When multimodal fusion helps: An ablation study of EEG-ECG fusion strategies for emotion recognition.
Journal:
Computers in biology and medicine
Published Date:
Aug 28, 2026
Abstract
Multimodal physiological signal fusion-particularly electroencephalography (EEG) and electrocardiography (ECG)-is widely assumed to improve emotion recognition over unimodal approaches. Yet the conditions under which fusion genuinely helps, and which fusion architecture is most robust, remain poorly characterised. We present a large-scale ablation study comparing five fusion strategies (EEG-only, ECG-only, concatenation, bidirectional cross-attention, and adaptive gated fusion) on the emotion datasets DREAMER and AMIGOS, and validate ECG signal-quality effects on the stress dataset WESAD. Using a lightweight deep learning architecture (PMMR-DL v3, ∼1.3M parameters) and leave-one-subject-out cross-validation (LOSO-CV) repeated over three random seeds, we evaluate macro F1-score and accuracy for binary emotion classification. Our results reveal pronounced dataset-dependent modality dominance: ECG dominates on DREAMER (F1 = 42.1% vs. EEG 28.6%), whereas EEG dominates on AMIGOS (F1 = 63.8% vs. ECG 13.8%). Critically, multimodal fusion does not consistently outperform the best unimodal baseline. On DREAMER, the strongest fusion strategy (adaptive gated, F1 = 37.1%) still underperforms ECG-only. On AMIGOS, all fusion conditions collapse toward the poor ECG performance (F1 17-21%) because the ECG recordings are noise-dominated. The adaptive gated strategy achieves the lowest inter-seed variance (std = 0.5 on DREAMER), making it the most reliable fusion mechanism when signal quality is adequate. These findings caution against blanket assumptions that multimodal fusion is universally beneficial, and provide practical guidance for modality selection in affective computing pipelines.
Authors
Keywords
No keywords available for this article.