MBDA: A modality-balanced framework with data augmentation and alignment for multimodal emotion recognition.

Cheng, Cheng; Shang, Ruisi; Wang, Zixu; Li, Huazhi; Jia, Ziyu · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

Multimodal Emotion Recognition (MER) aims to infer human emotional states by integrating complementary information from heterogeneous modalities. However, existing MER methods often suffer from modality imbalance, cross-modal misalignment, and limited data diversity, which hinder their robustness and generalization. To address these issues, we propose a Modality-Balanced framework with Data Augmentation and Alignment (MBDA), which integrates modality-aware augmentation, feature alignment, and counterfactual knowledge distillation into a unified framework in a progressive learning manner. MBDA boosts data diversity while preserving semantic consistency through modality-aware augmentation, enforces robust multi-level alignment across modalities, and adaptively rebalances modality contributions through counterfactual knowledge distillation. Experiments on the DEAP and SEED-IV datasets demonstrate that MBDA consistently outperforms state-of-the-art methods, achieving accuracies of 93.86%, 95.11%, 91.02%, and 92.66% on DEAP-A, DEAP-V, DEAP-AV, and SEED-IV, respectively.