Dis4DR: Disentangled Affective and Biometric Features for Multimodal Depression Recognition.

Pan, Yuchen; Shen, Wuxin; Yao, Hongxun · IEEE J Biomed Health Inform · 2026

basic_science · Level V

Where this comes from

Abstract

Depression recognition (DR) through facial images, audio signals, and text recordings has made significant progress. Recently, multimodal DR has shown advantages over single-modal approaches by integrating information from multiple sources. However, identity and gender information are inherently present in these signals alongside depression-related features, which can introduce ambiguity in the decision-making process. Additionally, collecting high-quality multimodal data is challenging, and existing methods often experience performance degradation when certain modalities are missing or incomplete. To tackle these challenges, we introduce Feature Disentanglement Network for DR, abbreviated as Dis4DR, a multimodal framework that combines feature disentanglement and privileged knowledge distillation to enhance DR. Our approach separates homogeneous and heterogeneous features within multimodal signals to better represent depression disorders. Moreover, identity and gender features are disentangled alongside depression-related features, allowing the model to selectively focus on the most relevant components while reducing the influence of unrelated information. To improve performance when input modalities are incomplete, we utilize knowledge distillation to transfer privileged knowledge from complete multimodal inputs to degraded scenarios, enhancing model robustness and generalization. We validate Dis4DR through experiments on the AVEC 2013, AVEC 2014, AVEC 2017, and AVEC 2019 datasets. Results demonstrate that Dis4DR consistently outperforms existing approaches, even when only a single modality is available. Furthermore, it surpasses its predecessor, Dis2DR, achieving up to a 3.27% relative reduction in prediction error.