ContextVecNet: A Context-Driven Multimodal Learning Framework for Depression Detection.

Tahir, Waleed Bin; Khalid, Shah; Alshahrani, Saied; Alharbi, Shuaa S; Alhasson, Haifa F · IEEE J Biomed Health Inform · 2025

basic_science · Level V

Where this comes from

Abstract

Depression is a major global health issue, and social media offers a rich source of data for early detection. However, existing multimodal approaches often fail to effectively capture the contextual relationships between text and images and tend to ignore temporal information, thereby limiting their predictive accuracy and reliability. To address these limitations, we propose ContextVecNet, a multimodal deep learning framework. It is built on a CLIP-based architecture enhanced by learnable context vectors, which are integrated into the text and image encoding branches. This design allows the model to dynamically learn and adapt to the specific visual and linguistic markers of depression present in social media data. The architecture further incorporates a cross-modal transformer with time-aware embeddings, which jointly models temporal dynamics and cross-modal interactions across user posts. Evaluated on a publicly available multimodal Twitter dataset, ContextVecNet achieves state-of-the-art performance, with an AUC of 0.9922 and an F1-score of 0.9619, significantly outperforming existing methods. An ablation study confirms that freezing the context vectors results in notable performance degradation, emphasizing their importance in depression detection.