Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42479522.
- Also identified by DOI 10.1109/JBHI.2026.3715892.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Speech-based depression detection (SDD) offers a non-invasive and scalable complement to conventional clinical assessment, but reliable detection remains challenging because depression-related speech markers are subtle, heterogeneous, and irregularly distributed over time. Although self-supervised learning (SSL) speech models provide informative multi-layer representations, most SSL-based SDD methods either select a single layer or collapse all layers into a weighted sum. Such strategies merge functionally different SSL layers into a single stream, limiting the model's ability to capture interactions between low-level acoustic markers and higher-level context. To address this limitation, we propose HAREN CTC, a hierarchical framework that learns two complementary SSL representation streams and connects them through semantic-conditioned cross-attention. This design enables the model to emphasize acoustic patterns that are informative under the surrounding semantic context. To address the sparse and irregular temporal distribution of depression-related markers, we further introduce an auxiliary Connectionist Temporal Classification (CTC) objective that weakly aligns temporal token representations with HuBERT-derived pseudo-token targets. Experiments on DAIC-WOZ and MODMA show that HAREN-CTC out performs strong baselines under both fixed-split bench mark and subject-level cross-validation settings, achieving Macro F1 scores of 0.81 and 0.82, respectively, in the fixed split setting and maintaining consistent gains across most metrics under cross-validation. These results suggest that modeling acoustic-semantic interactions can improve the robustness of speech-based depression assessment.