Scalable depression monitoring with smartphone speech using a multimodal benchmark and topic analysis.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 41764298.
- Also identified by DOI 10.1038/s41746-026-02486-9 and PMC identifier 12996298.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Objective, scalable biomarkers are needed for continuous monitoring of major depressive disorder. Smartphone-collected speech is promising, yet clinically useful signals remain elusive. We analyzed 3151 weekly voice diaries from 284 German-speaking adults (128 MDD, 156 controls) to predict Beck Depression Inventory (BDI) scores. Sentence-embedding models outperformed lexical and acoustic baselines: Qwen3-8B achieved MAE 4.65 and R<sup>2</sup> 0.34, and stacked generalization of multilingual-E5 with Qwen3-8B further improved performance (MAE 4.37, R<sup>2</sup> 0.41). Audio embeddings added little incremental value. In an MDD-only analysis, multilingual-E5 was the top single modality (MAE 6.74, R<sup>2</sup> 0.20). To aid interpretation, BERTopic uncovered six coherent themes; BDI scores were highest for "Distress & care", supporting clinical face validity. Together, LLM embeddings paired with lightweight topic analysis capture the dominant signal of depression severity in everyday speech and offer a scalable route to ecologically valid digital phenotyping.