The VOICE-DEP study protocol: Multimodal analysis of voice and speech during medical interviews to support diagnosis and longitudinal monitoring of major depressive disorder.

Zabalza-Zudaire, Maialen; Sayar-Beristain, Onintza; Fructos, Paula; Núñez, Favio Emanuel; Carpio, Franco Fabricio; García, Esmeralda; Ortiz, Álvaro; Ortuño, Felipe et al. · PLoS One · 2026

prospective_cohort · Level II

Where this comes from

Abstract

Major depressive disorder is a severe, recurrent and disabling condition. Although diagnosis and clinical monitoring are based on medical interviews and validated rating scales, voice and speech analysis may provide complementary digital biomarkers reflecting depressive severity and clinical evolution. However, current evidence remains limited by methodological heterogeneity, predominantly cross-sectional designs, limited longitudinal data and underrepresentation of non-English-speaking clinical populations. The aim of the VOICE-DEP study is to develop and formalize a standardized, reproducible and clinically grounded protocol for the multimodal analysis of voice and speech during medical interviews as a tool to support the diagnosis of depressive disorder and to assess whether speech-derived digital biomarkers change over time in parallel with clinical severity measures. VOICE-DEP is an observational, prospective, longitudinal pilot study of patients with major depressive disorder with a healthy control group, conducted in a hospital-based clinical setting in Spain. The study will include 25 adult patients with moderate or severe unipolar depression, with or without psychotic symptoms, and 50 healthy controls without a personal history of psychiatric disorders. Patients will be assessed at five time points: baseline (V0) and four monthly follow-up visits at 30, 60, 90 and 120 days. Healthy controls will be assessed once at baseline. The planned dataset comprises 175 voice recordings: 125 from patients and 50 from controls. At each assessment, the Montgomery-Asberg Depression Rating Scale related part of the medical interview, lasting approximately 10-30 minutes and including an initial free-speech segment, will be recorded using a standardized audio protocol. Acoustic (e.g., pitch, intensity), paralinguistic (e.g., speed rate, prosodic range) and linguistic (e.g., sentiment polarity, lexical diversity) features will be extracted and analyzed in relation to clinician-rated severity measures and self-reported symptoms. This protocol is expected to generate a clinically grounded Spanish-language longitudinal speech corpus and a transparent analytical framework for evaluating voice- and speech-derived digital biomarkers as complementary tools for depression assessment and monitoring.

Medical subject headings