On-premises open-source large language models for privacy-preserving multimodal depression screening.
Where this comes from
- Record sourced from PubMed, PMID 42379124.
- Also identified by DOI 10.1016/j.ijmedinf.2026.106577.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Narrative and speech data can provide valuable signals for depression screening, yet privacy and data-governance requirements often limit the use of closed-based models in clinical practice. In addition, existing large language model (LLM)-based approaches are largely text-centric, and multimodal integration of acoustic features and structured clinical variables remains limited. This study aimed to develop and externally validate a privacy-preserving multimodal depression screening prediction framework using open-source large language models that integrate sociodemographic information, emotion-memory narratives, and acoustic features. This study analyzed 3536 participants collected at Chonnam National University Hospital. Inputs combined sociodemographic and lifestyle variables, Korean transcripts of happy- and sad-memory narratives, and speech-derived extended Geneva minimalistic acoustic parameter set (eGeMAPS) features. To maintain prompt conciseness, statistically significant features were selected from the 88 eGeMAPS features extracted for each happy- and sad-memory narrative, with Mann-Whitney U tests conducted exclusively on the internal training split to prevent data leakage. Five open-source LLMs (Gemma-3-27B, Qwen-3-32B, Llama-3.3-70B, Phi4-14B, and gpt-oss-20b) were evaluated under zero-shot prompting, Chain-of-Thought prompting, and supervised fine-tuning. External validation used Extended Distress Analysis Interview Corpus (E-DAIC) (N = 275). Under zero-shot prompting, the best internal F1-score was 0.735 (Gemma-3-27B). Chain-of-Thought prompting improved Llama-3.3-70B (F1-score = 0.708) but reduced performance for other models. Supervised fine-tuning improved all models, yielding internal accuracies of 0.852 to 0.881 and F1-scores of 0.818 to 0.865 across five models, corresponding to F1 gains of 0.12 to 0.30 versus prompting-only approaches. In external validation, accuracy ranged from 0.764 to 0.822 and F1-score ranged from 0.683 to 0.807. This study suggests that multimodal open-source LLMs integrating clinical variables, narrative text, and acoustic features can support privacy-preserving depression screening in an on-premises setting. Supervised fine-tuning provided the most consistent performance improvements, and external validation supported robustness beyond the development cohort.