On-premises open-source large language models for privacy-preserving multimodal depression screening.

Kwon, Soonjun; Kim, Yihyun; Jhon, Min; Park, Jin-Hyun; Lim, Bahngtaik; Jeon, Eunkyoung; Kim, Jae-Min; Kim, Ju-Wan et al. · Int J Med Inform · 2026

Where this comes from

Abstract

Narrative and speech data can provide valuable signals for depression screening, yet privacy and data-governance requirements often limit the use of closed-based models in clinical practice. In addition, existing large language model (LLM)-based approaches are largely text-centric, and multimodal integration of acoustic features and structured clinical variables remains limited. This study aimed to develop and externally validate a privacy-preserving multimodal depression screening prediction framework using open-source large language models that integrate sociodemographic information, emotion-memory narratives, and acoustic features. This study analyzed 3536 participants collected at Chonnam National University Hospital. Inputs combined sociodemographic and lifestyle variables, Korean transcripts of happy- and sad-memory narratives, and speech-derived extended Geneva minimalistic acoustic parameter set (eGeMAPS) features. To maintain prompt conciseness, statistically significant features were selected from the 88 eGeMAPS features extracted for each happy- and sad-memory narrative, with Mann-Whitney U tests conducted exclusively on the internal training split to prevent data leakage. Five open-source LLMs (Gemma-3-27B, Qwen-3-32B, Llama-3.3-70B, Phi4-14B, and gpt-oss-20b) were evaluated under zero-shot prompting, Chain-of-Thought prompting, and supervised fine-tuning. External validation used Extended Distress Analysis Interview Corpus (E-DAIC) (N = 275). Under zero-shot prompting, the best internal F1-score was 0.735 (Gemma-3-27B). Chain-of-Thought prompting improved Llama-3.3-70B (F1-score = 0.708) but reduced performance for other models. Supervised fine-tuning improved all models, yielding internal accuracies of 0.852 to 0.881 and F1-scores of 0.818 to 0.865 across five models, corresponding to F1 gains of 0.12 to 0.30 versus prompting-only approaches. In external validation, accuracy ranged from 0.764 to 0.822 and F1-score ranged from 0.683 to 0.807. This study suggests that multimodal open-source LLMs integrating clinical variables, narrative text, and acoustic features can support privacy-preserving depression screening in an on-premises setting. Supervised fine-tuning provided the most consistent performance improvements, and external validation supported robustness beyond the development cohort.