Feelings behind words: A systematic review on how effective IS NLP-based assessment for mental health diagnosis in human studies.

Yulianti, Eka Putri; Eka Putri, Yossie Susanti; Keliat, Budi Anna; Hidayanto, Achmad Nizar · Int J Med Inform · 2026

systematic_review · Level I

Where this comes from

Abstract

Natural Language Processing (NLP)-based AI tools are increasingly used in mental health diagnostics, yet their real-world performance compared to traditional methods remains unclear. This systematic review evaluates the diagnostic accuracy, feasibility, and limitations of NLP tools for mental health assessment in human studies. Seventeen studies (2020-2025) were systematically reviewed across seven databases from February to March 2025. Key metrics included AUC (0.77-0.92), sensitivity (0.72-0.92), and specificity (0.68-0.96). Comparators included clinician evaluations with scales (59%) and standardised self-rated instruments (35%). LLMs outperformed traditional ML models in depression detection. Spoken and interactive models offered excellent performance. Mixed performance (23.5%) was observed in non-English contexts and asynchronous text analysis. Poor performance was observed in the stratification model. Studies predominantly used case-control designs (70.6%), with 94% originating from middle-to-high-income countries. NLP tools reduced clinical workload but struggled with cultural adaptability. NLP-based AI holds promise for mental health diagnostics in clinical settings, but it requires interactive models, clear clinical targets, diverse real-world data, broad cohort inclusion, balanced and standardised validation, and culturally adaptive models. Future research should prioritise rigorous randomised controlled trials (RCTs) to mitigate biases in AI-driven mental health care.

Medical subject headings