Feasibility and user evaluation of HopeBot: An LLM-powered conversational chatbot for depression screening.
prospective_cohort · Level II
Where this comes from
- Record sourced from PubMed, PMID 42348542.
- Also identified by DOI 10.1371/journal.pdig.0001446.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Clinician-administered Patient Health Questionnaire-9 (PHQ-9) interviews allow clarification of ambiguous responses but are resource-intensive and difficult to scale for routine use. Self-administered versions are widely adopted for depression screening yet offer little opportunity for interaction or clarification, which may limit engagement and scoring accuracy. We developed HopeBot, a conversational chatbot powered by a large language model (LLM) that delivers the PHQ-9 via text or voice, providing real-time clarification and safety guidance through a retrieval-augmented generation (RAG) layer drawing on validated psychological and helpline resources. The system aims to extend access to structured screening rather than replace clinician judgment. In a within-subject feasibility study, 132 adults from two countries completed both self-administered and chatbot-assisted PHQ-9 assessments followed by a 25-item evaluation survey. Chatbot and self-report scores showed high concordance (intraclass correlation coefficient = 0.92; median absolute difference = 1 point), indicating faithful replication of scoring. Of these participants, 75 completed a comparative feedback module (most with identical scores were not prompted for comparison); 71% (n = 53) reported greater confidence in chatbot-assisted scores, citing clearer structure, interpretive guidance, and a supportive tone. Mean usability ratings (0-10) were 8.4 for comfort, 7.7 for voice clarity, 7.6 for handling sensitive topics, and 7.4 for recommendation helpfulness; the latter varied significantly by employment status and prior experience with mental-health services. The self-administered PHQ-9 served as a pragmatic comparator reflecting real-world digital screening, allowing evaluation of whether the chatbot could faithfully reproduce the questionnaire's structure and scoring accuracy rather than diagnostic validity. These findings indicate that an LLM-powered conversational agent with RAG-based grounding can feasibly and acceptably administer the PHQ-9 with strong score concordance relative to self-report, suggesting potential as a scalable adjunct for early, low-burden depression screening. Further validation against clinician-administered assessments in real-world workflows is warranted. Trial registration ClinicalTrials.gov Identifier NCT06801925.