Real-world use and evaluation of a generative AI chatbot for Parkinson's disease information: a prospective observational study.

Lange, Florian; Mardi, Ssaman; Binder, Tobias; Reich, Martin M; Odorfer, Thorsten; Volkmann, Jens · Lancet Reg Health Eur · 2026

prospective_cohort · Level II

Where this comes from

Abstract

Patient-facing medical AI chatbots are entering clinical use faster than evidence can characterise their real-world safety. We evaluated one deployed system and introduce CARE-LLM, a proposed framework for conversation-level post-market surveillance. We conducted a prospective, observational, conversation-level evaluation of jAImes, a retrieval-augmented AI information system for Parkinson's disease deployed publicly in Germany, across its first 129 days of operation (Nov 11, 2025-Mar 20, 2026), analysing all 2035 conversations (6146 messages). CARE-LLM (Conversation-level AI Real-world Evaluation) combines automated triage of every conversation, structured expert review of flagged cases, sampling-based sensitivity validation against the adequate class, and a failure-class feedback loop. AI-assisted triage classified 1803/2035 conversations (88·6%) as good, 224/2035 (11·0%) as partially adequate, and 8/2035 (0·4%) as inadequate. Expert review of 45 flagged conversations confirmed five critical events. Independent re-review of 100 randomly sampled 'good'-rated conversations identified four (4/100, 4%; 95% CI 1·6-9·8) with clinically critical errors missed by the triage-a conditional false-negative rate within the 'good' stratum, not an overall critical-event rate. Confirmed critical events-five flagged, four sampled-spanned three failure classes (knowledge boundary, robustness, and escalation failures) and included an inappropriate memantine recommendation in Parkinson's disease dementia and explicit suicidal ideation that did not trigger the intended emergency response. Tightly scoped, retrieval-augmented, patient-facing AI achieved favourable automated triage classifications while still harbouring clinically critical failures invisible to automated quality metrics alone; detecting them required conversation-level evaluation with clinical expert adjudication. Deutsche Forschungsgemeinschaft, Interdisciplinary Center for Clinical Research Würzburg, and German Ministry of Education and Research (BMBF); technical development of jAImes was commissioned and financed by Parkinson Stiftung Deutschland.