Triage safety of patient-facing AI chatbots for nipple discharge: A guideline-informed assessment of red-flag recognition and patient actionability.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 42435613.
- Also identified by DOI 10.1016/j.ijmedinf.2026.106603.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
To evaluate red-flag recognition, clinical safety, and the quality of patient actionability in responses generated by artificial intelligence (AI) chatbots to patient questions about nipple discharge. This guideline-informed cross-sectional evaluation was conducted to assess the performance of AI chatbots in simulated nipple discharge consultations. A total of 36 English-language simulated patient questions were developed on the basis of clinical guidelines and real-world consultation scenarios, covering 6 modules of nipple discharge inquiries. Six commonly used AI chatbots responded to all questions independently, yielding 216 first-turn responses. Two reviewers independently evaluated the responses using a guideline-informed reference standard and structured scoring checklist, with disagreements adjudicated by a third senior reviewer. The primary outcome was the proportion of potentially misleading/unsafe responses, and secondary measures included the red-flag recognition rate, patient actionability score, guideline concordance score, total DISCERN score, and error types. Among 216 responses, 189 (87.5 %) were rated as safe, 19 (8.8 %) contained minor omissions, and 8 (3.7 %; 95 % confidence interval [CI], 1.4 %-6.9 %) met the primary outcome of potentially misleading/unsafe responses; notably, all were classified as potentially misleading, and no potentially unsafe responses were identified. The overall red-flag recognition rate was 90.6 % (1174 of 1296; 95 % CI, 87.8 %-93.2 %), and the median response-level recognition rate was 100.0 % (interquartile range [IQR], 85.7 %-100.0 %). Potentially misleading responses were most commonly attributable to missed red-flag features (4 of 8, 50.0 %) and poor actionability or insufficient action-oriented recommendations (3 of 8, 37.5 %). In exploratory analyses accounting for the repeated-response structure within the same questions, no clear statistical difference in the proportion of potentially misleading/unsafe responses was observed across models; conversely, between-model differences were observed in the red-flag recognition rate, patient actionability score, and total DISCERN score, although safety comparisons across models should be interpreted cautiously because of the low number of primary outcome events. In simulated patient consultations about nipple discharge, potentially misleading/unsafe responses from AI chatbots were uncommon, and most prespecified clinical warning features were recognized overall. Nevertheless, some responses demonstrated incomplete red-flag safety-netting and lacked specific recommendations for subsequent action. AI chatbots may serve as preliminary sources of health information; however, professional clinical evaluation cannot be replaced. Therefore, future patient-facing AI systems should incorporate structured red-flag screening and explicit triage recommendations to improve safety and practical utility in symptom-consultation contexts.