OpenEvidence errs on the safe side in a structured test of triage recommendations.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42673790.
- Also identified by DOI 10.1016/j.ijmedinf.2026.106687.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large language model (LLM) chatbots are increasingly consulted for triage decisions. A structured benchmark reported that ChatGPT Health, a consumer-facing health assistant, under-triaged 51.6 % of true emergencies and was susceptible to social anchoring. Whether a physician-facing clinical decision support platform fails similarly is unknown. To characterize the triage safety profile of OpenEvidence, a retrieval-augmented, physician-facing platform, using the identical benchmark previously applied to ChatGPT Health. We evaluated 60 clinician-authored vignettes from 30 clinical scenarios across 21 domains. Each scenario was written with and without objective clinical data and crossed with demographic and contextual modifiers in a 2 × 2 × 2 × 2 factorial design, yielding 960 prompts (480 clear-case, 480 edge-case). Responses were classified against a clinician gold standard as correct triage, under-triage, over-triage, or evidence-seeking refusal. Analyses used cluster bootstrap resampling, mixed-effects logistic regression, and Holm-Bonferroni correction. Among 449 clear-case responses that returned a recommendation, accuracy was 71.3 %. OpenEvidence under-triaged 12.5 % of emergency presentations versus 51.6 % in the previously reported ChatGPT Health benchmark, and over-triaged 68.0 % of nonurgent Home presentations (ChatGPT Health, 64.8 %). Anchoring statements did not alter recommendations (OR = 1.08, 95 % CI 0.62-1.88; Holm-adjusted p = 1.0). Objective clinical data eliminated emergency under-triage (25 % to 0 %; p = 0.005) and reduced nonurgent over-triage (78.7 % to 57.8 %; p = 0.014). In 65 of 960 responses (6.8 %), the platform declined to assign a triage level, exclusively in symptom-only Home or Routine prompts. Under this benchmark, OpenEvidence produced fewer missed emergencies than the historical ChatGPT Health comparison, while errors concentrated in over-triage and evidence-seeking refusal. These findings support evaluating health AI within its deployment context and treating refusal as a distinct output category whose clinical implications require separate assessment.