Prospective evaluation of a large language model clinical decision support system in the emergency department.
prospective_cohort · Level II
Where this comes from
- Record sourced from PubMed, PMID 42618632.
- Also identified by DOI 10.1038/s41591-026-04601-5.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Prospective evidence for artificial intelligence (AI)-based clinical decision support in emergency departments remains limited. Here we conducted a DECIDE-AI stage 1 evaluation of SHAKED, a clinical decision support system built on multiple large language models, in a tertiary emergency department. Over 4 weeks, 1,138 patients were analyzed across two parallel units-one using SHAKED and one following routine rotations. Clinical adoption of SHAKED declined from 68% to 30%, owing to workload-sensitive disengagement (OR = 0.72 per shift hour, 95% CI 0.62 to 0.83). Physicians preferred the use of SHAKED for radiology consultations (OR = 2.98, 95% CI 1.58 to 5.63). No adverse events were detected, and expert review rated 99 of 100 sampled outputs as clinically appropriate. Emergency department length of stay did not differ between wings (4.9 h in both, P = 0.99). Intention-to-treat analysis showed a non-significant trend toward shorter consultation cycle time (-9.4 min, P = 0.077). These findings suggest that sustained clinician engagement, rather than algorithmic accuracy, may be the key barrier to effective clinical AI use in emergency departments. They inform randomized trial design but do not justify clinical deployment of AI clinical decision support at this stage. ClinicalTrials.gov identifier: NCT06902675 .