Efficacy of large language models in detecting postoperative delirium from unstructured clinical notes: A retrospective cohort study.
retrospective_cohort · Level III
Where this comes from
- Record sourced from PubMed, PMID 41388138.
- Also identified by DOI 10.1038/s41746-025-02231-8 and PMC identifier 12816606.
- Licence recorded as CC BY-NC-ND.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Early identification of postoperative delirium (POD) remains challenging. This retrospective observational study compared the performance of large language models (LLMs), Llama-3-70B and GPT-4o, and physicians in predicting clinically significant POD, defined as either requiring antipsychotics or diagnosis of delirium by neurologists following consultation for delirium-related symptoms. The c-statistics of Llama-3-70B and GPT-4o were 0.74 and 0.76, respectively. LLMs showed higher sensitivity (Llama-3-70B, 0.900; GPT-4o, 0.868; physicians, 0.723) and lower specificity (0.463, 0.547, and 0.814, respectively) than physicians. Inter-rater agreement was almost perfect for both Llama-3-70B and GPT-4o (Fleiss' kappa = 0.852 and 0.854, respectively) but fair for physicians (0.219). Both LLMs detected clinically significant POD approximately one day earlier than physicians (Kaplan-Meier analysis, median time to diagnosis: Llama-3-70B, 34.5 h; GPT-4o, 37.5 h; physicians, 62.9 h; log-rank P < 0.001). The integration of LLMs as a complementary screening tool under physician supervision may improve the early, reproducible diagnosis of clinically significant POD.