Human versus machine: deciding on high-stakes surgery in possible Cauda Equina syndrome.
retrospective_cohort · Level III
Where this comes from
- Record sourced from PubMed, PMID 40348281.
- Also identified by DOI 10.1016/j.spinee.2025.05.026.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Cauda Equina Syndrome (CES) is a spine surgical urgency requiring prompt intervention to prevent neurological deficits. Accurate identification of CES cases needing urgent surgery is essential to avoid long-term sequelae. To evaluate the concordance between an AI language model (ChatGPT) and a Spinal Multidisciplinary Team (MDT) in recommending surgical intervention for suspected CES cases. Retrospective concordance analysis comparing surgical recommendations between ChatGPT and a Spinal MDT. Among 160 referrals presenting with red flags for possible CES, 10 cases were used to calibrate ChatGPT to specific clinical and diagnostic parameters, with the remaining 150 cases included in the primary analysis. The average patient age was 50.6 years (range 18-87), with a male-to-female ratio of 68:82. The primary outcome was the concordance rate between ChatGPT and the MDT in recommending surgery, evaluated through agreement rates and statistical analysis. Each of the 150 cases was presented as standardized slides including clinical history, imaging, and examination findings. Both the MDT and ChatGPT assessed the need for urgent surgery. Discordant cases (n=17) were further reviewed by 3 spinal surgeons blinded to prior decisions. ChatGPT and the MDT agreed on surgical recommendations in 133 out of 150 cases, achieving an 88.7% concordance (Cohen's Kappa = 0.764, p<.001). ChatGPT recommended surgery more frequently in the 17 discordant cases, but this difference was not statistically significant (McNemar's test statistic = 1.23, p=.46). Review by 3 independent surgeons reached consensus on 11 of the 17 discordant cases (64.7%), highlighting variability among experts; individual surgeons aligned with ChatGPT in 5 to 6 cases each (29.4%-35.3%). Substantial agreement between ChatGPT and the MDT suggests ChatGPT's comparable sensitivity in detecting surgical candidates in CES cases. Variability among surgeons on discordant cases underscores subjectivity in CES triage. ChatGPT may be a valuable adjunct in high-stakes clinical decision-making, though further validation and refinement are needed.
Medical subject headings
- Cauda Equina Syndrome
- Artificial Intelligence
- Clinical Decision-Making