Artificial Intelligence vs Human Clinicians in Esophagogastroduodenoscopy Appropriateness: A Comparative Study Using Clinical Vignettes.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 41257517.
- Also identified by DOI 10.14309/ajg.0000000000003837.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large language models are increasingly used in clinical decision-making, but their role in appropriateness-based indications is uncertain. We conducted a vignette-based Italian survey comparing esophagogastroduodenoscopy appropriateness by 5 artificial intelligence (AI) (ChatGPT-4.0, ChatGPT-4.5, Gemini, Claude AI, OpenEvidence) at 2 times (April and September 2025) with gastroenterologists, residents, and general practitioners, using American Society for Gastrointestinal Endoscopy and European Society of Gastrointestinal Endoscopy guidelines. A total of 135 physicians participated. AI performance varied over time: accuracies ranged from 50%-90% in April to 63%-80% in September, with ChatGPT-4.5 and ChatGPT-4.0 outperforming physicians in September. Temporal and prompt variability highlight the need for multirun, longitudinal evaluation before clinical adoption.
Medical subject headings
- Artificial Intelligence
- Endoscopy, Digestive System
- Clinical Decision-Making