Artificial Intelligence vs Human Clinicians in Esophagogastroduodenoscopy Appropriateness: A Comparative Study Using Clinical Vignettes.

Toppeta, Angelica; Corradi, Mattia; Mantia, Beatrice; Randazzo, Adelaide Maria; Schettino, Mario; De Lisi, Stefania; Dell'Era, Alessandra; Carmagnola, Stefania et al. · Am J Gastroenterol · 2026

cross_sectional · Level IV

Where this comes from

Abstract

Large language models are increasingly used in clinical decision-making, but their role in appropriateness-based indications is uncertain. We conducted a vignette-based Italian survey comparing esophagogastroduodenoscopy appropriateness by 5 artificial intelligence (AI) (ChatGPT-4.0, ChatGPT-4.5, Gemini, Claude AI, OpenEvidence) at 2 times (April and September 2025) with gastroenterologists, residents, and general practitioners, using American Society for Gastrointestinal Endoscopy and European Society of Gastrointestinal Endoscopy guidelines. A total of 135 physicians participated. AI performance varied over time: accuracies ranged from 50%-90% in April to 63%-80% in September, with ChatGPT-4.5 and ChatGPT-4.0 outperforming physicians in September. Temporal and prompt variability highlight the need for multirun, longitudinal evaluation before clinical adoption.

Medical subject headings