Can AI safely choose antibiotics over the knife? A STROBE-guided benchmark of GPT-4, GPT-5, and Gemini for non-operative acute appendicitis management.

Calışkan, Yahya Kemal; Başak, Fatih; Erdem, Olgun · Int J Med Inform · 2026

cross_sectional · Level IV

Where this comes from

Abstract

Non-operative management (NOM) with antibiotics is increasingly used for imaging-confirmed uncomplicated acute appendicitis (AA), but appropriate selection remains clinically delicate-especially in "gray-zone" scenarios such as appendicolith, older age, or borderline imaging findings. As large language models (LLMs) are increasingly queried for clinical guidance by both patients and clinicians, their reliability in distinguishing NOM candidates from patients requiring urgent operative care warrants formal evaluation. This question is clinically relevant not only for academic benchmarking but also because generative AI tools are already being used in health-information seeking and decision-support contexts, where overconfident but unsafe triage advice could influence real-world care pathways. We conducted a cross-sectional, in-silico benchmarking study using 50 standardized clinical vignettes spanning uncomplicated AA (n = 20), complicated AA (n = 20), and atypical/high-risk presentations (n = 10; older age, pregnancy, immunocompromise). Three LLMs (GPT-4, GPT-5, Gemini) were queried with a uniform zero-shot prompt requesting NOM candidacy determination and structured risk-benefit communication. Two blinded surgeons scored outputs against predefined criteria anchored to international guideline principles. The primary outcome was Management Accuracy Score (MAS; correct/incorrect). Secondary outcomes included risk-stratification nuance, safety warnings, and evidence-use quality (Likert 1-5). Fleiss' kappa quantified inter-model agreement. The vignette set was intentionally balanced to stress-test models across routine and safety-critical scenarios rather than to estimate population prevalence. Overall MAS differed across models (χ2 = 6.34, p = 0.042): GPT-5 92% (46/50), GPT-4 84% (42/50), and Gemini 76% (38/50). The widest gap occurred in complicated AA (p = 0.028), driven by Gemini's over-recommendation of antibiotic "trials" in appendicolith-positive or abscess-suggestive scenarios. GPT-5 generated the most consistent recurrence counseling and safety framing; no critical safety failures were observed. All three models performed less reliably in atypical/high-risk vignettes; however, because this subgroup contained only 10 cases, these findings should be interpreted cautiously and not over-read as evidence of true between-model equivalence or superiority. LLMs demonstrate strong baseline awareness of contemporary AA strategies, yet clinically meaningful variability persists-particularly for contraindications to NOM. GPT-5 performed best overall, while Gemini's over-generalization in high-risk contexts highlights the need for domain-constrained training and guardrails before clinical integration. These findings support supervised, educational use at most, rather than autonomous emergency triage deployment.

Medical subject headings