Guideline-based, but not error-free: Multilingual risks in AI-powered patient counseling on gallstones.

Erdem, Olgun; Canbak, Tolga; Acar, Aylin; Ceylan, Emine Maksude; Çakıt, Hakan; Başak, Fatih · Int J Med Inform · 2026

cross_sectional · Level IV

Where this comes from

Abstract

Patients increasingly use large language models (LLMs) for health information, yet the guideline concordance and safety of patient-facing outputs-particularly across languages-remain uncertain. We evaluated three widely used LLM platforms (web interfaces) and their underlying default models for gallstone-related counseling in Turkish and English. In this cross-sectional content analysis, 14 real-world, guideline-mappable patient questions were developed in Turkish and translated into semantically equivalent English. Each question was submitted once to ChatGPT (ChatGPT-4o mini), Gemini (Gemini 3-flash), and Perplexity (Sonar family; default free-tier routing at the time of testing) in both languages under standardized conditions, yielding 84 responses. Two blinded hepatobiliary surgeons independently rated each response using a prespecified 3-point guideline concordance scale (0-2) mapped to EASL 2016 gallstone guidelines and Tokyo Guidelines 2018 for acute cholecystitis; disagreements were adjudicated by a third surgeon. Within-model language differences were assessed with Wilcoxon signed-rank tests; between-model comparisons used Friedman tests. Full correctness (score = 2) was analyzed using Cochran's Q with McNemar post-hoc tests. Error types and response length were also examined. In English, model performance differed significantly, with ChatGPT and Gemini outperforming Perplexity (p < 0.01), while Turkish differences were not statistically significant. ChatGPT performed better in English than Turkish (p = 0.008). Error profiles were language-dependent: Turkish outputs more often showed under-explanation, whereas English outputs more frequently amplified risk. Perplexity demonstrated the highest overall error burden. . LLM responses to gallstone questions are often guideline-aligned but remain model- and language-sensitive, with clinically relevant safety risks. Multilingual evaluation standards are needed, and unsupervised reliance on LLMs for patient guidance-especially in low-resource languages-should be discouraged.

Medical subject headings