Guideline-based, but not error-free: Multilingual risks in AI-powered patient counseling on gallstones.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 41689953.
- Also identified by DOI 10.1016/j.ijmedinf.2026.106341.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Patients increasingly use large language models (LLMs) for health information, yet the guideline concordance and safety of patient-facing outputs-particularly across languages-remain uncertain. We evaluated three widely used LLM platforms (web interfaces) and their underlying default models for gallstone-related counseling in Turkish and English. In this cross-sectional content analysis, 14 real-world, guideline-mappable patient questions were developed in Turkish and translated into semantically equivalent English. Each question was submitted once to ChatGPT (ChatGPT-4o mini), Gemini (Gemini 3-flash), and Perplexity (Sonar family; default free-tier routing at the time of testing) in both languages under standardized conditions, yielding 84 responses. Two blinded hepatobiliary surgeons independently rated each response using a prespecified 3-point guideline concordance scale (0-2) mapped to EASL 2016 gallstone guidelines and Tokyo Guidelines 2018 for acute cholecystitis; disagreements were adjudicated by a third surgeon. Within-model language differences were assessed with Wilcoxon signed-rank tests; between-model comparisons used Friedman tests. Full correctness (score = 2) was analyzed using Cochran's Q with McNemar post-hoc tests. Error types and response length were also examined. In English, model performance differed significantly, with ChatGPT and Gemini outperforming Perplexity (p < 0.01), while Turkish differences were not statistically significant. ChatGPT performed better in English than Turkish (p = 0.008). Error profiles were language-dependent: Turkish outputs more often showed under-explanation, whereas English outputs more frequently amplified risk. Perplexity demonstrated the highest overall error burden. . LLM responses to gallstone questions are often guideline-aligned but remain model- and language-sensitive, with clinically relevant safety risks. Multilingual evaluation standards are needed, and unsupervised reliance on LLMs for patient guidance-especially in low-resource languages-should be discouraged.
Medical subject headings
- Gallstones
- Artificial Intelligence
- Counseling
- Multilingualism
- Practice Guidelines as Topic