Linguistic Disparities in Artificial Intelligence-Generated Patient Education for Total Hip Arthroplasty: A Pilot Study of Cross-Language Analysis of Leading Large Language Models.

Ali, Usman; Tareen, Hafsa Khan; Pedroza, Juan Antonio; Kanza, Syeda; Amjad, Fatima; Khan, Saifullah; García, José Antonio; Hasan, Asher et al. · JB JS Open Access · 2026

cross_sectional · Level IV

Where this comes from

Abstract

Large Language Models (LLMs) are increasingly used for health information, but concerns exist regarding performance disparities for non-English speakers, potentially exacerbating health inequities. Appropriate information is critical for patients with limited English proficiency undergoing orthopedic procedures such as total hip arthroplasty (THA). This pilot study evaluated differences in the clinical reliability of English and Spanish responses to common THA questions generated by leading LLMs. Three widely accessible LLMs (ChatGPT-4o, Gemini 2.0 Flash, and Microsoft Copilot) were evaluated using 10 standardized frequently asked questions on THA, posed in English and Spanish. Responses were independently graded by language-fluent medical experts using a 4-point rubric (1 = Unsatisfactory to 4 = Excellent) assessing clinical reliability and appropriateness. Nonparametric statistics, including Wilcoxon signed-rank, Kruskal-Wallis, and effect sizes (Cliff's Delta, η<sup>2</sup>), were used for comparisons. A statistically significant main effect of language was found (p = 0.014, η<sup>2</sup> = 0.151), indicating significantly lower clinical reliability scores for Spanish responses in all LLMs. A nonsignificant within-model score decline was observed across all 3 LLMs. Leading LLMs exhibit significant difference in clinical reliability when providing THA information, performing less reliably in Spanish compared with English. This linguistic gap suggests a potential risk for difference in response interpretation and could potentially worsen health inequities for Spanish-speaking populations. Efforts are needed to improve multilingual capabilities and manage biases in medical artificial intelligence (AI). Clinicians and patients should exercise caution when using LLMs for health information in languages other than English until cross-lingual reliability is demonstrably improved. This study highlights a significant linguistic disparity in AI-generated health information for THA. Improving LLMs' multilingual capabilities is essential to promote equitable access to reliable medical education and prevent the exacerbation of health inequities for non-English speaking patients. Level IV. See Instructions for Authors for a complete description of levels of evidence. This study evaluates LLMs in providing THA information in English and Spanish, revealing that Spanish responses are clinically less reliable. The findings highlight linguistic gap in AI healthcare tools, raising potential concerns for patient safety, and widening health inequities for non-English speakers.

Anatomy