Linguistic Disparities in Artificial Intelligence-Generated Patient Education for Total Hip Arthroplasty: A Pilot Study of Cross-Language Analysis of Leading Large Language Models.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 41589271.
- Also identified by DOI 10.2106/JBJS.OA.25.00207 and PMC identifier 12826221.
- Licence recorded as CC BY-NC-ND.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large Language Models (LLMs) are increasingly used for health information, but concerns exist regarding performance disparities for non-English speakers, potentially exacerbating health inequities. Appropriate information is critical for patients with limited English proficiency undergoing orthopedic procedures such as total hip arthroplasty (THA). This pilot study evaluated differences in the clinical reliability of English and Spanish responses to common THA questions generated by leading LLMs. Three widely accessible LLMs (ChatGPT-4o, Gemini 2.0 Flash, and Microsoft Copilot) were evaluated using 10 standardized frequently asked questions on THA, posed in English and Spanish. Responses were independently graded by language-fluent medical experts using a 4-point rubric (1 = Unsatisfactory to 4 = Excellent) assessing clinical reliability and appropriateness. Nonparametric statistics, including Wilcoxon signed-rank, Kruskal-Wallis, and effect sizes (Cliff's Delta, η<sup>2</sup>), were used for comparisons. A statistically significant main effect of language was found (p = 0.014, η<sup>2</sup> = 0.151), indicating significantly lower clinical reliability scores for Spanish responses in all LLMs. A nonsignificant within-model score decline was observed across all 3 LLMs. Leading LLMs exhibit significant difference in clinical reliability when providing THA information, performing less reliably in Spanish compared with English. This linguistic gap suggests a potential risk for difference in response interpretation and could potentially worsen health inequities for Spanish-speaking populations. Efforts are needed to improve multilingual capabilities and manage biases in medical artificial intelligence (AI). Clinicians and patients should exercise caution when using LLMs for health information in languages other than English until cross-lingual reliability is demonstrably improved. This study highlights a significant linguistic disparity in AI-generated health information for THA. Improving LLMs' multilingual capabilities is essential to promote equitable access to reliable medical education and prevent the exacerbation of health inequities for non-English speaking patients. Level IV. See Instructions for Authors for a complete description of levels of evidence. This study evaluates LLMs in providing THA information in English and Spanish, revealing that Spanish responses are clinically less reliable. The findings highlight linguistic gap in AI healthcare tools, raising potential concerns for patient safety, and widening health inequities for non-English speakers.
Anatomy
- hip