How accurately do large language models answer patient questions on anterior cruciate ligament tears? A comparative study.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 41253635.
- Also identified by DOI 10.1016/j.knee.2025.11.002.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large language models (LLMs) are increasingly used in the medical sector, raising questions about their reliability for patient education. With more LLMs becoming publicly available, it remains unclear whether meaningful performance differences exist between them. This is particularly relevant for anterior cruciate ligament (ACL) injuries, which mainly affect young, active individuals, those most likely to seek health advice from AI. This study aimed to evaluate and directly compare the accuracy of five leading LLMs in answering common patient questions about ACL tears. Fourteen commonly asked patient questions were identified in a systematic online search. Each question was submitted to five LLMs: ChatGPT-4, Gemini 2.0, Llama 3.1, DeepSeek-V3, and Grok3. Responses were assessed for accuracy by orthopedic consultants using a five-point Likert scale. Word count was recorded as a proxy for readability. Statistical analysis included ANOVA by Tukey's HSD post hoc test. All models achieved mean accuracy scores ≥3 (mostly accurate). DeepSeek (3.61) and Grok (3.59) demonstrated significantly higher mean accuracies than Llama (3.25; P < 0.05). ChatGPT and Gemini achieved mean scores of 3.48 and 3.52, respectively. Models generating longer responses, such as Grok and DeepSeek, tended to offer greater accuracy, whereas Llama produced the shortest and least accurate answers. All tested LLMs show promise for patient education regarding ACL injuries, but notable performance differences exist. Model choice is therefore critical. While all responses were evaluated by clinical experts, the lack of guideline-based validation highlights the need for further studies assessing both accuracy and patient comprehension.
Medical subject headings
- Anterior Cruciate Ligament Injuries
- Language
- Patient Education as Topic