Artificial Intelligence in Patient Education: A Comparative Evaluation of Chatbot Reliability After Total Hip Arthroplasty.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42296472.
- Also identified by DOI 10.1097/PHM.0000000000003061.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
This study aimed to evaluate and compare the responses of different artificial intelligence-based large language models (AI LLMs) to frequently asked patient questions regarding post-operative care following total hip arthroplasty (THA) across four domains: reliability, quality, accuracy, and readability. Twenty-seven commonly asked questions were identified through Google Trends and expert consensus, covering exercises, activities of daily living, and dislocation precautions. Responses generated by the three AI models (May 2-5, 2025) were assessed using the modified DISCERN scale, Global Quality Scale (GQS), Accuracy Scale, and Flesch Reading Ease Score (FRES). Inter-rater reliability was highest for Gemini (ICC = 0.71), while ChatGPT models demonstrated moderate-to-good reliability (ICC = 0.67-0.71). Gemini scored significantly higher in reliability (P < 0.001), whereas ChatGPT 3.5 achieved the greatest readability (median FRES = 55). Significant differences were observed among models across reliability, quality, and readability metrics. AI LLMs can serve as supplementary tools for patient education following THA; however, their use requires clinician oversight to ensure accuracy and safety. Future studies should explore patient-centered evaluations and hybrid approaches combining AI guidance with professional supervision in orthopedic rehabilitation.