ChatGPT provides high-quality answers to FAQs about high tibial osteotomy despite low inter-rater agreement.
case_series · Level V
Where this comes from
- Record sourced from PubMed, PMID 41323534.
- Also identified by DOI 10.1002/jeo2.70521 and PMC identifier 12661211.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
High tibial osteotomy (HTO) is frequently used to treat knee malalignment in younger patients. Given the rise in online health information-seeking behaviour, this study aimed to evaluate the quality of ChatGPT-generated responses to frequently asked questions (FAQs) about HTO and to assess the reliability of two scoring systems used by orthopaedic surgeons. Twenty-four FAQs were submitted to ChatGPT (GPT-4-turbo). Four orthopaedic surgeons independently rated the responses at two time points using: (1) a 4-point categorical scale (1 = excellent, 4 = poor), and (2) a 100-point numerical scale (0 = worst, 100 = best). Intra-observer reliability was assessed using weighted kappa (<i>κ</i>) and intraclass correlation coefficients (ICC); inter-observer agreement was measured using ICC values. Most responses were rated positively, with over 70% considered 'excellent' or requiring minimal clarification. Intra-observer agreement was variable, ranging from <i>κ</i> = 0.333 to 0.864 and ICC = 0.690-0.922. Inter-observer agreement was consistently low across both scales (ICC ≤ 0.390). ChatGPT responses to HTO-related FAQs were rated as high quality by most evaluators. However, the low inter-observer agreement highlights the need for standardised evaluation tools and suggests that expert oversight remains essential when integrating AI-generated content into patient education. Level V.