ChatGPT provides high-quality answers to FAQs about high tibial osteotomy despite low inter-rater agreement.

Akcaalan, Serhat; Şahbat, Yavuz; Loddo, Glauco; Erden, Tunay; Kocaoglu, Baris · J Exp Orthop · 2025

case_series · Level V

Where this comes from

Abstract

High tibial osteotomy (HTO) is frequently used to treat knee malalignment in younger patients. Given the rise in online health information-seeking behaviour, this study aimed to evaluate the quality of ChatGPT-generated responses to frequently asked questions (FAQs) about HTO and to assess the reliability of two scoring systems used by orthopaedic surgeons. Twenty-four FAQs were submitted to ChatGPT (GPT-4-turbo). Four orthopaedic surgeons independently rated the responses at two time points using: (1) a 4-point categorical scale (1 = excellent, 4 = poor), and (2) a 100-point numerical scale (0 = worst, 100 = best). Intra-observer reliability was assessed using weighted kappa (<i>κ</i>) and intraclass correlation coefficients (ICC); inter-observer agreement was measured using ICC values. Most responses were rated positively, with over 70% considered 'excellent' or requiring minimal clarification. Intra-observer agreement was variable, ranging from <i>κ</i> = 0.333 to 0.864 and ICC = 0.690-0.922. Inter-observer agreement was consistently low across both scales (ICC ≤ 0.390). ChatGPT responses to HTO-related FAQs were rated as high quality by most evaluators. However, the low inter-observer agreement highlights the need for standardised evaluation tools and suggests that expert oversight remains essential when integrating AI-generated content into patient education. Level V.