ChatGPT provides accurate and safe responses to patient questions on hip arthroscopy, while completeness remains variable: A systematic review and single-arm meta-analysis.
systematic_review · Level I
Where this comes from
- Record sourced from PubMed, PMID 41979367.
- Also identified by DOI 10.1002/ksa.70396 and PMC identifier 13266954.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Large language models, such as ChatGPT, are increasingly used by patients seeking information on hip arthroscopy (HAS) and femoroacetabular impingement (FAI). Despite their linguistic fluency, the accuracy, completeness and safety of procedure-specific patient information remain unclear. Although orthopaedic studies report variable performance across subspecialties, no systematic evaluation has specifically addressed HAS. PubMed, Embase, Scopus, CINAHL and Epistemonikos were searched to 10 February 2026 for studies evaluating ChatGPT responses to patient-oriented HAS or FAI questions. Randomized and non-randomized studies, observational cohorts and case series were eligible. Data on question sources, model versions, rating systems and performance domains (accuracy, relevance, completeness, safety, readability and clarity) were extracted. Heterogeneous rating scales were dichotomized into high- versus low-quality responses. Risk of bias was assessed using QUADAS-2 and ROBINS-I. Random-effects single-arm meta-analyses (REML) were conducted for each domain. Eight studies met eligibility criteria. Accuracy was high (pooled 88.6%). Relevance, safety, readability and clarity reached pooled values of 100% with low heterogeneity. Completeness was lower (83.8%) with moderate heterogeneity, mainly driven by early GPT-3.5 studies. Funnel plots showed no clear small-study effects, although interpretation was limited by the small number of studies. Risk of bias was predominantly high or moderate, largely due to non-systematic question selection and heterogeneous rating tools. Later models (GPT-4/4o and beyond) demonstrated higher performance compared with GPT-3.5. ChatGPT provides accurate, relevant, safe and clear responses to patient questions about HAS, while completeness shows moderate variability. Although LLMs appear promising as adjuncts to patient education, methodological limitations in the current evidence base underscore the need for expert clinical counselling and more rigorous, standardized evaluation frameworks. Level III systematic review and meta-analysis of non-randomized studies.
Medical subject headings
- Arthroscopy
- Patient Education as Topic
- Femoracetabular Impingement
- Hip Joint