Large language models are comparable with commonly used statistical software: A validation of GPT 5.1 for frequentist meta-analysis in orthopaedics.

Salzmann, Mikhail; Ramadanov, Nikolai; Prill, Robert; Hable, Robert; Becker, Roland · Knee Surg Sports Traumatol Arthrosc · 2026

other · Level V

Where this comes from

Abstract

The purpose of this study was to evaluate whether Chat Generative Pre-trained Transformer (ChatGPT; Version 5.1) can reproduce frequentist meta-analytic calculations with an accuracy comparable to established statistical software in orthopaedic research. In this methodological comparison study, data from two previously published orthopaedic meta-analyses with identical statistical architectures as reference standards were used. Between-study variance (τ<sup>2</sup>) was estimated using the Sidik-Jonkman method and uncertainty was quantified using the Hartung-Knapp adjustment for the random-effects models, while common-effect models assume τ<sup>2</sup> = 0. Original data extraction tables were provided to ChatGPT-5.1, which was instructed to perform the same analyses. ChatGPT-generated pooled mean differences, confidence intervals and heterogeneity statistics (I<sup>2</sup>, τ<sup>2</sup>, p values) were compared with verified reference results obtained using the meta and metafor packages in R. Across seven evaluated outcomes, ChatGPT-5.1 reproduced the direction of effects in all cases. Deviations compared with reference meta-analyses were classified as minor in three outcomes (43%), moderate in one outcome (14%) and major in three outcomes (43%). Agreement was highest in low-heterogeneity settings, whereas substantial deviations occurred in outcomes with pronounced between-study heterogeneity, particularly under random-effects models. ChatGPT-5.1 demonstrates emerging capability to approximate frequentist meta-analytic calculations, particularly in low-heterogeneity settings. However, its tendency to underestimate between-study variability and to deviate in complex random-effects scenarios limits its reliability as a standalone tool. At present, large language models may support exploratory analyses but cannot fully replace dedicated statistical software for meta-analyses in orthopaedic research. Level III.