Large language models are comparable with commonly used statistical software: A validation of GPT 5.1 for frequentist meta-analysis in orthopaedics.
Journal:
Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA
Published Date:
Mar 11, 2026
Abstract
PURPOSE: The purpose of this study was to evaluate whether Chat Generative Pre-trained Transformer (ChatGPT; Version 5.1) can reproduce frequentist meta-analytic calculations with an accuracy comparable to established statistical software in orthopaedic research. METHODS: In this methodological comparison study, data from two previously published orthopaedic meta-analyses with identical statistical architectures as reference standards were used. Between-study variance (τ2) was estimated using the Sidik-Jonkman method and uncertainty was quantified using the Hartung-Knapp adjustment for the random-effects models, while common-effect models assume τ2 = 0. Original data extraction tables were provided to ChatGPT-5.1, which was instructed to perform the same analyses. ChatGPT-generated pooled mean differences, confidence intervals and heterogeneity statistics (I2, τ2, p values) were compared with verified reference results obtained using the meta and metafor packages in R. RESULTS: Across seven evaluated outcomes, ChatGPT-5.1 reproduced the direction of effects in all cases. Deviations compared with reference meta-analyses were classified as minor in three outcomes (43%), moderate in one outcome (14%) and major in three outcomes (43%). Agreement was highest in low-heterogeneity settings, whereas substantial deviations occurred in outcomes with pronounced between-study heterogeneity, particularly under random-effects models. CONCLUSION: ChatGPT-5.1 demonstrates emerging capability to approximate frequentist meta-analytic calculations, particularly in low-heterogeneity settings. However, its tendency to underestimate between-study variability and to deviate in complex random-effects scenarios limits its reliability as a standalone tool. At present, large language models may support exploratory analyses but cannot fully replace dedicated statistical software for meta-analyses in orthopaedic research. LEVEL OF EVIDENCE: Level III.
Authors
Keywords
No keywords available for this article.