Evaluating large language model performance in Risk of Bias assessments: A cross-sectional validation study.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 42424352.
- Also identified by DOI 10.1371/journal.pone.0353155 and PMC identifier 13349098.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
To evaluate the reliability and diagnostic performance of ChatGPT-o3 in conducting Risk of Bias (RoB) assessments of randomized clinical trials (RCTs) using the Cochrane RoB 2.0 tool. This methodological validation study analyzed 50 RCTs sampled from 50 published meta-analyses. Each trial was independently assessed by the original systematic review authors (OSRAs), our masked human panel, and ChatGPT-o3. Structured prompts based on RoB 2.0 guidelines were used to elicit ChatGPT-o3 assessments. Agreement was evaluated using weighted Cohen's kappa and Gwet's AC2. Diagnostic performance was measured by sensitivity, specificity, and balanced accuracy, with human ratings as the reference. ChatGPT-o3 classified 34% of trials as high risk, compared with 22% by our panel, and 12% by the OSRAs. Agreement was modest (median κ: 0.33 with our panel; 0.14 with OSRAs). Overall Gwet's AC2 was 0.30. For detecting high-risk trials, ChatGPT-o3 achieved a sensitivity of 0.46, specificity of 0.69, and balanced accuracy of 0.57. For low-risk trials, its sensitivity was 0.47, specificity was 0.86, and balanced accuracy was 0.66. The results indicate that ChatGPT-o3 produced more conservative RoB ratings than human reviewers, identifying a greater percentage of trials as having a high RoB. While unsuitable to be used as a sole assessor, ChatGPT-o3 may serve as an adjunct tool to enhance the efficiency and consistency of RoB assessments in systematic reviews.
Medical subject headings
- Randomized Controlled Trials as Topic