Automating methodological quality assessment in orthopedic systematic reviews using large language models.
Where this comes from
- Record sourced from PubMed, PMID 42251603.
- Also identified by DOI 10.1177/10225536261459518.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
BackgroundSystematic reviews represent the foundation of evidence-based orthopedic practice, yet their methodological rigor relies heavily on accurate and consistent methodological quality assessment. This step remains time-consuming, labor-intensive, and prone to subjectivity. Recent advances in large language models (LLMs) suggest potential for automating parts of evidence synthesis.PurposeThis study examined whether LLMs can perform AMSTAR-1-based methodological quality assessment evaluations in orthopedic systematic reviews with accuracy comparable to human experts.MethodsTen sports medicine knee reviews were analyzed using three LLMs-GPT-4o, GPT-5, and GPT Consensus-and their binary responses were compared against expert AMSTAR-1 ratings from a published umbrella review (110 decisions). An external validation set of four reviews published between 2022 and 2025 was included to assess generalizability and safeguard against information leakage.ResultsAgreement with human reviewers reached 87% for GPT-4o, 89% for GPT-5, and 90% for GPT Consensus; all models achieved 84% agreement in the validation set. Concordance was strongest for structured, explicitly reported domains such as a priori design, literature search, and study characteristics, and lowest for judgment-based items including grey literature inclusion, publication bias, and conflict of interest.ConclusionsLLMs cannot yet replace human reviewers, they can serve as reliable adjunct tools to enhance efficiency, transparency, and reproducibility in systematic review workflows within orthopedic research.