Improvement of Clinical Practice Guideline Appraisal by Human Experts and AI Agents by Using Structured Guidance: Systematic Review, Meta-Analysis, and Validation Study.
meta_analysis · Level I
Where this comes from
- Record sourced from PubMed, PMID 42647856.
- Also identified by DOI 10.2196/96002.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Rehabilitation clinical practice guidelines (CPGs) have increased rapidly, but inconsistent methodological quality limits their implementation. Although Appraisal of Guidelines for Research and Evaluation II (AGREE II) and Reporting Items for Practice Guidelines in Health Care (RIGHT) provide standardized appraisal frameworks, their application is time-consuming. Large language model (LLM)-based AI agents may offer a scalable alternative with uncertain reliability. We evaluated rehabilitation CPGs' methodological and reporting quality and determined whether structured guidance improves human expert-AI agent agreement. We systematically reviewed English- and Chinese-language rehabilitation CPGs from Embase, Scopus, PubMed, China National Knowledge Infrastructure, Wanfang Data, National Institute for Health and Care Excellence, Scottish Intercollegiate Guidelines Network, and Guidelines International Network up to June 2026. Methodological and reporting quality were assessed using AGREE II and the RIGHT checklist. Factors associated with guideline quality were examined using regression and subgroup analyses. Two AI agents were compared with human consensus with and without a structured guideline appraisal workbook, followed by external validation using 6 anterior cruciate ligament reconstruction CPGs. We included 227 CPGs (163 English-language, 64 Chinese-language). After introducing a structured guideline appraisal workbook, agreement among human experts improved markedly-mean intraclass correlation coefficients (ICCs) increased from -0.09 to 0.66 to 0.84-0.92 across AGREE II domains. Overall guideline quality remained low, with 35.9% (SD 18,8%) applicability and 52% (SD 17.2%) stakeholder involvement. English-language guidelines outperformed Chinese-language guidelines in scope and purpose (mean 74.64, SD 15.4 vs mean 68.88, SD 13.9; P=.004) and applicability (mean 39.14, SD 18.6 vs mean 27.54, SD 16.7; P<.001). Backward-elimination logistic regression revealed external review as an associated process characteristic (odds ratio 20.39, 95% CI 4.66-89.27; P<.001). RIGHT assessments showed consistent reliability (ICC=0.80-0.88). Reporting was highest for basic information (70.5%) and lowest for funding, declaration, and management of interests (44.1%). Meta-analysis of RIGHT reporting rates showed lower reporting among Chinese-language than English-language guidelines (risk difference [RD] -0.07, 95% CI -0.13 to -0.02, 95% prediction interval [PI] -0.39 to 0.24) and among guidelines published before vs after RIGHT release (RD -0.19, 95% CI -0.26 to -0.13, 95% PI -0.56 to 0.17). Without additional guidance, agent-human agreement was moderate (ICC=0.608-0.629). The workbook improved agreement for both models, with DeepSeek-R1's increasing from 0.613 to 0.709 and o1-mini's from 0.629 to 0.687. In validation beyond rehabilitation, DeepSeek-R1 maintained stable agreement (ICC=0.711) and completed appraisals in 5.44 minutes compared to 11.18 minutes for humans. Rehabilitation CPGs, particularly Chinese-language CPGs, continue showing deficiencies in applicability and stakeholder involvement. LLM-based appraisal without structured guidance provides insufficient agreement. Structured guidance improved agent-human agreement, supporting AI-assisted guideline appraisal under human oversight. Although further validation across additional clinical specialties is needed, AI agents can serve as efficient assistants in guideline appraisal instead of replacing humans. Future synthesis requires human-AI integration guided by structured, expert-defined principles. PROSPERO CRD420251270676; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251270676.
Medical subject headings
- Practice Guidelines as Topic
- Artificial Intelligence