Development and performance validation of an automated GPT-based evaluation tool for the PEDro scale.
other · Level IV
Where this comes from
- Record sourced from PubMed, PMID 41771466.
- Also identified by DOI 10.1016/j.apmr.2026.02.011.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
To develop an automated GPT-PEDro evaluation tool (GPT-PEDro ET) and validate its performance (including intra-rater reliability and concurrent validity) for evaluating the methodological quality of RCTs using the PEDro scale. A psychometric validation study with a repeated measurements design. Research laboratory. One hundred and twenty-five RCTs on neurofacilitation interventions in stroke rehabilitation were retrieved from the PEDro database. The primary outcome was RCT quality scores evaluated by the GPT-PEDro ET at both total score and individual item levels across two rounds of evaluation. Intra-rater reliability was determined using the intraclass correlation coefficient (ICC) for total score and prevalence-adjusted and bias-adjusted kappa (PABAK) for individual items. Concurrent validity was assessed with ICC (total score) and PABAK (individual item), Bland-Altman analysis with 95% limits of agreement, and heteroscedasticity testing. The GPT-PEDro ET achieved almost perfect intra-rater reliability (total score ICC = 1.00; individual item PABAK = 0.94-1.00). The concurrent validity was moderate to high at both the total score level (ICCs = 0.83, 0.83) and the individual item level for all the items (PABAKs = 0.68-0.97). The results suggest that the GPT-PEDro ET may be a useful automated tool for evaluating the methodological quality of RCTs. It shows potential for reducing manual evaluation workload and supporting clinical and research applications.