Development and performance validation of an automated GPT-based evaluation tool for the PEDro scale.

Ju, Yu-Jeng; Wang, Yi-Ching; Chou, Fan; Hou, Wen-Hsuan; Chen, Shu-Mei; Hsieh, Ching-Lin · Arch Phys Med Rehabil · 2026

other · Level IV

Where this comes from

Abstract

To develop an automated GPT-PEDro evaluation tool (GPT-PEDro ET) and validate its performance (including intra-rater reliability and concurrent validity) for evaluating the methodological quality of RCTs using the PEDro scale. A psychometric validation study with a repeated measurements design. Research laboratory. One hundred and twenty-five RCTs on neurofacilitation interventions in stroke rehabilitation were retrieved from the PEDro database. The primary outcome was RCT quality scores evaluated by the GPT-PEDro ET at both total score and individual item levels across two rounds of evaluation. Intra-rater reliability was determined using the intraclass correlation coefficient (ICC) for total score and prevalence-adjusted and bias-adjusted kappa (PABAK) for individual items. Concurrent validity was assessed with ICC (total score) and PABAK (individual item), Bland-Altman analysis with 95% limits of agreement, and heteroscedasticity testing. The GPT-PEDro ET achieved almost perfect intra-rater reliability (total score ICC = 1.00; individual item PABAK = 0.94-1.00). The concurrent validity was moderate to high at both the total score level (ICCs = 0.83, 0.83) and the individual item level for all the items (PABAKs = 0.68-0.97). The results suggest that the GPT-PEDro ET may be a useful automated tool for evaluating the methodological quality of RCTs. It shows potential for reducing manual evaluation workload and supporting clinical and research applications.