Large language models for automated PRISMA 2020 adherence checking.

Kataoka, Yuki; So, Ryuhei; Banno, Masahiro; Tsujimoto, Yasushi; Takayama, Tomohiro; Yamagishi, Yosuke; Tsuge, Takahiro; Yamamoto, Norio et al. · Int J Med Inform · 2026

other

Where this comes from

Abstract

Evaluating adherence to PRISMA 2020 guideline remains a burden in the peer review process. However, there is a lack of shareable benchmarks for evaluating large language model (LLM) performance in this task. We constructed a copyright-aware benchmark of 108 Creative Commons-licensed systematic reviews. We first conducted parameter optimization using five SRs from the Suda dataset, then compared five checklist input formats (Markdown, JSON, XML, plain text, and manuscript-only control) using ten development-phase LLMs on ten further SRs from the Suda dataset, and finally validated the locked Markdown pipeline using nineteen LLMs on ten SRs from the Tsuge dataset as additional frontier models became available during the study period. Supplying structured PRISMA 2020 checklists yielded 78.7-79.7% accuracy versus 45.2% for manuscript-only input, with paired aggregate analyses showing that structured formats outperformed manuscript-only input while structured formats did not differ significantly from one another. In the validation sample, accuracy ranged from 68.5% to 86.0% with distinct sensitivity-specificity trade-offs. Using Qwen3-Max on the full dataset (n = 120), we achieved 95.1% sensitivity and 49.3% specificity. Structured checklist provision substantially improves LLM-based PRISMA assessment. However, given the observed proportion of false positives, human expert verification remains essential before editorial decisions.