Variation and inter-rater reliability of abstract reviewer scores for clinical vignettes submitted to a national hospitalist meeting.

Stephens, John R; Morgan, Andrew; Bellamy, Nelly; Raff, Evan J; Feldman, Leonard · J Hosp Med · 2026

cross_sectional · Level IV

Where this comes from

Abstract

Peer review of research products suffers from poor inter-rater reliability. Few studies examine whether this limitation generalizes to case reports. We conducted a cross-sectional analysis of peer reviews of clinical vignette abstracts submitted to a national hospitalist meeting in 2024 and 2025. Three randomly assigned reviewers scored each vignette on a 1-10 scale. We analyzed variation in scores across abstracts and reviewers and estimated inter-rater reliability via intraclass correlation coefficient (ICC). Two hundred twenty-one reviewers evaluated 1630 abstracts in 2024-2025. Abstract scores varied substantially: 384/1630 (23.6%) abstracts had a difference of 4 or more points (>2 standard deviations) between highest and lowest reviewer scores. Scores varied by reviewer: 2024 reviewer-level mean scores ranged 4.27-8.47 (standard deviation (SD): 0.70-2.80); 2025 scores ranged 4.06-8.59 (SD: 0.62-2.69). Inter-rater reliability was poor (ICC: 0.37). Adjusting final scores based on reviewer scoring tendencies changed the accept/reject category for 183 (11.2%) abstracts, suggesting opportunities for quality improvement.

Medical subject headings