Interobserver Variability Across Whole-Slide Imaging Systems.

Samueli, Benzion; Al-Ahmadie, Hikmat; Baine, Marina; Pulitzer, Melissa; Rosenblum, Marc; Alvi, Emaan; Mehrotra, Harshita; Pujadas, Elisabet et al. · Arch Pathol Lab Med · 2026

Where this comes from

Abstract

Digital pathology whole-slide images (WSIs) must be of adequate quality for use in diagnostics and the development of artificial intelligence algorithms. Inadequate slides should be re-scanned. There is insufficient data on what qualifies as WSI adequacy. To report on the degree of inter-cohort agreement when evaluating WSI quality and adequacy for use in primary diagnostics. The 60 validation slides were scanned by 3 scanners (Leica GT450, Hamamatsu NanoZoomer S360, and Huron TissueScope iQ). The resulting 180 WSIs were independently evaluated on a quality pass or fail basis by 12 observers representing 4 staff members of the following: in-house attending physicians, fellows, and imaging technicians. The Cohen's kappa was calculated overall and within each cohort, and the acceptance rate was compared between cohorts. For almost every cohort on each scanner, the Cohen's kappa was weak (0.21-0.4) or very poor (<0.2). Based on WSI pass rates, the data yielded different scanner rankings across cohorts, with the GT450, TissueScope iQ, and NanoZoomer S360 exhibiting the highest pass rates for technicians, fellows, and attendings, respectively. There is significant interobserver and inter-cohort variability in the rating of WSIs for suitability in diagnostic pathology. Setting uniform, reproducible standards is important for patient safety, creating a uniform level of care, efficiently re-scanning subpar slides, developing artificial intelligence models on images characteristic of those used in clinical practice, and accurately representing the quality of a scanner.