Development of an Artificial Intelligence Model to Aid in Measurement of Invasion, Comprehensive Histologic Subtyping, and Grading of Pulmonary Adenocarcinoma.

Boland, Jennifer M; Stetzik, Lucas; Roden, Anja C; Maleszewski, Joseph J; Jenkins, Sarah M; Kroneman, Trynda N; Yi, Eunhee S; Lo, Ying-Chun et al. · Mod Pathol · 2026

Where this comes from

Abstract

The World Health Organization classification of pulmonary adenocarcinoma is complex, posing challenges for pathological reporting. Key difficulties include assessing invasive size in lepidic-predominant tumors and performing comprehensive histologic subtyping. Although these evaluations inform tumor stage, grade, and prognosis, they are time consuming and subjective, leading to interobserver variability. Artificial intelligence (AI) may help streamline these tasks and improve consistency. One representative hematoxylin and eosin slide was selected from each of 100 resected pulmonary adenocarcinomas, which were divided into training (n = 35) and validation (n = 65) sets. Slides were scanned and uploaded to Aiforia for AI model creation. Annotations were completed on the training set by 6 expert pulmonary pathologists and used to train a nested AI model, which was used to evaluate whole slide images of the training and validation sets. Manual assessment of tumor size, invasive size, and comprehensive histologic subtyping was performed by 3 pulmonary pathologists. In both the training and validation sets, the mean and median difference between manual and AI estimations of tumor size was ≤1.3 mm and invasive size was ≤3 mm. The median and mean differences in invasive percentage were ≤15.3% in both the training and validation sets for all patterns except for acinar and lepidic. However, ranges were wide, indicating examples with substantial disagreement. Predominant pattern agreement between AI and each observer ranged from 65.7% to 71.4% in the training set and 45.3% to 54.7% in the validation set. Agreement in grade ranged from 77.1% to 88.6% in the training set and 62.5% to 67.2% in the validation set. There was moderate agreement in grade between the 3 pathologists in the training set and moderate to substantial agreement between AI and each observer. In the validation set, there was substantial agreement in grade between the 3 pathologists and moderate agreement between AI and each observer. Although the AI model shows promise and warrants further refinement, manual pathologist review and potential revision of AI assessments are necessary to ensure the diagnostic accuracy needed for clinical use because disagreement clearly occurs.

Medical subject headings