Analysis of Large Language Model Decision Making in Hormone Receptor-Positive/Human Epidermal Growth Factor Receptor 2-Negative Early Breast Cancer.

Buonaiuto, Roberto; Caltavituro, Aldo; Di Rienzo, Rossana; Grieco, Angela; Mangiacotti, Federica P; Longobardi, Alessandra; Cantile, Vincenza; Molinaro, Vittoria et al. · JCO Clin Cancer Inform · 2026

retrospective_cohort · Level III

Where this comes from

Abstract

To assess the ability of GPT-4o in adjuvant treatment decision making in hormone receptor-positive (HR+)/human epidermal growth factor receptor 2-negative (HER2-) early breast cancer by comparing its recommendations with those of clinicians including Oncotype DX data, and to explore its potential as a decision-support tool in routine clinical practice. We compared clinician and GPT-4o recommendations in patients tested with Oncotype DX in routine practice at the University of Naples Federico II (n = 607, cohort 1 [C1]) and within the prospective, multicenter PRO BONO study (n = 237, cohort 2 [C2]). Pre- and post-Oncotype DX treatment recommendations were categorized as chemotherapy (CT) + endocrine therapy (ET) or ET alone. Concordance between clinician and GPT-4o recommendations was assessed using agreement rates and Cohen's kappa. The accuracy of Oncotype DX results was evaluated using the AUC metric. The agreement between clinicians and GPT-4o in pretest recommendations was 68% (kappa, 0.381 [95% CI, 0.31 to 0.45], <i>P</i> < .001) in C1 and 70% (0.401 [95% CI, 0.29 to 0.52], <i>P</i> < .001) in C2. Before Oncotype DX, clinicians recommended CT more frequently than GPT-4o for C1 (58% <i>v</i> 38%) and C2 (53% <i>v</i> 43%). Post-test agreement increased to 93% (0.814 [95% CI, 0.76 to 0.87], <i>P</i> < .001) in C1 and 90% (0.741 [95% CI, 0.64 to 0.84], <i>P</i> < .001) in C2. The agreement between pre- and post-Oncotype DX treatment recommendations for clinicians was 56% and 63% versus 68% and 60% for GPT-4o in C1 and C2, respectively. GPT-4o showed higher accuracy in predicting low than high genomic risk in postmenopausal patients (87% <i>v</i> 43% in C1; 85% <i>v</i> 45% in C2, <i>P</i> < .001) and low versus intermediate and high risk in premenopausal patients in both cohorts (<i>P</i> < .001). The agreement between clinicians and GPT-4o in pretest recommendations was modest but improved post-test, highlighting the importance of multigene testing and the potential of large language models in clinical decision making.

Medical subject headings