Analysis of Large Language Model Decision Making in Hormone Receptor-Positive/Human Epidermal Growth Factor Receptor 2-Negative Early Breast Cancer.
retrospective_cohort · Level III
Where this comes from
- Record sourced from PubMed, PMID 41791000.
- Also identified by DOI 10.1200/CCI-25-00230 and PMC identifier 12986038.
- Licence recorded as CC BY-NC-ND.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
To assess the ability of GPT-4o in adjuvant treatment decision making in hormone receptor-positive (HR+)/human epidermal growth factor receptor 2-negative (HER2-) early breast cancer by comparing its recommendations with those of clinicians including Oncotype DX data, and to explore its potential as a decision-support tool in routine clinical practice. We compared clinician and GPT-4o recommendations in patients tested with Oncotype DX in routine practice at the University of Naples Federico II (n = 607, cohort 1 [C1]) and within the prospective, multicenter PRO BONO study (n = 237, cohort 2 [C2]). Pre- and post-Oncotype DX treatment recommendations were categorized as chemotherapy (CT) + endocrine therapy (ET) or ET alone. Concordance between clinician and GPT-4o recommendations was assessed using agreement rates and Cohen's kappa. The accuracy of Oncotype DX results was evaluated using the AUC metric. The agreement between clinicians and GPT-4o in pretest recommendations was 68% (kappa, 0.381 [95% CI, 0.31 to 0.45], <i>P</i> < .001) in C1 and 70% (0.401 [95% CI, 0.29 to 0.52], <i>P</i> < .001) in C2. Before Oncotype DX, clinicians recommended CT more frequently than GPT-4o for C1 (58% <i>v</i> 38%) and C2 (53% <i>v</i> 43%). Post-test agreement increased to 93% (0.814 [95% CI, 0.76 to 0.87], <i>P</i> < .001) in C1 and 90% (0.741 [95% CI, 0.64 to 0.84], <i>P</i> < .001) in C2. The agreement between pre- and post-Oncotype DX treatment recommendations for clinicians was 56% and 63% versus 68% and 60% for GPT-4o in C1 and C2, respectively. GPT-4o showed higher accuracy in predicting low than high genomic risk in postmenopausal patients (87% <i>v</i> 43% in C1; 85% <i>v</i> 45% in C2, <i>P</i> < .001) and low versus intermediate and high risk in premenopausal patients in both cohorts (<i>P</i> < .001). The agreement between clinicians and GPT-4o in pretest recommendations was modest but improved post-test, highlighting the importance of multigene testing and the potential of large language models in clinical decision making.
Medical subject headings
- Breast Neoplasms
- Erb-b2 Receptor Tyrosine Kinases
- Clinical Decision-Making
- Decision Making