Large Language Models in Multidisciplinary Decision-Making for Hepatopancreatobiliary Oncology: Retrospective Comparative Feasibility Study.
retrospective_cohort · Level III
Where this comes from
- Record sourced from PubMed, PMID 42753242.
- Also identified by DOI 10.2196/101137.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Hepatopancreatobiliary (HPB) malignancies require complex treatment planning that often relies on multidisciplinary team (MDT) discussions. Large language models (LLMs) have recently been explored for clinical decision support, but their performance within real-world multidisciplinary decision environments remains unclear. In particular, the stability of LLM-generated recommendations-that is, whether a model produces the same answer when given the same clinical input-has rarely been examined. This study aimed to evaluate the stability of treatment recommendations generated by contemporary LLMs when identical HPB cases are queried repeatedly, and their concordance with the treatment decisions reached at an institutional MDT conference. This retrospective study included consecutive cases discussed at a single-center HPB MDT conference between September 1, 2024, and August 31, 2025. Standardized clinical case summaries derived from preconference documentation were provided to 4 LLMs (GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5) through their consumer web interfaces. Each model recommended a treatment among predefined MDT treatment options, and identical queries were repeated 4 times in separate sessions. Stability was quantified as the discordance rate relative to the initial response and, without privileging any single query, as the mean pairwise agreement and Fleiss κ across the 4 iterations. Concordance with MDT decisions was assessed using both the initial and modal responses, together with Cohen κ and class-wise <i>F</i><sub>1</sub>-scores. A total of 107 MDT cases were analyzed. Stability differed significantly across models (<i>P</i>=.01). Gemini 3 Pro showed the lowest discordance rate (mean 12.8%, SD 2.3%) and the highest reference-free agreement (Fleiss κ=0.737), whereas GPT-4o showed the highest discordance rate (mean 30.2%, SD 6.5%) and the lowest agreement (Fleiss κ=0.430). Concordance with MDT decisions ranged from 48.6% to 72.9% using the initial response and from 66.3% to 74.5% using the modal response, and the highest-performing model differed between the 2 definitions. Class-wise <i>F</i><sub>1</sub>-score was consistently lower for surgery (0.400-0.520) than for chemotherapy (0.621-0.836). Complete discordance occurred in 17 of 107 (15.9%) cases and in none of the 31 anatomically unresectable cases (Fisher exact test, <i>P</i>=.003). Recurrent or on-treatment disease (adjusted odds ratio [OR] 5.40, 95% CI 1.65-17.68; <i>P</i>=.005), pancreatic tumor location (adjusted OR 7.37, 95% CI 2.03-26.78; <i>P</i>=.002), and low MDT agreement level (adjusted OR 10.33, 95% CI 1.54-69.38; <i>P</i>=.016) were independently associated with complete discordance. LLM-generated treatment recommendations demonstrated moderate alignment with MDT decisions in HPB oncology. Importantly, response stability varied substantially across models, indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools. These findings suggest that LLMs may serve as a reasoning-support layer in MDT-like decision environments, but their response stability must be systematically characterized before clinical integration.
Medical subject headings
- Large Language Models
- Liver Neoplasms
- Clinical Decision-Making
- Decision Making