Concordance Between a Generative Artificial Intelligence Model and a Hepatobiliary Multidisciplinary Team in Hepatocellular Carcinoma Management: A Retrospective Study.
retrospective_cohort · Level III
Where this comes from
- Record sourced from PubMed, PMID 42722940.
- Also identified by DOI 10.1245/s10434-026-20253-8.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The management of hepatocellular carcinoma (HCC) relies on decisions formulated by a multidisciplinary team (MDT). Generative artificial intelligence (AI) models are increasingly explored as decision-support tools. This study aimed to assess the level of agreement between AI-generated recommendation and expert MDT decisions in HCC management. A retrospective cohort of 100 HCC patients discussed at a hepatobiliary MDT between January 2023 and February 2025 was retrospectively included. Anonymized case summaries, excluding MDT conclusion, were submitted to a large language AI model (ChatGPT-4). Recommendations were categorized using the following four-level progressive concordance scale: (1) treatment aim (curative vs palliative), (2) treatment extent (locoregional vs systemic), (3) treatment modality (e.g., resection, ablation, embolization), and (4) specific technique. Concordance levels 1 and 2 were classified as low concordance (LC), whereas levels 3 and 4 were classified as high concordance (HC). In 81 % of the cases, HC between AI-generated and MDT recommendations was observed, with 53 % showing complete agreement on treatment technique. In 19 % of the cases, LC was identified and associated with older patient age (p = 0.001). In most discordant cases, discrepancies reflected the MDT's preference for more aggressive, tailored interventions. In HCC management, AI demonstrated a high level of agreement with expert MDT decisions, suggesting its potential role as a complementary decision-support tool. However, limitations persist for elderly patients and borderline clinical scenarios, in which individualized human judgment remains essential.