Assessing ChatGPT for Clinical Decision-Making in Radiation Oncology, With Open-Ended Questions and Images.

Chuang, Wei-Kai; Kao, Yung-Shuo; Liu, Yen-Ting; Lee, Cho-Yin · Pract Radiat Oncol · 2025

other · Level V

Where this comes from

Abstract

This study assesses the practicality and correctness of Chat Generative Pre-trained Transformer (ChatGPT)-4 and 4O's answers to clinical inquiries in radiation oncology, and evaluates ChatGPT-4O for staging nasopharyngeal carcinoma (NPC) cases with magnetic resonance (MR) images. A total of 164 open-ended questions covering representative professional domains (Clinical_G: knowledge on standardized guidelines; Clinical_C: complex clinical scenarios; Nursing: nursing and health education; and Technology: radiation technology and dosimetry) were prospectively formulated by experts and presented to ChatGPT-4 and 4O. Each ChatGPT's answer was graded as 1 (Directly practical for clinical decision-making), 2 (Correct but inadequate), 3 (Mixed with correct and incorrect information), or 4 (Completely incorrect). ChatGPT-4O was presented with the representative diagnostic MR images of 20 patients with NPC across different T stages, and asked to determine the T stage of each case. The proportions of ChatGPT's answers that were practical (grade 1) varied across professional domains (P < .01), higher in Nursing (GPT-4: 91.9%; GPT-4O: 94.6%) and Clinical_G (GPT-4: 82.2%; GPT-4O: 88.9%) domains than in Clinical_C (GPT-4: 54.1%; GPT-4O: 62.2%) and Technology (GPT-4: 64.4%; GPT-4O: 77.8%) domains. The proportions of correct (grade 1+2) answers (GPT-4: 89.6%; GPT-4O: 98.8%; P < .01) were universally high across all professional domains. However, ChatGPT-4O failed to stage NPC cases via MR images, indiscriminately assigning T4 to all actually non-T4 cases (κ = 0; 95% CI, -0.253 to 0.253). ChatGPT could be a safe clinical decision-support tool in radiation oncology, because it correctly answered the vast majority of clinical inquiries across professional domains. However, its clinical practicality should be cautiously weighted particularly in the Clinical_C and Technology domains. ChatGPT-4O is not yet mature to interpret diagnostic images for cancer staging.

Medical subject headings