Automated Classification of Radiation Oncology Safety Events Using Large Language Models: A Novel Approach to Streamline Reporting and Enable Retrospective Analysis.

Li, Qiongge; Liu, Jian; Li, Xing; Nie, Wei; Fan, Jiajin · Pract Radiat Oncol · 2026

retrospective_cohort · Level III

Where this comes from

Abstract

This study aims to evaluate the feasibility of using a large language model (LLM) to automate the classification of radiation oncology safety events and to outline a framework for future reporting systems that leverage artificial intelligence (AI)-assisted workflows. Retrospective safety event data were extracted from our institutional reporting system and processed using Python scripts for deidentification and formatting. GPT-5 (OpenAI) was accessed via an application programming interface and refined through iterative prompt engineering (a natural language process that does not modify model parameters) to classify incidents across multiple dimensions, including failure mode, severity, treatment type, discoverer's role, exclusion criteria, and occurring/discovering workflow stages. The model's performance was validated by three independent, blinded expert reviewers, with inter-rater agreement quantified using Cohen's κ. On a blinded 80-incident validation set, the model's classification fell within the experts' group-accepted answer in 67.9% of dimension-level comparisons on average (96.2% for the inclusion decision; 85.3% for treatment type and 76.6% for occurred workflow). Agreement between the model and individual reviewers (mean Cohen's κ = 0.31) was comparable to agreement among the reviewers themselves (mean κ = 0.42), indicating performance approaching that of an independent expert. Severity scoring showed the greatest variability for both the model and the human reviewers (model-reviewer κ = 0.14; inter-reviewer κ = 0.22), highlighting it as the primary area for improvement. The comparable model-expert and expert-expert agreement supported the application of the model to full data set analysis without further optimization. LLM-based classification demonstrates strong potential for automating retrospective safety event tagging and streamlining future reporting workflows. In a proposed future system, staff would only need to provide a detailed narrative description while the AI model performs classification and prompts for human validation when uncertainty arises. This hybrid approach can improve efficiency, consistency, and scalability in radiation oncology safety reporting. To our knowledge, this is the first study to apply an LLM for automated classification of radiation oncology safety events, introducing a proof-of-concept framework with potential for broader applicability for AI-assisted incident reporting and quality improvement.