AI-powered radiology report simplification in Arabic: A prospective evaluation of patient-perceived understandability and clinical safety.
prospective_cohort · Level II
Where this comes from
- Record sourced from PubMed, PMID 42546532.
- Also identified by DOI 10.1016/j.ijmedinf.2026.106637.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Radiology reports are written for clinicians, leaving the majority of patients unable to understand their own imaging findings. Large language models (LLMs) offer a means to simplify reports for patient-facing use; however, most implementations rely on cloud-based platforms that raise data privacy and regulatory concerns. Arabic-speaking populations - over 400 million people worldwide - remain substantially underserved in medical AI research, with no prospectively evaluated AI Arabic radiology report simplification system previously reported. This study aimed to conduct a preliminary prospective feasibility and safety evaluation of patient-perceived understandability outcomes and radiologist-assessed clinical safety of an on-site, institutionally governed AI system that generates simplified radiology reports in plain English and Arabic for Arabic-speaking outpatients. In this prospective single-center observational study conducted at King Faisal Specialist Hospital and Research Centre - Jeddah (KFSHRC-J) in January 2026, 98 adult outpatients (mean age 50.0 ± 14.9 years; 50 men, 48 women; IRB #2251488) reviewed three versions of their own radiology report in randomized order: a traditional radiologist report, an AI-simplified English report, and an AI-simplified Arabic report. The system used a two-stage pipeline: Qwen3-14B-FP8 with structured prompt engineering for English simplification (Stage 1), and a fine-tuned Hala-1.2B model for Arabic translation (Stage 2; 8,319 training samples). Patient-perceived understandability was measured on a 5-point Likert scale (1 = Very Easy; 5 = Very Difficult). Three board-certified radiologists independently assessed clinical safety using a standardized seven-domain rubric; inter-rater agreement was quantified using ICC(2,k). AI-simplified Arabic reports produced a median perceived understandability score of 1 (IQR 1-2) versus 5 (IQR 3-5) for traditional reports (p < 0.001; rank-biserial r = 0.97), with 94.9% (93/98) of patients preferring the Arabic simplified version. Safety evaluation demonstrated 96.9% (95/98) of reports safe for patient release, with a hallucination rate of 3.1% (n = 3; 1 unsafe, 2 safe by consensus). Inter-rater reliability was moderate by conventional thresholds (ICC 0.48-0.55) due to a ceiling effect from uniformly high safety scores (mean fidelity 4.8/5.0; 90.8% highest rating; SD = 0.3), with 91.2% absolute agreement on binary safety classification and Fleiss' kappa = 0.71 for safety. In this preliminary prospective feasibility evaluation, a locally deployed, clinician- and AI team-built system significantly improved patient-perceived understandability of radiology reports for Arabic-speaking patients, with a radiologist-assessed safety profile that is promising but requires further validation before broad clinical deployment. The hallucination rate of 3.1% (1.0% unsafe) is lower than published pooled benchmarks, though direct comparisons are limited by methodological heterogeneity across studies (pooled 7.2% across 38 simplification studies; 1.12% for fine-tuned models). To our knowledge, this is among the first prospectively evaluated AI systems for Arabic radiology report simplification, highlighting the potential of on-site AI deployment as a safe, privacy-preserving approach for multilingual radiology communication, warranting further investigation of its contribution to health equity.