Performance of large language models and clinical decision support in perioperative management of oral anticoagulants.

Çalışkan, Mehmet Uğur; Sarıbaş, Halenur; Keskin, Gökhan; Kertmen, Ömer; Çakmak, Abdulkadir; Özbay, Yılmaz · Int J Med Inform · 2026

other

Where this comes from

Abstract

Perioperative anticoagulant management is critical because of the competing risks of ischemia and bleeding. Large language models (LLMs) and clinical decision support (CDS) have shown substantial advances and are increasingly being applied across various domains of cardiology. However, to date, no studies have specifically evaluated the performance of LLMs combined with CDS in perioperative anticoagulant management. Two cardiologists developed 40 guideline-based clinical scenarios involving patients receiving oral anticoagulants scheduled for non-cardiac surgery, including 20 direct oral anticoagulant (DOAC) and 20 warfarin scenarios. Three commonly used LLMs (ChatGPT 5.2, Gemini 3.0 Pro, and DeepSeek V3.2) were evaluated in two phases: baseline performance (Phase 1) and performance after integration of structured CDS tables (Phase 2). Responses were assessed for guideline concordance on a per-scenario basis, requiring correct recommendations across all predefined domains. Model performances were compared using Cochran's Q test, with post hoc Dunn-Bonferroni correction, and within-model comparisons were performed using McNemar's test. In Phase 1, guideline-concordant management was achieved in 60% of scenarios by ChatGPT-5.2, 55% by Gemini 3.0 Pro, and 50% by DeepSeek V3.2, with no significant difference among models (p = 0.180). Following CDS augmentation in Phase 2, scenario-level accuracy improved significantly for all models (all p < 0.05), increasing to 80% for ChatGPT-5.2, 100% for Gemini 3.0 Pro, and 85% for DeepSeek V3.2. Comparative analysis in Phase 2 demonstrated a significant difference among models (p = 0.009), driven primarily by superior performance of Gemini 3.0 Pro compared with ChatGPT-5.2 (adjusted p = 0.009). Baseline performance of LLMs in perioperative anticoagulant management was modest and insufficient to replace clinical judgment. However, integration of structured CDS tools resulted in marked and clinically meaningful improvements in guideline-concordant performance across all evaluated models.