On-premise medical AI agents for reliable clinical decision-making.

Zhang, Li; Wölflein, Georg; Ferber, Dyke; Liang, Junhao; Carrero, Zunamys I; Wu, Xuewei; Vibert, Julien; Clusmann, Jan et al. · Nat Med · 2026

other · Level V

Where this comes from

Abstract

Autonomous clinical artificial intelligence (AI) agents powered by large language models (LLMs), meaning systems that can complete a diagnostic workflow without continuous human input, are increasingly capable of supporting complex reasoning and decision-making. Clinical translation, however, remains limited by two unmet requirements: institutionally governed deployment and reliable decision-time uncertainty estimation. Here we developed and evaluated a fully on-premise clinical agent that couples local operational control with a multi-perspective reliability framework to support selective autonomy. Across two Medical Information Mart for Intensive Care IV (MIMIC-IV)-derived benchmarks, the agent achieved 90.04% accuracy on a seven-disease task and 83.8% accuracy on a four-disease task, approaching a cloud baseline on the primary benchmark. To assess decision-time reliability, we quantified internal-likelihood, language-based and behavioral-stability measures for diagnosis and reasoning. Diagnostic behavioral consistency provided the strongest discrimination of correctness (area under the curve (AUC) = 0.860) and remained informative under stress testing (AUC = 0.875). At a consistency threshold of 0.90, 49.4% of cases were retained at 98.9% diagnostic accuracy. These findings support a practical framework for institutionally governed clinical agents in which decision-time reliability signals identify a lower-risk subset for autonomous handling and defer the remainder for review.