Avoiding common failures in AI for health and medicine.
review · Level V
Where this comes from
- Record sourced from PubMed, PMID 42753694.
- Also identified by DOI 10.1016/j.cell.2026.07.025.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The widespread adoption of artificial intelligence (AI) in healthcare necessitates reliable AI systems. Reliability refers to clinically acceptable performance across contexts such as patient and clinician populations and over time under real-world deployment conditions. This review synthesizes common reliability failure modes in predictive and generative AI systems, which include erroneous model outputs, clinically unjustified performance differences across patient populations, and performance degradation under changing deployment conditions. We examine approaches to addressing these failures and critically assess the empirical support and practical limits of current solutions, in both predictive and generative AI settings. We argue that current technical solutions are often insufficient to address the challenges that lead to unreliable AI, thus motivating the need for lifecycle-aware evaluation, continuous monitoring, and governance.
Medical subject headings
- Artificial Intelligence