Why almost all ML models for medicine are wrong-and what we need for evidence-based medical AI.
editorial · Level V
Where this comes from
- Record sourced from PubMed, PMID 42308951.
- Also identified by DOI 10.1016/j.ijmedinf.2026.106538.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Machine learning (ML) models are increasingly proposed to support clinical decision-making, yet their evidentiary basis remains weaker than their publication volume suggests. This editorial argues that the problem is not only translational, regulatory, or infrastructural, but methodological. Many medical ML pipelines rely on uncertain ground truths, optimize performance around clinically irrelevant thresholds, report unstable or prevalence-dependent metrics, neglect calibration and uncertainty, and lack rigorous external and temporal validation. These weaknesses produce optimistic estimates that do not reliably anticipate performance in heterogeneous clinical settings. We call for an evidence-based medical AI grounded in more reliable annotation practices, explicit modeling of uncertainty, clinically meaningful threshold selection, calibration and decision-utility analyses, robustness testing, external validation on independent datasets, and post-deployment monitoring. The editorial also invites authors, reviewers, users, and vendors to adopt stricter standards so that predictive models can become credible, accountable, and clinically useful tools in everyday practice, rather than merely publishable artifacts.