Deep learning models for acute kidney injury prediction: multi-center external validation and evaluation under simulated continuous monitoring conditions.
retrospective_cohort · Level III
Where this comes from
- Record sourced from PubMed, PMID 42103942.
- Also identified by DOI 10.1038/s41746-026-02722-2.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Acute kidney injury (AKI) is a common hospital complication with substantial morbidity and mortality. Deep learning models for AKI prediction show strong development-cohort performance, but single-point evaluation fails to capture behaviour under continuous monitoring. We conducted a multi-centre retrospective study using electronic health records from three cohorts (n = 157,323 admissions): National Health Insurance Service Ilsan Hospital (development), Chuncheon Sacred Heart Hospital, and MIMIC-IV (external validation). Three deep learning architectures (LSTM-Attention, Masked CNN, ITE-Transformer) and two baselines (XGBoost, logistic regression) were developed across 0-, 48-, and 72-h horizons, with an online simulation framework generating predictions at 12-h intervals before onset. Deep learning substantially outperformed baselines externally (AUROC 0.956-0.963 vs. 0.630-0.686). Online simulation revealed that 0-h models exhibited "clinical faithfulness"-consistent AUROC improvement as onset approached (Mann-Kendall significant in 15/15 combinations)-whereas longer horizons showed unstable trajectories. Notably, the highest single-point AUROC model (Masked CNN, 0.961) had the worst deployment profile (NNE 17.6-564), while ITE-Transformer (AUROC 0.924) achieved the most favourable alert burden (NNE 1.5-2.4). Deployment-oriented evaluation should complement conventional metrics for continuous monitoring models.