Tipping the balance: impact of class imbalance correction on the performance of clinical risk prediction models.
other
Where this comes from
- Record sourced from PubMed, PMID 42533626.
- Also identified by DOI 10.1093/jamia/ocag127.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Machine-learning-based clinical risk prediction models are increasingly used to support decision-making in healthcare. While class-imbalance correction techniques are commonly applied to address rare outcomes, their impact on probabilistic calibration remains insufficiently understood. This study evaluated the effect of widely used resampling strategies on both discrimination and calibration across real-world clinical prediction tasks. Ten clinical datasets spanning diverse medical domains and including over 600 000 patients were analyzed. Multiple machine-learning model families were evaluated. Models were trained on original data and using 3 1:1 class-imbalance correction strategies (synthetic minority oversampling technique, random undersampling, and random oversampling). Performance was assessed on held-out data using discrimination and calibration metrics. Resampling had no positive impact on predictive performance. Changes in area under the receiver operating characteristic curve (ROC-AUC) and precision-recall AUC were small and inconsistent (ROC-AUC: -0.002 to -0.01; PR-AUC: -0.10 to -0.03), with no method showing systematic improvement. In contrast, calibration was consistently degraded. Resampled models showed higher Brier scores (increase 0.029-0.080) and marked deviations in calibration intercept and slope, indicating distorted predicted risks despite preserved ranking performance. Across diverse clinical datasets, resampling primarily altered the implicit class prior learned during training, leading to miscalibration when models were evaluated. The consistent dissociation between discrimination and calibration highlights that rank-based metrics alone are insufficient for evaluating clinical utility. Gains from imbalance correction can typically be reproduced by threshold adjustment without distorting predicted probabilities. Common 1:1 class-imbalance correction techniques do not improve discrimination and may substantially degrade calibration, limiting their suitability for clinical risk prediction where accurate probabilities are essential.