Sample size calculation for training ensemble machine learning models on health data.

Mitsakakis, Nicholas; Liu, Dan; Walters, Thomas; El Emam, Khaled · Patterns (N Y) · 2026

other · Level V

Where this comes from

Abstract

Health research studies often suffer from small sample sizes, and training machine learning (ML) models requires large datasets. There is a dearth of literature on determining the adequate sample size for using ML models. We developed an empirically derived sample size calculator for ensemble ML models: random forests and two gradient-boosted decision trees (light gradient boosting machine [LGBM] and extreme gradient boosting [XGBoost]). This predicts the sample size required to achieve a pre-defined level of prognostic performance with a certain probability. Prognostic performance is defined as the sample area under the ROC curve (ROC-AUC) relative to the optimal model trained on the full (population) dataset. Our calculator's accuracy was compared to three common heuristics and a statistical approach to sample size calculation. For example, the median relative error sample size prediction was 25% to achieve 85% of the optimal performance with 90% certainty for LGBM. Our model has significantly better accuracy than other methods for tree-based ensemble ML models.