Sample size calculation for training ensemble machine learning models on health data.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42328195.
- Also identified by DOI 10.1016/j.patter.2026.101498 and PMC identifier 13280678.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Health research studies often suffer from small sample sizes, and training machine learning (ML) models requires large datasets. There is a dearth of literature on determining the adequate sample size for using ML models. We developed an empirically derived sample size calculator for ensemble ML models: random forests and two gradient-boosted decision trees (light gradient boosting machine [LGBM] and extreme gradient boosting [XGBoost]). This predicts the sample size required to achieve a pre-defined level of prognostic performance with a certain probability. Prognostic performance is defined as the sample area under the ROC curve (ROC-AUC) relative to the optimal model trained on the full (population) dataset. Our calculator's accuracy was compared to three common heuristics and a statistical approach to sample size calculation. For example, the median relative error sample size prediction was 25% to achieve 85% of the optimal performance with 90% certainty for LGBM. Our model has significantly better accuracy than other methods for tree-based ensemble ML models.