Machine learning versus logistic regression for differentiating Kawasaki disease from febrile illnesses: methodological and performance comparisons.
systematic_review · Level I
Where this comes from
- Record sourced from PubMed, PMID 42673789.
- Also identified by DOI 10.1016/j.ijmedinf.2026.106689.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Differentiating Kawasaki disease (KD) from other febrile illnesses remains challenging because of overlapping clinical and laboratory features. This study systematically evaluated the methodological quality, risk of bias, applicability, and predictive performance of machine learning (ML) and logistic regression-family (LR-family) models. PubMed, Web of Science, and Embase were searched from January 1, 2006, to December 31, 2025, for studies developing or validating ML or LR-family models for differentiating KD from other febrile illnesses. Risk of bias and applicability were assessed using PROBAST + AI. The required minimum sample size for each study was formally estimated, and exploratory meta-analyses, subgroup analyses, and leave-one-out sensitivity analyses were performed. Twenty-eight studies (10 ML and 18 LR-family) were included. Common methodological limitations included inadequate sample size, retrospective design, inadequate handling of missing data, limited external validation, and poor calibration reporting; no ML study reported calibration metrics. Given the high risk of bias across all studies, pooled areas under the receiver operating characteristic curve (AUCs) were interpreted as exploratory quantitative summaries, showing no significant difference in internal validation performance between ML and LR-family models (0.95 [95% CI 0.91-0.97] vs. 0.92 [95% CI 0.89-0.94]; P = 0.222). Current evidence is limited by high risk of bias, substantial heterogeneity, and important methodological limitations, and limited independent external validation evidence precluded a reliable comparison of the external validation performance of ML and LR-family models. Accordingly, current prediction models are not yet ready for routine clinical decision-support deployment.