Performance of Machine Learning Models in Predicting Outcomes After ACDF: A Systematic Review and Meta-Analysis of 443 000 Patients.

P Mitre, Lucas; Sanglard, Rômulo S; Brown, Ethan D L; Ghanekar, Shaila D; Serrato, Paul; Elsamadicy, Aladine A · Global Spine J · 2026

meta_analysis · Level I

Where this comes from

Abstract

Study DesignSystematic review and meta-analysis.ObjectiveDirect head-to-head comparison of machine learning models aiming to predict outcomes in Anterior Cervical Discectomy and Fusion (ACDF) is necessary because existing studies typically evaluate algorithms in isolation, using heterogeneous datasets, features, and performance metrics, which limits interpretability and prevents meaningful comparison of predictive performance.MethodsWe conducted a systematic review and meta-analysis according to PRISMA guidelines, searching PubMed, Embase, Web of Science and Cochrane through December 2024. We identified 26 studies (n = 443 445 patients) that developed ML models for ACDF outcomes. Algorithms were categorized into five classes by taxonomy: Logistic regression, tree-based, boosting ensembles, kernel methods and NNs. Pooled ML models' AUCs and accuracy were extracted and estimated via random-effects inverse-variance model.ResultsOverall discrimination ranged from 0.59 for major complications to 0.81 for adjacent-level disease. Logistic regression led in predicting unfavorable discharge (AUC 0.76), readmission/reintervention (0.68) and cost of care (0.83). Boosting ensembles excelled in predicting thromboembolic events (AUC 0.74; 0.68-0.80; <i>P</i> < 0.0001). Neural networks achieved the highest discrimination for opioid prescription (AUC 0.80; 0.75-0.85; <i>P</i> = 0.02) and adjacent-level disease (AUC 0.81; 0.72-0.91; <i>P</i> < 0.01). Kernel methods delivered an exceptional AUC of 0.97 (0.96-0.97) for adjacent-level fusion but underperformed for other outcomes (AUC 0.43-0.49). Decision-tree and mixed-ensemble approaches demonstrated intermediate performance for various outcomes (AUC range 0.54-0.75).ConclusionLogistic regression and gradient-boosting models offer robust, generalizable discrimination across diverse ACDF outcomes. Neural networks and kernel methods showed endpoint-specific strengths. These data support prospective validation and rapid integration of ML-driven risk calculators into perioperative workflows.