Methodological quality and performance of artificial intelligence and machine learning models for preoperative risk prediction in plastic surgery: A systematic review.
systematic_review · Level I
Where this comes from
- Record sourced from PubMed, PMID 42537558.
- Also identified by DOI 10.1016/j.bjps.2026.07.015.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Artificial intelligence (AI) and machine learning (ML) are increasingly being applied to preoperative risk prediction in plastic surgery; however, the methodological quality and clinical readiness of these models are yet to be systematically evaluated. This systematic review assessed the quality, risk of bias, and predictive performance of AI/ML preoperative risk prediction models in plastic surgery using the PROBAST+AI framework. Five databases were searched from inception through October 2025. Ten studies met the inclusion criteria, encompassing autologous breast reconstruction (n = 2), alloplastic breast reconstruction (n = 5), head and neck reconstruction (n = 1), burn surgery (n = 1), and aesthetic surgery (n = 1). Random forest was the most frequently used algorithm (n = 4), followed by neural networks (n = 2), deep forest with RUSBoost (n = 1), support vector machine (n = 1), and logistic regression (n = 1). AUC ranged from 0.66 to 0.82 among the 8 studies reporting discrimination. Critical methodological limitations were identified: only 2 studies (20%) performed external validation, 5 of 7 development studies (71.4%) had events per variable <10 indicating inadequate sample size, and 7 studies (70%) did not report model calibration. Pre-reconciliation inter-rater reliability across 102 paired domain-level ratings yielded a Cohen's kappa of 0.240 and Prevalence-Adjusted Bias-Adjusted Kappa of 0.039, consistent with published benchmarks for PROBAST-based systematic reviews. All discrepancies were resolved via structured consensus. Current AI/ML models for preoperative risk prediction in plastic surgery demonstrate variable performance and substantial methodological limitations that preclude clinical implementation. Multi-institutional prospective validation studies with rigorous methodology are needed before clinical adoption.