Underperformance of Machine Learning Models Predicting Readmission and Prolonged Length of Stay Following Total Knee Arthroplasty Among Underrepresented Populations: Fairness Analysis and Mitigation Strategies.

Shimizu, Michelle R; Raza, Marium; Xiao, Pengwei; Li, Zhijun; Freeman, Isaiah; Kwon, Young-Min · J Arthroplasty · 2026

retrospective_cohort · Level III

Where this comes from

Abstract

While machine learning (ML) models demonstrate high predictive accuracy, recent studies reveal that ML models underperform for smaller subcohorts such as racial and ethnic minorities, suggesting inherent ML biases that may exacerbate health disparities. To assume a "one-size-fits-all" approach perpetuates inequities in decision-making for underrepresented groups. This study assessed validated ML model "fairness" for total knee arthroplasty (TKA) outcomes and explored bias mitigation strategies to enhance equitable prediction performance. There were four ML models that were developed and validated to predict readmission and prolonged lengths of stay (pLOS) following TKA. Bias assessment was performed using protected attributes (age, sex, race, and ethnicity), and various fairness metrics were evaluated. There were three bias mitigation strategies incorporated and trialed for each algorithm, and the most effective strategy for a given protected attribute was integrated with the original model and reassessed based on the fairness metrics. The random forest model had the best predictive performance (Readmission<sub>AUC</sub> = 0.98; pLOS<sub>AUC</sub> = 0.92). Inferior predictive equality for readmission was demonstrated in women (0.47 versus men), non-White (0.75 versus White), and Hispanic or LatinoX (0.53 versus not Hispanic/LatinoX) subcohorts. Significant differences in all, but accuracy equality metrics for pLOS were unveiled across underprivileged cohorts. Mitigation strategies were effective in both ML models; however, trade-offs were observed between the fairness metrics. Our study highlights the "unfairness" or underperformance of current validated ML models in predicting readmission and pLOS in smaller subcohorts after TKA. While mitigation techniques improved fairness metrics, no approach fully eliminated bias, underscoring the need to carefully consider the trade-offs of each strategy. Our findings suggest substantial efforts should be made to correct potential bias in subcohorts before clinical algorithms are published or utilized as clinical decision-support tools to ensure patient equity.

Anatomy