Explainable diabetes prediction using a stacked ensemble framework.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42743221.
- Also identified by DOI 10.1371/journal.pone.0352313.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Diabetes is a chronic disease that significantly increases the risk of serious complications such as cardiovascular disorders and kidney failure. Early detection through predictive modeling can lead to timely interventions and significantly improve patient health outcomes. Several machine learning approaches have been proposed for predicting diabetes, but the main focus has been on improving prediction accuracy, while interpretability has received limited attention. To address this gap, we present a robust and explainable machine learning framework based on a stacked ensemble model that uses Random Forest, Support Vector Machine, and Gradient Boosting as base learners and Catboost as the meta-learner. The model was trained on the PIMA Indians Diabetes dataset using a preprocessing pipeline that included standard scaling, analysis of variance (ANOVA)- F-score-based feature selection, and class balancing with the synthetic minority oversampling technique (SMOTE). The proposed ensemble model outperformed the latest methods with an accuracy of 86%. We integrated explainable AI techniques such as Local Interpretable Model-Agnostic Explanations (LIME) and Shapley Additive Explanation (SHAP) to enhance transparency, which provide both local and global interpretability by identifying the most influential features contributing to each prediction, thus supporting more informed and trustworthy decision-making in healthcare applications.
Medical subject headings
- Diabetes Mellitus