Cardiovascular risk prediction and influencing predictors identification among Bangladeshi individuals using machine learning algorithms and association rule mining.
other
Where this comes from
- Record sourced from PubMed, PMID 41056324.
- Also identified by DOI 10.1371/journal.pone.0333913 and PMC identifier 12503297.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Cardiovascular disease (CVD) encompasses a group of disorders that affect the heart and blood vessels, making it one of the leading causes of death globally, including in Bangladesh. Applying predictive modeling for the early identification and detection of CVD holds significant promise for saving lives by enhancing prediction precision through machine learning algorithms. Therefore, this study aimed to predict high-risk individuals for CVD using machine learning algorithms and identify its influencing predictors by association mining rules among individuals in Bangladesh. This study utilized the most recent Bangladesh Demographic and Health Survey (BDHS) 2022 data, which encompassed 2,221 respondents. A Boruta-based feature selection method is employed to determine the important features associated with the high risk of CVD. Different machine learning algorithms, including logistic regression, Naïve Bayes, artificial neural network, random forest, and extreme gradient boosting (XGB), are adopted to predict the high-risk individuals for CVD in the training dataset. The predictive performance of the models is evaluated using accuracy, precision, recall, F1-score, and area under the curve (AUC) in the testing set. Additionally, the most significant rules are analyzed using the association mining technique to identify the influencing predictors of high risk of CVD. The Boruta method indicated that age, residence, marital status, wealth, having an air conditioner (AC), and body mass index (BMI) are important predictors of high risk of CVD. The XGB-based predictive model achieves impressive performance compared to other models, with an accuracy of 68.22%, precision of 69.70%, F1-score of 79.54%, and AUC of 0.721. The association rules identified that being aged 65 or older, living in an urban area, having the richest wealth status, having AC, and being widowed are the influencing predictors of high risk of CVD. This study emphasizes the potential of XGB in predicting high-risk individuals for CVD and enhances the investigation of key factors contributing to CVD risk in this population, thereby facilitating the development of targeted prevention strategies that can effectively mitigate the high CVD risk.
Medical subject headings
- Cardiovascular Diseases
- Machine Learning
- Data Mining