Linear and non-linear feature selection for survival modelling of ischemic heart disease in patients with diabetes: Insights from electronic health records.
retrospective_cohort · Level III
Where this comes from
- Record sourced from PubMed, PMID 41856423.
- Also identified by DOI 10.1016/j.jbi.2026.105020.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Cardiovascular disease, particularly ischemic heart disease (IHD), is a major cause of morbidity and mortality in patients with type 2 diabetes mellitus (T2DM). Accurate risk prediction is essential, yet the influence of linear versus non-linear feature selection and survival modelling on performance and feature interpretability remains insufficiently explored. We analysed 12,281 patients with T2DM from a university hospital EHR, followed for up to 15 years. The outcome was incident IHD, defined by ICD-10 codes. After variance thresholding and multicollinearity filtering, 263 predictors were retained. Features were selected using univariate Cox regression, Random Survival Forest (RSF), and their consensus. Seven survival models (Cox, Ridge Cox, Weibull, RSF, Gradient-Boosted Survival [GBS], Survival SVM, XGBoost) were trained using five-fold cross-validation. Performance was assessed using the concordance index (C-index) and Integrated Brier Score (IBS). Feature importance stability, Spearman rank correlations, and top-20 feature contributions were compared across models. RSF, Weibull, GBS, and SSVM achieved the highest discrimination (C-index up to 0.78) with comparable calibration, whereas XGBoost consistently performed poorest (C-index 0.66-0.68). Linear models produced stable, diffuse feature importance profiles, while non-linear models concentrated importance on a narrower set of dominant predictors. Across all approaches, cardiovascular comorbidities (I10, I50, I15, I11) remained the most influential predictors. Linear models ensured stability and interpretability, whereas non-linear methods enhanced discrimination and calibration but increased variability. Combining linear and non-linear feature selection provided complementary insights for EHR-based risk prediction of IHD in T2DM.