Ensemble Feature Selection for Microarray Data Classification.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41231697.
- Also identified by DOI 10.1109/JBHI.2025.3631336.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Microarray data classification is challenged by high dimensionality and small sample sizes, causing feature selection instability. Traditional ensemble feature selection methods struggle to balance diversity and quality effectively. We propose a novel Ensemble Feature Selection Method (EFSM) that introduces a feature mapping diversity metric to generate a robust candidate pool. EFSM first generates a diverse candidate pool of feature selectors by leveraging randomized neural networks to create multiple non-linear feature mappings (views) of the original data. Its core innovation is an ensemble pruning technique formulated as an optimization problem that jointly maximizes both the predictive accuracy of individual selectors and their pairwise diversity. We simplify this NP-hard problem by converting it into a Semi-Definite Programming (SDP) problem and deriving a novel bound for efficient solution. Finally, the rankings from the pruned ensemble are aggregated using the Borda count method. Extensive experiments on 15 biological datasets demonstrate that EFSM outperforms nine state-of-the-art feature selection methods across popular classifiers, achieving superior and stable performance for high-dimensional data analysis.