Classification and feature selection algorithms for multi-class CGH data.
Where this comes from
- Record sourced from PubMed, PMID 18586749.
- Also identified by DOI 10.1093/bioinformatics/btn145 and PMC identifier 2718623.
- Licence recorded as CC BY-NC.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Recurrent chromosomal alterations provide cytological and molecular positions for the diagnosis and prognosis of cancer. Comparative genomic hybridization (CGH) has been useful in understanding these alterations in cancerous cells. CGH datasets consist of samples that are represented by large dimensional arrays of intervals. Each sample consists of long runs of intervals with losses and gains. In this article, we develop novel SVM-based methods for classification and feature selection of CGH data. For classification, we developed a novel similarity kernel that is shown to be more effective than the standard linear kernel used in SVM. For feature selection, we propose a novel method based on the new kernel that iteratively selects features that provides the maximum benefit for classification. We compared our methods against the best wrapper-based and filter-based approaches that have been used for feature selection of large dimensional biological data. Our results on datasets generated from the Progenetix database, suggests that our methods are considerably superior to existing methods. All software developed in this article can be downloaded from http://plaza.ufl.edu/junliu/feature.tar.gz.
Medical subject headings
- Algorithms
- Chromosome Mapping
- Gene Dosage
- Oligonucleotide Array Sequence Analysis
- Pattern Recognition, Automated
- Sequence Alignment
- Sequence Analysis, DNA