Optimization of gene set annotations via entropy minimization over variable clusters (EMVC).
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 24574114.
- Also identified by DOI 10.1093/bioinformatics/btu110 and PMC identifier 4058919.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
MOTIVATION: Gene set enrichment has become a critical tool for interpreting the results of high-throughput genomic experiments. Inconsistent annotation quality and lack of annotation specificity, however, limit the statistical power of enrichment methods and make it difficult to replicate enrichment results across biologically similar datasets. RESULTS: We propose a novel algorithm for optimizing gene set annotations to best match the structure of specific empirical data sources. Our proposed method, entropy minimization over variable clusters (EMVC), filters the annotations for each gene set to minimize a measure of entropy across disjoint gene clusters computed for a range of cluster sizes over multiple bootstrap resampled datasets. As shown using simulated gene sets with simulated data and Molecular Signatures Database collections with microarray gene expression data, the EMVC algorithm accurately filters annotations unrelated to the experimental outcome resulting in increased gene set enrichment power and better replication of enrichment results. AVAILABILITY AND IMPLEMENTATION: http://cran.r-project.org/web/packages/EMVC/index.html.
Medical subject headings
- Algorithms
- Gene Expression Profiling
- Molecular Sequence Annotation