CompMap: a reference-based compression program to speed up read mapping to related reference sequences.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 25282641.
- Also identified by DOI 10.1093/bioinformatics/btu656.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Exhaustive mapping of next-generation sequencing data to a set of relevant reference sequences becomes an important task in pathogen discovery and metagenomic classification. However, the runtime and memory usage increase as the number of reference sequences and the repeat content among these sequences increase. In many applications, read mapping time dominates the entire application. We developed CompMap, a reference-based compression program, to speed up this process. CompMap enables the generation of a non-redundant representative sequence for the input sequences. We have demonstrated that reads can be mapped to this representative sequence with a much reduced time and memory usage, and the mapping to the original reference sequences can be recovered with high accuracy. CompMap is implemented in C and freely available at http://csse.szu.edu.cn/staff/zhuzx/CompMap/. xiaoyang@broadinstitute.org Supplementary data are available at Bioinformatics online.
Medical subject headings
- Data Compression
- Genome, Human
- High-Throughput Nucleotide Sequencing
- Sequence Alignment
- Sequence Analysis, DNA
- Software