Compression of DNA sequence reads in FASTQ format.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 21252073.
- Also identified by DOI 10.1093/bioinformatics/btr014.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Modern sequencing instruments are able to generate at least hundreds of millions short reads of genomic data. Those huge volumes of data require effective means to store them, provide quick access to any record and enable fast decompression. We present a specialized compression algorithm for genomic data in FASTQ format which dominates its competitor, G-SQZ, as is shown on a number of datasets from the 1000 Genomes Project (www.1000genomes.org). DSRC is freely available at http:/sun.aei.polsl.pl/dsrc.
Medical subject headings
- Data Compression
- Sequence Analysis, DNA