TagDust--a program to eliminate artifacts from next generation sequencing data.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 19737799.
- Also identified by DOI 10.1093/bioinformatics/btp527 and PMC identifier 2781754.
- Licence recorded as CC BY-NC.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Next-generation parallel sequencing technologies produce large quantities of short sequence reads. Due to experimental procedures various types of artifacts are commonly sequenced alongside the targeted RNA or DNA sequences. Identification of such artifacts is important during the development of novel sequencing assays and for the downstream analysis of the sequenced libraries. Here we present TagDust, a program identifying artifactual sequences in large sequencing runs. Given a user-defined cutoff for the false discovery rate, TagDust identifies all reads explainable by combinations and partial matches to known sequences used during library preparation. We demonstrate the quality of our method on sequencing runs performed on Illumina's Genome Analyzer platform. Executables and documentation are available from http://genome.gsc.riken.jp/osc/english/software/. timolassmann@gmail.com.
Medical subject headings
- Computational Biology
- Sequence Analysis, DNA
- Sequence Analysis, RNA
- Software