Classification of DNA sequences using Bloom filters.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 20472541.
- Also identified by DOI 10.1093/bioinformatics/btq230 and PMC identifier 2887045.
- Licence recorded as CC BY-NC.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
New generation sequencing technologies producing increasingly complex datasets demand new efficient and specialized sequence analysis algorithms. Often, it is only the 'novel' sequences in a complex dataset that are of interest and the superfluous sequences need to be removed. A novel algorithm, fast and accurate classification of sequences (FACSs), is introduced that can accurately and rapidly classify sequences as belonging or not belonging to a reference sequence. FACS was first optimized and validated using a synthetic metagenome dataset. An experimental metagenome dataset was then used to show that FACS achieves comparable accuracy as BLAT and SSAHA2 but is at least 21 times faster in classifying sequences. Source code for FACS, Bloom filters and MetaSim dataset used is available at http://facs.biotech.kth.se. The Bloom::Faster 1.6 Perl module can be downloaded from CPAN at http://search.cpan.org/ approximately palvaro/Bloom-Faster-1.6/ henrik.stranneheim@biotech.kth.se; joakiml@biotech.kth.se Supplementary data are available at Bioinformatics online.
Medical subject headings
- Algorithms
- Metagenome
- Sequence Analysis, DNA