ntHash: recursive nucleotide hashing.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 27423894.
- Also identified by PMC identifier 5181554.
- Licence recorded as CC BY-NC.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Hashing has been widely used for indexing, querying and rapid similarity search in many bioinformatics applications, including sequence alignment, genome and transcriptome assembly, k-mer counting and error correction. Hence, expediting hashing operations would have a substantial impact in the field, making bioinformatics applications faster and more efficient. We present ntHash, a hashing algorithm tuned for processing DNA/RNA sequences. It performs the best when calculating hash values for adjacent k-mers in an input sequence, operating an order of magnitude faster than the best performing alternatives in typical use cases. ntHash is available online at http://www.bcgsc.ca/platform/bioinfo/software/nthash and is free for academic use. hmohamadi@bcgsc.ca or ibirol@bcgsc.caSupplementary information: Supplementary data are available at Bioinformatics online.
Medical subject headings
- Algorithms
- Nucleotides