RabbitMash: accelerating hash-based genome analysis on modern multi-core architectures.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 32845281.
- Also identified by DOI 10.1093/bioinformatics/btaa754.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Mash is a popular hash-based genome analysis toolkit with applications to important downstream analyses tasks such as clustering and assembly. However, Mash is currently not able to fully exploit the capabilities of modern multi-core architectures, which in turn leads to high runtimes for large-scale genomic datasets. We present RabbitMash, an efficient highly optimized implementation of Mash which can take full advantage of modern hardware including multi-threading, vectorization and fast I/O. We show that our approach achieves speedups of at least 1.3, 9.8, 8.5 and 4.4 compared to Mash for the operations sketch, dist, triangle and screen, respectively. Furthermore, RabbitMash is able to compute the all-versus-all distances of 100 321 genomes in <5 min on a 40-core workstation while Mash requires over 40 min. RabbitMash is available at https://github.com/ZekunYin/RabbitMash. Supplementary data are available at Bioinformatics online.
Medical subject headings
- Algorithms
- Software