cl-dash: rapid configuration and deployment of Hadoop clusters for bioinformatics research in the cloud.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 26428290.
- Also identified by DOI 10.1093/bioinformatics/btv553 and PMC identifier 4708102.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
: One of the solutions proposed for addressing the challenge of the overwhelming abundance of genomic sequence and other biological data is the use of the Hadoop computing framework. Appropriate tools are needed to set up computational environments that facilitate research of novel bioinformatics methodology using Hadoop. Here, we present cl-dash, a complete starter kit for setting up such an environment. Configuring and deploying new Hadoop clusters can be done in minutes. Use of Amazon Web Services ensures no initial investment and minimal operation costs. Two sample bioinformatics applications help the researcher understand and learn the principles of implementing an algorithm using the MapReduce programming pattern. Source code is available at https://bitbucket.org/booz-allen-sci-comp-team/cl-dash.git. hodor_paul@bah.com.
Medical subject headings
- Algorithms
- Biomedical Research
- Computational Biology
- Genomics
- Information Storage and Retrieval
- Search Engine
- Software