Biospark: scalable analysis of large numerical datasets from biological simulations and experiments using Hadoop and Spark.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 27663493.
- Also identified by DOI 10.1093/bioinformatics/btw614 and PMC identifier 6276899.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Data-parallel programming techniques can dramatically decrease the time needed to analyze large datasets. While these methods have provided significant improvements for sequencing-based analyses, other areas of biological informatics have not yet adopted them. Here, we introduce Biospark, a new framework for performing data-parallel analysis on large numerical datasets. Biospark builds upon the open source Hadoop and Spark projects, bringing domain-specific features for biology. Source code is licensed under the Apache 2.0 open source license and is available at the project website: https://www.assembla.com/spaces/roberts-lab-public/wiki/Biospark CONTACT: eroberts@jhu.eduSupplementary information: Supplementary data are available at Bioinformatics online.
Medical subject headings
- Computational Biology
- Computer Simulation
- Software