SeQuiLa: an elastic, fast and scalable SQL-oriented solution for processing and querying genomic intervals.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 30428005.
- Also identified by DOI 10.1093/bioinformatics/bty940.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Efficient processing of large-scale genomic datasets has recently become possible due to the application of 'big data' technologies in bioinformatics pipelines. We present SeQuiLa-a distributed, ANSI SQL-compliant solution for speedy querying and processing of genomic intervals that is available as an Apache Spark package. Proposed range join strategy is significantly (∼22×) faster than the default Apache Spark implementation and outperforms other state-of-the-art tools for genomic intervals processing. The project is available at http://biodatageeks.org/sequila/. Supplementary data are available at Bioinformatics online.
Medical subject headings
- Genomics
- Software