AlignBucket: a tool to speed up 'all-against-all' protein sequence alignments optimizing length constraints.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 26231432.
- Also identified by DOI 10.1093/bioinformatics/btv451.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The next-generation sequencing era requires reliable, fast and efficient approaches for the accurate annotation of the ever-increasing number of biological sequences and their variations. Transfer of annotation upon similarity search is a standard approach. The procedure of all-against-all protein comparison is a preliminary step of different available methods that annotate sequences based on information already present in databases. Given the actual volume of sequences, methods are necessary to pre-process data to reduce the time of sequence comparison. We present an algorithm that optimizes the partition of a large volume of sequences (the whole database) into sets where sequence length values (in residues) are constrained depending on a bounded minimal and expected alignment coverage. The idea is to optimally group protein sequences according to their length, and then computing the all-against-all sequence alignments among sequences that fall in a selected length range. We describe a mathematically optimal solution and we show that our method leads to a 5-fold speed-up in real world cases. The software is available for downloading at http://www.biocomp.unibo.it/∼giuseppe/partitioning.html. giuseppe.profiti2@unibo.it. Supplementary data are available at Bioinformatics online.
Medical subject headings
- Algorithms
- Computational Biology
- Databases, Protein
- Proteins
- Sequence Alignment
- Software