Large-scale structure-informed multiple sequence alignment of proteins with SIMSApiper.
Where this comes from
- Record sourced from PubMed, PMID 38648741.
- Also identified by DOI 10.1093/bioinformatics/btae276 and PMC identifier 11099654.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
SIMSApiper is a Nextflow pipeline that creates reliable, structure-informed MSAs of thousands of protein sequences faster than standard structure-based alignment methods. Structural information can be provided by the user or collected by the pipeline from online resources. Parallelization with sequence identity-based subsets can be activated to significantly speed up the alignment process. Finally, the number of gaps in the final alignment can be reduced by leveraging the position of conserved secondary structure elements. The pipeline is implemented using Nextflow, Python3, and Bash. It is publicly available on github.com/Bio2Byte/simsapiper.
Medical subject headings
- Software
- Proteins
- Sequence Alignment
- Sequence Analysis, Protein