Sequence alignment using machine learning for accurate template-based protein structure prediction.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 31197318.
- Also identified by DOI 10.1093/bioinformatics/btz483.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Template-based modeling, the process of predicting the tertiary structure of a protein by using homologous protein structures, is useful if good templates can be found. Although modern homology detection methods can find remote homologs with high sensitivity, the accuracy of template-based models generated from homology-detection-based alignments is often lower than that from ideal alignments. In this study, we propose a new method that generates pairwise sequence alignments for more accurate template-based modeling. The proposed method trains a machine learning model using the structural alignment of known homologs. It is difficult to directly predict sequence alignments using machine learning. Thus, when calculating sequence alignments, instead of a fixed substitution matrix, this method dynamically predicts a substitution score from the trained model. We evaluate our method by carefully splitting the training and test datasets and comparing the predicted structure's accuracy with that of state-of-the-art methods. Our method generates more accurate tertiary structure models than those produced from alignments obtained by other methods. https://github.com/shuichiro-makigaki/exmachina. Supplementary data are available at Bioinformatics online.
Medical subject headings
- Proteins
- Sequence Alignment
- Sequence Analysis, Protein