A new representation for protein secondary structure prediction based on frequent patterns.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 16940325.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
A new representation for protein secondary structure prediction based on frequent amino acid patterns is described and evaluated. We discuss in detail how to identify frequent patterns in a protein sequence database using a level-wise search technique, how to define a set of features from those patterns and how to use those features in the prediction of the secondary structure of a protein sequence using support vector machines (SVMs). Three different sets of features based on frequent patterns are evaluated in a blind testing setup using 150 targets from the EVA contest and compared to predictions of PSI-PRED, PHD and PROFsec. Despite being trained on only 940 proteins, a simple SVM classifier based on this new representation yields results comparable to PSI-PRED and PROFsec. Finally, we show that the method contributes significant information to consensus predictions. The method is available from the authors upon request.
Medical subject headings
- Algorithms
- Models, Chemical
- Models, Molecular
- Protein Structure, Secondary
- Proteins
- Sequence Alignment
- Sequence Analysis, Protein