A trimodal protein language model enables advanced protein searches.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41039041.
- Also identified by DOI 10.1038/s41587-025-02836-0.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
ProTrek unifies protein sequence, structure and natural language function in a trimodal language model through contrastive learning, enabling comprehensive searches between any two modalities, including within modality. ProTrek surpasses current alignment tools (for example, Foldseek and MMseqs2) in speed and accuracy for identifying functionally related proteins. Computational and wet-lab experimental validations show that the ProTrek server ( www.search-protrek.com ), with precomputed embeddings for over 5 billion proteins, efficiently processes and analyzes large-scale protein repositories.