Mapping the space of protein binding sites with sequence-based protein language models.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40576205.
- Also identified by DOI 10.1093/bioinformatics/btaf284 and PMC identifier 12208174.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Binding sites are the key interfaces that determine a protein's biological activity, and therefore common targets for therapeutic intervention. Techniques that help us detect, compare, and contextualize binding sites are hence of immense interest to drug discovery. Here, we present an approach that integrates protein language models with a 3D tessellation technique to derive rich and versatile representations of binding sites that combine functional, structural, and evolutionary information with unprecedented detail. We demonstrate that the associated similarity metrics induce meaningful pocket clusterings by balancing local structure against global sequence effects. The resulting embeddings are shown to simplify a variety of downstream tasks: they help organize the 'pocketome' in a way that efficiently contextualizes new binding sites, construct performant druggability models, and define challenging train-test splits for believable benchmarking of pocket-centric machine-learning models. A Python package that implements the EPoCS method is freely available at https://github.com/tugceoruc/epocs.
Medical subject headings
- Proteins