Contrastive learning enables epitope overlap predictions for targeted antibody discovery.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41726099.
- Also identified by DOI 10.1016/j.patter.2025.101419 and PMC identifier 12921510.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Computational epitope prediction remains an unmet need for therapeutic antibody development. We present three complementary approaches for predicting epitope relationships from antibody sequences. First, by analyzing approximately 18 million antibody pairs targeting around 250 protein families, we establish that over 70% of heavy-chain complementarity-determining region 3 (CDRH3) sequence identity among antibodies sharing both V genes reliably predicts overlapping epitopes. Second, we develop a supervised contrastive fine-tuning framework for antibody large language models that enriches embeddings with epitope information. Applied to SARS-CoV-2 receptor-binding-domain antibodies, this approach achieves 97% total accuracy in predicting high levels of structural overlap. Third, we create AbLang-PDB, a generalized model achieving 5-fold improvement in average precision over sequence-based methods and correlating strongly with epitope overlap (<i>ρ</i> = 0.81). Experimental validation with HIV-1 antibody 8ANC195 shows that 70% of selected candidates demonstrate HIV-1 specificity and 50% compete for binding. These models provide powerful tools for epitope-targeted antibody discovery while demonstrating contrastive learning's efficacy for encoding epitope information.