H3BERTa: A CDR-H3-specific language model for antibody repertoire analysis.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42453687.
- Also identified by DOI 10.1016/j.patter.2026.101561 and PMC identifier 13366522.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Antibodies are central to immune defense and therapeutic design, yet predicting functional sequences remains challenging. Deep learning models trained on full variable regions often struggle due to sparse experimental data, signal dilution from conserved framework residues, and extreme diversity of hypervariable loops. The heavy-chain complementarity-determining region 3 (CDR-H3) is the most variable segment, shaping antigen specificity and immune diversity. Here, we present H3BERTa, a language model trained solely on CDR-H3 sequences to assess the extent of biological information encoded by this region. H3BERTa embeddings recapitulate immunologically relevant features, including J-gene usage, inferred B cell maturation state, and antigen binding. Pseudo-perplexity profiles enable repertoire analysis, distinguishing healthy from human immunodeficiency virus type 1 (HIV-1)-derived sequences and suggesting measurable immune response signatures. These embeddings support classifiers for broadly neutralizing antibodies using limited labeled data, highlighting utility for antibody discovery. CDR-H3 alone encodes a rich immunological signal, which H3BERTa captures, offering a focused tool for repertoire analysis and antibody engineering.