GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41537246.
- Also identified by DOI 10.1093/bioinformatics/btag014 and PMC identifier 12866627.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).
Medical subject headings
- Algorithms
- Histocompatibility Antigens Class I
- Data Mining
- Computational Biology
- Sequence Analysis, Protein