PLiCat: decoding protein-lipid interactions by large language model.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41378883.
- Also identified by DOI 10.1093/bib/bbaf665 and PMC identifier 12696715.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Protein-lipid interactions are essential for many cellular processes, as proteins associate with diverse lipid molecules to exert distinct functions. However, existing approaches are limited in discriminating among lipid categories, leaving a gap in our understanding of lipid-binding specificity. Recent advances in protein language models have opened new possibilities for discovering novel sequence insights. Here, we introduce PLiCat (Protein-Lipid interaction Categorization tool), a sequence-based framework designed to predict the categories of lipids that interact with proteins. PLiCat employs a hybrid deep learning architecture that integrates ESMC with BERT, enabling accurate and interpretable classification across eight major lipid categories. Through attribution analysis, PLiCat uncovers the sequence-encoded lipid-binding signatures and highlights residues contributing to binding specificity. Notably, PLiCat shows potential for identifying lipid-binding sites hidden in protein sequences. Furthermore, PLiCat can be used to assess the impact of pathogenic mutations on lipid-binding events. Collectively, PLiCat provides the first computational tool for lipid category prediction from protein sequence alone, offering valuable insights into lipid recognition mechanisms, and with promising applications in guiding rational protein design. The PLiCat source code and processed datasets are available at https://github.com/Noora68/PLiCat.
Medical subject headings
- Proteins
- Lipids
- Software
- Computational Biology
- Lipid Metabolism