PLiCat: decoding protein-lipid interactions by large language model.

Dong, Feitong; Wu, Jingrou · Brief Bioinform · 2025

basic_science · Level V

Where this comes from

Abstract

Protein-lipid interactions are essential for many cellular processes, as proteins associate with diverse lipid molecules to exert distinct functions. However, existing approaches are limited in discriminating among lipid categories, leaving a gap in our understanding of lipid-binding specificity. Recent advances in protein language models have opened new possibilities for discovering novel sequence insights. Here, we introduce PLiCat (Protein-Lipid interaction Categorization tool), a sequence-based framework designed to predict the categories of lipids that interact with proteins. PLiCat employs a hybrid deep learning architecture that integrates ESMC with BERT, enabling accurate and interpretable classification across eight major lipid categories. Through attribution analysis, PLiCat uncovers the sequence-encoded lipid-binding signatures and highlights residues contributing to binding specificity. Notably, PLiCat shows potential for identifying lipid-binding sites hidden in protein sequences. Furthermore, PLiCat can be used to assess the impact of pathogenic mutations on lipid-binding events. Collectively, PLiCat provides the first computational tool for lipid category prediction from protein sequence alone, offering valuable insights into lipid recognition mechanisms, and with promising applications in guiding rational protein design. The PLiCat source code and processed datasets are available at https://github.com/Noora68/PLiCat.

Medical subject headings