TSEDTA: a transformer-based neural network with SMILES transformer and ESM2 embeddings for drug-target binding affinity prediction.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42105214.
- Also identified by DOI 10.1093/bioinformatics/btag298 and PMC identifier 13211980.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Drug-target binding affinity (DTA) prediction plays a vital role in drug repositioning. The emergence of large language models (LLMs) has introduced new perspectives for predicting DTA. Herein, we present TSEDTA, a Transformer-based neural network with SMILES Transformer and ESM2 embeddings for predicting DTA. It leverages pre-trained LLMs (SMILES Transformer and ESM2) to extract deep evolutionary representations from drug SMILES and protein sequences. The representations are directly fused with raw sequence embeddings and processed via dual Transformer encoders to capture complex local and global dependencies. The experiments demonstrate that TSEDTA outperforms ten advanced models on the Davis and KIBA datasets, and seven on the BindingDB dataset. Ablation studies show that incorporating LLM embeddings significantly improves the performance of TSEDTA. Furthermore, a practical case study demonstrates its real-world applicability. Ultimately, TSEDTA provides a highly accurate, robust tool for DTA prediction, offering new insights into the application of LLMs for DTA tasks. The source code and data are available at: https://github.com/SunXu24Math/TSEDTA. The version of record is archived in Zenodo with the DOI: 10.5281/zenodo.19103249.
Medical subject headings
- Neural Networks, Computer
- Proteins
- Computational Biology