DDA-BERT: end-to-end training for data-dependent acquisition mass spectrometry-based proteomics.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42045233.
- Also identified by DOI 10.1038/s41467-026-72246-6.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Peptide-spectrum match (PSM) rescoring is critical for accurate peptide identification in data-dependent acquisition (DDA)-based proteomics. Existing rescoring frameworks typically combine search-engine scores with heuristic or learned auxiliary features to refine PSM ranking and confidence estimation. Although recent approaches incorporate deep learning-derived representations of spectra, retention time, or ion mobility, the final decision stage still commonly relies on separately trained shallow classifiers, constraining the expressive capacity of the overall scoring framework. Here, we introduce DDA-BERT, a transformer-based end-to-end deep learning model trained with ~271 million PSMs from 11 species. DDA-BERT consistently outperforms existing tools across species-specific benchmarks, achieving 2.24%-269.35%, 3.73%-141.46%, 5.53%-45.64%, and 3.68%-62.77% increases in peptide identifications on human, yeast, Drosophila, and Arabidopsis datasets, respectively. The model retains high sensitivity in trace-level proteomics samples. On HLA immunopeptidomics data, DDA-BERT further increases peptide identifications by 4.14%-87.47%. The main limitations of DDA-BERT include the requirement for GPU-based computing and the need for substantial, diverse training datasets to achieve optimal model performance. This study introduces an alternative DDA rescoring approach and establishes a methodological foundation for scalable, AI-driven peptide identification in DDA proteomics.