Carafe enables high quality in silico spectral library generation for data-independent acquisition proteomics.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41198693.
- Also identified by DOI 10.1038/s41467-025-64928-4 and PMC identifier 12592563.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Data-independent acquisition (DIA)-based mass spectrometry is becoming an increasingly popular mass spectrometry acquisition strategy for carrying out quantitative proteomics experiments. Most of the popular DIA search engines make use of in silico generated spectral libraries. However, the generation of high-quality spectral libraries for DIA data analysis remains a challenge, particularly because most such libraries are generated directly from data-dependent acquisition (DDA) data or are from in silico prediction using models trained on DDA data. In this study, we introduce Carafe, a tool that generates high-quality experiment-specific in silico spectral libraries by training deep learning models directly on DIA data. We demonstrate the performance of Carafe on a wide range of DIA datasets, where we observe improved fragment ion intensity prediction and peptide detection relative to existing pretrained DDA models. To make Carafe more accessible to the community, we integrate Carafe into the widely used Skyline tool.
Medical subject headings
- Proteomics
- Software