Language model-guided anticipation and discovery of mammalian metabolites.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41535467.
- Also identified by DOI 10.1038/s41586-025-09969-x and PMC identifier 12960238.
- Licence recorded as CC BY-NC-ND.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Despite decades of study, large parts of the mammalian metabolome remain unexplored<sup>1</sup>. Mass spectrometry-based metabolomics routinely detects thousands of small molecule-associated peaks in human tissues and biofluids, but typically only a small fraction of these can be identified, and structure elucidation of novel metabolites remains challenging<sup>2-4</sup>. Biochemical language models have transformed the interpretation of DNA, RNA and protein sequences, but have not yet had a comparable impact on understanding small molecule metabolism. Here we present an approach that leverages chemical language models<sup>5-7</sup> to anticipate the existence of previously uncharacterized metabolites. We introduce DeepMet, a chemical language model that learns from the structures of known metabolites to anticipate the existence of previously unrecognized metabolites. Integration of DeepMet with mass spectrometry-based metabolomics data facilitates metabolite discovery. We harness DeepMet to reveal several dozen structurally diverse mammalian metabolites. Our work demonstrates the potential for language models to advance the mapping of the mammalian metabolome.
Medical subject headings
- Mammals
- Metabolomics
- Models, Chemical