Task-specific pre-training for molecular property prediction.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41608985.
- Also identified by DOI 10.1093/bib/bbag010 and PMC identifier 12853129.
- Licence recorded as CC BY-NC.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Molecular property prediction is a critical task in computational chemistry and drug discovery. While deep learning has advanced this field, the increasing complexity of models contrasts with the scarcity of labeled data, leading to severe overfitting and limited generalization. In this paper, we propose TasProp, a task-specific pre-training strategy for molecular property prediction, particularly for the scenarios with small labeled datasets. To learn a robust molecular representation, TasProp first projects both labeled and unlabeled data into a unified latent space. Then, we introduce a task-specific contrastive loss that aligns closely with the final prediction task and apply it to the labeled data. This contrastive loss encourages the model to learn more cohesive and distinguishable molecular representations corresponding to property categories, which in turn, enhances the model's performance on downstream property prediction tasks. Additionally, we propose a novel data augmentation method, accompanied by a theoretical analysis, to mitigate the challenge of labeled data scarcity. With the task-specific pre-training and augmented data, TasProp outperforms the state-of-the-art methods on many molecular property prediction tasks, including three publicly available datasets and two curated datasets related to anesthesiology. Furthermore, we provide an interactive web resource to facilitate model exploration and application, allowing users to easily predict the properties of input molecules online.
Medical subject headings
- Deep Learning
- Drug Discovery
- Computational Biology