Carcinogenicity prediction via multi-task learning of cross-organ representations with attention mechanisms.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42242680.
- Also identified by DOI 10.1093/bib/bbag296.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Cancer is caused by the uncontrolled growth and division of abnormal cells. In industrialized societies, chemical exposure is a leading cause of cancer. Since certain compounds induce cancer by damaging genes or affecting cellular metabolism, studying carcinogens is essential. However, previous studies used separate models for each organ and failed to capture carcinogenic features shared across organs, limiting generalization. Thus, this study developed a multi-task learning framework to predict organ-specific carcinogenicity in the liver, lung, stomach, and breast. This framework consisted of a shared layer and task-specific layers. The shared layer contains a graph attention network layer to make atom-level representations, along with parallel fully connected layers designed for each task combination. The resulting shared representations are passed to task-specific layers to predict organ-specific carcinogenicity. The training process followed stepwise learning, whereby the model was first trained using partially labeled data to capture cross-organ representations and determine initial weights. In the second step, fully labeled data for all organs were used for final training. The proposed multi-task model achieved superior performance in the liver, lung, and stomach tasks. Notably, it recorded the highest area under the receiver operating characteristic curve in the stomach task (0.7636), outperforming the single-task model (0.7055) and all comparative models (0.5527-0.7418). The highest area under the precision-recall curve was observed in the liver task (0.9646), surpassing the single-task model (0.9505) and all comparative models (0.9373-0.9621). We further analyzed molecules with high predicted carcinogenicity and identified critical substructures using an attention mechanism. This research can contribute to predicting organ-specific carcinogenicity of candidate chemicals in the early stages of drug development.