Biologically-informed regional subset analysis with CatBoost for robust tissue-of-origin prediction.
Where this comes from
- Record sourced from PubMed, PMID 41343608.
- Also identified by DOI 10.1371/journal.pone.0337106 and PMC identifier 12677570.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Accurate identification of cancer tissue/cell of origin (TOO/COO) is critical for diagnosis and treatment; yet existing whole-genome approaches demand extensive computation and often struggle with sparse mutation signals. Here, we introduce an informative regional subset framework that selects a small number of biologically and statistically significant 1Mbp genomic intervals to train a CatBoost prediction model. On a benchmark of 137 whole-genome samples across six cancer types, our method achieved a 4% gain in melanoma accuracy (from 88.0% using all 2,128 regions to 92.0% with 300 regions), a 4.4% gain in multiple myeloma (87.0% with 600 regions), and perfect (100%) accuracy in high-mutation cancers such as esophageal adenocarcinoma and glioblastoma with as few as 50 informative regions. When extended to 934 PCAWG samples spanning 14 cancer lineages, the same limited regional subsets matched or improved whole-genome performance, reaching up to 100% accuracy in gastrointestinal, skin, and brain cancers, demonstrating exceptional scalability. Our approach not only reduces computational burden and enhances interpretability but also provides a robust, generalizable tool for precision oncology and the diagnosis of cancers of unknown primary.
Medical subject headings
- Neoplasms