Semi-supervised Omics Factor Analysis (SOFA) disentangles known and latent sources of variation in multi-omic data.
Where this comes from
- Record sourced from PubMed, PMID 42624833.
- Also identified by DOI 10.1038/s41467-026-74694-6.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
A fundamental design pattern in biomolecular studies is to assay the same set of samples (organisms, tissue biopsies, or individual cells) by multiple different 'omics assays. Group Factor Analysis (GFA) and its adaptation to high-dimensional settings, Multi-Omics Factor Analysis (MOFA), are widely used as a first-line approach to analyze such data and are effective in detecting patterns of correlation, organize them into so-called latent factors, and identify common and assay-specific factors. However, in many applications, a subset of the found factors just rediscovers already known covariates (e.g., disease subtypes, environmental covariates) while others may represent genuine novelty.Here, we present Semi-supervised Omics Factor Analysis (SOFA), a method that incorporates known covariates into the model upfront and focuses the factor discovery on novel sources of variation. We show SOFA's effectiveness for discovering novel patterns by applying it to cancer, brain development and heart failure multi-omic data sets.
Medical subject headings
- Multiomics