Spatial histology and gene-expression representation and generative learning via online self-distillation contrastive learning.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40618351.
- Also identified by DOI 10.1093/bib/bbaf317 and PMC identifier 12229093.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Spatial transcriptomics quantifies spatial molecular profiles alongside histology, enabling computational prediction of spatial gene expression distribution directly from whole slide images. Inspired by image-to-text alignment and generation, we introduce Magic, a self-training contrastive learning model designed for histology-to-gene expression prediction. Magic (i) employs contrastive learning to derive shared embeddings for histology and gene expression while utilizing a momentum-based module to generate pseudo-targets to reduce the impact of noise; and (ii) leverages a transformer-based decoder to predict the expression of 300 genes based on histological features. Trained on 75 760 spots from 56 breast cancer slices and validated on 11 026 spots from five independent slices, Magic outperforms existing methods in aligning and generating histology-gene expression data, achieving a 10% improvement over the second-best approach. Furthermore, Magic demonstrates robust generalization, effectively predicting gene expression in colorectal cancer samples and The Cancer Genome Atlas (TCGA) datasets through zero-shot learning. Notably, Magic's predicted gene expression captures interpatient differences, highlighting its strong potential for clinical applications.
Medical subject headings
- Breast Neoplasms
- Gene Expression Profiling
- Machine Learning
- Computational Biology
- Transcriptome