scGenoByte: a GenoByte embedding transformer with biological priors for cell type annotation.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42407121.
- Also identified by DOI 10.1093/bib/bbag369.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Effective cell representation learning is crucial for accurate cell annotation and the deciphering of cellular heterogeneity in single-cell RNA sequencing (scRNA-seq) analysis. Current foundation models have achieved superior performance compared with traditional methods. However, due to data sparsity and the complexity of model, existing methods often compromise by selecting highly variable genes or filtering for nonzero expressions, which discard potentially significant genes. Thus, modeling the complete transcriptome for cell representation remains computationally challenging; we present scGenoByte, a unified framework designed to enhance cell representation learning through biologically informed full-gene modeling. To enable efficient modeling of the full transcriptome, we design GenoBytes, biologically coherent units that are constructed by leveraging biological priors in terms of protein-protein interaction network and gene paralogy network. Furthermore, considering that the information of protein and pathway is critical for analyzing cell functions and representation, scGenoByte encapsulates biological priors by harmonizing GenoByte embeddings with protein representations and leveraging an auxiliary task of pathway activity prediction to impose pathway-guided regularization. Extensive results on eight datasets have shown that scGenoByte achieves better performance than competing methods, which confirms the efficacy of combining full-gene context with biological priors.