Sparse autoencoders reveal interpretable cell-type programs in single-cell foundation model representations.

Kendiukhov, Ihor · J Biomed Inform · 2026

basic_science · Level V

Where this comes from

Abstract

Single-cell foundation models such as scGPT learn rich representations of cellular identity, yet the biological programs encoded in their internal activations remain opaque. We investigate whether sparse autoencoders (SAEs), a mechanistic interpretability technique from AI safety research, can decompose these representations into sparse, biologically interpretable features. We extract residual-stream activations from all 12 transformer layers of a pre-trained scGPT model processing 1000 human immune cells from the Tabula Sapiens atlas. We train SAEs with dictionary size M=2,048 at multiple sparsity levels (λ∈{1,3,10}) and evaluate recovered features using cell-type classification (AUROC), gene set enrichment (Fisher's exact test, FDR <0.05), and comparison with PCA baselines. Biological foundation models require substantially stronger L<sub>1</sub> regularisation (λ≥1) than language models (λ≈0.01-0.1) to achieve genuine sparsity. Appropriately regularised SAEs achieve L<sub>0</sub>≈48-54 active features while maintaining R<sup>2</sup>>0.76. Later-layer SAE features recover biologically coherent programs aligned with annotated cell types, with 64% of alive features receiving significant gene set annotations at layer 11 (λ=3). We observe a sparsity-dead-feature trade-off: at λ=10, up to 66% of dictionary elements become inactive. Mechanistic interpretability methods developed for large language models transfer productively to biological foundation models, but require domain-specific calibration. SAEs provide a principled approach to understanding what single-cell foundation models learn about cellular identity, with potential applications in model auditing and biological discovery.