Probing genomic language models: Nucleotide Generative Pretrained Transformer and the role of pretraining in learned representations.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41632593.
- Also identified by DOI 10.1093/bib/bbag011 and PMC identifier 12866925.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The value and nature of the representations learned during the pretraining of genomic language models (gLMs) remain actively debated. We introduce Nucleotide Generative Pretrained Transformer (GPT), a decoder-only transformer with single-nucleotide tokenization, to dissect the role of pretraining. Through experiments varying repetitive element (RE) weights during pretraining (0.0-1.0), comparative finetuning against random initialization, linear probing of internal representations, and sparse autoencoder (SAE)-based interpretability, we evaluated the impact of pretraining and how REs in genomic data influence model learning. Models with moderate RE downweighting (0.5) consistently achieved optimal performance across seven genomic classification tasks, with pretrained models providing substantial performance gains over baselines. SAE feature annotation via sequence alignment revealed substantial RE-associated patterns in the pretrained model internal representations, suggesting that REs-which comprise 30%-60% of mammalian genomes-may dominate the pretraining objective. Our findings support the utility of pretraining and underscore the need for pretraining strategies that better accommodate repetitive sequences across the genome while also fostering the learning of less common but biologically important representations. This study highlights a key challenge for gLMs: ensuring that models broadly learn functional genomic syntax beyond simply recognizing ubiquitous repeats.
Medical subject headings
- Genomics
- Nucleotides
- Models, Genetic