Zero-shot benchmarking of RNA language models in structural, functional, and evolutionary learning.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41789565.
- Also identified by DOI 10.1093/bib/bbag098 and PMC identifier 12963973.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
RNA language models (LMs) are increasingly applied to RNA structure and function analysis, yet their intrinsic representational capacities remain poorly characterized. Here, we present a standardized zero-shot evaluation of 21 RNA LMs, with representative DNA LMs included as reference controls. Three complementary tasks-attention-based RNA secondary structure prediction, embedding-based RNA classification, and mutational fitness estimation from sequence likelihoods-are evaluated without downstream fine-tuning. Our results reveal substantial variability across models and clear trade-offs between structural, functional, and evolutionary representations. RNA-specific, noncoding RNA-enriched pretraining is crucial for capturing structural information, while evolutionary signals from multiple sequence alignments substantially boost performance. Although model scaling yields gains, architectural and objective choices critically influence performance across task categories. Together, this study provides a foundational benchmark, highlights inherent challenges in learning unified RNA representations, and offers insights for developing next-generation RNA foundation models.
Medical subject headings
- RNA
- Evolution, Molecular