ADAR-GPT: A continually fine-tuned language model for predicting A-to-I RNA editing sites.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41512020.
- Also identified by DOI 10.1073/pnas.2529073123 and PMC identifier 12798952.
- Licence recorded as CC BY-NC-ND.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Adenosine-to-inosine (A-to-I) RNA editing by ADAR enzymes shapes transcript fate and underpins emerging RNA editing therapeutics, yet predicting which adenosines are edited remains difficult. We introduce ADAR-GPT, a model-agnostic fine-tuning framework that adapts a GPT-class language model to classify editing at candidate sites using sequence context in standardized 201 nt windows with the target adenosine explicitly marked. We train and evaluate on GTEx liver data ([Formula: see text] samples) at a clinically relevant 15% editing threshold, using a two-stage continual fine-tuning approach where lower thresholds serve as curriculum data to progressively sharpen decision boundaries. Using sequence data, ADAR-GPT demonstrates competitive or superior performance when benchmarked against established computational approaches, including convolutional and foundation model architectures, achieving a better balance of recall, precision, and specificity alongside stronger operating-curve metrics. The approach is reproducible and portable across GPT backbones without architectural changes. Beyond accurate site classification, ADAR-GPT provides practical adenosine scoring to prioritize experimental targets and inform guide RNA design, with a framework adaptable to new datasets and model architectures.
Medical subject headings
- RNA Editing
- Adenosine
- Inosine
- Adenosine Deaminase
- Computational Biology