Direct high-throughput deconvolution of non-canonical bases via nanopore sequencing and bootstrapped learning.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40739090.
- Also identified by DOI 10.1038/s41467-025-62347-z and PMC identifier 12311055.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The discovery of non-canonical bases (NCBs) and development of synthetic xeno-nucleic acids (XNAs) has spawned interest in many applications in viral genomics, synthetic biology and DNA storage. However, inability to do high-throughput sequencing of NCBs has been a significant limitation. We demonstrate that XNAs with NCBs can be robustly sequenced on a MinION system ( > 2.3×10<sup>6</sup> reads/flowcell) to obtain significantly distinct signals from controls (median fold-change >6×). To enable AI-model training, we synthesized and sequenced a complex pool of 1,024 NCB-containing oligonucleotides with varied 6-mer contexts and high purity ( > 90%). Bootstrapped models assisted in data preparation, and data augmentation with spliced reads provided high context diversity, enabling learning of generalizable models to decipher NCB-containing sequences with high accuracy ( > 80%) and specificity (99%). These results highlight the versatility of nanopore sequencing for interrogating unusual nucleic acids, and the potential to transform the study of genetic material beyond those that use canonical bases.
Medical subject headings
- Nanopore Sequencing
- High-Throughput Nucleotide Sequencing
- Sequence Analysis, DNA
- Nucleic Acids