Leveraging basecaller's move table to generate a lightweight k-mer model for nanopore sequencing analysis.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40085000.
- Also identified by DOI 10.1093/bioinformatics/btaf111 and PMC identifier 11964489.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Nanopore sequencing by Oxford Nanopore Technologies (ONT) enables direct analysis of DNA and RNA by capturing raw electrical signals. Different nanopore chemistries have varied k-mer lengths, current levels, and standard deviations, which are stored in "k-mer models." In cases where official models are lacking or unsuitable for specific sequencing conditions, tailored k-mer models are crucial to ensure precise signal-to-sequence alignment, analysis and interpretation. The process of transforming raw signal data into nucleotide sequences, known as basecalling, is a fundamental step in nanopore sequencing. In this study, we leverage the move table produced by ONT's basecalling software to create a lightweight de novo k-mer model for RNA004 chemistry. We demonstrate the validity of our custom k-mer model by using it to guide signal-to-sequence alignment analysis, achieving high alignment rates (97.48%) compared to larger default models. Additionally, our 5-mer model exhibits similar performance as the default 9-mer models another analysis, such as detection of m6A RNA modifications. We provide our method, termed Poregen, as a generalizable approach for creation of custom, de novo k-mer models for nanopore signal data analysis. Poregen is an open source package under an MIT license: https://github.com/hiruna72/poregen.
Medical subject headings
- Nanopore Sequencing
- Software
- Sequence Analysis, DNA
- Nanopores
- Sequence Analysis, RNA