Efficient reconstruction of full-length RNA isoforms using ISAtools and large-scale PacBio circular consensus sequencing data.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42531061.
- Also identified by DOI 10.1093/bib/bbag403.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Accurate reconstruction and quantification of full-length RNA isoforms remain challenging in long-read RNA sequencing due to sequencing artifacts, complex splicing, and incomplete annotations. Although Pacific Biosciences circular consensus sequencing (PacBio CCS) provides high-fidelity long reads, scalable and annotation-flexible analysis frameworks remain limited. Here, we present ISAtools, an efficient framework specifically designed for PacBio CCS data. ISAtools introduces a splice site chain representation that unifies read alignments and transcript annotations into a compact format for scalable isoform reconstruction. It further integrates unsupervised density-based clustering for transcription start and end site detection and a truncation-aware quantification strategy. Benchmarking on simulated, SIRV spike-in controls, and biological datasets shows that ISAtools accurately reconstructs both annotated and novel isoforms with reliable splice structures and transcript boundaries while maintaining high computational efficiency. ISAtools processes up to 80 million reads in ~21 min using ~6 GB memory and scales to nearly 1 billion reads in under 4 h with ~8 GB memory, supporting large-scale PacBio Iso-Seq studies.
Medical subject headings
- RNA Isoforms
- Sequence Analysis, RNA
- Software