TranscriptClean: variant-aware correction of indels, mismatches and splice junctions in long-read transcripts.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 29912287.
- Also identified by DOI 10.1093/bioinformatics/bty483 and PMC identifier 6329999.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Long-read, single-molecule sequencing platforms hold great potential for isoform discovery and characterization of multi-exon transcripts. However, their high error rates are an obstacle to distinguishing novel transcript isoforms from sequencing artifacts. Therefore, we developed the package TranscriptClean to correct mismatches, microindels and noncanonical splice junctions in mapped transcripts using the reference genome while preserving known variants. Our method corrects nearly all mismatches and indels present in a publically available human PacBio Iso-seq dataset, and rescues 39% of noncanonical splice junctions. All Python and R scripts used in this paper are available at https://github.com/dewyman/TranscriptClean.
Medical subject headings
- Genome
- INDEL Mutation
- Software