Full-length isoform constructor (FLIC) - a tool for isoform discovery based on long reads.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41028964.
- Also identified by DOI 10.1093/bioinformatics/btaf551 and PMC identifier 12771368.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Advances in high-throughput sequencing have illuminated the complexity of transcriptome landscape in eukaryotes. An inherent part of this complexity is the presence of multiple isoforms generated by the alternative splicing and the use of alternative transcription start and polyadenylation sites. However, currently available tools have limited capacity to infer full-length isoforms. We developed a new pipeline, FLIC (full-length isoform constructor). FLIC is based on the long-read transcriptome data and integrates several key features: (1) utilizing biological replicate concordance to filter out noise and artifacts; (2) employing peak calling to precisely identify transcription start and polyadenylation sites; (3) enabling robust isoform reconstruction with minimal reliance on existing annotations. We evaluated FLIC using a dedicated set of real and simulated data of Arabidopsis thaliana cDNA sequencing. Results demonstrate that FLIC accurately reconstructs known and novel isoforms, outperforming existing tools, especially in the absence of reference annotations. A direct comparison with CAGE, currently regarded as the gold standard for transcription start site identification, shows that FLIC is equally accurate, while being much less time-consuming. Thus, FLIC provides a valuable tool for comprehensive transcript characterization, particularly for non-model organisms or when dealing with incomplete or inaccurate annotations. FLIC is available at https://github.com/albidgy/FLIC.