Repeat- and error-aware comparison of deletions.
other
Where this comes from
- Record sourced from PubMed, PMID 25979471.
- Also identified by DOI 10.1093/bioinformatics/btv304.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The number of reported genetic variants is rapidly growing, empowered by ever faster accumulation of next-generation sequencing data. A major issue is comparability. Standards that address the combined problem of inaccurately predicted breakpoints and repeat-induced ambiguities are missing. This decisively lowers the quality of 'consensus' callsets and hampers the removal of duplicate entries in variant databases, which can have deleterious effects in downstream analyses. We introduce a sound framework for comparison of deletions that captures both tool-induced inaccuracies and repeat-induced ambiguities. We present a maximum matching algorithm that outputs virtual duplicates among two sets of predictions/annotations. We demonstrate that our approach is clearly superior over ad hoc criteria, like overlap, and that it can reduce the redundancy among callsets substantially. We also identify large amounts of duplicate entries in the Database of Genomic Variants, which points out the immediate relevance of our approach. Implementation is open source and available from https://bitbucket.org/readdi/readdi roland.wittler@uni-bielefeld.de or t.marschall@mpi-inf.mpg.de Supplementary data are available at Bioinformatics online.
Medical subject headings
- Algorithms
- Computational Biology
- Genetic Variation
- Repetitive Sequences, Nucleic Acid
- Sequence Analysis, DNA
- Sequence Deletion
- Software