Finding low-complexity DNA sequences with longdust.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41826798.
- Also identified by DOI 10.1093/bioinformatics/btag112 and PMC identifier 13003316.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Low-complexity (LC) DNA sequences are compositionally repetitive sequences that are often associated with spurious homologous matches and variant calling artifacts. While algorithms for identifying LC sequences exist, they either lack concise mathematical definition of complexity or are inefficient with long or variable context windows. Longdust is a new algorithm that efficiently identifies long LC sequences including centromeric satellite and tandem repeats with moderately long motifs. It defines string complexity by statistically modeling the k-mer count distribution with the parameters: the k-mer length, the context window size and a threshold on complexity. Longdust exhibits high performance on real data and high consistency with existing methods. https://github.com/lh3/longdust.
Medical subject headings
- Algorithms
- Sequence Analysis, DNA
- Software
- DNA