The gift of novelty: repeat-robust k-mer-based estimators of mutation rates.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42412834.
- Also identified by DOI 10.1093/bioinformatics/btag234.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Estimating mutation rates between evolutionarily related sequences is a central problem in molecular evolution. Due to the rapid expansion of datasets, modern methods avoid costly alignment and instead compare sketches of sets of constituent k-mers. While these methods perform well on many sequences, they are not robust to highly repetitive sequences such as centromeres. We present three new estimators that are robust to the presence of repeats. The estimators are applicable in different settings, depending on whether count information is available from zero, one, or both sequences. We evaluate our estimators empirically using highly repetitive alpha satellite sequences. Each estimator performs best within its class, and our strongest estimator outperforms all other tested estimators. Our software is open-source and freely available at https://github.com/medvedevgroup/Accurate_repeat-aware_kmer_based_estimator.
Medical subject headings
- Mutation Rate
- Repetitive Sequences, Nucleic Acid
- Software
- Sequence Analysis, DNA