TADA: taxonomy-aware dataset aggregator.
Where this comes from
- Record sourced from PubMed, PMID 38060257.
- Also identified by DOI 10.1093/bioinformatics/btad742 and PMC identifier 10733731.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
The profusion of sequenced genomes across the bacterial and archeal domains offers unprecedented possibilities for phylogenetic and comparative genomic analyses. In general, phylogenetic reconstruction is improved by the use of more data. However, including all available data is (i) not computationally tractable, and (ii) prone to biases, as the abundance of genomes is very unequally distributed over the biological diversity. Thus, in most cases, subsampling taxa to build a phylogeny is necessary. Currently, though, there is no available software to perform that handily. Here we present TADA, a taxonomic-aware dataset selection workflow that allows sampling across user-defined portions of the prokaryotic diversity with variable granularity, while setting constraints on genome quality and balance between branches. TADA is implemented as a snakemake workflow and is freely available at https://github.com/emilhaegglund/TADA.
Medical subject headings
- Software
- Genome