Negative dataset selection impacts machine learning-based predictors for multiple bacterial species promoters.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40152247.
- Also identified by DOI 10.1093/bioinformatics/btaf135 and PMC identifier 11993300.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Advances in bacterial promoter predictors based on machine learning have greatly improved identification metrics. However, existing models overlooked the impact of negative datasets, previously identified in GC-content discrepancies between positive and negative datasets in single-species models. This study aims to investigate whether multiple-species models for promoter classification are inherently biased due to the selection criteria of negative datasets. We further explore whether the generation of synthetic random sequences (SRS) that mimic GC-content distribution of promoters can partly reduce this bias. Multiple-species predictors exhibited GC-content bias when using CDS as a negative dataset, suggested by specificity and sensibility metrics in a species-specific manner, and investigated by dimensionality reduction. We demonstrated a reduction in this bias by using the SRS dataset, with less detection of background noise in real genomic data. In both scenarios DNABERT showed the best metrics. These findings suggest that GC-balanced datasets can enhance the generalizability of promoter predictors across Bacteria. The source code of the experiments is freely available at https://github.com/maigonzalezh/MultispeciesPromoterClassifier.
Medical subject headings
- Machine Learning
- Promoter Regions, Genetic
- Bacteria