Genome-wide modelling of plant transcription factor binding captures regulatory variants associated with phenotypic traits.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42236716.
- Also identified by DOI 10.1038/s41467-026-73634-8 and PMC identifier 13234004.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
The sequence-specific recognition of cis-regulatory elements (CRE) by transcription factors (TF) propagates genotype information to phenotypes. Understanding how genetic variation affects gene regulation remains limited by the diversity and complexity of CRE interactions. Here, we address this challenge using an explainable multi-label deep learning model trained on A. thaliana DNA-binding data to capture how CRE sequence, their broader sequence context, and syntax influence TF occupancy. Once trained, the model annotates cistrome-wide TF-binding sites and uncovers condition-specific regulatory syntax. By integrating genomic and GWAS data from A. thaliana, our approach predicts differential TF-binding and identifies regulatory gene variants within quantitative trait loci. Experimental validation highlights the link between cis-regulatory variation, gene expression, and phenotypic outcomes. Finally, applying our model to untargeted DNA binding assays in Z. mays under heat-stress conditions demonstrates its potential to characterize condition-responsive TF binding in phylogenetically distant crops.
Medical subject headings
- Transcription Factors
- Arabidopsis
- Genome, Plant
- Plant Proteins