Efficient and accurate framework for genome-wide gene-environment interaction analysis in large-scale biobanks.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40157913.
- Also identified by DOI 10.1038/s41467-025-57887-3 and PMC identifier 11955004.
- Licence recorded as CC BY-NC-ND.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Gene-environment interaction (G×E) analysis elucidates the interplay between genetic and environmental factors. Genome-wide association studies (GWAS) have expanded to encompass complex traits like time-to-event and ordinal traits, which provide richer phenotypic information. However, most existing scalable approaches focus only on quantitative or binary traits. Here we propose SPAGxE<sub>CCT</sub>, a scalable and accurate framework for diverse trait types. SPAGxE<sub>CCT</sub> fits a genotype-independent model and employs a hybrid strategy including saddlepoint approximation (SPA) for accurate p value calculation, especially for low-frequency variants and unbalanced phenotypic distributions. We extend SPAGxE<sub>CCT</sub> to SPAGxEmix<sub>CCT</sub>, which accounts for population stratification and is applicable to multi-ancestry or admixed populations. SPAGxEmix<sub>CCT</sub> can further be extended to SPAGxEmix<sub>CCT-local</sub>, which identifies ancestry-specific G×E effects using local ancestry. Through extensive simulations and real data analyses of UK Biobank data, we demonstrate that SPAGxE<sub>CCT</sub> and SPAGxEmix<sub>CCT</sub> are scalable to analyze large-scale study cohort, control type I error rates effectively, and maintain power.
Medical subject headings
- Genome-Wide Association Study
- Gene-Environment Interaction
- Biological Specimen Banks