SPLENDID incorporates continuous genetic ancestry in biobank-scale data to improve polygenic risk prediction across diverse populations.
Where this comes from
- Record sourced from PubMed, PMID 42736360.
- Also identified by DOI 10.1038/s41592-026-03235-2.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Polygenic risk scores are widely used in disease risk stratification, but their accuracy varies across different ancestries. Recent methods leverage multi-ancestry data to improve accuracy in under-represented populations but require the labeling of individuals by ancestry. This poses practical challenges, given that clinical decisions are typically not based on ancestry, and many individuals may not fit into a pre-specified ancestry group. Here we propose SPLENDID, a penalized regression framework for large-scale individual-level data that models genetic ancestry as a continuum to produce a single prediction model without any ancestry labels. In extensive simulations and analyses in the All of Us Research Program (n = 224,364) and UK Biobank (n = 340,140), we show that SPLENDID significantly improved prediction accuracy over existing methods, particularly for non-European and admixed ancestries. SPLENDID stands as a valuable tool for robust risk prediction across diverse populations, reduced health disparities in genetic research, and fairer clinical implementation.