A one-shot, lossless algorithm for cross-cohort learning in mixed-outcomes analysis.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41040961.
- Also identified by DOI 10.1016/j.patter.2025.101321 and PMC identifier 12485519.
- Licence recorded as CC BY-NC-ND.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
In cross-cohort studies, integrating diverse datasets is essential and challenging due to cohort-specific variations, distributed data storage, and privacy concerns. Traditional methods often require data pooling or harmonization, which can reduce efficiency and limit the scope of cross-cohort learning. We introduce mixWAS, a one-shot, lossless algorithm that efficiently integrates distributed electronic health record (EHR) datasets via summary statistics. Unlike existing approaches, mixWAS preserves cohort-specific covariate associations and supports simultaneous mixed-outcome analyses. Simulations demonstrate that mixWAS outperforms conventional methods in accuracy and efficiency across various scenarios. Applied to EHR data from seven cohorts in the US, mixWAS identified 4,530 significant cross-cohort genetic associations among traits such as blood lipids, BMI, and circulatory diseases. Validation with an independent UK EHR dataset confirmed 97.7% of these associations, underscoring the algorithm's robustness. By enabling lossless cross-cohort integration, mixWAS improves the precision of multi-outcome analyses and expands the potential for actionable insights in healthcare research.