A probabilistic graphical model for estimating selection coefficients of nonsynonymous variants from human population sequence data.

Zhao, Yige; Lan, Tian; Zhong, Guojie; Hagen, Jake; Pan, Hongbing; Chung, Wendy K; Shen, Yufeng · Nat Commun · 2025

basic_science · Level V

Where this comes from

Abstract

Accurately predicting the effect of missense variants is important in discovering disease risk genes and clinical genetic diagnostics. Commonly used computational methods predict pathogenicity, which does not capture the quantitative impact on fitness in humans. We develop a method, MisFit, to estimate missense fitness effect using a graphical model. MisFit jointly models the effect at a molecular level ( <math xmlns="http://www.w3.org/1998/Math/MathML"><mi>d</mi></math> ) and a population level (selection coefficient, <math xmlns="http://www.w3.org/1998/Math/MathML"><mi>s</mi></math> ), assuming that in the same gene, missense variants with similar <math xmlns="http://www.w3.org/1998/Math/MathML"><mi>d</mi></math> have similar <math xmlns="http://www.w3.org/1998/Math/MathML"><mi>s</mi></math> . We train it by maximizing probability of observed allele counts in 236,017 individuals of European ancestry. We show that <math xmlns="http://www.w3.org/1998/Math/MathML"><mi>s</mi></math> is informative in predicting allele frequency across ancestries and consistent with the fraction of de novo mutations in sites under strong selection. Further, <math xmlns="http://www.w3.org/1998/Math/MathML"><mi>s</mi></math> outperforms previous methods in prioritizing de novo missense variants in individuals with neurodevelopmental disorders. In conclusion, MisFit accurately predicts <math xmlns="http://www.w3.org/1998/Math/MathML"><mi>s</mi></math> and yields new insights from genomic data.

Medical subject headings