Enhancing validation of case-control omics signatures through "minimalist" single-subject analysis (N-of-1 trials): proof of concept in sepsis.

Wilson, Liam S; Pouladi, Nima; Nelson, Rachel F; Middleton, Elizabeth A; Tolley, Neal D; Shabanian, Mahdieh; Kenost, Colleen; Campbell, Robert A et al. · J Am Med Inform Assoc · 2026

case_control · Level III

Where this comes from

Abstract

To evaluate if a single-subject study (S3) design, utilizing paired transcriptome samples from the same patient (eg, "sepsis" vs "recovered"), can replicate transcriptomic signatures from small case-control studies, addressing challenges in patient accrual for rare or sub-stratified diseases. We generated a sepsis gene signature (SGS) comprising 300 differentially expressed genes (DEGs; FDR < 5%) from a human sepsis case-control cohort using general linear models (GLMs). Reproducibility of SGS was assessed through three approaches applied to sub-sampled independent datasets: single-subject analyses (N-of-1-MixEnrich), anticipated to perform better; conventional paired-sample GLM analyses; and a traditional case-control GLM analysis. SGS reproducibility in GLM analyses was inconsistent at smaller cohort sizes (∼80% reproducibility; n = 5) but stabilized at cohort sizes >6. Remarkably, the single-subject-study approach consistently reproduced SGS in each of the 18 subjects individually (100% reproducibility; n = 1). Conventional GLMs are not designed for single-subject or small cohort analyses due to their dependence on larger samples to mitigate variable dispersion and human heterogeneity. In contrast, S3 methods enhance statistical power by: reducing multiple testing through gene set aggregation, emphasizing concordant changes in pathway activity rather than exact molecular consistency, and exploiting paired samples from the same individual. This proof-of-concept demonstrates that S3 designs effectively validate gene expression signatures derived from case-control studies, highlighting their potential in research or clinical trials constrained by small sample sizes. However, further validation and computational simulation are needed to demonstrate scalability to other conditions and sensitivity to validation subject variations from the "average subject" of discovery cohorts.