IRS: iterative reference selection improves normalization of microbiome sequencing data.

Shi, Yiming; Liu, Lili; Chen, Jun; Wylie, Kristine M; Wylie, Todd N; Park, Sung Hee; Zhou, Ruiwen; Cao, Yin et al. · Brief Bioinform · 2026

Where this comes from

Abstract

Microbiome studies often seek to determine how the absolute abundances of individual taxa change across biological conditions, yet sequencing read counts are sample-specific scaled representations of those abundances. Because sampling depth can differ across samples, fold changes calculated directly from sequencing read counts do not generally represent absolute-abundance fold changes. Normalization methods attempt to account for these between-sample differences in sampling depth, but their accuracy depends on the reference used. In particular, total-sum scaling uses all taxa as the reference and can introduce compositional bias. Reference-based methods instead rely on taxa that are stable across conditions, but contamination of the reference set by differentially abundant (DA) taxa can distort sampling-depth estimation and downstream inference. Here, we present iterative reference selection (IRS), a robust normalization method that iteratively screens and refines a candidate reference set to exclude DA taxa. By deriving a clean reference set, IRS accurately captures between-sample differences in sampling depth and recovers absolute-abundance fold changes. Benchmarking using simulations and datasets with experimental absolute quantification shows that IRS outperforms standard scaling and existing reference-based methods in controlling false discovery rates while maintaining power.

Medical subject headings