Correcting collection bias in comparative studies of diversity.
other
Where this comes from
- Record sourced from PubMed, PMID 42448349.
- Also identified by DOI 10.1098/rsif.2026.0163.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Comparative analyses of diversity in human populations often rely on historical, archival and observational data, which are systematically shaped by uneven documentation, preservation and research attention. Such collection bias can distort comparisons across populations, regions or time periods, obscuring genuine patterns of variation or generating spurious ones. Standard approaches that equalize sample size or sampling effort implicitly assume comparable sampling completeness, which is an assumption rarely satisfied in human and historical datasets. Here, we evaluate coverage-based standardization, adapted from ecological diversity estimation, as a general framework for comparative diversity analysis under biased and incomplete observation. Using population-level simulations, we show that coverage-based approaches reliably recover true diversity relationships across multiple, qualitatively distinct mechanisms of collection bias, whereas sample-size-based methods yield systematically distorted inferences. We illustrate the utility of this framework with an application to a large historical cultural dataset, where correcting for uneven documentation substantially refines inferred diversity patterns. By conditioning comparisons on sampling completeness rather than raw sample size, coverage-based standardization offers a principled and broadly applicable solution for comparative research in any domain where observation is biased, incomplete or historically contingent.