Dataset decay and the problem of sequential analyses on open datasets.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 32425159.
- Also identified by DOI 10.7554/eLife.53498 and PMC identifier 7237204.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Open data allows researchers to explore pre-existing datasets in new ways. However, if many researchers reuse the same dataset, multiple statistical testing may increase false positives. Here we demonstrate that sequential hypothesis testing on the same dataset by multiple researchers can inflate error rates. We go on to discuss a number of correction procedures that can reduce the number of false positives, and the challenges associated with these correction procedures.
Medical subject headings
- Data Interpretation, Statistical
- Datasets as Topic
- Information Dissemination