Model-based analysis of sample index hopping reveals its widespread artifacts in multiplexed single-cell RNA-sequencing.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 32483174.
- Also identified by DOI 10.1038/s41467-020-16522-z and PMC identifier 7264361.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Index hopping is the main cause of incorrect sample assignment of sequencing reads in multiplexed pooled libraries. We introduce a statistical model for estimating the sample index-hopping rate in multiplexed droplet-based single-cell RNA-seq data and for probabilistic inference of the true sample of origin of hopped reads. We analyze several datasets and estimate the sample index hopping probability to range between 0.003-0.009, a small number that counter-intuitively gives rise to a large fraction of phantom molecules - the fraction of phantom molecules exceeds 8% in more than 25% of samples and reaches as high as 85% in low-complexity samples. Phantom molecules lead to widespread complications in downstream analyses, including transcriptome mixing across cells, emergence of phantom copies of cells from other samples, and misclassification of empty droplets as cells. We demonstrate that our approach can correct for these artifacts by accurately purging the majority of phantom molecules from the data.
Medical subject headings
- Algorithms
- Artifacts
- High-Throughput Nucleotide Sequencing
- Models, Statistical
- RNA
- Single-Cell Analysis