FlowGrid enables fast clustering of very large single-cell RNA-seq data.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 34289014.
- Also identified by DOI 10.1093/bioinformatics/btab521.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Scalable clustering algorithms are needed to analyze millions of cells in single cell RNA-seq (scRNA-seq) data. Here, we present an open source python package called FlowGrid that can integrate into the Scanpy workflow to perform clustering on very large scRNA-seq datasets. FlowGrid implements a fast density-based clustering algorithm originally designed for flow cytometry data analysis. We introduce a new automated parameter tuning procedure, and show that FlowGrid can achieve comparable clustering accuracy as state-of-the-art clustering algorithms but at a substantially reduced run time for very large single cell RNA-seq datasets. For example, FlowGrid can complete a one-hour clustering task for one million cells in about five min. https://github.com/holab-hku/FlowGrid. Supplementary data are available at Bioinformatics online.
Medical subject headings
- Gene Expression Profiling
- Single-Cell Gene Expression Analysis