Scalable data harmonization for single-cell image-based profiling with CytoTable.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42130941.
- Also identified by DOI 10.1016/j.patter.2026.101514 and PMC identifier 13161684.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
High-content imaging (HCI) involves the automated acquisition and quantitative analysis of cell phenotypes from microscopy images. These studies often rely on screening, which can involve thousands of chemical or genetic perturbations that produce terabytes of microscopy data. To extract meaningful biological insights, these data must be processed into quantitative features through a technique known as image-based profiling. A major analytical bottleneck is curating the high-dimensional, single-cell data derived from various image-analysis tools. These datasets suffer from inconsistent schemas, inefficient file formats, and undocumented ontological relationships. These challenges reduce reproducibility and slow progress in downstream applications. To solve these issues, we introduce CytoTable, a software package for harmonizing single-cell image-based profiling. CytoTable enables modular, portable, and cross-language data integration through a robust, reproducible, and scalable engine that harmonizes single-cell readouts from multiple image-analysis tools, preparing for feature integration with software in the Cytomining ecosystem such as Pycytominer.