How far are we from the era of big data in transcriptomics? Lessons from the bacterial data in GEO.

Escobedo-Muñoz, A S; Carmona-Campos, Diego; Trapaga, Armando G G; Freyre-González, Julio A · Brief Bioinform · 2025

review · Level V

Where this comes from

Abstract

The Gene Expression Omnibus (GEO) is the largest functional genomics repository, including ~5 million entries related to the main transcriptomic technologies: microarrays and RNA-seq. This amount of data has the potential to be reused in large-scale meta-analysis, such as those in bacterial systems biology, where the landscape of biological conditions is wider and more diverse than any individual experiment alone. Notwithstanding the accelerated growth in RNA-seq experiments, microarray still accounts for ~48% of bacterial transcriptomic entries in GEO, highlighting the need to revalue this data. Therefore, in this work, we assess the current state of bacterial microarray and RNA-seq data and metadata. We report diverse inconsistencies in both the GEO metadata documentation and community usage, limiting the automated access to biological context essential for high-throughput analysis interpretation. Additionally, while access to and analysis of RNA-seq data are topics widely discussed by the community, microarray data processing and normalization present challenges that need to be addressed for the proper data integration into large-scale reanalysis. Thus, we delve into the availability and processability of bacterial microarray data in GEO, showing a complex panorama where the lack of standard formats limits our reusability potential to at least 44% of the ~45 000 microarray entries. We conclude that GEO transcriptomic data and metadata should be viewed as valuable resources that require ongoing revision and maintenance. Finally, we propose a series of guidelines to enhance the Findability, Accessibility, Interoperability, and Reusability of GEO, thereby taking a step forward into the era of big data.

Medical subject headings