Towards a universal JPEG lossless recompression foundation model for pathology images: A transformer context modeling approach.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42269197.
- Also identified by DOI 10.1016/j.media.2026.104152.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Lossless recompression of JPEG images remains fundamentally constrained by the limited modeling capacity of traditional context-mixing entropy estimators, yielding suboptimal compression ratios. Recently, CNN-based learned recompression methods have demonstrated improved entropy modeling by exploiting the strong representational capacity of deep networks. However, their reliance on local convolutional operations restricts long-range dependency modeling and limits generalization across diverse image domains. In this study, we introduce a Universal Pathology JPEG Lossless Recompression Foundation Model (ULRFM), a transformer-based architecture explicitly designed to build long-range contextual dependencies within JPEG DCT coefficient streams. Leveraging a large-scale pathology dataset comprising more than nine million image tiles across multiple cancers and multiple organs, we systematically investigate the effects of model capacity and data quantity on lossless recompression performance. Extensive experiments demonstrate that ULRFM substantially outperforms existing CNN-based learned recompression approaches in both compression efficiency and cross-distribution generalization. ULRFM provides a maximum file size reduction of 34.13% relative to the original JPEG format, highlighting its potential to markedly alleviate the growing storage burden in digital pathology infrastructures.