Diversity-driven MG-MAE: Multi-granularity representation learning for non-salient object segmentation.

Yu, Chengjin; Zhang, Bin; Xu, Chenchu; Ruan, Dongsheng; Wang, Rui; Liu, Huafeng; Li, Xiaohu; Li, Shuo · Med Image Anal · 2026

basic_science · Level V

Where this comes from

Abstract

Masked Autoencoders (MAEs) have grown increasingly prominent as a powerful self-supervised learning paradigm. They are capable of effectively leveraging inherent image prior information and are gaining traction in the field of medical image analysis. However, their application to feature representations of the non-salient objects, such as microvasculature, accessory organs, and early-stage tumors-is fundamentally limited by dimensional collapse problem, which diminishes feature diversity critical for non-salient structure discrimination. To address this, we propose a Multi-Granularity Masked Autoencoder (MG-MAE) framework for feature diversity learning: (1) We extend the conventional MAE into a multi-granularity framework, a global branch reconstructs global pixels, with a local branch recovering Histogram of Oriented Gradients (HOG) features, enabling hierarchical representation of both coarse-grained and fine-grained patterns; (2) Critically, in the local branch, a diversity-enhanced loss function incorporating Nuclear Norm Maximization (NNM) constraint to explicitly mitigate feature space collapse through orthogonal embedding regularization; and (3) A Dynamic Weight Adjustment (DWA) strategy that dynamically prioritizes hard-to-reconstruct regions via entropy-driven gradient modulation. Comprehensive evaluations across five clinical benchmarks-CCTA139, BTCV, LiTS, ACDC, and MSD Pancreas Tumour datasets-demonstrate that MG-MAE achieves statistically significant improvements in Dice Similarity Coefficient (DSC) scores for non-salient object segmentation, outperforming state-of-the-art methods. The code is available at https://github.com/zhangbbin/mgmae.

Medical subject headings