Sparse mixture of experts-driven multimodal degraded image fusion.
Where this comes from
- Record sourced from PubMed, PMID 42612307.
- Also identified by DOI 10.1016/j.neunet.2026.109497.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Multi-modal image fusion aims to integrate complementary information from different modal images to improve image quality and performance in visual tasks. However, existing methods face challenges such as the loss of key feature information in degraded scenarios, difficulty in adapting to multi-task fusion, and excessive computational overhead. To address these issues, this paper proposes an efficient fusion method. First, a staged collaborative optimization architecture is designed to decouple encoder pre-training from fusion layer fine-tuning, thereby enhancing single-modal feature representation and cross-modal semantic alignment. Second, a multi-scale heterogen eous hybrid expert architecture is proposed, integrated with a sparse activation mechanism, which dynamically selects the most relevant Top-K experts, significantly reducing redundant computations. Finally, a dynamic fusion paradigm selection mechanism is constructed, which adaptively selects the optimal fusion path based on input feature differences. We conducted extensive qualitative and quantitative experiments on the visible and infrared image fusion (VIF) Dataset and the Harvard Medical Dataset, validating the superior performance of this method on both datasets. Our method consistently outperforms existing SOTA methods in the core dimensions of 8 objective metrics on 6 key test datasets, especially showing superior performance in infrared-visible scenarios with low light, haze and noise interference, as well as noise-degraded medical imaging scenarios.