A DINO-based progressive semantic enhanced infrared and visible image fusion network.
Where this comes from
- Record sourced from PubMed, PMID 41520573.
- Also identified by DOI 10.1016/j.neunet.2025.108527.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Infrared and visible image fusion aims to integrate complementary information from two source images into a single fused image with rich detail. However, most existing fusion methods focus on visual appearance and pay little attention to the semantic requirements of downstream applications. Although some semantic driven approaches enhance the semantic content of fused images, they rely on labelled data that contain only limited semantic target information.To address this limitation, this paper proposes a DINO-based progressive semantic enhanced infrared and visible image fusion network (DPSEF). DINO is a self supervised model that learns representations from large volumes of unlabelled images and exhibits powerful spatial semantic clustering capabilities. We exploit DINO to extract fine grained spatial semantic features as prior knowledge, and then introduce a semantic enhanced fusion module (SEFM) that progressively injects these semantic priors into the fusion network. This mechanism guides the model to focus on target relevant regions and generates high quality fused images that combine rich semantic and detailed information, thereby meeting the needs of subsequent high level vision tasks.Extensive experiments demonstrate that DPSEF produces fused images whose visual quality significantly exceeds that of mainstream algorithms. Qualitative and quantitative analyses further confirm the strong potential of DPSEF in high level vision applications. Moreover, additional experiments on multi focus image fusion validate the generality and robustness of the proposed network.