DMDNet: Dual-branch multi-modal deep fusion network for V-D-T salient object detection.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41544496.
- Also identified by DOI 10.1016/j.neunet.2026.108579.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
In the multi-modal salient object detection task, depth or thermal features are often directly fused with visible feature during the encoding stage, which directly results in the fused encoder features containing a large amount of noise information and reducing the accuracy of detection. To address this challenge, in this paper, we propose a novel dual-branch multi-modal deep fusion network (DMDNet) where visible image serves as one branch, and depth and thermal images serve as another branch to achieve multi-modal feature fusion in the decoder phase. In the encoder phase, we apply two types of backbone networks to three modalities to ensure sufficient information extraction and design the modal interaction (MI) module to dig the complementarity between depth and thermal features. In the decoder phase, we propose the multi-scale feature perception (MFP) module and region optimization (RO) module in succession to mine and optimize the saliency region. After that, we introduce the dual-branch fusion (DF) module to integrate multi-modal feature in the bottom-to-top manner for generating final saliency map. DMDNet achieves superior performance on the VDT-2048 dataset, as verified by comprehensive experimental results.
Medical subject headings
- Neural Networks, Computer
- Deep Learning
- Pattern Recognition, Automated