Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41379913.
- Also identified by DOI 10.1109/TPAMI.2025.3642842.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Image fusion aims to blend complementary information from diverse sensing modalities, yet most current methods lack robustness in complex fusion scenarios and cannot flexibly accommodate user intent. We present DiTFuse, the first Diffusion-Transformer (DiT) framework for instruction-driven, dynamic fusion control. Guided by natural-language instructions, DiTFuse flexibly blends multimodal content to enable hierarchical and fine-grained control over fusion dynamics. The training phase employs a multi-degrade-mask-image-modeling (M3) strategy, so the network jointly learns cross-modal alignment, modality-invariant restoration, and task-aware feature selection without relying on ideal reference images. A curated, multi-granularity instruction dataset further equips the model with interactive fusion capabilities. DiTFuse unifies infrared-visible, multi-focus, and multi-exposure fusion-as well as text-controlled refinement and downstream tasks-within a single architecture. Experiments on public IVIF, MFF, and MEF benchmarks confirm superior quantitative and qualitative performance, sharper textures, and better semantic retention. The model also supports multi-level user control and zero-shot generalization to other multiimage fusion scenarios, including instruction-conditioned segmentation.