Semantic consistency-aware pseudo-temporal framework for multimodal remote sensing image segmentation.
Where this comes from
- Record sourced from PubMed, PMID 42259113.
- Also identified by DOI 10.1016/j.neunet.2026.109187.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Multimodal remote sensing semantic segmentation, as a fundamental task in intelligent perception and geospatial analysis, enables fine-grained land-cover recognition and robust understanding of complex scenes. However, existing approaches face two major challenges when dealing with ultra-high-resolution imagery: the conventional sliding window-based local modeling leads to spatial context fragmentation and semantic discontinuity, making it difficult to capture long-range contextual dependencies across regions; and multimodal feature fusion often lacks unified semantic constraints, resulting in modality inconsistency and semantic drift. These limitations severely hinder the generalization ability of multimodal segmentation models in complex environments. To address these issues, we propose a semantic consistency-aware pseudo-temporal multimodal segmentation framework that jointly models cross-region contextual dependencies and cross-modal semantic complementarity in ultra-high-resolution imagery. Built upon the pretrained Segment Anything Model (SAM), the framework comprises three key innovations: a pseudo-temporal input construction strategy that employs the random walk and image transformations to dynamically organize static image patches into spatially continuous "patch frame" sequences, explicitly modeling cross-region contextual relationships; a cross-modal temporal interaction module that integrates pyramid-based multi-scale feature fusion and global cross-frame attention to achieve deep multimodal alignment and temporal dependency modeling; and a prompt-guided decoder that leverages semantic similarity constraints to enhance class separability and structural consistency. Extensive experiments on the ISPRS Vaihingen and Potsdam datasets demonstrate that the proposed framework significantly outperforms existing methods in multi-class semantic segmentation, validating the effectiveness of our proposed method for ultra-high-resolution multimodal remote sensing segmentation.