S<sup>2</sup>ANet: Semantic-spatial driven alignment salient object detection network in UAV-based unregistered RGB-T image.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42492100.
- Also identified by DOI 10.1016/j.neunet.2026.109410.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Detecting salient objects in complex traffic environments remains a challenging task for ground-based systems. Recently, unmanned aerial vehicles (UAVs) have emerged as an effective solution by offering flexible perspectives and the ability to capture complementary RGB and thermal images. However, direct fusion of these modalities often introduces artifacts caused by spatial misalignment and semantic discrepancies. To overcome these challenges, we propose a Semantic-Spatial Driven Alignment Network (S<sup>2</sup>ANet) for salient object detection in UAV-based unregistered RGB-T imagery. The proposed network performs progressive semantic and spatial alignment through a Semantic-Spatial Alignment (SSA) module, achieving precise cross-modal registration. After alignment, three specialized components are designed to enhance feature representation: the Deep Positional Awareness (DPA) module extracts accurate positional cues via self-attention; the Cross-Hierarchical Contextual Interaction (CCI) module captures both intra- and inter-feature dependencies; and the Multi-scale Detail Perception (MDP) module refines fine-grained details through multi-receptive-field convolutions and spatial attention. Finally, a Dual-Cascade Feature Fusion (DFF) module integrates positional, contextual, and detailed information to generate high-quality saliency maps. Extensive experiments validate that S<sup>2</sup>ANet achieves accurate feature alignment and superior detection performance in UAV-based RGB-T scenarios.