MGNet: RGBT tracking via cross-modality cross-region mutual guidance.

Zhang, Jianming; Yang, Jing; Qin, Yu; Xiao, Zhu; Wang, Jin · Neural Netw · 2025

basic_science · Level V

Where this comes from

Abstract

Compared to single modal object tracking, the main challenge in RGBT tracking lies in effectively fusing features from both modalities. However, many existing methods neglect the dependence of distinct regions from different modalities, instead only considering that of identical regions, which fails to capture the cross-modal cross-regional relationships. In other words, they do not leverage the mutual guidance between different regions of different modalities. To address this limitation, we propose a novel RGBT tracking network, MGNet, which employs dual-stage attention and multi-scale feature fusion. The network includes the Cross-modality Cross-region Dual-stage Attention (CCDA) module and the Multi-scale Intra-region Feature Fusion (MIFF) module. The CCDA module processes features in two stages to preserve the unique features of identical region of different modalities, and then achieves mutual guidance across them. Specifically, in the first stage, features from different regions of different modalities are combined into a mixed representation, maintaining the distinct features of each region. In the second stage, attention mechanisms are applied to the mixed representation, facilitating cross-modality cross-region mutual guidance. Additionally, the MIFF module can perceive feature changes at multiple scales, ensuring effective fusion within each region. Our method achieves superior performance on three RGBT benchmark datasets (GTOT, RGBT234, and LasHeR) while running at 75 FPS, demonstrating both high accuracy and real-time performance.

Medical subject headings