Visual tracking with unified relation modeling and masked appearance learning.

Gong, Xiaomei; Zhang, Yi; Liu, Yanli · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

Recently, the Transformer-based tracking frameworks begin to prevail in visual tracking with huge success. However, both the two-stream and one-stream pipelines have some intrinsic flaws due to their structural characteristics. Under such case, a novel and effective visual tracker (dubbed RMATrack) is proposed in this paper. Firstly, a unified relation modeling scheme has been developed to realize flexible relation computation between the template and the search images. It selects the salient region of search image in a learnable manner and calculate the cross-relation between template and search images based on the state of the object (size, shape, disappear or not). Meanwhile, a target-aware representation learning method has been developed to extract the target-specific feature. Finally, to avoid the excessive reliance on the intra-frame spatial clues while ignoring the inter-frame temporal relations during the reconstruction process, a temporal reinforcement strategy is presented. Discarding partial spatial clues makes our tracker focus on the inter-frame temporal information so as to capture the appearance changes. Extensive experiments have been conducted on 5 mainstream datasets, in which our tracker has achieved appealing results among the state-of-the-art methods while meeting real-time demand.

Medical subject headings