Visual tracking with unified relation modeling and masked appearance learning.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41145051.
- Also identified by DOI 10.1016/j.neunet.2025.108197.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Recently, the Transformer-based tracking frameworks begin to prevail in visual tracking with huge success. However, both the two-stream and one-stream pipelines have some intrinsic flaws due to their structural characteristics. Under such case, a novel and effective visual tracker (dubbed RMATrack) is proposed in this paper. Firstly, a unified relation modeling scheme has been developed to realize flexible relation computation between the template and the search images. It selects the salient region of search image in a learnable manner and calculate the cross-relation between template and search images based on the state of the object (size, shape, disappear or not). Meanwhile, a target-aware representation learning method has been developed to extract the target-specific feature. Finally, to avoid the excessive reliance on the intra-frame spatial clues while ignoring the inter-frame temporal relations during the reconstruction process, a temporal reinforcement strategy is presented. Discarding partial spatial clues makes our tracker focus on the inter-frame temporal information so as to capture the appearance changes. Extensive experiments have been conducted on 5 mainstream datasets, in which our tracker has achieved appealing results among the state-of-the-art methods while meeting real-time demand.
Medical subject headings
- Machine Learning
- Learning
- Eye-Tracking Technology