Not all regions are equal: Spatially adaptive representation learning for efficient visual object tracking.

Zhang, Zicheng; Lin, Shan; Xu, Hongke; Liu, Zhanwen; Li, Dejun; Wang, Longguang · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

The sparsity of information in natural images poses great challenges for trackers to strike a balance between accuracy and efficiency. Existing methods commonly process all regions in the template and search area equally without considering their difference. As a result, considerable redundant computation is involved and limited inference efficiency is achieved. To remedy this, in this paper, we argue that not all regions are equal during the representation learning for visual object tracking. Specifically, we develop a sparse mask Transformer (SMTransformer) that is able to achieve spatially adaptive representation learning. Particularly, a deformable patch embedding module is constructed to adapt the receptive field to focus on the object in the template. In addition, sparse mask module is developed to dynamically identify regions with low object existence probabilities, thereby reducing the search region progressively for higher computational efficiency. With these two modules, our SMTransformer can significantly reduce the redundant computation for superior efficiency while maintaining high performance. Extensive experiments are conducted on a wide range of benchmark datasets and the results demonstrate the state-of-the-art performance of the proposed SMTransformer against previous methods in terms of both accuracy and efficiency.