A Multi-scale Spatial-temporal Attention Model for Person Re-identification in Videos.

Zhang, Wei; He, Xuanyu; Yu, Xiaodong; Lu, Weizhi; Zha, Zhengjun; Tian, Qi · IEEE Trans Image Process · 2019

Where this comes from

Abstract

In this paper, we propose a novel deep neural network based attention model to learn the representative local regions from a video sequence for person re-identification. Specifically, we propose a multi-scale spatial-temporal attention (MSTA) model to measure the regions of each frame in different scales from the perspective of whole video sequence. Compared to traditional temporal attention models, MSTA focuses on exploiting the importance of local regions of each frame to the whole video representation in both spatial and temporal domains. A new training strategy is designed for the proposed model by incorporating the image-to-image mode with the videoto- video mode. Extensive experiments on benchmark datasets demonstrate the superiority of the proposed model over state-ofthe- art methods.