STAU: A SpatioTemporal-Aware Unit for Video Prediction and Beyond.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40402714.
- Also identified by DOI 10.1109/TPAMI.2025.3572735.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Video prediction aims to predict future frames by modeling the complex spatiotemporal dynamics in videos. However, most existing methods only model the temporal information and the spatial information for videos in an independent manner but have not fully explored the correlations between both terms. In this paper, we propose a SpatioTemporal-Aware Unit (STAU) for video prediction and beyond by exploring the significant spatiotemporal correlations in videos. On the one hand, the motion-aware attention weights are learned from the spatial states to help aggregate the temporal states in the temporal domain. On the other hand, the appearance-aware attention weights are learned from the temporal states to help aggregate the spatial states in the spatial domain. In this way, the temporal information and the spatial information can be greatly aware of each other in both domains, during which, the spatiotemporal receptive field can also be greatly broadened for more reliable spatiotemporal modeling. Experiments are not only conducted on video prediction tasks (deterministic and stochastic), but also another task beyond video prediction, the early action recognition task. Experimental results show that the proposed STAU can achieve satisfactory performance on all tasks compared with other methods.