Video-Causal State Space Model (VCSSM) for Video Analysis.

Korban, Matthew; Youngs, Peter; Acton, Scott T · IEEE Trans Image Process · 2026

Where this comes from

Abstract

We propose a Video Causal State Space Model (VCSSM) for video understanding. VCSSM integrates a learned causal graph into a latent state-space dynamics model, enabling the explicit modeling of cause-and-effect relationships over time. By representing underlying video factors as latent states linked by a directed acyclic graph (DAG), VCSSM achieves interpretable temporal modeling and improved robustness to distribution shifts. The state-space formulation ensures efficient sequential prediction, while the causal structure yields clear reasoning about interventions and counterfactuals. We validate VCSSM on standard action recognition benchmarks, including the large-scale Kinetics-700 dataset, as well as HMDB-51, UCF-101,HAR, and CoPhy, where our method significantly outperforms state-of-the-art approaches in accuracy. In particular, VCSSM excels in generalizing to new contexts and facilitating causal reasoning about video events. The proposed framework demonstrates superior performance and interpretability compared to prior methods, marking a substantial step forward in video understanding.