Video-Causal State Space Model (VCSSM) for Video Analysis.
Where this comes from
- Record sourced from PubMed, PMID 42726632.
- Also identified by DOI 10.1109/TIP.2026.3730842.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
We propose a Video Causal State Space Model (VCSSM) for video understanding. VCSSM integrates a learned causal graph into a latent state-space dynamics model, enabling the explicit modeling of cause-and-effect relationships over time. By representing underlying video factors as latent states linked by a directed acyclic graph (DAG), VCSSM achieves interpretable temporal modeling and improved robustness to distribution shifts. The state-space formulation ensures efficient sequential prediction, while the causal structure yields clear reasoning about interventions and counterfactuals. We validate VCSSM on standard action recognition benchmarks, including the large-scale Kinetics-700 dataset, as well as HMDB-51, UCF-101,HAR, and CoPhy, where our method significantly outperforms state-of-the-art approaches in accuracy. In particular, VCSSM excels in generalizing to new contexts and facilitating causal reasoning about video events. The proposed framework demonstrates superior performance and interpretability compared to prior methods, marking a substantial step forward in video understanding.