CSVSUF: A Deep Unfolding Framework for Compressive Spectral Video Sensing.

Li, Zhilin; Wang, Han; Duan, Jizhong; Li, Baihua; Liu, Yu · IEEE Trans Image Process · 2026

basic_science · Level V

Where this comes from

Abstract

Spectral videos (SVs) capture spatio-temporal-spectral information from dynamic scenes, but their acquisition traditionally requires expensive and complex systems, motivating the development of compressive spectral video sensing (CSVS). It typically employs the coded aperture snapshot spectral imager (CASSI) to acquire compressed measurements, from which SVs are reconstructed via model-driven or learning-based algorithms. However, two major limitations remain in current CASSI-based reconstruction methods: i) conventional model-driven algorithms rely on iterative optimization, which limits their representational capacity in complex scenes and results in slow reconstruction; ii) existing deep learning-based approaches overlook the joint modeling of spatial, temporal, and spectral correlations, failing to fully exploit the multi-dimensional dependencies. Hence, we propose a principled compressive spectral video sensing unfolding framework (CSVSUF) in a CASSI system for spectral video reconstruction. Moreover, we develop a novel spatio-temporal-spectral prior-learning Transformer (STS-PLT) to capture the multi-dimensional correlations within each unfolding stage. By treating STS-PLT as a Gaussian denoiser for the prior term in CSVSUF, we establish a deep unfolding-based method for CSVS. Extensive experiments demonstrate that our method consistently outperforms existing approaches in both reconstruction accuracy and visual quality, validating the benefit of combining physics-guided modeling with deep prior learning in CSVS. Code is available at https://github.com/zli1024/CSVSUF.