CSVSUF: A Deep Unfolding Framework for Compressive Spectral Video Sensing.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42275332.
- Also identified by DOI 10.1109/TIP.2026.3700920.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Spectral videos (SVs) capture spatio-temporal-spectral information from dynamic scenes, but their acquisition traditionally requires expensive and complex systems, motivating the development of compressive spectral video sensing (CSVS). It typically employs the coded aperture snapshot spectral imager (CASSI) to acquire compressed measurements, from which SVs are reconstructed via model-driven or learning-based algorithms. However, two major limitations remain in current CASSI-based reconstruction methods: i) conventional model-driven algorithms rely on iterative optimization, which limits their representational capacity in complex scenes and results in slow reconstruction; ii) existing deep learning-based approaches overlook the joint modeling of spatial, temporal, and spectral correlations, failing to fully exploit the multi-dimensional dependencies. Hence, we propose a principled compressive spectral video sensing unfolding framework (CSVSUF) in a CASSI system for spectral video reconstruction. Moreover, we develop a novel spatio-temporal-spectral prior-learning Transformer (STS-PLT) to capture the multi-dimensional correlations within each unfolding stage. By treating STS-PLT as a Gaussian denoiser for the prior term in CSVSUF, we establish a deep unfolding-based method for CSVS. Extensive experiments demonstrate that our method consistently outperforms existing approaches in both reconstruction accuracy and visual quality, validating the benefit of combining physics-guided modeling with deep prior learning in CSVS. Code is available at https://github.com/zli1024/CSVSUF.