Beyond Stepwise Modeling: Towards a Unified Contextual Reasoning Framework for Hyperspectral Video Object Tracking.

Chen, Yuzeng; Yuan, Qiangqiang; Xie, Hong; Su, Xin; Tang, Yuqi; Guan, Renxiang; Liu, Li; Liu, Xinwang et al. · IEEE Trans Image Process · 2026

Where this comes from

Abstract

Hyperspectral video data provide complementary contextual cues across spectral, spatial, and temporal dimensions for modeling object dynamics under challenging conditions. Many existing hyperspectral video object tracking (HVOT) approaches organize spatial-spectral and temporal modeling in successive stages, leaving room for closer interaction among video-level contextual cues. To address this, we propose HucrTrack, a unified contextual reasoning framework for HVOT trained by parameter-efficient fine-tuning (PEFT). HucrTrack forms synchronized hyperspectral and false-color representations from each hyperspectral cube and enhances spatial-spectral features through a weight-shared dual-representation backbone with unified contextual cue modeling. To effectively leverage contextual dynamics, we design a unified contextual reasoning module (UCRM) composed of three key components: memory dynamics unit (MDU), contextual injection unit (CIU), and selective retrieval unit (SRU). Specifically, MDU maintains a frame-wise dynamic memory via Mamba's hidden states; CIU hierarchically integrates this memory into the spectral-spatial backbone features; and SRU selectively retrieves relevant contextual information to reinforce the tracking representation. In contrast to representative stepwise designs, HucrTrack enables concurrent, unified reasoning over all three dimensions within one recurrent process. Extensive experiments on ten benchmarks demonstrate that HucrTrack compares favorably with existing trackers in both robustness and generalization.