Multimodal-guided prototype calibration and temporal coherence-aware hybrid matching for few-shot action recognition.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41962365.
- Also identified by DOI 10.1016/j.neunet.2026.108930.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Few-shot action recognition methods achieve impressive performance by learning discriminative features and designing temporal alignment strategies. However, these methods suffer from three significant challenges in distinguishing similar classes: (a) spatiotemporal and motion features are complementary and essential, yet the latter is frequently overlooked; (b) excessive reliance on unimodal video data leaves multimodal information underexplored, and (c) the temporal distribution of actions is prone to intra-class temporal offsets and inter-class local similarity. To overcome these challenges, we propose a multimodal-guided prototype calibration and temporal coherence-aware hybrid matching (MGTH), which integrates four innovative components: a motion-enhanced temporal aggregation module (MTAM), a text-guided prototype calibration module (TPCM), a video-text adapter objective (VTA), and a temporal coherence-aware hybrid matching (TCH). The MTAM encodes complementary spatiotemporal and motion features without any 3D convolution or optical flow calculations. The TPCM fully utilizes video-text information to optimize video prototypes. Meanwhile, the VTA maximizes the similarity between video features and corresponding textual representations. The TCH strategy alleviates metric bias caused by intra-class temporal offsets and inter-class local similarity through a hybrid matching mechanism, and strengthens this mechanism's ability to distinguish similar actions by applying temporal coherence regularization to the input video. Additionally, we extend the proposed MGTH to more challenging tasks, including cross-domain few-shot action recognition and zero-shot action recognition. Experimental results on multiple benchmark datasets demonstrate that MGTH achieves state-of-the-art performance, confirming the superiority of our approach.