An Episode Memory-guided Dual-stage Framework for Long-Form Video Temporal Grounding.
Where this comes from
- Record sourced from PubMed, PMID 42340901.
- Also identified by DOI 10.1109/TIP.2026.3705206.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Video temporal grounding (VTG) aims to localize video moments that are semantically related to a given natural language query. In spite of recent progress in short-form videos, research on VTG in long-form videos (e.g., hours long) remains highly demanded yet underexplored. Existing methods predominantly adopt sliding window-based or multi-scale anchor-based strategies to generate temporal proposals, which require time-consuming post-processing or are independent of video content, thereby limiting their performance and efficiency. To address this dilemma, in this paper, we propose an episode memory-prompted (EMP) two-stage framework for temporal grounding in long-form videos. Specifically, the first stage generates a set of dynamic episode memories, which explicitly summarize various activities occurring throughout the lengthy video. An unsupervised memory learning paradigm is formulated by imposing discriminability and diversity constraints, eliminating the reliance on additional activity-instance annotations. Then, in the second stage, based on the supplement of frame-level detailed content and the guidance of a language query, the augmented memory prompts function as anchors for efficiently regressing the refined boundaries of the target video moment. Extensive experimental results on two public long-form video data sets, i.e., MAD and Ego4d, validate that the proposed EMP framework saves more than 8.5% trainable parameters and 13.9% FLOPs, while still achieving comparable performance with existing methods.