Weakly and Single-Frame Supervised Temporal Sentence Grounding With Gaussian-Based Contrastive Proposal Learning.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41406257.
- Also identified by DOI 10.1109/TPAMI.2025.3644900.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Temporal sentence grounding aims to localize moments corresponding to language queries from videos. As labeling the temporal boundaries is cumbersome and subjective, the weakly-supervised methods have received increasing attention. Most of the existing weakly-supervised methods generate proposals by sliding windows, which are content-independent and of low quality. Besides, they train their model to distinguish positive visual-language pairs from negative ones randomly collected from other videos, ignoring the highly confusing video segments within the same video. In this paper, we propose Contrastive Proposal Learning (CPL) to overcome the above limitations. Specifically, we use multiple learnable asymmetric Gaussian functions to generate both positive and negative proposals within the same video. Then, we propose a controllable, easy-to-hard negative proposal mining strategy to collect negative samples within the same video, which enables CPL to distinguish highly confusing scenes. Finally, we propose an extension of the proposal generation algorithm to explore the use of low-cost single-frame annotation and achieve a balance between annotation burden and grounding performance. Our CPL can be applied to both MIL-based and reconstruction-based mainstream frameworks and achieves state-of-the-art performance on Charades-STA, ActivityNet Captions, and DiDeMo datasets.