Weakly and Single-Frame Supervised Temporal Sentence Grounding With Gaussian-Based Contrastive Proposal Learning.

Zheng, Minghang; Huang, Yanjie; Chen, Qingchao; Peng, Yuxin; Liu, Yang · IEEE Trans Pattern Anal Mach Intell · 2025

basic_science · Level V

Where this comes from

Abstract

Temporal sentence grounding aims to localize moments corresponding to language queries from videos. As labeling the temporal boundaries is cumbersome and subjective, the weakly-supervised methods have received increasing attention. Most of the existing weakly-supervised methods generate proposals by sliding windows, which are content-independent and of low quality. Besides, they train their model to distinguish positive visual-language pairs from negative ones randomly collected from other videos, ignoring the highly confusing video segments within the same video. In this paper, we propose Contrastive Proposal Learning (CPL) to overcome the above limitations. Specifically, we use multiple learnable asymmetric Gaussian functions to generate both positive and negative proposals within the same video. Then, we propose a controllable, easy-to-hard negative proposal mining strategy to collect negative samples within the same video, which enables CPL to distinguish highly confusing scenes. Finally, we propose an extension of the proposal generation algorithm to explore the use of low-cost single-frame annotation and achieve a balance between annotation burden and grounding performance. Our CPL can be applied to both MIL-based and reconstruction-based mainstream frameworks and achieves state-of-the-art performance on Charades-STA, ActivityNet Captions, and DiDeMo datasets.