VLG-RMOT: Calibration-before-control for end-to-end referring multi-object tracking.

Zhao, Hongshen; Dai, Ming; Xie, Fei; Zhang, Wenkang; Yao, Yuncong; Yang, Wankou · Neural Netw · 2026

Where this comes from

Abstract

Referring multi-object tracking (RMOT) aims to detect and track all objects that satisfy a natural-language expression in video. End-to-end RMOT relies on object queries for detection, association, and temporal updating, so language must function as both an alignment cue and a control signal for evolving queries. Existing methods have advanced visual-linguistic fusion and query interaction; however, alignment alone does not ensure a scene-discriminative language signal as candidate objects, their appearance, and their relations change across frames. We therefore propose VLG-RMOT, a calibration-before-control framework that adapts language to the current visual scene before using it for query control. Specifically, cross-modal semantic calibration performs bidirectional visual-linguistic calibration with learnable channel-wise residual scaling. The resulting scene-conditioned language guides two complementary stages: contextual query synthesis establishes a shared referring prior before visual decoding, while text-aware gating modulates query states after visual cross-attention. On Refer-KITTI, Refer-KITTI-V2, and Refer-BDD, VLG-RMOT achieves the best HOTA among the compared end-to-end methods. Under a matched ablation protocol, learnable calibration improves HOTA by 1.97, 1.40, and 2.98 points over raw-language control, an unscaled residual, and fixed scaling, respectively; scaling-vector statistics and token-level visualization provide representation-level evidence of channel-selective and scene-dependent adaptation.