C<sup>3</sup>aptioner: Improving change captioning by leveraging momentum cross-view and cross-modality contrastive learning.

Deng, Lin; Kang, Borui; Zhong, Yuzhong; Wang, Maoning; Zhang, Jianwei · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

The primary goal of change captioning is to identify subtle visual differences between two similar images and express them in natural language. Existing research has been significantly influenced by the task of vision change detection and has mainly concentrated on the identification and description of visual changes. However, we contend that an effective change captioner should go beyond mere detection and description of what has changed. Two additional aspects are crucial: 1) retaining significant and unique semantic elements that persist across both images, and 2) forging a robust link between visual cues and their concomitant descriptive linguistic elements. This paper addresses these challenges by presenting the C<sup>3</sup>aptioner, which seamlessly incorporates dual momentum contrastive learning objectives into change captioning. Our model architecture consists of intra-image and inter-image Transformer encoders for visual feature extraction, complemented by unimodal language and multimodal decoders. Specifically, we introduce a cross-view contrastive learning objective to capture essential invariant features by aligning cross-view representations with a momentum-updated queue of negative samples, addressing the challenge of viewpoint variations. Additionally, our cross-modality contrastive learning objective aligns and interacts visual and textual modalities using a separate momentum-maintained queue, resolving the modality gap that hampers existing methods. This dual contrastive approach enables C<sup>3</sup>aptioner to model both changed and unchanged elements while establishing strong vision-language correspondence, resulting in more contextually rich and human-like descriptions. Extensive experiments across five distinct datasets confirm that our approach achieves state-of-the-art performance, with particularly significant improvements in challenging scenarios involving extreme viewpoint changes. Source code is available at https://github.com/DenglinGo/C-3aptioner.

Medical subject headings