MTCM: Multi-context temporal consistent modeling for referring video object segmentation.
Where this comes from
- Record sourced from PubMed, PMID 40532521.
- Also identified by DOI 10.1016/j.neunet.2025.107701.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Referring Video Object Segmentation (RVOS) focuses on segmenting objects in a video that is based on a provided text description. With recent advancements in transformers, many transformer-based RVOS methods have emerged to enhance interactions between the two modalities. However, these methods often struggle with temporal modeling due to issues with query consistency and limited context awareness. Query inconsistency could result in unstable masks that switch between different objects in the middle of the video, and insufficient context consideration could cause incorrect object segmentation due to a poor alignment with the textual description. To overcome the above challenges, we propose the Multi-context Temporal Consistency Module (MTCM), which integrates an Aligner and a Multi-Context Enhancer (MCE). The Aligner enhances query consistency by filtering out noise and aligning queries, while the MCE selects text-relevant queries through comprehensive context analysis. We applied MTCM to four distinct models, achieving performance improvements across all of them, including a J&F score of 47.6 on the MeViS dataset. The code is available in https://github.com/Choi58/MTCM.
Medical subject headings
- Video Recording
- Neural Networks, Computer
- Image Processing, Computer-Assisted
- Pattern Recognition, Automated