Hierarchical knowledge-guided reasoning for text-based person re-identification.

Zeng, Ruigeng; Ma, Wentao; Zhou, Tongqing; Zhao, Shan; Mao, Xinjun; Liu, Jie · Neural Netw · 2025

basic_science · Level V

Where this comes from

Abstract

Masked language modeling (MLM) has expanded the exploration of text-image person re-identification (TIReID) tasks from coarse-granularity to fine-grained alignment. Whereas, we note that vanilla MLM picks random tokens for visual-to-token reasoning, which could fail the intention of semantic visual-textual alignment by indistinguishably focusing on all the sub-words. This work proposes to leverage the inherent hierarchical scene graph knowledge in each text for guiding token masking and enhancing cross-modal representation in TIReID, thus relieving the pitfall of blind visual-textual alignment. The proposed framework, Hierarchical Knowledge-Guided Reasoning (HKGR), parses object-level, attribute-level, and relation-level masking according to phrase knowledge constructions and explicitly lets the training of a dedicated encoder focus on the visual-to-token reasoning of these highlighted tokens. In addition, we propose a Multi-Grained Semantic Alignment (MGA) module, which leverages the token selection method and image-text similarity distribution constraint to further facilitate the semantic alignment between image and text at both coarse-grained and fine-grained levels. Experimental results demonstrate that our HKGR framework achieves state-of-the-art (SoTA) performance on three public benchmark datasets at all evaluation metrics. We believe that the knowledge-guided idea is beneficial to other multi-modal research communities, including cross-modal retrieval and visual question answering. Code is available at https://github.com/Ray-Zhen/HKGR.git.

Medical subject headings