Multimodal-guided self-distillation for unified person search.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42700682.
- Also identified by DOI 10.1016/j.neunet.2026.109581.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Person search is challenging due to limitations in identity representation. Existing methods rely on one-hot encoding, ignoring semantic relationships among pedestrians. This leads to a fragmented feature space and reduces generalization ability, especially in large-scale scenarios with a significant proportion of unlabeled identities. For instance, in the CUHK-SYSU dataset, 72.7% of pedestrians lack identity annotations, limiting the effectiveness of supervised learning. To address these issues, we propose a novel Multimodal-Guided Self-Distillation (MGSD) method for Unified Person Search that leverages multimodal textual descriptions and self-distillation to enhance pedestrian representation learning. Specifically, we introduce three key innovations: (1) Multimodal LLM-Assisted Text Generation (MLTG) to provide fine-grained semantic context beyond discrete identity labels, enabling the model to capture inter-person relationships based on clothing attributes, appearance features, and environmental cues; (2) Semantic Structural Consistency Constraint (SSCC) to impose global structural constraints on the feature space, ensuring that distinct identities remain separable while preserving semantic similarities among visually similar individuals; and (3) Multimodal-Aware Self-Distillation Framework (MSDF), where the learnable visual encoder is progressively aligned with the pre-trained CLIP multimodal encoder, improving robustness to variations in illumination, occlusion, and background clutter. Extensive experiments demonstrate that our method significantly enhances retrieval accuracy and generalization, achieving state-of-the-art performance with an mAP of 56.1% on the PRW dataset while maintaining computational efficiency for large-scale real-world applications.