Robust zero-shot learning with ambiguous labels via visual-semantic alignment and dynamic disambiguation.

Li, Jiangnan; Yan, Xiaowen; Huang, Linqing; Fan, Jinfu · Neural Netw · 2026

Where this comes from

Abstract

Zero-shot Learning (ZSL) has garnered significant attention for its ability to recognize unseen classes without requiring any visual instances of those classes. However, existing methods typically assume clean labels, overlooking real-world label noise and ambiguity, which can substantially degrade performance. To address this limitation, we study robust zero-shot learning with ambiguous labels and propose a unified framework, termed Dynamic Visual-semantic Alignment (DVSA), which task-specifically integrates bidirectional visual-semantic alignment, attribute-level Mutual Information (MI) regularization, and dynamic label disambiguation. Specifically, we employ a bidirectional attention-based visual-semantic alignment module to encourage mutual calibration between visual features and attribute prototypes, thereby improving cross-modal correspondence under ambiguous supervision. In addition, DVSA introduces attribute-level contrastive optimization guided by MI to enhance the dependency between discriminative attributes and their semantically consistent counterparts, leading to more separable embeddings. Moreover, a dynamic label disambiguation mechanism progressively refines supervision signals through iterative noise correction while maintaining semantic consistency, helping narrow the instance-label semantic gap during training. These components work together within an ambiguity-aware learning framework, enabling more reliable transfer to unseen classes when training labels are ambiguous. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness and robustness of the proposed framework.