Unified semantic space learning for cross-modal retrieval.

Zhu, Jie; Liu, Jianan; Wu, Shufang; Zhang, Feng · Neural Netw · 2025

Where this comes from

Abstract

With the increasing amount of multimodal data on the Internet, cross-modal retrieval has gradually become a hot research topic and has achieved significant progress, especially since graph convolutional networks were introduced. Most methods based on graph convolutional networks tend to focus on incorporating the correlations among samples and the correlations among labels into the common representations, but neglect the correlations among the semantic contents. Moreover, the semantic similarity between instances and semantic contents is also underutilized. To address these issues, we propose a Unified Semantic Space Learning (USSL) method, which not only explores the correlations of the semantic contents but also maps images, texts, labels, and multi-labels into a unified semantic space, facilitating the calculation of similarities between samples and between samples and semantic contents. To fully explore the correlations of the semantic contents, we construct a label-multi-label graph and learn the correlations of the semantic contents in a data-driven manner using our proposed Group Semantic Sharing Graph Convolutional Network. Furthermore, we propose an isomorphic InfoNCE loss to bridge the heterogeneity gap between the samples and semantic contents, along with an intra-modality InfoNCE loss and an inter-modality InfoNCE loss to maintain the semantic and structural consistencies of the learned modality-invariant common representations. Through comparative experiments on three representative cross-modal datasets, we have demonstrated the superiority of our proposed method.

Medical subject headings