Multimodal Image Representation Learning With Limited Visual-Tactile Data.

Qiu, Liuxiang; Da, Hui; Liu, Wenxi; Niu, Yuzhen; Wang, Hanli; Zhao, Tiesong · IEEE Trans Image Process · 2026

basic_science · Level V

Where this comes from

Abstract

Previous multimodal visual-tactile image representation learning (VTL) methods have achieved significant success in object understanding through large-scale training data. However, obtaining sufficient training data is often infeasible, and the above methods struggle to effectively focus on discriminative visual and tactile features with limited data, resulting in degraded performance. To solve the above issue, we introduce a new task called visual-tactile image representation learning with limited data (VTL-L), which better facilitates real-world applications. To address the challenges of limited data and modality discrepancy in the VTL-L task, we propose a novel multi-order feature enhancement-based, alignment-free fusion network (MOA-Net). First, we introduce a multi-order feature enhancement (MFE) module to hierarchically strengthen the detailed and structural representation by aggregating the low- and high-order topological information. This approach can effectively reduce the attention noise and obtain discriminative features with limited data. Then, we propose the alignment-free visual-tactile fusion (AVTF) module to achieve representative spatial and channel features and perform the cross-modality fusion without alignment, which efficiently mitigates the modality discrepancy. Finally, we develop a dual counterfactual intervention (DCI) loss to jointly optimize fused visual-tactile feature and probability distributions, thereby improving the performance of the MOA-Net in the VTL-L task. Extensive experiments demonstrate the superiority of the proposed method across three types of tasks on four datasets under diverse limited-data settings (source code available at: https://github.com/liuxiangqiu007/MOA-Net).