Leveraging VLMs for MUDA: Category-specific prompt with multi-modal interactive LoRA.

Yang, Jianing; He, Xihuai; Li, Xueqiong; Huang, Wanrong; Liu, Hengzhu; Tan, Huibin · Neural Netw · 2026

other

Where this comes from

Abstract

Multi-Source Unsupervised Domain Adaptation (MUDA) aims to leverage labeled data from multiple distinct source domains and unlabeled data from the target domain in training a model that can adapt to the target domain. Many classical methods have been proposed and utilized, nevertheless, alongside the emergence and progression of pre-trained Visual Language Models (VLMs), there is currently a dearth of effective methods based on VLMs to address the MUDA problem. To address this issue, we have constructed a novel CLIP-model-based framework for MUDA, which introduces a method that integrates category-specific prompts with a multimodal Low-Rank (LoRA) matrix adaptation approach. Our method employs learnable, class-specific prompts to extract and learn shared knowledge, while incorporating a multimodal LoRA matrix to acquire domain-specific knowledge. Furthermore, we introduce a modality interaction mechanism to foster the interplay between fine-tuning parameters across different modalities. Extensive experiments have demonstrated that our approach yields significant improvements on standard image classification benchmark datasets.