Leveraging VLMs for MUDA: Category-specific prompt with multi-modal interactive LoRA.
other
Where this comes from
- Record sourced from PubMed, PMID 42308835.
- Also identified by DOI 10.1016/j.neunet.2026.109249.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Multi-Source Unsupervised Domain Adaptation (MUDA) aims to leverage labeled data from multiple distinct source domains and unlabeled data from the target domain in training a model that can adapt to the target domain. Many classical methods have been proposed and utilized, nevertheless, alongside the emergence and progression of pre-trained Visual Language Models (VLMs), there is currently a dearth of effective methods based on VLMs to address the MUDA problem. To address this issue, we have constructed a novel CLIP-model-based framework for MUDA, which introduces a method that integrates category-specific prompts with a multimodal Low-Rank (LoRA) matrix adaptation approach. Our method employs learnable, class-specific prompts to extract and learn shared knowledge, while incorporating a multimodal LoRA matrix to acquire domain-specific knowledge. Furthermore, we introduce a modality interaction mechanism to foster the interplay between fine-tuning parameters across different modalities. Extensive experiments have demonstrated that our approach yields significant improvements on standard image classification benchmark datasets.