Semantic clustering under resource constraints via Bayesian low-rank adaptation.

Yu, Minda; Ye, Xulun · Neural Netw · 2026

Where this comes from

Abstract

Image clustering has long relied on visual feature similarity, which limits its capacity to meet user-defined semantic needs. Recent advances in Vision-Language Models (VLMs) offer a path to extract semantically meaningful textual descriptions from images, enabling semantics-aware clustering. However, in resource-constrained environments, small-scale VLMs suffer from noisy generation and semantic inconsistency, compromising the quality of text-based representations. We propose a user-guided multimodal semantic clustering framework that demonstrates clear advantages over prior methods in low-resource, user-defined semantic clustering scenarios. Our method first utilizes a visual question answering style VLM to generate natural language descriptions of images, and then leverages a small-scale large language model to perform semantic classification according to user-defined queries. To mitigate the impact of noisy textual outputs, we introduce a token-level confidence estimation mechanism, assigning adaptive weights to visual and textual features for robust fusion. Furthermore, we design a Bayesian low-rank adaptation alignment module with automatic relevance determination priors, enabling cross-modal projection with structured sparsity and improved semantic consistency. Extensive experiments on public benchmarks demonstrate that our approach achieves consistent improvements in accuracy and robustness under user-defined semantic clustering tasks, while maintaining competitive performance on standard category clustering, with substantially reduced computational overhead compared to large-scale vision-language models. The framework provides a promising direction for controllable, interpretable, and scalable clustering in practical multimodal settings.