Semantic clustering under resource constraints via Bayesian low-rank adaptation.
Where this comes from
- Record sourced from PubMed, PMID 42085836.
- Also identified by DOI 10.1016/j.neunet.2026.109021.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Image clustering has long relied on visual feature similarity, which limits its capacity to meet user-defined semantic needs. Recent advances in Vision-Language Models (VLMs) offer a path to extract semantically meaningful textual descriptions from images, enabling semantics-aware clustering. However, in resource-constrained environments, small-scale VLMs suffer from noisy generation and semantic inconsistency, compromising the quality of text-based representations. We propose a user-guided multimodal semantic clustering framework that demonstrates clear advantages over prior methods in low-resource, user-defined semantic clustering scenarios. Our method first utilizes a visual question answering style VLM to generate natural language descriptions of images, and then leverages a small-scale large language model to perform semantic classification according to user-defined queries. To mitigate the impact of noisy textual outputs, we introduce a token-level confidence estimation mechanism, assigning adaptive weights to visual and textual features for robust fusion. Furthermore, we design a Bayesian low-rank adaptation alignment module with automatic relevance determination priors, enabling cross-modal projection with structured sparsity and improved semantic consistency. Extensive experiments on public benchmarks demonstrate that our approach achieves consistent improvements in accuracy and robustness under user-defined semantic clustering tasks, while maintaining competitive performance on standard category clustering, with substantially reduced computational overhead compared to large-scale vision-language models. The framework provides a promising direction for controllable, interpretable, and scalable clustering in practical multimodal settings.