FreeScene: Generative-enhanced visual modeling for generalized zero-shot image retrieval with scene sketches.
Where this comes from
- Record sourced from PubMed, PMID 42476090.
- Also identified by DOI 10.1016/j.neunet.2026.109384.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
In this paper, we study zero-shot sketch-based image retrieval (ZS-SBIR) in the scene-level scenario, with two key challenges compared to prior art: (i) fine-grained scene-level (FG-SL) matching, which requires precise multi-object correspondence in both visual appearance and spatial configuration between the scene sketch and target image; (ii) generalized zero-shot (GZS) retrieval, which arises from the natural mixture of seen and unseen object categories in scene content, preventing a clear separation of unseen object categories for testing. Existing ZS-SBIR methods focus on instance-level retrieval and fail to interpret complex scene compositions, while FGSL-SBIR methods struggle to generalize in zero-shot settings when unseen and seen object categories are intermingled within intricate layouts. To address this, we propose FreeScene, a novel framework that captures semantic and structural cues from both scene sketches and images, while leveraging the generative capabilities of Multimodal Large Language Models (MLLMs) and diffusion models to enhance visual representations and zero-shot generalization. Specifically, we distill high-level scene semantics from MLLM-generated textual embeddings, and then employ the semantic-centroid confidence strategy to confidence-weight the alignment of these embeddings with visual features, thereby injecting external semantic priors. To capture complex spatial layouts, we introduce a hybrid CNN-ViT backbone that injects CNN layers into ViT via an adaptive attention mechanism for multi-scale structure encoding. A generative feature correlation module further optimizes semantic and structural representations via the diffusion-based denoising process. Extensive experiments demonstrate that FreeScene significantly outperforms state-of-the-art methods, establishing a new and effective benchmark for zero-shot learning in scene-level SBIR.