FDSRM: A Feature-Driven Style-Agnostic Foundation Model for Sketch-Less Facial Image Retrieval.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41289103.
- Also identified by DOI 10.1109/TNNLS.2025.3633075.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Sketch-less facial image retrieval (SLFIR) framework efficiently retrieves target images with minimal strokes through human-computer interaction, thus overcoming the traditional model's reliance on high-quality sketch images. However, the variability in sketching styles and the randomness of stroke placement during the drawing process pose challenges in matching target images. To address this issue, we propose a feature-driven foundation model for sketch-less facial image retrieval (FDSRM), which is designed to be independent of the sketch style and comprises two core components: the feature observer and the adaptive fusion adapter (AFA). First, to address the diversity of sketch styles, we design the feature observer module (FOM). It employs multiple experts focused on extracting key features and semantic information common to various sketch styles and the target image. This helps the model to precisely identify crucial features for effective matching in stylistically diverse sketches. Second, to address the randomness of stroke placement, we introduce prior knowledge of sketching and, in conjunction with the AFA component, dynamically learn and adjust the fusion strategy of sketches and text based on the current state of sketch strokes. This enables more accurate and targeted feature fusion throughout the sketching process. Furthermore, we train a facial image-text alignment pretraining (FAIP) model on a large-scale facial dataset and use it as the backbone of FDSRM, which significantly improved the model's robustness to unknown facial features. Extensive experiments demonstrate that our method exhibits significant advantages in terms of accuracy in early retrieval and system generalization capabilities. Even without additional auxiliary information, it outperforms state-of-the-art methods in both qualitative and quantitative measures in multistyle application scenarios.