Multi-condition guided diffusion model for face sketch-to-photo synthesis.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42480163.
- Also identified by DOI 10.1016/j.neunet.2026.109406.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Facial sketch synthesis is important for cross-modal face analysis and digital forensics, yet existing models often suffer from structural distortions and identity inconsistency under limited paired training data. Traditional methods, primarily based on generative adversarial networks, often suffer from training instability and insufficient detail reconstruction. To address these limitations, this study proposes a diffusion-based framework with a stage-wise multi-condition guidance mechanism that enhances both structural and textural fidelity. Specifically, (i) during the downsampling phase, semantic segmentation features are fused to guide the model with accurate structural information, such as facial part locations; and (ii) during upsampling, a hybrid cross-attention mechanism is employed to integrate coarse image textures with denoised noise, refining fine-grained details. Additionally, we incorporate the Vision Transformer within the U-Net backbone to better capture global contextual information in low-frequency regions, further enhancing image realism. Experiments on multiple benchmark datasets demonstrate strong overall performance in SSIM and FSIM, indicating improved structural similarity and feature-level fidelity under the evaluated settings. Moreover, our method obtains competitive LPIPS results, indicating improved perceptual similarity. We further compare with recent image-to-image translation and diffusion-based baselines and observe competitive performance in both visual coherence and identity preservation.