Identity-style preserving image composition with domain-adaptive diffusion.
Where this comes from
- Record sourced from PubMed, PMID 42623764.
- Also identified by DOI 10.1016/j.neunet.2026.109493.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
In generative image compositing, effectively achieving cross-domain and same-domain composition while balancing semantic fidelity and visual coherence remains a key challenge. Specifically, current research faces three critical limitations: 1) existing research on image composition has rarely focused on domain-adaptive solutions, leaving this promising direction with substantial room for further exploration. 2) difficulty balancing foreground-background style consistency and foreground identity alignment, leading to perceptual inconsistencies; and (3) inadequate background retention in existing models, leading to boundary artifacts and potential degradation of background structural integrity. To address these issues, this paper proposes an identity-preserving, style-consistent domain-adaptive diffusion-based image composition model with targeted solutions. First, we integrate regional identity anchoring and domain-aware style alignment to align foreground identity in the latent space, and adopt domain-specific adaptive AdaIN to adjust foreground-background style consistency, resolving the lack of domain-adaptive methods and the style-identity balance dilemma. Further, two complementary optimization pathways refine this balance to enhance semantic fidelity and visual coherence. Finally, a refined mask fusion mechanism eliminates boundary artifacts and improves background preservation, mitigating structural loss. Extensive experiments show that our method consistently outperforms state-of-the-art approaches in both same-domain and cross-domain scenarios: compared to the best previous methods, it reduces LPIPS<sub>bg</sub> by 19.1% (same-domain) and 20.1% (cross-domain), improves CLIP<sub>I</sub> by 0.11% (same-domain), and increases CSD by 0.67% (cross-domain), demonstrating superior structural fidelity, semantic consistency, and visual realism. These advances establish a new benchmark for domain-adaptive image compositing and offer a practical solution for real-world applications (e.g., e-commerce visual editing), while providing a generalizable framework for generative models beyond image composition.