Object-centric image editing via position-structure guided diffusion.

Si, Qi; Chen, Xiangrui; Wang, Bo; Zhang, Zhao; Zhao, Mingbo; Yang, Yun; Zhang, Haijun · Neural Netw · 2026

Where this comes from

Abstract

Fine-grained object-centric editing in complex scenes, while preserving contextual integrity, remains a persistent challenge. The core difficulties arise from two sources: (1) inaccurate object localization stemming from cross-attention misalignment and inter-object interference in diffusion models, where imperfect attention correspondence frequently yields incomplete or misplaced edits; and (2) the reliance of mask-conditioned diffusion on random Gaussian noise for generating content within edited regions, which affords limited control over precise object placement. To tackle these issues, this paper proposes a training-free framework grounded in latent diffusion models. Concretely, we introduce a latent space optimization strategy that refines cross-attention maps to disentangle object representations and achieve accurate spatial alignment, dynamically adjusting attention weights across distinct objects to suppress mutual interference. Furthermore, we design a region-aware fusion mechanism to safeguard background structure and content during editing, adaptively blending the edited latent features with the original background information to prevent structural distortion. Experimental evaluations on public benchmarks demonstrate that the proposed method consistently outperforms state-of-the-art approaches, delivering clear gains in both structural fidelity and semantic coherence.