Object-centric image editing via position-structure guided diffusion.
Where this comes from
- Record sourced from PubMed, PMID 42142411.
- Also identified by DOI 10.1016/j.neunet.2026.109058.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Fine-grained object-centric editing in complex scenes, while preserving contextual integrity, remains a persistent challenge. The core difficulties arise from two sources: (1) inaccurate object localization stemming from cross-attention misalignment and inter-object interference in diffusion models, where imperfect attention correspondence frequently yields incomplete or misplaced edits; and (2) the reliance of mask-conditioned diffusion on random Gaussian noise for generating content within edited regions, which affords limited control over precise object placement. To tackle these issues, this paper proposes a training-free framework grounded in latent diffusion models. Concretely, we introduce a latent space optimization strategy that refines cross-attention maps to disentangle object representations and achieve accurate spatial alignment, dynamically adjusting attention weights across distinct objects to suppress mutual interference. Furthermore, we design a region-aware fusion mechanism to safeguard background structure and content during editing, adaptively blending the edited latent features with the original background information to prevent structural distortion. Experimental evaluations on public benchmarks demonstrate that the proposed method consistently outperforms state-of-the-art approaches, delivering clear gains in both structural fidelity and semantic coherence.