PDMark: Plausibly deniable watermarking for diffusion models.

Xu, Zikai; Liu, Bin; Li, Weihai; Zhang, Lijunxian; Yu, Nenghai · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

Diffusion watermarking has become an important technique for copyright protection and provenance tracing in AI-generated content. However, most existing methods rely on deterministic extraction, where the embedded watermark is expected to be revealed once the corresponding key or extractor is available. This design becomes problematic under coercive disclosure, where revealing a watermark may expose sensitive identity or provenance information. In this paper, we propose PDMark, a plausibly deniable watermarking method for diffusion models. To the best of our knowledge, PDMark is the first framework that explicitly studies deniable watermarking for diffusion-generated images under coercive disclosure. PDMark is a train-free latent-based framework, and it constructs a single watermark representation from the real and fake messages. The representation is then mapped to a watermarked latent through Gaussian-consistent sampling. During extraction, the disclosed key determines which watermark is recovered, allowing the content owner to disclose a fake extraction key that produces a plausible fake watermark while the undisclosed watermark remains protected. Extensive experiments and theoretical analysis demonstrate that PDMark realizes deniability while maintaining high extraction accuracy, competitive image fidelity, and robustness under various attacks.