Adversarial discriminant attack on text-to-image diffusion models.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41707457.
- Also identified by DOI 10.1016/j.neunet.2026.108716.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Despite advancements in concept-erased diffusion models, the persistent risk of generating Not-Safe-For-Work (NSFW) content in text-to-image tasks remains a critical challenge. To expose vulnerabilities in these models, some existing works designs attack method from generation perspective, which attempts to constraint the similarity between generate images and specific inappropriate images. However, generating visually similar images does not necessarily imply that the NSFW content has been successfully reconstructed, so the effectiveness of existing attack methods remains limited. To address this limitation, we propose Adversarial Discriminant Attack (ADAtk), a novel method designed to expose vulnerabilities in concept-erased diffusion models. Unlike existing attacks that focus on generation, ADAtk adopts a more intuitive discriminative perspective, aiming to generate images that are classified as inappropriate. By optimizing the likelihood of producing NSFW content, ADAtk crafts adversarial perturbations in the model's latent space, thereby guiding the reconstruction of NSFW concepts (e.g., nudity) aligned with the target discriminant class. Experimental results show that ADAtk can achieve an over 90% success rate in bypassing current internal security mechanisms, exposing critical limitations in existing concept-erasure techniques. These findings provide essential insights for improving the safety and reliability of text-to-image generation systems, paving the way for more secure generative AI models. Warning: This paper includes model outputs that may be considered offensive.
Medical subject headings
- Computer Security
- Neural Networks, Computer