Adversarial discriminant attack on text-to-image diffusion models.
Journal:
Neural networks : the official journal of the International Neural Network Society
Published Date:
Feb 12, 2026
Abstract
Despite advancements in concept-erased diffusion models, the persistent risk of generating Not-Safe-For-Work (NSFW) content in text-to-image tasks remains a critical challenge. To expose vulnerabilities in these models, some existing works designs attack method from generation perspective, which attempts to constraint the similarity between generate images and specific inappropriate images. However, generating visually similar images does not necessarily imply that the NSFW content has been successfully reconstructed, so the effectiveness of existing attack methods remains limited. To address this limitation, we propose Adversarial Discriminant Attack (ADAtk), a novel method designed to expose vulnerabilities in concept-erased diffusion models. Unlike existing attacks that focus on generation, ADAtk adopts a more intuitive discriminative perspective, aiming to generate images that are classified as inappropriate. By optimizing the likelihood of producing NSFW content, ADAtk crafts adversarial perturbations in the model's latent space, thereby guiding the reconstruction of NSFW concepts (e.g., nudity) aligned with the target discriminant class. Experimental results show that ADAtk can achieve an over 90% success rate in bypassing current internal security mechanisms, exposing critical limitations in existing concept-erasure techniques. These findings provide essential insights for improving the safety and reliability of text-to-image generation systems, paving the way for more secure generative AI models. Warning: This paper includes model outputs that may be considered offensive.
Authors
Keywords
No keywords available for this article.