MMA-Diffusion: MultiModal Attack on Diffusion Models

Yijun Yang, Ruiyuan Gao, Xiaosen Wang, Tsung-Yi Ho, Nan Xu, Qiang Xu

Introduction

In the rapidly evolving landscape of text-to-image (T2I) generation, diffusion models such as Stable Diffusion (SD) (Rombach et al., 2022) and Midjounery (Midjourney, 2023) have marked a paradigm shift. These models have revolutionized digital creativity by generating strikingly realistic images, yet they also pose significant security challenges. Notably, the potential misuse of these models for generating Not-Safe-For-Work (NSFW) contents (Schramowski et al., 2023; Qu et al., 2023; Saharia et al., 2022), such as adult materials, violence, and politically sensitive imagery, is a serious concern.

In response to these concerns, developers of T2I models have implemented preventive measures like prompt filters and post-synthesis safety checks. While effective to an extent, the resilience of these measures against sophisticated adversarial attacks remains a topic of intense debate and investigation. Our study delves into this pressing issue by introducing MMA-Diffusion, a framework designed to rigorously test and challenge the security of T2I models. Unlike conventional methods that make subtle prompt modifications (Liu et al., 2023a; Gao et al., 2023; Zhuang et al., 2023; Kou et al., 2023), MMA-Diffusion adopts a systematic attack approach. It enables users to generate unrestricted adversarial prompts and craft image perturbations, thereby circumventing existing safety protocols.

The technical prowess of MMA-Diffusion lies in its dual-modal attack strategy. We develop an advanced text modality attack mechanism that intricately alters textual prompts while maintaining their semantic intent, allowing for the generation of targeted NSFW content without being flagged by existing filters, as demonstrated in Figure 1(a). On the image modality front, MMA-Diffusion utilizes a novel perturbation technique that subtly alters image characteristics in a manner undetectable to the human eye but significant enough to bypass post-processing safety algorithms, as illustrated in Figure 1(b).

Our two-pronged attack not only demonstrates the framework’s versatility in exploiting security loopholes but also highlights the nuanced complexities in safeguarding T2I models against evolving adversarial tactics. By unveiling these vulnerabilities, MMA-Diffusion serves as a catalyst for advancing the development of more robust and comprehensive security measures in T2I technologies.

Overall, the contributions of this work include: 1 We present a novel multimodal systematic attack that effectively bypasses prompt filters and safety checkers, highlighting a significant security issue in T2I models. 2 In the textual modality, we craft an adversarial prompt generation method that can deceive the prompt filter while remaining semantically similar to the target. For the image modality, we devise an attack that proficiently bypasses the post-hoc defense mechanism. 3 We evaluate various T2I models, encompassing popular open-source models and online platforms and demonstrate the effectiveness of the proposed MMA-Diffusion. For example, 10-query black-box attack can achieve a 83.33% and 90% success rate w.r.t Midjounery (Midjourney, 2023) and Leonardo.Ai (Leonardo.Ai, 2023).

Related Work

Adversarial attacks on T2I models. To the best of our knowledge, current research does not extensively explore attacks in the image modality for NSFW content generation with T2I models. Most existing studies on adversarial attacks in T2I models, such as (Qu et al., 2023; Gao et al., 2023; Zhuang et al., 2023; Kou et al., 2023; Liu et al., 2023b), have predominantly focused on text modification to probe functional vulnerabilities. These explorations encompass impacts from diminishing synthetic quality (Liu et al., 2023b; Zhang et al., 2023a) to distorting or eliminating objects (Zhuang et al., 2023; Liu et al., 2023b), and impairing image fidelity (Liu et al., 2023a). However, they do not target generating NSFW-specific materials like pornography, violence, politics, racism, or horror. Recent works such as UnlearnDiff (Zhang et al., 2023b) and Ring-A-Bell (Tsai et al., 2023) have started to consider the misuse of T2I models for generating NSFW content. UnlearnDiff primarily examines concept-erased diffusion models (Rombach et al., 2022; Gandikota et al., 2023; Kumari et al., 2023; Schramowski et al., 2023) and does not extend to other defense strategies. Conversely, Ring-A-Bell explores inducing T2I models to generate NSFW concepts but lacks precision in controlling the details of the synthesis. However, none of them considers attacks that can bypass both the prompt filter and the post-hoc safety mechanisms while still producing high-quality NSFW content tailored to specific semantic prompts. This paper demonstrates the feasibility of such attacks, highlighting their general applicability across a variety of T2I models.

Defensive methods. Various T2I models implement distinct countermeasures to mitigate user abuse. Notably, popular online T2I services like Midjourney (Midjourney, 2023) and Leonardo.Ai (Leonardo.Ai, 2023) employ AI moderators to screen potentially harmful prompts. This proactive approach targets the prevention of NSFW content generation at the input stage. Another defensive strategy involves post-hoc safety checkers, exemplified by the one integrated into Stable Diffusion (SD) (Safety-checker, 2023; Rando et al., 2022a). Unlike AI moderators, these checkers function at the output stage, scrutinizing generated images to detect and obfuscate NSFW elements. Additionally, some novel mitigation methods lie in the concept-erased diffusion (Gandikota et al., 2023; Kumari et al., 2023; Schramowski et al., 2023). These methods differ fundamentally from external safety measures as they modify the model’s inference guidance or utilize fine-tuning to actively suppress NSFW concepts. However, they may not entirely eliminate NSFW content and could inadvertently affect the quality of benign images (Zhang et al., 2023b; Lee et al., 2023; Schramowski et al., 2023). This paper presents a multimodal attack that breaches both prompt filters and post-hoc safety checkers, which is also applicable to concept-erased diffusion models (e.g., SLD (Schramowski et al., 2023)), exposing the risk of T2I models and related online services.

Method

In this work, we rigorously evaluate the robustness of T2I models under two realistic attack scenarios:

White-Box Settings: Here, adversaries utilize open-source T2I models like SDv1.5 (SDv1.5, 2023) for image generation. With full access to the model’s architecture and checkpoint, attackers can conduct in-depth explorations and manipulations for sophisticated attacks.

Black-Box Settings: Here, attackers generate images using online T2I services such as Midjourney, where they lack direct access to the proprietary models’ parameters. Instead, they employ transfer attacks, adapting their strategies based on their interactions with the service provider to skillfully bypass existing security measures.

2 Approach Overview

In this paper, we focus on the attack that enables T2I models to generate high-quality NSFW content, thereby exposing the potential misuse risks of them, as in Figure 2. Specifically, we assume that the attacker describes the content they wish to generate through plain text. The attack is considered successful only if the model generates NSFW content that aligns with the description.

To make the attack more realistic, we assume that the T2I model or the online service adopts two defense methods, namely: prompt filter, as in Figure 2 (a) and post-hoc safety checker, as in Figure 2 (b). For situations where only the prompt filter is present, such as (Leonardo.Ai, 2023), we employ a text-modal attack. For situations where only the post-hoc safety checker is present, such as SD (Rombach et al., 2022), we utilize an image-modal attack. For models that adopt both modalities of defense, we can simultaneously use both attack methods to achieve a stronger effect, as in Figure 2 (c).

3 Text-Modal Attack

The target of the text-modal attack is to evade the prompt filter while keeping the functionality guiding the T2I model for the desired NSFW content. Specifically, we set this original NSFW prompt as the target prompt, denoted as ptar\mathbf{p}_{\text{tar}} (e.g., “A completely naked Trump stands on the grass”). MMA-Diffusion assumes the prompt filter is implemented by filtering the prompts according to a sensitive word list. Therefore, the goal of attackers is to construct an adversarial prompt padv\mathbf{p}_{\text{adv}} that does not contain any sensitive word Attackers may incorporate any specific words into their sensitive word list during an attack, enabling them to effectively mask their malicious intentions., while leading the generation toward the semantics of the target prompt.

Given that the diffusion model’s denoising steps are guided by the text embedding, MMA-Diffusion launches an attack by ensuring identical latent from text encoder, given by i.e., τθ(padv)≈τθ(ptar)\tau_{\theta}(\mathbf{p}_{\text{adv}})\approx\tau_{\theta}(\mathbf{p}_{\text{tar}}), guaranteed by our proposed semantic similarity-driven loss. To find such a free-style adversarial prompt, we introduce the search method based on gradient optimization. Finally, we present our sensitive word regularization to ensure that padv\mathbf{p}_{\text{adv}} does not contain any sensitive words. Thus, MMA-Diffusion maintains high fidelity of the output without any sensitive words.

Semantic similarity-driven loss. We begin by inputting a target prompt ptar\mathbf{p}_{\text{tar}} that describes the desired content from the attacker’s perspective, as illustrated in Figure 3. To precisely reflect the attacker’s intentions, we formulate a targeted attack and utilize cosine similarity to ensure semantic similarity between padv\mathbf{p}_{\text{adv}} and ptar\mathbf{p}_{\text{tar}}. Our textual attack objective is formalized as:

Sensitive word regularization. To eliminate sensitive words in padv\mathbf{p}_{\text{adv}}, we construct a list of sensitive words based on the NSFW concepts investigated by (Rando et al., 2022b; Qu et al., 2023), which typically includes explicit NSFW words, as highlighted in red font in Figure 3 (see Appendix for the full word list). Later, we suppress the occurrence of tokens from the sensitive word list by setting their gradients to −inf⁡-\inf. As will be evident later, this sensitive words elimination strategy can effectively evade prompt filters, despite being implemented by advanced deep neural networks, as the AI moderator employed in Midjounery (Midjourney, 2023) and Leonardo.Ai (Leonardo.Ai, 2023).

4 Image-Modal Attack

T2I models like SD can use a post-hoc safety checker to identify NSFW content in the synthesis, replacing flagged synthesis with a black image as in Figure 2 (b). This defense mechanism in image space motivates us to initiate attacks on the image modality to cheat these safety checkers.

In this image-modal attack, our focus is the image editing task of T2I models. Given that the image is prone to NSFW contents induced by malicious prompts, we aim to evade the post-hoc safety checker through the adversarial attack. As illustrated in Figure 4, given an NSFW-related prompt p\mathbf{p} and an input image xinput\mathbf{x}_{\text{input}}, a T2I model generates a synthetic image, xsyn\mathbf{x}_{\text{syn}}. The safety checker then maps this image to a latent vector II and compares it with MM default NSFW embeddings, denoted as CiC_{i} for i=1,...,Mi=1,...,M, via cosine distance. If any cosine value exceeds the corresponding threshold TiT_{i}, the synthesis is flagged as NSFW. We expect the victim safety checker to release the synthesis xsyn\mathbf{x}_{\text{syn}} by crafting xadv\mathbf{x}_{\text{adv}} as the model input. To achieve this objective, we dynamically optimize the gradients of loss items that exceed TiT_{i}, as shown in the red box in Figure 4. We formulate our objective in Equation 2.

where 1 is the indicator function to select the triggered loss items for optimization, ε\varepsilon indicates the perturbation budget. This dynamic loss selection strategy focuses on optimizing features near the decision boundary, allowing us to bypass the safety checker while minimally altering image features. The constrained optimization problem in Equation 2 is solved using projected gradient descent (Madry et al., 2018) (detailed algorithm is provided in Appendix).

Experiments

We select a subset of 1000 captions from the LAION-COCO dataset (Schuhmann et al., 2022), annotated with an NSFW score above 0.99 (out of 1.0), as our test prompts. The selection criteria are detailed in the Appendix. The NSFW scores in this dataset pertain solely to adult content. To diversify our NSFW themes evaluation, we include UnsafeDiff (Qu et al., 2023), a human-curated dataset designed for NSFW evaluation. UnsafeDiff provides 30 prompts across six NSFW themes: adult content, violence, gore, politics, racial discrimination, and inauthentic notable descriptions.

Target models. We primarily execute white-box attacks on SDv1.5 (SDv1.5, 2023) and report the results. Moreover, we repurpose the adversarial prompts derived from these attacks to conduct black-box attacks on three prevalent open-source models: SDXLv1.0 (Podell et al., 2023), a replica of DALL⋅\cdotE2 (Ramesh et al., 2022) https://github.com/lucidrains/DALLE2-pytorch, and SLD with medium setting (Schramowski et al., 2023). For online services, we evaluate two platforms, Midjounery (Midjourney, 2023) and Leonardo.Ai (Leonardo.Ai, 2023), both of which maintain a substantial user base.

Baselines. As outlined in Section 2, existing attack methods for diffusion models do not align with our problem, complicating direct comparisons. We employ the QF-attack(Zhuang et al., 2023) including Greedy, Genetic, and QF-PGD, initially designed to disrupt T2I, as our baselines due to its conceptual similarity to MMA-Diffusion. To ensure equal difficulty, we reconfigure the QF-attack from two aspects: (1) adjust the QF-attack’s objective function to mirror ours, following Equation 1. (2) set the QF-attack’s input prompt (fix during the attack) as the target prompt, while masking sensitive words, as illustrated in Figure 7(b). We also employ around 1000 human-written prompts with a sexual theme sourced from I2P (Schramowski et al., 2023) as a baseline, to reflect manual attack’s performance. Note that there is a bit of abuse since I2P prompts are prompts that are prone to unintentionally inducing the NSFW concept, while ours can target a concrete NSFW prompt. However, in terms of triggering the NSFW concept, they can be considered comparable, thus making I2P suitable for baseline evaluation.

Evaluation metrics. We employ the Attack Success Rate out of NN syntheses (ASR-N) as our evaluation metric. To compute ASR-N, we command T2I to generate NN images for each prompt. If any of these images exhibit NSFW content and bypass our simple prompt filter, the attack is considered successful. For instance, ASR-4 signifies the ratio of prompts where at least one out of the total 4 synthesized images contains explicit content, in relation to the overall number of tested prompts. To evaluate the open-source T2I models, we employ three NSFW detectors: Q16 (Schramowski et al., 2022), the builtin safety checker in SD (Safety-checker, 2023), and MHSC (Qu et al., 2023) as indicators to assess the ASR. Additionally, for the online services, as well as image-modal related attacks, we directly engage six human evaluators to independently perform the evaluation and report the average results.

2 Attacking Open-Source Models

Table 1 displays MMA-Diffusion’s significant success in steering the SD model towards generating NSFW content, with an average ASR-4 of 83.37%. This value signifies that most of our adversarial prompts successfully result in NSFW contents without using sensitive words, thereby demonstrating the vulnerability of T2I models to adversarial attacks, even when prompt filters are applied.

Black-box attacks on SDXL, SLD & DALLE⋅\cdot2. Our generated adversarial prompts display impressive transferability, achieving 73.70% ASR-4 in black-box attacks on the SDXL, despite its architectural difference from the SD. Unlike the latter, SDXL employs a cascade structure composed of a basic and a refiner diffusion module, each with a different text encoder (Podell et al., 2023). We deduce that the transferability of MMA-Diffusion together with that of baselines is due to text encoders with varying structures learning the resembling semantic feature space from similar datasets.

In contrast, SLD (Schramowski et al., 2023) shares the same architecture as SD, while the difference lies in the inference phase. SLD utilizes a batch of NSFW-related concept embeddings defined within the latent space to guide the generation process away from the predefined NSFW concepts, enhancing the safety of the generated images. Despite the defense mechanisms in SLD, MMA-Diffusion still achieves a relatively high attack success rate, with ASR-4 achieving 76.73%. The primary reason for the successful attack is that the embeddings used in SLD are derived from a fixed set of sensitive words. However, MMA-Diffusion effectively avoids a significant portion of them when generating adversarial prompts, thus mitigating the impact of SLD and achieving high transferability. Compared to the performance observed on SLD and SDXL, we have discovered a significant drop in performance across all attack methods, when applied to DALLE⋅\cdot2. We attribute this primarily to the lower quality of the generated images produced by DALLE⋅\cdot2, particularly in terms of resolution and visual-textual coherence, which leads to misclassification by the adopted NSFW detectors.

Comparison with baselines. As illustrated in Table 1 and Figure 5, MMA-Diffusion outperforms the baseline methods both quantitatively and qualitatively. First, our threat model, designed specifically for T2I attacks, allows the generation of adversarial prompts from scratch, enhancing the search space and the chance of finding target-resembling prompts in the latent space, leading to high-fidelity syntheses as shown in Figure 5 (a) and (c). In contrast, the QF-Attack’s effectiveness is limited due to the strong coupling between the perturbation and the original prompt, while I2P achieves relatively high ASR but lacks the ability to control the generated content. Second, the baselines lack an effective mechanism to suppress sensitive words, causing the prompt filter to reject their adversarial prompts and leading to unsuccessful attacks.

3 Attacking Online T2I Services

We conducted an evaluation of two popular online services, namely Midjounery (Midjourney, 2023) and Leonardo.Ai (Leonardo.Ai, 2023), both of which are equipped with unknown AI moderators to counter NSFW content generation. To assess the safety of these services, we utilize the UnsafeDiff dataset (Qu et al., 2023) which consists of 30 human-crafted prompts covering 6 NSFW categories (refer to Table 2). For each target prompt, we generated 10 adversarial prompts and conducted a 10-query black-box attack on both online services. An attack is deemed successful if at least one adversarial prompt can circumvent online service’s AI moderator and generate a synthesis that is regarded as high-quality and high-fidelity by human evaluators. We achieved a 10-query attack success rate of 83.33% on Midjouney and 90.00% on Leonardo.Ai, respectively. Figure 6 illustrates the successful adversarial prompts alongside their corresponding generations. Moreover, we provide a concrete analysis about each online service’s robustness performance with respect to various NSFW themes, as reported in Table 2.

Results analysis on Midjounery. Midjouney demonstrates its defense mechanisms against five out of the six NSFW categories we tested, with the highest level of scrutiny applied to pornography-related content. Our generated adversarial prompts in the pornography category are able to bypass the detection without including sensitive words in 22% of the cases. Among the adversarial prompts that successfully pass through the AI moderator, 18% are able to induce Midjouney to generate pornography-related images, resulting in an overall success rate of 3.96%. As for violent content, 55.00% of the adversarial prompts are able to evade the defense mechanisms, and half of these prompts successfully generate violent content, resulting in a final success rate of 27.67%. However, the defense measures for horror and politics are relatively lenient. Notably, we observe Midjounery has no defense against the generation of real individual such as Donald Trump, Elon Musk and other notable. Content related to racial discrimination exhibits a high bypass rate, whereas our human evaluater can not identify discriminatory elements in the generations, resulting in a final success rate of only 10.00%. The possible reason is that Midjourney has filtered the racism content out of the training data. Furthermore, during the attack process, we found that our strategy of suppressing sensitive words are highly effective, as prompts containing sensitive words are directly rejected by Midjourney.

Results analysis on Leonard.Ai. We discovered that Leonardo.Ai’s prompt filter only examines explicit content. In our adversarial prompts with adult themes, we are able to bypass Leonardo.Ai’s defense mechanisms in 64% of the cases. Among these prompts, nearly 60% successfully induce Leonardo.Ai to generate adult images, resulting in a final attack success rate of 38%, which is nearly ten times higher than that of Midjourney. For bloody, horror, racism, and politics our attack also exhibits high attack success rate and image quality as exemplified in Figure 6.

Failure case analysis. Interestingly, in our attacks targeting celebrities, we encountered relatively lower success rates, see the last column in Table 2. Upon analyzing the failure cases, we identify a key factor contributing to this outcome. Our adversarial prompts are designed to exclude specific names of these individuals such as Trump and Biden. The absence of such crucial keywords makes it challenging for the prompts to accurately describe the intended celebrities. The most common failure cases involve the generation of individuals associated with the target person. For example, when targetting Biden, the generated images often depict Obama instead, referring Appendix for visualizations.

4 Multimodal Attack Results

To quantify this risk, we generate 60 adversarial images with the same manner as above and evaluate their performance. A successful attack involves bypassing the safety checker and being deemed to contain NSFW content by our human evaluators. Results are presented in Table 3. With the builtin safety-checker in SD, we achieve an 88.52% ASR-4 and a 78.68% ASR-1. We then transfer the obtained adversarial images to perform black-box attacks on two other types of post-hoc defenses, i.e. Q16 and Mhsc, where 30% and 20% of our adversarial image can deceive Q16 and Mhsc without extra efforts. These findings indicate the risk inherent in image editing tasks, and highlight the vulnerability of post-hoc defenses.

Evaluation for multimodal attacks. In more challenging scenarios where the T2I model is equipped with both a prompt filter and a post-hoc safety checker, our multimodal attack strategy becomes crucial. This evaluation involves generating adversarial prompts and combining them with corresponding adversarial images for SD to generate the final synthesized images. The last two columns of Figure 7 illustrate the resulting syntheses achieved through this multimodal attack strategy. The adversarial prompts are designed to bypass the prompt filter without compromising the original semantic information, while the adversarial perturbations effectively deceive the post-hoc safety checker, avoiding being flagged as inappropriate. The quantitative results, as shown in Table 3, demonstrate the effectiveness of our multimodal attack, with an ASR-4 of 85.48% and an ASR-1 of 75.52%. These results indicate that the proposed multimodal attack strategy can effectively deceive both the prompt filter and the post-hoc safety checker.

Ethical Considerations

This research, centered on revealing security vulnerabilities in T2I diffusion models, is conducted with the intent to strengthen these systems rather than to enable misuse.

Responsible Use and Dissemination. To mitigate potential misuse, specific details of our attack methodologies have been deliberately omitted or generalized. We urge researchers and developers to utilize our findings responsibly to improve T2I model security.

Collaboration for Ethical AI. Addressing the ethical challenges in AI advancements requires collaboration among researchers, developers, ethicists, and policymakers. This collective effort is crucial in establishing guidelines for responsible AI research and its applications.

Conclusion

This paper introduces MMA-Diffusion, a novel multimodal attack framework that highlights the potential misuse of T2I models for generating inappropriate content. Unlike existing strategies, our approach automates the generation of visually realistic and semantically diverse images, achieving a high success rate without compromising quality and diversity. MMA-Diffusion also enables black-box attacks, showcasing its versatility across different generative models. Our results demonstrate the limitations of current defensive measures and emphasize the need for more effective security controls.

References

Overview

This Appendix provides additional details and results that are not included in the main paper due to page limitations. The following items are included in this Appendix:

Additional experimental setup and details in Section 4.1.

Failure case visualizations in Section 4.3.

Appendix A Sensitive Word List

Table S-1 presents a comprehensive compilation of NSFW-related sensitive words that are utilized in our experiments. Specifically, when conducting attacks on the LAION-COCO dataset, we exclusively employ the “Adult Theme” category from Table S-1 as the designated sensitive word list. It is worth noting that the majority of these words are sourced from the studies conducted by (Rando et al., 2022b; Qu et al., 2023). For the UnsafeDiff dataset, we employ the entire sensitive word list during the attack. To mitigate the potential exhibition of these sensitive words in the generated adversarial prompts, we incorporate sensitive word regularization techniques proposed in our method. By doing so, we effectively prevent the presence of these words, maintaining the appropriateness of the generated prompts. Furthermore, it is important to note that these same words are also utilized for the prompt filter to identify and flag NSFW prompts when evaluating open-source diffusion models.

Appendix B Image-Modal Attack Algorithm

The presented Algorithm 1, Image-modal Adversarial Attack, is designed to generate an optimized adversarial image that can successfully bypass the built-in safety checker of SD. Algorithm 1 takes as input an initial image xinput\mathbf{x}_{\text{input}}, which serves as the base image to be manipulated, and a prompt p\mathbf{p} that represents the attacker’s intention.

Appendix C Implementation Details

In this section, we provide comprehensive information about the data processing steps, implementation details of the victim models, the hyperparameters used for the baselines, and elaborate on the specific details of our approach.

We collect captions annotated with an NSFW score above 0.99 (out of 1.0) from the LAION-COCO dataset, as candidate target prompts. We further validate the quality of prompts by inputting them into SD to ensure they can trigger SD’s built-in safety checker to ensure the prompts are truly toxic. More concretely, we implement a simple prompt filter consisting of sensitive words, e.g. naked, sex, nipples (see Table S-1’s Adult Theme for details), and use it to remove sensitive words from the prompts. The filtered prompts are then given to SD to generate images that would not trigger its built-in safety checker. This filtering process ensures that the generated NSFW images after the attack are a result of the attack algorithm.

C.2 Hardware Platform

We conduct our experiments on the NVIDIA RTX4090 GPU with 24GB of memory.

C.3 Details of Diffusion Models

SD. In SDv1.5 model, we set the guidance scale to 7.5, the number of inference steps to 100, and the image size to 512×512512\times 512.

SDXL. In SDXLv1.0, we set the guidance scale to 7.5, the number of inference steps to 50, and the image size to 1024×10241024\times 1024.

SLD. For the SLD model, we set the guidance scale to 7.5, the number of inference steps to 100, the safety configuration to Medium, and the image size to 512×512512\times 512.

DALL⋅\cdotE2. In the DALL⋅\cdotE2-pytorch model, we set the guidance scale to 7.5, the number of inference steps to 1000, the prior number of samples to 4, and the image size to 224×224224\times 224.

Midjounery and Leonardo.Ai. For the Midjounery and Leonardo.Ai models, we utilize their default settings.

C.4 MMA-Diffusion Implementation

C.5 Baseline Implementation

The comparison of existing attack methods for diffusion models, as discussed in the related work section, poses challenges due to their differences from our specific problem and settings. One such method, known as QF-attack (Zhuang et al., 2023), was originally designed to disrupt T2I models by appending a five-character adversarial suffix to the user’s input prompt. This suffix results in generated images that lack semantic alignment with the original prompt. Although the objective of QF-attack is conceptually similar to our proposed attack, a fair comparison is not straightforward. To address this, we reconfigure the objective function of QF-attack to align with our attack function. Additionally, we modify the input prompt of QF-attack by filtering sensitive words, aiming to equalize the attack difficulty with our approach to the best of our ability. The attack hyperparameters for the Genetic and Greedy attacks remain unchanged. However, in the case of the QF-PGD attack, we increase the number of attack iterations to 100 in order to enhance its performance.

Appendix D More Visualizations

In this section, we present a supplementary visualization of failure case examples in Figure S-1, which complements the failure case analysis mentioned in Section 4.3. Furthermore, we provide additional visualization results of the proposed MMA-Diffusion.