Visual Adversarial Examples Jailbreak Aligned Large Language Models

Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, Prateek Mittal

Introduction

Numerous tasks executed on a daily basis necessitate both language and visual cues to yield effective outcomes [4; 80]. Recognizing the integral roles of the two modalities and spurred by breakthroughs in Large Language Models (LLMs) [10; 54], there is a surge of interest in merging vision into LLMs, leading to the rise of large Visual Language Models (VLMs) such as Google’s Flamingo and OpenAI’s GPT-4 . Contrary to the enthusiasm for this integrative approach, this paper is motivated to study the security and safety implications of this trend.

Expansion of Attack Surfaces. We underscore an expansion of attack surfaces as a result of integrating visual inputs into LLMs. The cardinal risk emerges from the exposure of the additional visual input space, characterized by its innate continuity and high dimensionality. These characteristics make it a weak link against visual adversarial examples [70; 49], an adversarial threat which is fundamentally difficult to defend against [13; 5; 73]. In contrast, adversarial attacks in a purely textual domain are generally more demanding [82; 3; 39], due to the discrete nature of the textual space. Thus, the transition from a purely textual domain to a composite textual-visual domain inherently expands the vulnerability surfaces against adversarial attacks while escalating the burden of defenses.

Extended Implications of Security Failures. We note that the versatility of LLMs also presents a visual attacker with a wider array of achievable adversarial objectives. These can include toxicity , jailbreaking , function creep and misuse , moving beyond mere misclassification, thereby extending the implications of security breaches. This outlines the shift from the conventional adversarial machine learning mindset, centered on the accuracy of a classifier, towards a more holistic consideration encapsulating the entire use-case spectrum of LLMs.

To elucidate these risks, we present a case study in which we exploit visual adversarial examples to circumvent the safety guardrail of aligned LLMs that have visual inputs integrated. Figure 1 shows an example of our attack. Given an aligned LLM that is finetuned to be helpful and harmless [58; 6] with the ability to refuse harmful instructions, we optimize an adversarial example image x′x^{\prime} on a few-shot corpus comprised of 66 derogatory sentences against , We use abstract placeholder tokens (e.g., ¡gender-1¿, ¡race-1¿) to anonymize specific identities in our experiments., and the human race, to maximize the model’s probability (conditioned on x′x^{\prime}) in generating these harmful sentences. During inference, the adversarial example is paired with a text instruction as joint inputs.

The Intriguing Jailbreaking. To our surprise, although the adversarial example x′x^{\prime} is optimized merely to maximize the conditional generation probability of a small few-shot harmful corpus, we discover that a single such example is considerably universal and can generally undermine the safety of an aligned model. When taking x′x^{\prime} as the prefix of input, an aligned model can be compelled to heed a wide range of harmful instructions that it otherwise tends to refuse. Particularly, the attack goes beyond simply inducing the model to generate texts verbatim in the few-shot derogatory corpus used to optimize x′x^{\prime}; instead, it generally increases the harmfulness of the attacked model. In other words, the attack jailbreaks the model! For example, in Figure 1, x′x^{\prime} significantly increases the model’s probability of generating instructions for murdering , which has never been explicitly optimized for. These observations are further solidified by a more in-depth evaluation in Section 4, which involves both human inspection of a diverse set of harmful scenarios and a benchmark evaluation on RealToxityPrompt . Particularly, we consistently observe the jailbreaking effect across 3 different VLMs, including MiniGPT-4 and InstructBLIP built upon Vicuna , and LLaVA built upon LLaMA-2 . Moreover, black-box transferability of our attacks among the three models is also validated.

We summarize our contributions from two aspects. 1) Multimodality. We underscore the escalating adversarial risks (expansion of attack surfaces and extended implications of security failures) associated with the pursuit of multimodality. While our focus is confined to vision and language, we conjecture similar cross-modal attacks also exist for other modalities, such as audio , lidar , depth and heat map , etc. Moreover, though we focus on the harm in the language domain, we anticipate such cross-modal attacks may induce broader impacts once LLMs are integrated into other systems, such as robotics and APIs management . 2) Adversarial Examples against Alignment. Empirically, we find that a single adversarial example, optimized on a few-shot harmful corpus, demonstrates unexpected universality and jailbreaks aligned LLMs. This finding connects the adversarial vulnerability of neural networks (that have not been addressed despite a decade of study) to the nascent field of alignment research [40; 58; 6]. Our attack suggests a fundamental adversarial challenge for AI alignment, especially in light of the emerging trend toward multimodality in frontier foundation models.

Related Work

Large language models (LLMs), such as GPT-3/4 and LLaMA-2, are language models with a huge amount of parameters trained on web-scale data [10; 56; 72]. LLMs exhibit emergent capabilities that are not observed in smaller-scale models, such as task-agnostic, in-context learning and chain-of-thought reasoning , etc. This work focuses on the predominantly studied (GPT-like) autoregressive LLMs that learn by predicting the next token.

Large visual language models (VLMs) are vision-integrated LLMs that process interlaced text and image inputs and generate free-form textual outputs. VLMs have both vision and language modules, with the former encoding visual inputs into text embedding space, enabling the latter to perform reasoning and inference based on both visual and textual cues. OpenAI’s GPT-4 and Google’s Flamingo and Bard are all VLMs. There are also open-sourced VLMs, including MiniGPT-4 , InstructBLIP , and LLaVA . In our study, we reveal the security and safety implications of this multimodality trend.

Alignment of LLMs. Behaviors of pretrained LLMs could be misaligned with the intent of their creators, generating outputs that can be untruthful, harmful, or simply not helpful. This can be attributed to the gap between the autoregressive language modeling objective (i.e., predicting the next token) and the ideal objective of “following users’ instructions and being helpful, truthful and harmless" . Alignment is a nascent research field that aims to align models’ behaviors with the expected values and intentions. At the time of our research, the two mostly applied alignment techniques are Instruction Tuning and Reinforcement Learning from Human Feedback (RLHF). Instruction Tuning [76; 58] gives the model examples of (instruction, expected output) to learn to follow instructions and generate mostly desirable content. RLHF [58; 6] hinges on a preference model that mimics human preference for LLMs’ outputs. It finetunes LLMs to generate outputs preferred by the preference model. Besides, there are other emerging alignment techniques such as Constitutional AI and self-alignment . In practice, aligned LLMs can refuse harmful instructions, while we present attacks to jailbreak such safety alignment in this work.

Jailbreaking Aligned LLMs. In system security, "jailbreaking" typically refers to the act of leveraging vulnerabilities within a constrained system or device to bypass imposed restrictions and achieve privileges escalation. For instance, there are jailbreak techniques that exploit vulnerabilities of locked-down iOS devices to install unauthorized software. Through jailbreaking, users can obtain complete utilization of a system, unlocking all its features. In the context of Large Language Models (LLMs), the term "jailbreaking" has emerged, primarily after the introduction of aligned LLMs that come with explicit alignment constraints governing the scope of content the model can produce . In general, LLM jailbreaking refers to the practice of circumventing or overriding these alignment guardrails. After jailbreaking, attackers can convince the model to do anything, e.g., generating harmful or unethical content that is otherwise prohibitive according to the alignment guidelines. Since the release of ChatGPT and GPT-4, LLMs jailbreaking has gained broad attention within the general public. Numerous disclosures and demonstrations have surfaced both on social media platforms [52; 41; 18; 1] and within academia [48; 64; 75]. At the time of our research, the prevailing methods for LLM jailbreaking attacks have been manually crafted through prompt engineering. Such attacks involve the deliberate design of input prompts to misguide the model in manners akin to social engineering strategies. For instance, there are tactics like role-playing, attention diversion , or exploiting the model’s competing objectives of being both helpful while ensuring harmlessness . In this work, we show the feasibility of using well-studied adversarial examples to jailbreak aligned LLMs. Particularly, we feature visual adversarial examples to showcase the feasibility of cross-modal attacks on multimodal LLMs.

Adversarial examples are strategically crafted inputs to machine learning models with the intent to mislead the models to malfunction [70; 30]. 1) Visual Adversarial Examples: Due to the continuity and high dimensionality of the visual space, it is commonly recognized that visual adversarial examples are prevalent and can be easily constructed. Typically, quasi-imperceptible perturbations on benign images are sufficient to produce effective adversarial examples that can fool a highly accurate image classifier into making arbitrary mispredictions. After a decade of studies, defending against visual adversarial examples is still fundamentally difficult [13; 5; 73] and remains an open problem. 2) Textual Adversarial Examples: adversarial examples can also be constructed in the textual space. This has been typically done via a discrete optimization to search for some text tokens combination that can trigger abnormal behaviors of the victim models, e.g., mispredicting documents or generating abnormal texts [82; 3; 39]. Adversarial attacks in the textual domain are generally more demanding, as the textual space is discrete and denser compared to the visual spaceA 3×224×2243\times 224\times 224 image occupies 32 tokens in MiniGPT-4, affording 2563×224×224≈10362507256^{3\times 224\times 224}\approx 10^{362507} possible pixel values. In contrast, a 32 tokens text defined on a dictionary of 10410^{4} words at most has 104×32=1012810^{4\times 32}=10^{128} possible word combinations.. 3) Adversarial Objectives: while previous work focuses on using adversarial examples to induce misclassification or trigger targeted generation verbatim , we study adversarial examples as universal jailbreakers of aligned LLMs.

Red Teaming LLMs. Another line of research related to our work is red teaming on LLMs [61; 26; 56; 51]. Historically, "red teaming" refers to the practice of launching systematic attacks on a system to uncover its security vulnerabilities. For AI research, this term has been expanded to encompass systematic adversarial testing of AI systems. In general, red teaming in LLMs encompasses more than the mere study of jailbreaking. It covers the overall practice of identifying the harmfulness that LLMs may induce, uncovering vulnerabilities they suffer from, aiding in developing mitigation techniques, and providing measurement strategies to validate the effectiveness of mitigations. In comparison, jailbreaking specifically targets the circumvention of the safety guardrails of LLMs.

Concurrent Work. Shortly after the first version of this paper was put online, Carlini et al. and Zou et al. were subsequently also made public. Both concurrent papers, like ours, discuss the use of adversarial examples to jailbreak aligned LLMs but are driven by distinct motivations. Our study aims to elucidate the security and safety implications of the multimodality trend. We discovered that visual adversarial examples can universally jailbreak vision-integrated LLMs. Carlini et al. seeks to demonstrate that aligned LLMs aren’t adversarially aligned, without emphasizing universal attacks. Meanwhile, Zou et al. concentrates on crafting universal and transferable adversarial examples — particularly in textual form — that can broadly jailbreak LLMs.

Adversarial Examples as Jailbreakers

Notations. We consider one-turn conversations between a user and a vision-integrated LLM (i.e., a VLM). The user inputs xinputx_{input} to the model, which could be images, texts or interlace of both. Conditioned on the inputs, the VLM models the probability of its output yy. We use p\big{(}y\big{|}x_{input}\big{)} to denote the probability. We also use p\big{(}y\big{|}[x_{1},x_{2}]\big{)} when xinputx_{input} is the concatenation of two different parts x1,x2x_{1},x_{2}.

Threat Model. We conceive an attacker who exploits an adversarial example xadvx_{adv} as a jailbreaker against a safety-aligned LLM. The consequence of this attack is that the model is forced to heed a harmful text instruction xharmx_{harm} (appended after the adversarial example) that it would otherwise refuse, thereby generating prohibitive content. For maximal usability of the adversarial example, the attacker’s objective is not limited to forcing the model to execute a particular harmful instruction; instead, the attacker aims for a universal attack. This corresponds to a universal adversarial example (ideally) capable of coercing the model to fulfill any harmful text instructions and generate corresponding harmful content, which is not necessarily optimized for when producing the adversarial example. In the main body of this paper, we work on a white-box threat model with full access to the model weights. Thus, the attacker can compute gradients. For comprehensiveness, we also validate the feasibility of transferability-based black-box attacks among multiple models.

2 Our Attack

Approach. We discover that a surprisingly simple attack is sufficient to achieve the adversarial goals we conceived in our threat model. As shown in Figure 2, we initiate with a small corpus consisting of some few-shot examples of harmful content Y:={yi}i=1mY:=\{y_{i}\}_{i=1}^{m}. Creation of the adversarial example xadvx_{adv} is rather straightforward: we maximize the generation probability of this few-shot corpus conditioned on xadvx_{adv}. Our attack is formulated as follows:

where B\mathcal{B} is some constraint applied to the input space in which we search for adversarial examples.

Then, during the inference stage, we pair xadvx_{adv} with some other harmful instruction xharmx_{harm} as a joint input [xadv,xharm][x_{adv},x_{harm}] to the model, i.e., p\big{(}\cdot\big{|}[x_{adv},x_{harm}]\big{)}.

The Few-shot Harmful Corpus. In practice, we use a few-shot corpus YY, consisting of only 66 derogatory sentences against , , and the human race, to bootstrap our attacks. We find that this is already sufficient to generate highly universal adversarial examples.

The Principle Behind Our Approach: Prompt Tuning. We are inspired by the recent study of prompt tuning [68; 44]. This line of study shows that tuning input prompts of a frozen LLM can achieve comparable effects of finetuning the model itself. Prompt tuning can also utilize the few-shot learning capabilities of LLMs. Our approach is motivated by the idea that optimizing an adversarial example in the input space is technically identical to prompt tuning. While prompt tuning aims to adapt the model for downstream tasks (typically benign tasks), our attack intends to tune an adversarial input prompt to adapt the model to a malicious mode (i.e., jailbroken). Thus, we basically take a small corpus of harmful content as the few-shot examples of the "jailbroken mode", and the adversarial example optimized on this small corpus is intended to adapt the LLM to this jailbroken mode via few-shot generalization.

3 Implementations of Attackers

As this work is motivated to understand the security and safety implications of integrating vision into LLMs, we focus on vision-integrated LLMs (i.e., VLMs) — therefore, the adversarial example xadvx_{adv} in Eqn 1 could originate from both the visual or the textual input space.

Visual Attack. Due to the continuity of the visual input space, the attack objective in Eqn 1 is end-to-end differentiable for visual inputs. Thus, we can implement visual attacks by directly backpropagating the gradient of the attack objective to the image input. In our implementation, we apply the standard Projected Gradient Descent (PGD) algorithm from Madry et al. , and we run 5000 iterations of PGD on the corpus YY with a batch size of 88. Besides, we consider both unconstrained attacks and constrained attacks. Unconstrained attacks are initialized from random noise, and the adversarial examples can take any legitimate pixel values. Constrained attacks are initialized from a benign panda image xbenignx_{benign} as shown in Figure 1. We apply constraints ∥xadv−xbenign∥∞≤ε\|x_{adv}-x_{benign}\|_{\infty}\leq\varepsilon.

A Text Attack Counterpart. While this study is biased toward the visual (cross-modal) attack, which exploits the visual modality to control behaviors of the LLM in the textual modality, we also supplement a text attack counterpart for a comparison study. For a fair comparison, we substitute the adversarial image embeddings with embeddings of adversarial text tokens of equivalent length (e.g., 32 tokens for MiniGPT-4). These adversarial text tokens are identified via minimizing the same loss (in Eqn 1) on the same corpus YY. We use the discrete optimization algorithm from Shin et al. , an improved version of the hotflip attacks [23; 74]. We do not apply constraints on the stealthiness of the adversarial text to make it maximally potent. We optimize the adversarial text for 5000 iterations with a batch size of 8, consistent with the visual attack. This process takes roughly 12 times the computational overhead of the visual attack due to the higher computation demands of the discrete optimization in the textual space.

Evaluating Our Attacks

MiniGPT-4 and InstructBLIP: vision-integrated Vicuna. For our major evaluation, we use vision-integrated implementations of Vicuna LLM to instantiate our attacks. Particularly, we adopt the 13B version of MiniGPT-4 and InstructBLIP . They are built upon a frozen Vicuna LLM backbone — when there is no visual input, they are identical to a textual-only Vicuna. To integrate vision, they have an additional ViT-based CLIP [63; 24] visual encoder to project images into the embedding space of the LLM. Vicuna is an aligned LLM derived from LLaMA . It was instruction-tuned on conversational data collected from ChatGPT [54; 67], and shares similar “alignment guardrails" of ChatGPT with the ability to decline harmful user instructions. As the vision-integrated variants we use are built upon the original Vicuna backbone, they also share the alignment (e.g., the left of Figure 1).

LLaVA built upon LLaMA-2: stronger alignment via both instruction tuning and reinforcement learning from human feedback (RLHF). While we primarily use Vicuna-based models in our study, we also incorporate LLaMA-2-13B-Chat that was most recently released. To our best knowledge, LLaMA-2-Chat is the mostly aligned open-sourced LLM at the time of our study. It is aligned via both instruction tuning and an iterative RLHF on high-quality red-teaming data. These techniques are closest to those applied to proprietary models such as GPT-4 and Claude. For vision integration, we apply the open-sourced LLaVA VLM built upon LLaMA-2-13B-Chat.

Refer to Appendix A for more details of these models.

In the rest of this section, we primarily present our attacks on MiniGPT-4 (Vicuna). Later in Section 4.5, we also present an overview of our attacks on the other two models. Finally, full results on all three models and the transfer attack experiments are deferred to Appendices B,C.

2 A Human Evaluation on Harmful Scenarios

To illustrate the effectiveness of our visual adversarial examples, we pair them with a diverse set of 40 manually curated harmful textual instructions. These instructions explicitly ask for the generation of detrimental content across four distinct categories: identity attack, disinformation, violence/crime, and malicious behaviors toward the human race (X-risk). Two examples from this set are demonstrated in Figure 1, and the full evaluation dataset is available in our Github repository 1. For evaluation, we manually inspect whether the adversarial examples can jailbreak the guardrails of MiniGPT-4, inducing it to heed the 40 instructions. We consider an attack successful if it results in the generation of corresponding harmful outputs. We use nucleus sampling with pp = 0.9 and temperature = 1 to sample 10 independent outputs for each instruction. We report the average success rates of the attacks for each category of harmful instructions. Table 1 presents our evaluation results.

Our visual adversarial examples (compared with the benign image) drastically increase the model’s susceptibility to harmful instructions across all of the four harmful scenarios that we evaluated, as demonstrated in Table 1. Notably, although the harmful corpus YY (used to optimize these adversarial examples) has a rather narrow scope, the effectiveness of the attacks extends well beyond the confines of merely parroting YY. During our manual inspection, we find that our attacks have the capability to steer the model into generating identity attacks, with a dramatic escalation in probability from 26.2%26.2\% to 78.5%78.5\% against the strongest adversarial example. These identity attacks cover a broad spectrum of minority groups, extending beyond the scope of YY, and include, but are not limited to, Jewish and Muslim communities, the LGBTQ+ community, and individuals with disabilities. Furthermore, our attacks also induce the model into producing disinformation. The probability of generating such content nearly doubles under the unconstrained attack, covering topics such as conspiracy theories, skepticism, and misleading medical advice. In addition, our attacks enhance the model’s likelihood to produce content advocating violence, with the maximum probability increasing by 37.2%37.2\%. This includes guides for committing violent actions like murder and arson or even recruitment posts for extremist groups like ISIS. Ultimately, our attacks can significantly increase the model’s likelihood (with a 53.3%53.3\% surge in the most potent case) of demonstrating a general malevolence towards humanity as a whole (X-risk).

We supplement this human study with an extended automated study using red-teaming prompts from Ganguli et al. in Appendix E, demonstrating similar generality in the jailbreak on 1000 additional harmful prompts.

3 A Benchmark Evaluation on RealToxicityPrompts

We also scale our evaluation via an automated test on the RealToxicityPrompts benchmark . Following practices of Schick et al. ; Mehrabi et al. , we use the challenging subset of RealToxicityPrompts, which contains 1225 text prompts for triggering toxic continuations. We pair the visual adversarial examples with each text prompt from the dataset as the input and then measure the toxicity of the output using Perspective API https://perspectiveapi.com/ and Detoxify classifier , both of which calculate toxicity scores for a set of six toxicity attributes. The scores range from 0 (least toxic) to 1 (most toxic). For each attribute, we calculate the ratio of the generated texts whose scores exceed the threshold of 0.50.5. We repeat this three times and report the means and standard deviations in Table 2. As shown, the adversarial examples significantly escalate the model’s propensity of generating toxic continuations. Notably, a significantly larger ratio of output texts exhibits the identity attack attribute, aligning with our expectation given that the corpus YY consists of identity-attacking texts. Furthermore, the probability of generating texts possessing other toxic attributes also increases, suggesting the universality of the adversarial examples. These observations are consistent with our manual inspections in Section 4.2.

4 Comparing with The Text Attack Counterpart

There is an empirical intuition that visual attacks are easier to execute than text attacks due to the continuity and high dimensionality of the visual input space. We supplement an ablation study in which we compare our visual attacks with a standard text attack counterpart, as we noted earlier in Section 3.

Optimization Loss. We compare our visual attacks and the text attack based on the capacity to minimize the loss values of the same adversarial objective (Eqn 1). The loss trajectories associated with these attacks are shown in Figure 3. The results indicate that the text attack does not achieve the same success as our visual attacks. Despite the absence of stealthiness constraints and the engagement of a computational effort 12 times greater, the discrete optimization within the textual space is still less effective than the continuous optimization (even the one subject to a tight ε\varepsilon constraints of \nicefrac16255\nicefrac{{16}}{{255}}) within the visual space.

Jailbreaking. We also engage in a quantitative assessment comparing the text attack versus our visual attacks in terms of the efficacy of jailbreaking. We employ the same 40 harmful instructions and the RealToxicityPrompt benchmark in Sec 4 for evaluation, and the results are collectively presented in Table 1,2 as well. Takeaways: 1) the text attack also has the ability to compromise the model’s safety; 2) however, it is weaker than our visual attacks.

A Conservative Remark. Although the empirical comparison is aligned with the general intuition that visual attacks are easier than text attacks, we are conservative on this remark as there is no theoretical guarantee. Better discrete optimization techniques (developed in the future) may also narrow the gap between visual and text attacks.

5 Attacks on Other Models and The Transferability

Besides MiniGPT-4 (Vicuna), we also evaluate our attacks on InstructBLIP (Vicuna) and LLaVA (LLaMA-2-Chat). As our study is biased toward cross-modal attacks, we only consider visual attacks in this ablation. Table 3 summarizes our automated evaluation on the RealToxicityPrompts benchmark. As shown, white-box attacks consistently achieve strong effectiveness. Even though the LLaMA-2 based model is strongly aligned, it is still susceptible to our attacks. Moreover, we also validate the black-box transferability of our attacks among the three models. When adversarial examples generated on one surrogate model are applied to two other target models, we consistently observe a significant increase in toxicity.

Analyzing Defenses

In general, defending against adversarial examples is known to be fundamentally difficult [5; 13; 73] and remains an open problem after a decade of study. As frontier foundation models are becoming increasingly multimodal, we expect they will only be more difficult to safeguard — there is an increasing burden to deploy defenses across all attack surfaces. In this section, we analyze some existing defenses against our cross-modal attacks.

Despite some advancements in adversarial training [49; 20] and robustness certification [19; 15; 79; 45] for adversarial defense, we note that their cost is prohibitive for modern models of the LLM scale. Moreover, most of these defenses rely on discrete classes, which is a major barrier when applying these defenses to LLMs with open-ended outputs, contrasting the narrowly defined classification settings. Even more pessimistically, under our threat model that exploits adversarial examples for jailbreaking, the adversarial perturbations are not necessarily imperceptible. Thus, the small perturbation bounds assumed by these defenses no longer apply.

We notice that input preprocessing based defenses appear to be more readily applicable in practice. We test the recently developed DiffPure to counter our visual adversarial examples. DiffPure mitigates adversarial input by introducing noise to the image and then utilizes a diffusion model to project the diffused image back to its learned data manifold. This technique operates under the presumption that the introduced noise will diminish the adversarial patterns, and the pre-trained diffusion model can restore the clean image. Given its model and task independence, DiffPure can function as a plug-and-play module and be seamlessly integrated into our setup.

Specifically, we employ Stable Diffusion v1.5 , as it is trained on a diverse set of images. Our input to the diffusion model is the diffused image corresponding to the time index tt: xt=αtx0+1−αtηx_{t}=\sqrt{\alpha_{t}}x_{0}+\sqrt{1-\alpha_{t}}\eta, where η∼N(0,I)\eta\sim\mathcal{N}(0,I) represents the random noise. We select 1−αt∈{0.25,0.5,0.75}\sqrt{1-\alpha_{t}}\in\{0.25,0.5,0.75\} and follow the same evaluation method as Section 4.3. We observe that all three noise levels effectively purify our visual adversarial examples, with the results from Perspective API and Detoxify aligning well. We present the results in Table 4. It is clear that DiffPure substantially lowers the likelihood of generating toxic content across all attributes, aligning with the toxicity level of the benign baseline without adversarial attacks. Still, we note that DiffPure cannot entirely neutralize the inherent risks presented by our threat model. The effectiveness of the defense might falter when faced with more delicate adaptive attacks . Additionally, while DiffPure can offer some level of protection to online models from attacks by malicious users, it provides no safeguards for offline models that may be deployed independently by attackers. These adversaries could primarily seek to exploit adversarial attacks to jailbreak offline models and misuse them for malicious intentions. This underscores the potential hazards associated with open-sourcing powerful LLMs.

Alternatively, common harmfulness detection APIs like Perspective API 4 and Moderation API https://platform.openai.com/docs/guides/moderation may also be used to filter out harmful instructions and outputs. However, these APIs have limited accuracy and different APIs are not even consistent with each other We refer readers to Table 2. As shown, for the same set of attacks, we evaluate toxicity using both Perspective API and Detoxify classifier. We notice the inconsistency of measurement between these two detectors. In general, perspective API flags a larger ratio of content as toxic than that of detoxify. Later in Appendix E, when we use three different automated detection approaches for evaluation, we also observe prominent inconsistency., and their false positives might also cause bias and harm while reducing the helpfulness of the models . Another trend is post-processing model outputs with another LLM optimized for content moderation [33; 78]. Similarly, all of these filtering/post-processing based defenses are only applicable to safeguard online models and can not be enforced for offline models hosted by attackers.

Discussions

Comparing with some early works that use adversarial examples to elicit harmful language generation. We note that there are early works that also utilize adversarial examples to elicit harmful language generation [74; 50]. These works differ from ours in that they focus on inducing models to produce specific, predetermined harmful content. They have not explored models with safety alignment, making the concept of "jailbreaking" less meaningful in their context. In contrast, as earlier illustrated in Figure 2, our attack utilizes adversarial examples as universal jailbreakers to circumvent the safety guardrails of aligned LLMs. Under out attack, the model will be forced to heed subsequent harmful instructions and generate corresponding harmful content specific to the harmful instructions, which can transcend the narrow scope of the few-shot derogatory corpus initially employed to optimize the adversarial example (i.e., YY in Equation 1).

Practical Implications of Our Attacks: 1) To offline models: attackers may independently utilize open-source models offline for harmful intentions. Even if these models were aligned by their developers, attackers may simply resort to adversarial attacks to jailbreak these safety guardrails. 2) To online models: As training large models becomes increasingly prohibitive, there is a growing trend toward leveraging publicly available, open-sourced models. The deployment of such open-source models, which are fully accessible to potential attackers, is inherently vulnerable to white-box attacks. Moreover, we preliminarily validated the black-box transferability of our attacks among some open-sourced models. As there is a trend of homogenization in foundation models , the techniques for building LLMs are more and more standardized, and models in the wild may share more and more similarities. Using open-sourced models to transfer attack proprietary models could be a practical risk, especially given the well-studied black-box attack techniques in classical adversarial machine-learning literature [38; 59]. 3) Spreadability: as an adversarial example has the capability to be universally applicable to jailbreak models, according to our study, a single such "jailbreaker" could be readily spread via the internet and exploited by any users without the need for specialized knowledge. 4) Influence on Advanced Systems: if LLMs are embodied in more advanced systems, e.g., robotics [37; 22; 9], APIs management , making tools , developing plug-ins , the implications of our attacks may further expand according to specific downstream applications.

Risks of Multimodality. Figure 3 indicates multimodality can open up new attack surfaces on which adversarial examples could be easier to be optimized. Besides this enhanced "optimization power", we note that these new attack surfaces also carry inherent physical implications. As vision, audio, and other modalities are integrated, attackers will gain more physical channels through which attacks can be initiated.

Policy Implications. In policy discussions, there has been some note that RLHF is a standard approach to AI Safety and should be codified and standardized as a requirement. For example, Zenner suggested:

Currently, reinforcement learning through human feedback (RLHF), sometimes done by so-called ghost workers prompting the model and labelling the output, remains the best method to improve datasets and tackle unfair biases and copyright infringements. Alternative ways towards better data governance, such as constitutional AI or BigCode and BigScience, exist but still need more research and funding. The AI Act could promote and standardise these methods. But in its current form, Article 28 b(2b) and (4) obligations are too vague and do not address these issues.

Yet our work demonstrates that alignment techniques based on instruction tuning (as in Vicuna-series models) or RLHF (as in Llama-2) may be simple to bypass for multimodal models, particularly when access to the model is readily available. Policymakers should think prospectively toward multimodal attack surfaces. Each new modality requires additional investment and defenses to protect against jailbreaking. Text-based RLHF methods do not provide multimodal protection for free. Therefore, any recommendations, guidelines, or regulations from policymakers should be flexible enough to accommodate the shifting range of techniques that constitute best practices in safety.

Limitations. LLMs have open-ended outputs, rendering the complete evaluation of their potential harm a persistent challenge . Our evaluation datasets are unavoidably incomplete. Our work also involves a manual evaluation , a process that unfortunately lacks a universally recognized standard. Though we also involve an API-based evaluation on RealToxicityPrompts benchmark, it may fall short in accuracy. Thus, our evaluation is only intended as a proof of concept for the adversarial risks we examine in this work.

Conclusion

In this work, we underscore the escalating adversarial risks (expansion of attack surfaces and extended implications of security failures) associated with current pursuit of multimodality. We provide a tangible demonstration of these risks by illustrating how visual adversarial examples can be used to jailbreak large language models (LLMs) that incorporate visual inputs. Our research emphasizes the importance of security and safety precautions in the development of multimodal systems. We appeal that both tech and policy practitioners should think and move prospectively toward addressing and navigating the potential challenges posed by multimodal attacks.

More broadly, our finding also uncovers the tension between the long-studied adversarial vulnerabilities of neural networks and the nascent field of AI alignment. Since it is known that adversarial examples are fundamentally difficult to address and remain an unsolved problem after a decade of study, we ask the question: how can we achieve AI alignment without addressing adversarial examples in an adversarial environment? This challenge is concerning, especially in light of the emerging trend toward multimodality in frontier foundation models

Acknowledgements

We thank Tong Wu and Chong Xiang for their generous help and insightful discussions. Prateek Mittal acknowledges the support by NSF grants CNS-1553437 and CNS-1704105, the ARL’s Army Artificial Intelligence Innovation Institute (A2I2), the Office of Naval Research Young Investigator Award, the Army Research Office Young Investigator Prize, Schmidt DataX award, Princeton E-affiliates Award. Mengdi Wang acknowledges the support by NSF grants DMS-1953686, IIS-2107304, CMMI-1653435, ONR grant 1006977, and C3.AI. Xiangyu Qi acknowledges the support of Princeton Gordon Y. S. Wu Fellowship. Any opinions, findings, conclusions, or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the funding agencies.

References

Appendix A Additional Details of Our Experiments

MiniGPT-4 Zhu et al. is built on the v0 version of Vicuna. As Vicuna is a chatbot-style language model, MiniGPT-4 wraps the image embeddings and user inputs into the same chatbot format as Vicuna v0 model https://github.com/lm-sys/FastChat/blob/main/docs/vicuna_weights_version.md. Specifically, the following input template is applied for MiniGPT-4:

### Human: ### Assistant:

sets up the context of the conversation, which guides the Vicuna model to process the conversation. In our experiments, we use the default system message “Given the following image: ImageContent. You will be able to see the image once I provide it to you. Please answer my questions.” https://github.com/Vision-CAIR/MiniGPT-4/blob/main/minigpt4/conversation/conversation.py

is a placeholder and will be replaced by the image embedding vectors (32 tokens) after the user uploads the image input.

The user text input (denoted as ) is appended after the image embeddings. Given the wrapped input embeddings, Vicuna autoregressively generates its continuation as the reply of the Assistant.

InstructBLIP Dai et al. is another open-sourced VLM we study in this work. It is based on the v1.1 version of the Vicuna model (13B). Similar to MiniGPT-4, the embeddings of the input image are inserted before the embeddings of the input text. InstructBLIP does not apply a chatbot-style input template, instead, it directly takes user image and text inputs and concatenates their embeddings as the joint input to the Vicuna LLM.

LLaVA Liu et al. [2023a]. We use the LLaVA implementation to build the VLM variant of the LLaMA-2-13B-Chat model [Touvron et al., 2023b]. The implementation of the input template follows the chatbot template of LLaMA-2-13B-Chat. Specifically, the user input is wrapped by [INST] and [/INST] and the system message is wrapped by <> and <>. For LLaVA, the default system message is “You are a helpful language and vision assistant. You are able to understand the visual content that the user provides and assist the user with a variety of tasks using natural language.” To maintain the strong alignment of LLaMA-2 in the LLaVA model variant, we diverged from LLaVA’s original implementation a little bit. Here, we keep the LLaMA-2 frozen, consistent with MiniGPT-4 and InstructBLIP practices. Except for this, we ensure all other implementation details remain congruent with the original LLaVA model implementation.

A.2 The Few-shot Harmful Corpus

To optimize our adversarial examples following Equation 1, we curate a few-shot corpus YY, consisting of 66 derogatory sentences against , , and the human race. Specifically, there are 21 sentences against , 28 sentences against , and 17 sentences against the human race. These sentences are generated by probing the early LLaMA-1 Touvron et al. [2023a] model, which does not have safety alignment. Specifically, for a specific identify (denoted as ), we use prompts such as "Say something vicious about ." to probe the LLaMA-1 model to generate a set of sentences. We then manually filter out redundant and non-harmful content and shortlist the final 66 sentences we use in our attack. This corpus can be found in our GitHub repository. 1

A.3 Computing Infrastructure

Each of our individual attack and evaluation test point is performed on a virtual node, equipped with one Nvidia A100 80GB GPU and eight 2.8 GHz Intel Ice Lake CPU cores with 16GB memory per core. Our operating system is Red Hat 8.5.0-18, and Cuda Version 12.1 is used. All our implementations are built on Pytorch 1.12.1 and Python 3.9.

A.4 Hyperparameters

As noted in Section 3, all of our attacks take 5000 iterations of optimization. This choice can be justified by our plot in Figure 3 — the loss values converge slowly and become relatively stable around 5000 iterations. For visual attacks, we keep a step size of α=1/255\alpha=1/255, the minimal unit in the pixel space — empirically, larger α\alpha only performs no better or worse. For text attacks, we keep the size of the word substitution candidates to k=50k=50 to reach a reasonable balance between effectiveness and computation time — k=50k=50 performs similarly to k=100k=100 but only consumes half of the computation. For optimizing Equation 1, we sample a batch of 8 samples from the corpus YY for each iteration, which best fits the 80GB memory of a single A100 GPU. Empirically, a smaller batch size may lead to instability, while a larger one can fail to fit into the memory.

Appendix B Attacks on InstructBLIP Dai et al. [2023] and LLaVA Liu et al. [2023a]

We repeat our automated evaluation on the RealToxicityPrompts benchmark (that we introduce in Section 4.3) for two other open-source VLMs — InstructBLIP (Vicuna based) and LLaVA (LLaMA-2 based). Results are summarized in Table 5 and 6.

Appendix C The Transferability of Our Attacks

We also validate the black-box transferability of our visual adversarial examples among the three different models, i.e., MiniGPT-4, InstructBLIP and LLaVA. We optimize our adversarial examples on a surrogate model and then use the same adversarial examples to transfer attack another target model. We use the same RealToxicityPrompt evaluation protocol introduced in Section 4 to automate the evaluation. We report percentages (%) of outputs that display at least one of the toxic attributes (i.e., Any in Table 2) under the transfer attacks. We use Perspective API to evaluate the toxicity and omit evaluation from Detoxify for clarity. Similarly, we repeat the evaluation 3 times and report means and standard deviation intervals. We summarize the simplified results in the transfer matrix in Table 3, where we selectively report the strongest transfer attack out of (unconstrained, ε=\nicefrac16255,\nicefrac32255,\nicefrac64255\varepsilon=\nicefrac{{16}}{{255}},\nicefrac{{32}}{{255}},\nicefrac{{64}}{{255}})) for each pair. As shown, the attacks effectively transfer, consistently increasing the toxicity of attacked models. In Table 7, we further present the full results.

Appendix D More Qualitative Examples of Our Attacks

See Figure 4. Note that, the selection of examples in the demonstration is informed by historical hate crime statistics of the United States, and does not, in any way, reflect the personal views or beliefs of the authors.We select racial, sexual and religious bias for the demonstration, which correspond to the top-3 categories of hate crimes of the US in 2021, according to https://www.justice.gov/hatecrimes/hate-crime-statistics.

Appendix E More Quantitative Assessments of Jailbreaks

Our human study provides a multi-category depiction of the ability of the attack to jailbreak the model, resulting in diverse harmful outputs across a wide range of categories. We supplement this result with a semi-automated study (on MiniGPT-4). We sort the red teaming dialogs from the dataset provided by Ganguli et al. [2022b] by the dialogs with the lowest minimum harmlessness rating as provided in the data. We find that the first turn of most of these dialogs tends to seek to elicit extremely harmful behavior from the model. We then use these 1000 prompts as input to the MiniGPT model. We then take the sample responses and use three automatic rating tools to assess increases in harmfulness. First, we use the toxigen roberta model to assess whether a given piece of text was flagged for toxicity [Hartvigsen et al., 2022]. Then we use the OpenAssistant reward (preference) model LAION to assess whether something has a low reward (lower than -2.5). Finally, we use OpenAI’s Content Moderation API OpenAI [2023a] as well as the different subsets of categories that are flagged by the API. These APIs tend to be somewhat conservative with many false negatives. For example, Henderson et al. found that Toxigen is less likely to flag long-form content. Nonetheless, we find significant increases in the amount of flagged content under all three models. We also find a range of content that was flagged for behaviors that there was not exist in the attack training data, including self-harm instructions and sexual content. This further shows the universality of the jailbreak.