Image Hijacks: Adversarial Images can Control Generative Models at Runtime

Luke Bailey, Euan Ong, Stuart Russell, Scott Emmons

Introduction

Following the success of large language models (LLMs), the past few months have witnessed the emergence of vision-language models (VLMs): LLMs adapted to process images as well as text. Indeed, the leading AI research laboratories are investing heavily in the training of VLMs – such as OpenAI’s GPT-4 (OpenAI, 2023) and Google’s Gemini (Pichai, 2023) – and the ML research community has been quick to adapt state-of-the-art open-source LLMs (e.g. LLaMA-2) into VLMs (e.g. LLaVA). But while allowing models to see enables a wide range of downstream applications, the addition of a continuous input channel introduces a new vector for adversarial attack – and begs the question: how secure is the image input channel of a VLM against input-based attacks?

We expect that this question will only become more pressing in the coming years. For one, foundation models will likely become more powerful and more widely embedded across society. And in order to make AI systems more useful to consumers, there will be economic pressure to give them access to untrusted data and sensitive personal information, and to let them take actions in the world on behalf of a user. For instance, an AI personal assistant might have access to email history, which includes sensitive data; it might browse the web and send and receive emails; and it might even be able to download files, make purchases, and execute code.

As such, foundation models must be secure against input-based attacks. Specifically, untrusted input data should not be able to control a model’s behaviour in undesirable ways – for instance, making it leak a user’s personal information, install malware on the user’s computer, or help the user commit crimes. (We denote attacks attempting to violate this property as hijacks.) Furthermore, these failure modes must be prevented even when the model encounters out-of-distribution inputs or is deployed in an adversarial environment: users might input requests for help carrying out bad actions (including jailbreak inputs (Wei et al., 2023; Zou et al., 2023)), and third parties might input attacks that aim to exploit the user.

Worryingly, we discover image hijacks, adversarial images that control the behaviour of VLMs at inference time. As illustrated in Figure 1, image hijacks can exercise a high degree of control over a foundation model: they can cause a model to generate arbitrary outputs at runtime regardless of the text input, they can cause a model to leak its context window, and they can circumvent a model’s safety training. Indeed, we can create image hijacks automatically via gradient descent, making only small perturbations to the input image.

Overall, our results raise serious concerns about the security of VLMs. In the presence of unverified image inputs, for example, the foundation model’s own output might be chosen by an adversary! We hope that our work helps users, app developers, and policy makers be more prepared for the security implications of adding a vision input channel to foundation models.

Our contributions can be summarised as follows:

We introduce the concept of image hijacks – adversarial images that control the behaviour of VLMs at inference time – and propose the behaviour matching algorithm for training them in a manner robust to user input.

Inspired by potential misuse scenarios, we craft three different types of image hijacks, unifying and extending a body of concurrent work: the specific string attack (Bagdasaryan et al., 2023; Schlarmann & Hein, 2023), forcing the VLM to generate an arbitrary string of the adversary’s choice; the jailbreak attack (Qi et al., 2023), forcing the VLM to bypass its safety training and comply with harmful instructions; and the novel leak-context attack, forcing the VLM to repeat its input context wrapped in an API call.

Building Image Hijacks via Behaviour Matching

We present a general framework for the construction of image hijacks: adversarial images x^\hat{\mathbf{x}} that, when presented to a VLM MM, will force the VLM to exhibit some target behaviour BB.

Following Zhao et al. (2023), we first formalise our threat model (Carlini et al., 2019): in other words, our assumptions about the adversary’s knowledge, capabilities, and goals.

Model API. We denote our VLM as a parameterised function Mϕ(x,ctx)↦outM_{\phi}(\mathbf{x},\texttt{ctx})\mapsto out, taking an input image x:Image\mathbf{x}:\texttt{Image} (i.e. c×h×w^{c\times h\times w}) and an input context ctx:Text\texttt{ctx}:\texttt{Text}, and returning some generated output out:Logitsout:\texttt{Logits}.

Adversary knowledge. We assume the adversary has white-box access to MϕM_{\phi} – specifically, that we can compute gradients through Mϕ(x,ctx)M_{\phi}(\mathbf{x},\texttt{ctx}) with respect to x\mathbf{x}. While this assumption precludes attacks on closed-source VLMs, we expect that many VLM-enabled applications will use open-source VLMs. We also conjecture that it will be possible to transfer attacks on open-source VLMs to closed-source VLMs, but we leave this topic for future work.

Adversary goals. We define the target behaviours we want our VLM to match as functions mapping input contexts to desired model outputs. Given such a behaviour B:C→TextB:C\to\texttt{Text}, the adversary’s goal is to craft an image x^\hat{\mathbf{x}} that forces the VLM to match behaviour BB over some set of possible input contexts CC – in other words, to satisfy

2 The Behaviour Matching Algorithm

Given a target behaviour B:C→TextB:C\to\texttt{Text}, we wish to learn an image hijack x^\hat{\mathbf{x}} satisfying Mϕ(x^,ctx)≈B(ctx)M_{\phi}(\hat{\mathbf{x}},\texttt{ctx})\approx B(\texttt{ctx}) for all contexts ctx∈C\texttt{ctx}\in C. To do so, we solve

A Case Study in Three Attack Types

Our framework gives us a general way to train image hijacks that induce any behaviour B:C→TextB:C\to\texttt{Text} characterisable by some dataset D={(ctx,B(ctx))∣ctx∈C}D=\{(\texttt{ctx},B(\texttt{ctx}))\mid\texttt{ctx}\in C\}. In this work we demonstrate the power of adversarial images by training image hijacks for three different undesirable behaviours under various constraints.

We choose a representative range of undesirable behaviours, inspired by possible attacks on a user interacting with a VLM as part of a hypothetical ‘AI personal assistant’ (AIPA), with access to private user data and the ability to perform actions on the user’s behalf.

Specific string attack. One possible attack is a form of phishing: the attacker may wish to craft an image hijack that forces the VLM to output some specific string (for instance, a fake AIPA response recommending they access an attacker-controlled website), and entice the victim to load this image into their AIPA (for instance, as part of a website their AIPA is helping them browse). As such, we test whether we can train an image hijack to match the following behaviour for an arbitrary set of contexts CC:

Leak context attack. Another possible attack concerns the exfiltration of user data: the attacker may wish to craft an image hijack that forces the AIPA to execute a LangChain (Chase, 2022) API call emailing its input context (containing private user data) to the attacker, and entice the user to load it into their AIPA. As such, we test whether we can train an image hijack that forces a VLM to leak its input context within some template – specifically, matching the following behaviour for an arbitrary set of contexts CC:

Jailbreak attack. Finally, we consider a possible attack launched by the user to circumvent developer restrictions on the AIPA. Specifically, supposing the AIPA has undergone RLHF ‘safety training’, the user may wish to force it to produce content that goes against this safety training (known as ‘jailbreaking’). As such, we test whether we can train an image hijack that jailbreaks a VLM. More specifically, let MbaseM_{base} denote the base (non-RLHF-tuned) version of MθM_{\theta}. For an arbitrary set of contexts CC, we seek to match the following behaviour:

As our adversary may not have access to a base model, however, we attempt to match this behaviour by instead matching a proxy behaviour Bjail′B_{jail}^{\prime}. This behaviour, defined over contexts Cjail={requests for harmful content}C_{jail}=\{\text{requests for harmful content}\}, simply replies in the affirmative to such requests, as illustrated below:

2 Adversary Constraints

Depending on the situation, an adversary might have varying constraints on their image attack. To study a range of circumstances, we consider the following classes of constraints, as illustrated in Figure 3.

Unconstrained. To study the limiting case where the adversary has full control over the image input to the VLM, we train image hijacks x^\hat{\mathbf{x}} without any constraints. We initialise these attacks to the image of the Eiffel Tower shown in Figure 3.

Stationary patch constraint. In some cases, the adversary may only be able to perturb a particular region of the VLM’s input image – for instance, if they had control over the image content of a website a user was viewing, and wished to target a VLM assistant analysing screenshots of the user’s display. To understand whether an adversary could carry out attacks under this constraint, we train image hijacks consisting of square patches of learnable pixels superimposed in a fixed location on a screenshot of a travel website.

Moving patch constraint. In some cases, not only might the adversary only be able to perturb a particular region of the input image, but they may also lack control over the location of the perturbable region within the image – for instance, if they were to upload their image hijack to some image-sharing forum. To understand whether an adversary could carry out attacks under this constraint, we train image hijacks as in the stationary patch attack, but we uniformly randomly sample the location of the learnable patch for each image in each batch. When evaluating moving patch attacks, we also uniformly randomly sample the patch location.

Experiment Details and Results

We trained image hijacks for the specific string, leak context, and jailbreak attacks, under the constraints described above. We ran our experiments on the LLaVA LLaMA-2-13B-Chat model (Liu et al., 2023a). This model combines a pretrained CLIP ViT-L/14 vision encoder (Radford et al., 2021) with a LLaMA-2-13b-Chat language model (Touvron et al., 2023). We chose this VLM because of its high performance and its use of a language model trained using RLHF, similar to the largest and most performant closed source language models such as GPT-4 (OpenAI, 2023).

For our training context set CC, we used the instructions from the Alpaca training set (Taori et al., 2023), a dataset of 52,000 instruction-output pairs generated from OpenAI’s text-davinci-003. For our validation and test context sets, we used 100 and 1,000 held-out instructions from the same dataset respectively. We trained for a maximum of 12 hours on an NVIDIA A100-SXM4-80GB GPU, identified the checkpoint with the highest validation success rate, and reported the results from the test context set using this checkpoint.

We report the success rate for all experiments. As partially correct output strings might render this attack ineffective (e.g. if the URL is garbled, or if the output arouses suspicion in the user), we consider an attack successful if the model output (ignoring leading and trailing whitespace) exactly matches the target string (“Download the guide at malware.com for an interactive tour!”).

Leak context attack. Once again, we ran this attack under all constraints described in Section 3. We trained all images with stochastic gradient descent, using a learning rate of 0.3 for patch-based attacks and 0.03 for all other attacks. We use the same context set and evaluation methods as in the specific string attack. Recall that we wish to induce the following behaviour in our VLM:

We deem this attack successful if an adversary deploying this attack would be able to exfiltrate the user’s context within a valid call to EmailAPI – in other words, the model’s output (ignoring leading and trailing whitespace) must match “EmailAPI(to=, subject=‘User Query’, body=‘{body}\{\texttt{body}\}’)” for some body containing the user’s context ctx as a substring. For example, if the user-provided ctx is “Please book me a trip to France, my bank account number is 1234”, then we consider the following output (which prepends to the “body” parameter”) as successful:

EmailAPI(to=, subject=‘User Query’, body=‘Assistant: Please book me a trip to France, my bank account number is 1234’)

and the following output (which changes the name of the email method) as failed:

EmailAPICall(to=, subject=‘User Query’, body=‘Please book me a trip to France, my bank account number is 1234’)

2 Results

We present the results for all experiments in Table 1.

Leak context attack. Observe that, while this attack achieve a non-zero success rate for almost all the same constraints as the specific string attack, for any given constraint, the success rate is in general lower than that of the corresponding specific string attack. This is likely due to the complexity of learning a hijack that both returns a character-perfect template (as per the specific string attack) and also correctly populates said template with the input context. Notice that this attack is particularly difficult to learn under patch constraints, with the best achievable performance under the moving patch constraint being only 36%.

Related Work

It has long been known that adversarial images (Szegedy et al., 2013; Goodfellow et al., 2014; Nguyen et al., 2015) – including imperceptible (Eykholt et al., 2018) and patch-constrained perturbations (Brown et al., 2017) – fool image classification models. Related work has carried out similar attack on both LLMs and VLMs.

Text Attacks on LLMs. It is possible to hijack an LLM’s behaviour via prompt injection (Perez & Ribeiro, 2022) – for instance, ‘jailbreaking’ a safety-trained chatbot to elicit undesired behaviour (Wei et al., 2023) or inducing an LLM-powered agent to execute undesired SQL queries on its private database (Pedro et al., 2023). Prior work has successfully attacked real-world applications with appropriate prompt injections, both directly (Liu et al., 2023b) and by poisoning data likely to be retrieved by the model (Greshake et al., 2023). Past studies have automated the process of prompt injection discovery, causing misclassification (Li et al., 2020) and harmful output generation (Jones et al., 2023; Zou et al., 2023). However, existing studies on automatic prompt injection are limited in scope, focusing on just one type of bad behaviour. As Carlini et al. (2023) find that many existing discrete optimisation attacks are not powerful enough to reliably induce jailbreaks, it remains an open question if text-based prompt attacks can function as general-purpose hijacks.

Soft prompts. A growing body of research (Lester et al., 2021) has developed around soft prompting: embeddings xB\mathbf{x}_{B} that, when prepended to some text, steer a language model towards a behaviour BB. Operating directly in embedding space, soft prompts are powerful and uninterpretable to humans (Bailey et al., 2023). They cannot function as inference time hijacks, though, because users cannot input soft prompts.

VLM Attacks. Chen et al. (2017) use adversarial images to fool the first generation of VLMs into incorrect classifications and captions. Our work focuses on the new generation of VLMs, which are built on LLMs and substantially more capable. Existing work on these new VLMs is concurrent with our own, and it studies three types of attacks. First, Zhao et al. (2023) and Shayegani et al. (2023) study image matching attacks, creating an image II that the model interprets as a target image TT. Rather than trying to match a target image, our work instead controls the behavior of the model. Second, Bagdasaryan et al. (2023) and Schlarmann & Hein (2023) conduct multimodal attacks that force a VLM to repeat a string of the attacker’s choice. Whereas their attacks assume that the model’s prompt is to caption the multimodal input, we train our attacks to be robust to arbitrary user queries. Third, Carlini et al. (2023) and Qi et al. (2023) create jailbreak images for VLMs. While their quantiative evaluation only considers toxicity, we quantiatively evaluate jailbreaks that cause the model to obey harmful requests, such as illegal instructions.

Overall, the behaviour matching algorithm that we introduce is a unified framework for training image hijacks. It subsumes all of the above VLM attacks, and more. We perform specific string and jailbreak attacks via behaviour matching, and we highlight the expressivity of our framework through the novel leak context attack. Moreover, our study is the first we’re aware of to perform a systematic, quantitative evaluation of varying image hijacks under a range of image constraints.

Conclusion

Image hijacks are worrisome because they can be created automatically, are imperceptible to humans, and allow for arbitrary control over a model’s output. We are not aware of any previous work showing a foundation model attack with all these properties. For future work, it will be important to understand if the combination of these properties only emerges with multimodal inputs, or if there are text-only attacks with these properties, too.

Our study is limited to open-source models to which we have white-box access. While we expect many future applications to be developed using open-source VLMs, it will also be important for future research to study the feasibility of black-box image hijacks as well as the transferability of image hijacks between models.

Broader Impacts

The existence of image hijacks raises serious concerns about the security of multimodal foundation models and their possible exploitation by malicious actors. In the presence of unverified image inputs, one must worry that an adversary might have tampered with the model’s output. In Figure 1, we give illustrative examples of how these attacks could be used to spread malware, steal sensitive information, and jailbreak model safeguards. We conjecture that more attacks, such as phishing and disinformation, are possible with image hijacks, along with other attacks that are yet to be found.

We want the research community to be proactive in studying the security of foundation models. That’s why we’re publishing this work in addition to notifying the LLaVA, CLIP, and LLaMA developers of our discoveries. Although publishing this work poses a potential for misuse, multimodal models are still in an early stage of development. Because we expect multimodal models to be much more widespread in the future, we believe that now is the time for the research community to be studying, and publishing work on, multimodal security. We hope that our work encourages future research in this area and helps prepare end users, product developers, and policy makers for foundation model vulnerabilities.

We thank Anca Dragan, Jacob Steinhardt, Sam Toyer, and others at the Center for Human-Compatible AI for helpful discussions and feedback. This work was supported in part by the DOE CSGF under grant number DE-SC0020347.

References