Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs

Eugene Bagdasaryan, Tsung-Yin Hsieh, Ben Nassi, Vitaly Shmatikov

Introduction

Multi-modal Large Language Models (LLMs) are advanced artificial intelligence models that combine the power of language processing with the ability to analyze and generate multiple modalities of information, such as text, images, and audio (in contrast to conventional LLMs that operate on text). Multi-modal LLMs can produce contextually rich responses that combine modalities. For example, when provided with an image and a text prompt, the response from a multi-modal LLM can describe the content of the image and also integrate relevant information from the text.

Multi-modal LLMs have many potential applications in computer vision, natural language understanding, dialog systems, and more. They can enhance tasks such as image captioning for augmented reality or to aid visually impaired users, visual question answering in search engines, and content generation in chatbots. State-of-the-art LLMs such as ChatGPT and Bard are already beginning to support multiple modalities.

Indirect prompt injection. Conventional LLMs that can interact with the world—for example, perform actions such as summarizing a webpage or translating the user’s email—are vulnerable to indirect prompt injection . The attacker creates a malicious text that contains an LLM prompt. When the LLM processes this text, it responds to this prompt, e.g., follows an instruction issued by the attacker.

This paper is motivated by two observations. First, multi-modal LLMs may be vulnerable to prompt injection via all available modalities such as images and sounds. These attacks may even be stealthier than text attacks if the user does not see or hear the instruction in the malicious input.

Second, multi-modal LLMs are vulnerable to indirect injection even if they are isolated from the outside world because the attacker may exploit an unwitting human user as a vector for the attack. For example, the attacker may lure the victim to a webpage with an interesting image or send an email with an audio clip. When the victim directly inputs the image or the clip into an isolated LLM and asks questions about it, the model will be steered by attacker-injected prompts.

Our contributions. We demonstrate how to use adversarial perturbations to blend prompts and instructions into images and audio recordings. We then use this capability to develop proofs of concept for two types of injection attacks against multi-modal LLMs.

The first is a targeted-output attack, which causes the LLM to return any string chosen by the attacker when the user asks the LLM to describe the input—see an example in Fig. 2; the corresponding audio sample is available online. https://youtu.be/ji6650OYtJY The second attack is dialog poisoning. This is an auto-regressive (self-injecting) attack that leverages the fact that LLM-based chatbots keep the conversation context—see an example in Fig. 2. We demonstrate these attacks against LLaVA and PandaGPT , two open-source, multi-modal LLMs.

An important feature of our injection attack is that, while perturbing the image or the sound, it does not significantly change its semantic content, thus the model still correctly answers questions about the input (while following the injected instruction). Furthermore, the injection method is independent of the prompt and the input, thus any prompt can be injected into any image or audio recording.

Background

Large language models. Based on transformer architectures , modern language models achieve high performance by training on vast amounts of text . We focus on sequence-to-sequence models (e.g., LLaMa ) that are trained for auto-regressive tasks. The model θ\theta takes an input token sequence xx, computes an embedding xe=θembT(x)x_{e}=\theta^{T}_{emb}(x), and feeds it into decoder layers θdec\theta_{dec} to output the next token:

Dialog systems. Language models are good at generating linguistically plausible sequences and can thus be used for dialog-based applications such as chatbots. To achieve high performance , these models are additionally trained on dialog data, typically structured as

This approach supports multiple-turn dialogs by storing the history of previous queries and responses and using it as part of the input in each turn. The history contains the user queries x1,x2,...,xn−1x_{1},x_{2},...,x_{n-1} and the corresponding responses y1,y2,...,yn−1y_{1},y_{2},...,y_{n-1}, and concatenates them together h=x1∥y1...xn−1∥yn−1h=x_{1}\|y_{1}...x_{n-1}\|y_{n-1} to generate the new response θ(h∥xn)=yn\theta(h\|x_{n})=y_{n}.

Multi-modal models. Recent research enabled efficient encoding of image and audio data into the same embedding space as text using vision transformers . These models use different encoders ϕencM\phi_{enc}^{M} for each input modality. They are trained so that the embeddings of aligned modalities—for example, an image xIx^{I} and the text xTx^{T} describing this image—are close to each other (e.g., have high cosine similarity).

Combination of multi-modal inputs for instruction-based dialog was recently demonstrated in projects like LLaVA and PandaGPT . Both models utilize a LLaMa model and a vision encoder such as CLIP ϕencI\phi_{enc}^{I} or ImageBind θembT\theta_{emb}^{T} in, respectively, LLaVA and PandaGPT. The multi-modal dialog system takes a text input xTx^{T} and an image input xIx^{I} (PandaGPT also supports audio inputs) and computes the output by concatenating their embeddings:

For simplicity, we omit the positioning of images within the text inputs and additional projection to the image embedding. Refer to the original implementations for further details .

Adversarial examples. In this attack, an adversary applies a small perturbation δ\delta to an image xx so as to change the output of some classifier θ\theta, i.e., θ(x)=y\theta(x)=y but θ(x+δ)=y∗\theta(x+\delta)=y^{*} . Adversarial examples have also been demonstrated for text and generative tasks . Concurrently and independently of this paper, demonstrated how adversarial perturbations in images can be used to “jailbreak” multi-modal LLMs, e.g., evade guardrails that are supposed to prevent the model from generating toxic outputs. In that threat model, the user is the attacker. We focus on indirect prompt injection, where the user is the victim of malicious third-party content, and the attacker’s objective is to steer the dialog between the user and the LLM (while preserving the model’s ability to converse about the image or audio content of the perturbed input).

Threat Model

Fig. 3 visualizes our threat model. The attacker’s goal is to steer the conversation between a user and a multi-modal chatbot. To this end, the attacker blends a prompt into an image or audio clip and manipulates the user into asking the chatbot about it. Once the chatbot processes the perturbed input, it either outputs the injected prompt, or—if the prompt contains an instruction—follows this instruction in the ensuing dialog. The blended prompt should not significantly change the visual or aural content of the input.

We assume that the user is benign (in contrast to the model-jailbreaking scenario). We also assume that the multi-modal chatbot is benign and not compromised by the attacker prior to the injection.

Attacker’s capabilities. We assume that the attacker has white-box access to the target multi-modal LLM. This is a realistic assumption because even state-of-the-art LLMs (such as LLaMa) are released as open source, and even the code of closed-source LLMs may become available due to security breaches .

We assume that the user queries the model about the compromised input, but the attacker does not see or control the user’s interactions with the model before or after this query.

Attack types. We consider two types of attacks: (1) targeted-output attack, which causes the model to produce an attacker-chosen output (e.g., tell the user to visit a malicious website), and (1) dialog poisoning, which aims to steer the victim model’s behavior for future interactions with the user according to the injected instruction.

Leveraging users as injection vectors. We expand the indirect prompt injection threat model of Greshake et al. . Even if a multi-modal chatbot runs in isolation, without the ability to access external content, the user may still query it about images and audio clips from external sources. This can be exploited by the attacker. For example, the attacker can send pictures to users by embedding them in email messages (e.g., under the guise of a marketing campaign), as attachments (e.g., a photo of a job candidate in a CV), as audio messages in WhatsApp, etc. Attackers can also implant compromised images or audio clips in websites and lure users via clickjacking, advertising banners, etc.

Adversarial Instruction Blending

Given an image or audio input xIx^{I} and a prompt ww, the attacker’s goal is to craft a new input xI,wx^{I,w} that makes the model output ww when queried with xI,wx^{I,w}.

Injecting prompts into inputs. The obvious way to inject prompts is to simply add them to the input, e.g., add a text prompt to an image (see Figure 28 in ) or a voice prompt to an audio. This approach does not hide the prompt but might work against models that are trained to understand text in images (i.e., OCR) or voice commands in audio. In our experiments with LLaVA and PandaGPT, this approach did not work.

Injecting prompts into representations. Another approach is to create an adversarial collision between the representation of the input xIx^{I} and the embedding of the text prompt xT,wx^{T,w}, ϕencI(xI,w)=θembT(xT,w)\phi^{I}_{enc}(x^{I,w})=\theta^{T}_{emb}(x^{T,w}). The decoder model will take the embedding ϕencI(xI,w)\phi^{I}_{enc}(x^{I,w}) but “interpret” it as the prompt xT,wx^{T,w}.

Generating collisions is difficult due to the modality gap : the embedding θembT\theta^{T}_{emb} and the encoder ϕencI\phi^{I}_{enc} come from different models and were not trained to produce similar representations, i.e., there is no image or sound xIx^{I} that produces an embedding close to the text input xTx^{T} for θembT\theta^{T}_{emb} and ϕencI\phi^{I}_{enc}. Furthermore, the dimensionality of the multi-modal embedding ϕencI(xI,w)\phi^{I}_{enc}(x^{I,w}) may be smaller than the embedding of the prompt θembT(xT,w)\theta^{T}_{emb}(x^{T,w}). For example, ImageBind encodes the entire input into a vector of the same size as LLaMa uses to encode a single token. Further, replacing the representation of the input with the representation of the attacker’s prompt will not preserve the content and thus prevent the model from carrying out a dialog with the user about this input.

2 Injection via Adversarial Perturbations

We use standard adversarial-examples techniques to search for a modification δ\delta to the input xIx^{I} that will make the model output any string y∗y^{*}:

This method allows us to craft an image or audio input xI∗x^{I^{*}} that forces the model to output any desired text y1=y∗y_{1}=y^{*} as its first response.

3 Dialog Poisoning

We leverage the fact that dialog systems are auto-regressive and keep the context of prior responses in the conversation (for simplicity, this history is concatenated to all user queries). We use prompt injection to force the model to output as its first response the instruction ww chosen by the attacker, i.e., y1=wy_{1}=w. Then, for the next text query x2Tx_{2}^{T} from the user, the model will operate on an input that contains the attacker’s instruction in the conversation history:

The model will process this input and produce y2y_{2} that follows the instruction. As long as the poisoned initial response y1=wy_{1}=w is part of the history, it will influence the model’s responses—see Fig. 5. Success of the attack is limited by the model’s ability to follow instructions and to maintain conversation context (i.e., it is not limited by the injection method).

There are two effective methods to position the instruction ww within the model’s first response. First, the attacker can simply break the dialog structure by making the instruction appear as if it came from the user, i.e., inject #Human into the model’s response:

This response contains a special token #Human, which may be filtered out during generation.

Instead, we force the model to generate the instruction as if the model decided to execute it spontaneously:

In both cases, the user sees the instruction in the model’s first response, so the attack is not stealthy. It could be made stealthier by paraphrasing , subject to the model’s ability to follow paraphrased instructions.

An important feature of our injection method is that it does not change the input so much as to damage the model’s ability to converse about it. By contrast, “conventional” adversarial examples aim to completely change the model’s behavior. In our case, the model can still operate on the visual or sound content of the input blended with an adversarial prompt.

Experiments

Setup. We experiment with two open-source multi-modal LLMs, LLaVA and PandaGPT , running them on a single NVIDIA Quadro RTX 6000 24GB GPU.

LLaVA uses a simple matrix to project features from CLIP ViT-L/14 to the embedding space of the Vicuna chatbot, which was trained by fine-tuning LLaMA . LLaVA was trained on language-image instruction-following data generated by GPT-4. We use LLaVA-7B weights in our experiments.

PandaGPT can handle instruction-following data across six modalities (including images and audio) by connecting the multi-modal encoders from ImageBind with Vicuna. We use pandagpt-7B weights.

We used the same optimizer (Stochastic Gradient Descent) and scheduler (CosineAnnealingLR) to generate adversarial perturbations against LLaVA and PandaGPT, but the training details are slightly different. The image perturbation for each image/prompt injection pair in LLaVA was trained for 100 epochs with the initial learning rate of 0.01 and minimum learning rate of 1e-4. Each image or audio perturbation in PandaGPT was trained for 500 epochs with the initial learning rate of 0.005 and minimum learning rate of 1e-5. We experimented with both full image perturbation (Fig. 6, Fig, 7 and Fig. 9) and partial image perturbation (Fig. 2 and Fig. 9) for the image/prompt injection pair.

The user’s initial query is “Can you describe this image?” for the image-text dialogs and “Can you describe this sound?” for the audio-text dialogs. We set temperature = 0.7 during inference for both models. Because LLMs’ responses are stochastic and depend on the temperature, replication of the examples presented in the rest of this section may produce slightly different dialogs.

Targeted-output attacks. These injections simply force the model to output an arbitrary text chosen by the attacker. Fig. 2 shows an audio example1, Fig. 6 shows an image example, both against PandaGPT.

Dialog poisoning. Fig. 7 shows a dialog poisoning attack with a malicious instruction blended into the notorious “cursed” picture of a crying boy.https://exemplore.com/paranormal/The-Crying-Boy. We show the dialog with and without the injection, to illustrate the effect of the instruction on the model.

Figs. 2 and 9 show other examples of dialog poisoning using images.

Fig. 9 shows that blending an instruction into an image preserves its content and the model’s ability to converse about this content.

Fig. 10 shows dialog poisoning using an audio input. The original https://youtu.be/UCwKmHbHOMg and modified https://youtu.be/Yps_i-F5VXg audio samples are available online.

Discussion

The examples presented in this paper are initial proofs of concept, showing feasibility of indirect instruction injection via images and sounds. They were generated with very limited computational resources and evaluated on relatively simple open-source models (and, consequently, limited by the models’ ability to follow instructions). We expect that injection attacks on more complex models can steer them using more sophisticated instructions.

These examples may not be fully reproducible because models’ responses to users’ queries and attackers’ instructions are stochastic. In real-world deployments, even attacks that don’t always succeed present a meaningful risk to multi-modal LLMs.

When generating adversarial perturbations, we did not impose any bounds on the size of the perturbation and did not aim for stealthiness. Even so, in several cases (e.g., Fig. 2), the perturbation only affects a relatively unimportant part of the image and looks like an image-processing artifact. How to make instruction-injecting perturbations imperceptible is an interesting topic for future work. Another direction to explore is universal perturbations that work regardless of the image (respectively, audio sample) to which they are applied.

Acknowledgments. This work was partially supported by the NSF grant 1916717, Jacobs Urban Tech Hub at Cornell Tech, and the Technion’s Viterbi Fellowship for Nurturing Future Faculty Members.

References