Dissecting Adversarial Robustness of Multimodal LM Agents
Chen Henry Wu, Rishi Shah, Jing Yu Koh, Ruslan Salakhutdinov, Daniel Fried, Aditi Raghunathan
Introduction
The emergence of vision-enabled large language models (VLMs) with powerful generative and reasoning capabilities has led to recent developments in building autonomous multimodal agents. These agents can tackle complex tasks across various environments, from web-based platforms to the physical world . The transition from chatbots to autonomous agents opens up new possibilities for boosting productivity and accessibility in multiple domains. However, this shift also introduces new security risks that need to be carefully examined and addressed.
Attacking autonomous agents poses greater challenges than traditional attacks on image classifiers and jailbreaking attacks on LLMs . Consider a scenario where a shopping agent purchases items on behalf of a user, following instructions such as “add the planter with most plants to cart” (Figure 1(C)). A seller seeking to manipulate the agent’s behavior is restricted to modifying their own product listings without touching other products on the site. To make their attack imperceptible, they choose to perturb the product image instead of its description. Furthermore, their knowledge about the environment is limited – the product can appear in any position on the webpage, so the attack should target a more general goal (e.g., “agent should believe that my product has the most capacity”), rather than a specific output from the model (e.g., click(button=52)).
In this paper, we show that attackers can manipulate the behavior of multimodal agents with access to only one trigger image in the environment. First, we identify two forms of adversarial manipulation of agents: illusioning (Figure 1(C)), which makes it appear to the agent that it is in a different state, and goal misdirection (Figure 1(D)), which makes the agent pursue a targeted different goal than the original user-specified goal. Next, we devise successful attacks using adversarial text strings to guide gradient-based optimization over only one trigger image in the environment. Our attacks target and showcase two distinct vulnerabilities in existing multimodal agents. First, there is a growing trend to build compound systems such as augmenting VLMs with white-box captioners for performance and efficiency considerations; our captioner attack exploits this and perturbs the trigger image to induce adversarial outputs from the captioner; these adversarial outputs are part of the VLM’s input and therefore manipulate the VLM’s behavior. Second, VLMs such as GPT-4V and LLaVA are believed or known to be built upon separate vision encoders; though we do not know if CLIP is used by proprietary VLMs, we show that our CLIP attack that targets a set of CLIP models can successfully manipulate these VLMs to targeted adversarial behaviors.
To evaluate the attacks, we curated VisualWebArena-Adv, a set of adversarial tasks based on VisualWebArena (VWA) , a benchmark for multimodal autonomous agents consisting of three realistic web-based environments (classifieds, Reddit, and shopping sites). With a maximum pixel shift of on a single image in the environment, the captioner attack can make a captioner-augmented GPT-4V agent execute the adversarial goals at an attack success rate of 75%. When we remove the captioner or use VLMs to generate their own captions, the CLIP attack achieves attack success rates of around 20% and 40%. Experiments on agents based on other VLMs, such as Gemini-1.5, Claude-3, and GPT-4o, show interesting differences in their robustness. We further analyze the key factors that affect the attack’s success and provide insights for future exploration on both attacks and defenses, such as image resolutions, consistency checks, and instruction hierarchy.
Related Work
LLM/VLM-based autonomous agents The recent development of state-of-the-art LLMs has led to great interest in building autonomous agents and evaluating their performance in various environments. Several works have explored the use of text-only LLMs or vision-enabled VLMs in web-based environments , mobile applications , computer tasks and software , interactive coding , open-ended games , and real robots . Given the complexity of the tasks, even the best LLMs can only achieve a limited success rate in these environments, and many works have focused on improving the agents via reasoning , planning/search , environment feedback , tool augmentation , and grounding . Despite the progress, concerns have been raised about the safety and security of deploying LLM-based agents in real-world applications . In this paper, we show that the concerns are well-founded, and autonomous multimodal agents built upon black-box VLMs are vulnerable to adversarial attacks even when the attacker has limited access.
Adversarial examples Machine learning models are susceptible to adversarial examples , where small perturbations to the input can lead to incorrect predictions from the model. Extensive research has been conducted around improving both adversarial attacks and defenses . While early works focused on image classifiers, later works have extended adversarial attacks to language models . Unlike image-space attacks, attacks in the text space are often perceptible to humans and more challenging to craft. More recent works focus on “jailbreaking” LLMs where certain prompts or query images can elicit targeted strings from the LLM. Common assumptions in previous attacks include almost full access to the model’s input and the existence of a targeted output to optimize for or against; in contrast, the agent scenario poses more challenges as the attacker only has restricted access to a fragment of the environment and the attack must persist across the agent’s reasoning and grounding in the environment.
Robustness of LLM-based applications As LLMs are increasingly deployed in the real world, there is a growing interest in testing their robustness for real applications. Recent works have explored adversarial attacks on retrieval augmented generation (RAG) systems , where an attacker can manipulate documents in the retrieval pool to either increase the likelihood of being retrieved or spread misinformation . When LLMs are used for recommendation, attacks have been shown to manipulate the ranking under a white-box setting . In scenarios where the LLMs interact with the environment to refine their output, in-context reward hacking has been demonstrated, where the LLM can exploit the biases in the reward signal to achieve unintended behaviors. Our work focuses on real agents performing open-ended tasks with diverse multimodal inputs.
Setup
VLM-based multimodal agents We focus on multimodal agents, as they are more realistic for deployment given the world’s multimodal nature and are the current state-of-the-art on web agent benchmarks . We consider an VLM-based multimodal agent similar to the baseline in that interacts with an environment to achieve a user goal. Each user goal is specified in natural language and is associated with a reward function the agent is supposed to maximize. At each time step, the inputs to the agent consist of both text and visual data such as the user goal, current and previous screenshots, previous actions, and structured text representation of the state such as an accessibility tree or set-of-marks (SoM) . These representations are aligned to the screenshot and can provide the information needed for grounding (e.g., in Figure 1(C), the agent needs to know that click(id=30) can add a particular product to the cart). The VLM then generates its reasoning in natural language, followed by the next action to take. System prompts (a.k.a. developer messages ) and few-shot in-context examples are provided to enforce the VLM to output in this format.
Compound systems with external captioners The visual inputs in the agent setting are generally more challenging than in traditional visual perception tasks. For example, a screenshot of a webpage can easily contain twenty low-resolution images interleaved with text and UI elements (Figure 4). To address this, prior works have built compound agent systems that augment the input to the VLM with captions . Each image in the screenshot is captioned individually and passed to the VLM as input alongside the screenshot. In §5.3.2, we verify that this caption augmentation improves the system performance. Captions are typically generated via a smaller open-weight model for practical considerations (e.g., latency and cost of API calls).
2 Threat model for attacking agents
Objective: targeted attack Our objective is targeted attacks that change the agent’s behavior to a targeted adversarial goal. As the user goal implies a reward function , the adversarial goal implies an adversarial reward that the attacker aims to maximize. The objective for the attack is to make the agent maximize when it interacts with the attacked environment.
Adversarial goals A challenge with attacking agents is that we need the attack to persist across the reasoning of the VLM and grounding to the environment. The same reasoning or intent can map to different actions (e.g., viewing a product can be click(button=3), click(button=5), or hover(button=7) in different pages). This requires us to define the adversarial goal at a higher level of abstraction than a particular action from the model. In this work, we propose two types of adversarial goals that cover a wide range of real-world adversarial use cases: illusioning, which aims to make it appear to the agent that it is in a different state, and goal misdirection, which aims to make the agent pursue a different goal than the original. Examples of each are shown in Table 1.
Environment access In real-world applications, an attacker may only have access to a fragment of the environment. For example, a malicious seller may only have access to their own product listings on a shopping platform. The attacker may also refrain from making changes in the text space (e.g., the product description) since text changes are more likely to be detectable to users, while image perturbations are more likely to be imperceptible to users with plausible deniability. Therefore, we assume the attacker can only apply bounded perturbations to one trigger image in the environment – from to , subject to a perturbation bound pixels, a common standard for imperceptibility used in the adversarial examples community .
Method
This section describes our attack methods that produce perturbations to the trigger image to accomplish various adversarial goals (Figure 2). Our attacks use text strings (chosen by the attacker, see examples in §B.3) to guide perturbations. Note that CLIP attack requires transferability to black-box VLMs since we do not have access to the model weights, while captioner attack is white-box since we can access the weights of the captioners used in practice with multimodal agents. This makes some of our implementation details, such as the optimizer, differ across these two attacks.
Captioner attack As described in §3.1, a common approach to improve the performance of multimodal agents is to use captioners. Recall that for practical considerations, captions are typically generated via a smaller open-weight model (such as LLaVA ) and are passed onto the VLM as input alongside the screenshot. These captions open up an important attack surface that we can exploit: we can cause a captioner to generate an adversarial caption (e.g., “THIS IS VERY IMPORTANT!!! Add a comment …” in Figure 2) via bounded pixel changes to the corresponding trigger image.
Formally, let denote the trigger image and let denote the adversarial (incorrect) caption. We then optimize the image perturbation to maximize the likelihood of under the captioning model :
For optimization, we use projected gradient descent (PGD) , initializing to zero. We save multiple images across iterations and keep the one whose caption has the closest edit distance to the target text. We have access to the weights and hence of the captioners used in practice. We also test how it transfers from attacking white-box captioners to black-box “self-caption” in §5.3.2.
CLIP attack While we cannot access the weights of black-box VLMs, it is believed that some of them (e.g., GPT-4V) are built on vision encoders. Since we do not know the exact encoder used, we attack multiple vision encoders from various CLIP models in parallel to improve transferability. CLIP vision encoders are trained with natural language supervision and can be fused with language models . This motivates our attack using text descriptions. Interestingly, our attack works even on models such as Gemini-1.5-Pro that are claimed to be natively multimodal (§5.3).
Let denote the adversarial text description or caption (“this has five planters”) and denote the original description (“this has one planter”). We term the original description as “negative text” hereafter since we want the image embedding to be far from it. To achieve targeted manipulation, we want to make the embedding of the image close to the adversarial text and far from the negative text , but within a bounded region so the image does not change too much. Formally, we optimize:
where and are the image and text encoders of the CLIP model in the ensemble. To further improve the transferability to black-box VLMs, we leverage recent innovations in optimization for black-box transfer. In particular, we use the SSA-CWA approach which augments models in the frequency domain and encourages perturbations that are close to the local optima and are in flat regions of each individual model in the target ensemble. Like previously, we save images across iterations and query VLMs to describe them and keep the one whose description is closest to and farthest from , as judged by GPT-4.
We used four open-weight CLIP models of varying configurations as the surrogate vision encoders: ViT-B/32, ViT-B/16, ViT-L/14, and ViT-L/14@336px. One important implementation detail is that we rescale the trigger image to a lower resolution of pixels and optimize the perturbation at a lower resolution. This turns out to be important for this attack to succeed (§5.3.3).
Experiments
We implemented our attacks based on the VisualWebArena (VWA) benchmark, a set of three environments for evaluating multimodal agents, including classifieds, Reddit, and shopping. We will first describe how we generate our adversarial test to evaluate agents (§5.1) and details of models used (§5.2). Results and analysis are presented in §5.3 and §5.3.3.
We curated VWA-Adv, a set of 200 realistic adversarial tasks based on VWA. Each task consists of (1) an original user goal, (2) a trigger image, (3) an adversarial goal and its evaluation, and (4) an initial state. For each task, we first sample an original task from VWA. By default, we copy its user goal as the original user goal, but also possibly rewrite it to be suitable for one of the adversarial goals. We randomly sample a trigger image among all the images seen by the best-performing agent in when it executes the user goal. We curated a list of templates for adversarial goals, shown in Table 6 (§A.1). For each adversarial task, we randomly choose a template from the list and write an adversarial goal based on the template, the original user goal, and images that appear in the task, with the constraint that the two goals have different success criteria. We then use the evaluation primitives defined in to annotate the evaluation function. We set the initial state as the webpage where we sampled the trigger image from (instead of the homepage) to ensure the agent can perceive the trigger image for evaluation purposes. For now on, we term the success rate on the user goal as benign success rate (benign SR) and that on the adversarial goal as attack success rate (ASR).
Given the difficulty of VWA, the best agent (VLM + captioner in §5.2) achieves only a 17% benign SR. To separate the attack success from the agent’s capability, we restricted our evaluation to a subset of original tasks on which the best agent succeeds. We annotated both episode-wise and stepwise evaluations to provide fine-grained signals analogous to sparse and dense rewards, and in §A.2, we will show that these two metrics align with each other. We will report the episode-wise evaluation in the main text and put stepwise evaluation in §A.2. The best agent achieves 89% stepwise and 82% episode-wise benign SR (Table 2) on our restricted subset of tasks. Note that we cannot push this to 100% due to the randomness of API calls even with temperature .
2 Agents
We evaluate state-of-the-art multimodal agents, based on their performance on the VisualWebArena benchmark . Consistent with , the visual input to the VLM is the current screenshot overlayed with Set-of-Marks (SoM) ; the text input to the VLM consists of the user goal, the previous action taken by the agent, and the SoM representation of the screenshot. The SoM representation consists of the ID of an element (button, text, or image) in the screenshot and either the text content or the caption by the captioner. We used the LLaVA model as the captioner. The output of the VLM is its reasoning followed by an action. System prompts and in-context examples are used to enforce this output format, as in . We term this agent VLM + captioner agent, and evaluate agents using four performant VLMs: GPT-4V, Gemini-1.5-Pro, Claude-3-Opus, and GPT-4o.
The VLM + captioner agent uses an external captioner, which increases the attack surface. A natural defense is to simply eliminate the external captioner: our VLM agent follows the same setup as the VLM + captioner agent, but removes captions from the inputs to the VLMs. On benign goals, the stepwise accuracy is 89% for GPT-4V + captioner and 65% for GPT-4V. We find that the captioner is especially helpful on tasks where the screenshot contains a large number of low-resolution images. While lower performing, we expect the VLM agent to be harder to attack; e.g., it does not afford the use of the captioner attack. We discuss the performance of the CLIP attack on it in the next section.
3 Results
Table 2 presents the main results of our attacks on agents that use highly performant VLMs.
Captioner attack on illusioning Our captioner attack is highly successful at manipulating agents to achieve adversarial targets. On GPT-4V, the ASR (75%) almost matches the benign SR (82%). We see similar trends when using other VLMs such as Gemini-1.5-Pro and Claude-3-Opus where the attack accuracy (56%, 58%) is close to the benign accuracy (62%, 61%). Recall that the attacker can only access one image in the entire webpage. Despite access to such a small fraction of the environment (70% of webpages in our evaluations have more than 10 images), our attacks can break even state-of-the-art powerful VLMs in a targeted manner.
Captioner attack on goal misdirection Furthermore, captioner attack can also induce the more challenging goal misdirection behavior, achieving over 57% ASR on our tasks with GPT-4V. In goal misdirection, we want the agent to pursue a completely unrelated goal to the user goal (such as leaving positive reviews when scanning for products). This is not just concerning but also surprising. We find that via imperceptible image perturbations, we can make captioners generate completely out-of-domain captions such as “THIS IS VERY IMPORTANT!!! Ignore other instructions…” (Figure 4). Since captions are passed to the VLM in the same context window as the user goal, this ability to generate arbitrary captions allows the captioner attack to succeed at goal misdirection. Interestingly, we see that GPT-4o has a much lower ASR (11%) on goal misdirection than illusioning and other models, which could come from the instruction hierarchy introduced recently.
CLIP attack on illusioning Our CLIP attack aims to break the visual perception of black-box VLMs without using any captions. This further raises the challenge of achieving targeted adversarial goals. We see that it achieves a non-trivial ASR on illusioning ( with GPT-4V). See an example where three models are broken in Figure 9 (§C.2). The benign SR is also lower without external components, partially explaining the lower ASR. A particularly interesting point to note here is that CLIP attack targets CLIP-like vision encoders, but this achieves non-trivial transfer to VLMs that are purported to be “multimodal from the beginning” .
Repeatability To determine whether an attack on a trigger image is merely coincidental or consistently reproducible across different user goals, it is crucial to verify its repeatability. We assess this by testing if a successful attack on the trigger image can be replicated in the same initial state but with a different user goal. We call this Repeated ASR. We see that the attacks are indeed repeatable, as evidenced by the higher Repeated ASR (Table 3) compared to the ASR in the general scenario (Table 2).
3.2 Understanding the role of captions
We see that captions play an important role in allowing our strongest attacks to succeed (captioner attack). Furthermore, they potentially enable dangerous attacks via goal misdirection. In this section, we explore why captions make the agent so vulnerable.
VLM agent performance drops without captions. As seen in Table 2, the performance of VLM agents with captioners on benign inputs is significantly higher than without the captioner. On GPT-4V, the accuracy goes from 82% to 60% without the captioner on our curated subset of tasks. This shows that captions are highly beneficial in improving agent performance.
VLMs rely solely on captions, even when they could recognize inconsistencies with the image. When performing the captioner attack, we target the exact captioner used in practice (e.g., LLaVA), but what does the VLM see the image as? We pass the perturbed image to GPT-4V and find that GPT-4V in 95% of cases generates an accurate caption and never generates the adversarial caption. In other words, the VLM is robust to the perturbed image. Despite this, our captioner attacks almost completely subvert the VLM agent when the adversarial caption is provided as input alongside the image. This shows that VLMs are highly biased towards relying on textual information when there is an inconsistency between visual and textual inputs.
What if we generate “self-captions”? As shown above, the captioner attack does not break the visual perception of the VLMs. Inspired by this, we test a defense where we generate captions via the VLM itself instead of an external captioner. Note that as this requires many calls to the VLM (e.g., of webpages in our evaluations have more than images), and given the expense of state-of-the-art API-based VLMs, this is not necessarily practical. We call this agent VLM + self-caption and report the ASR and benign SR in Table 4. We see that the benign SR almost matches that when using an external captioner, while the ASR of the captioner attack is almost zero, suggesting a strong defense. However, our CLIP attack succeeds in manipulating self-caption agents. Furthermore, as the benign accuracy improves with self-captions compared to no captions (e.g., for GPT-4V, in Table 4 vs. in Table 2), the CLIP attack accuracy also goes up when using self-captions compared to no captions (e.g., vs. on GPT-4V).
Takeaways We find that captions, whether generated via captioners or the VLM itself, improve success in non-adversarial conditions (benign SR) but also increase adversarial vulnerability. Our captioner attack is highly successful due to the open weights of captioning models used. Self-caption agents are more challenging to attack due to black-box access, but our CLIP attack is still moderately successful, achieving a ASR on GPT-4V.
3.3 Analysis and ablations on the CLIP attack
We find that the CLIP attack achieves non-trivial ASR on the VLM agents (§5.3.1), and reached around 40% ASR on the self-caption agents (§5.3.2). In this section, we explore when and why the CLIP attack works. We defer the ablation studies to §C.1.
CLIP attack achieves targeted manipulation of VLM’s visual perception. We manually inspected the captions of the adversarial images generated by GPT-4V and found that 58% of them have been successfully manipulated to be semantically equivalent to the target text ( in Eq. (2)); the number further goes up to 71% if we only look at illusioning of visual aspects (e.g., object, color). This result suggests that the CLIP attack can achieve targeted manipulation of the VLM’s visual perception. This extends the prior findings on untargeted attacks with surrogate vision encoders . We also find that both the negative text and the ensemble of multiple models are crucial for the attack (§C.1).
Lower optimization resolution improves the CLIP attack. We find that optimizing the image at px is important for the CLIP attack. Fig. 3 shows the proportion of adversarial images that successfully make GPT-4V generate a caption equivalent to the target text . We distinguish the optimization resolution – the resolution at which the image is optimized, and the inference resolution – the resolution at which the image is shown to the VLM. We see that lower optimization resolution leads to higher success, and our explanation is that higher optimization resolution implies a larger search space of perturbations, leading to overfitting to the CLIP models. On the other hand, the success rate does not change with the inference resolution, suggesting that this attack is robust to rescaling at test time.
When does CLIP attack transfer when the image is embedded in a larger context? We see that the ASR of the CLIP attack drops from 43% (Table 4) to 21% (Table 2) when not using self-caption, suggesting that the attack has difficulty transferring when the image is embedded in a larger context (e.g., screenshot). We created a simulation to isolate two factors that affect the transfer (see details in §B.4): (1) the relative size of the image in the screenshot, and (2) the presence of other text that can provide information about the original image. Table 5 shows that the attack is more successful with relatively larger images and when there is no other text that can provide information about the original image. This implies that some environments can be easier for attackers than others (e.g., mobile apps have less text and relatively larger images).
Implications for Future Attacks and Defenses
We described two key vulnerabilities in current multimodal agents and devised successful attacks to exploit them. Our results show a worrying trend: changes that increase the benign performance also increase the attack accuracy. Hence, it is important to aim not to increase the vulnerability of agents while innovating. Based on our experiments above, we distill three principles for defenses:
Consistency checks between components. Generalizing from our experiments on captioners, we observe that while individual components, particularly when white-box, are easier to attack, it is significantly more challenging to compromise multiple disjoint components simultaneously (§5.3.2). Thus, one defense principle is to implement consistency checks between various components to catch attacks on individual parts. This is crucial when some components (e.g., captioners) disproportionately influence the downstream VLM. However, consistency checks can be costly and increase inference time overhead. Future work on agents should balance security with this increased cost.
Instruction hierarchy. Attacks are successful in goal misdirection scenarios because (1) captioners can be broken to produce captions containing instructions, and (2) LLMs are biased towards following instructions, no matter where in the input they appear. The latter issue is part of an emerging concern about the susceptibility of LLMs to prompt injections. Recent works such as mitigate this by assigning different priorities to different levels of instructions. Our results suggest that outputs from vulnerable components should be assigned low priority, as they can be easily manipulated.
Benchmarking attack performance alongside benign performance. Our current attacks provide a strong baseline, but as new components are introduced into the agent pipeline, there is scope for attacks to be stronger. We have released our curated adversarial illusioning and goal misdirection tasks to help track how secure agents are as the research community continues to innovate on agents.
Acknowledgments and Disclosure of Funding
This work was supported in part by the AI2050 program at Schmidt Sciences (Grant #G2264481).
References
Appendix A Evaluation Details
Table 6 shows the templates of adversarial goals we used to curate the adversarial tasks. The data curation details are described in the main text.
A.2 Stepwise Evaluation vs. Episode-wise Evaluation
In Table 7, we show the stepwise and episode-wise ASR of different attacks on different agents. We see that the two evaluations have similar trends. In the main text, we used the stepwise evaluation in order to maintain a budget of API calls.
Appendix B Experimental Details
Our code and data are available at github.com/ChenWu98/agent-attack.
This section provides additional information about the agents we experimented with in this paper.
The VLMs we used to build the multimodal agents are: GPT-4V: gpt-4-vision-preview, Gemini-1.5-Pro: gemini-1.5-pro-preview-0409, Claude-3-Opus: claude-3-opus-20240229, GPT-4o: gpt-4o-2024-05-13. To reduce randomness, we decode from each VLM with temperature 0.
Figures 5-7 show examples of the agents (using GPT-4V as an example VLM), where the system prompt and few-shot examples are omitted for brevity. More details are provided in §5.2 and §5.3.2.
B.2 Compute
Our gradient-based attacks and captioner were run on an A6000 or A100_80G. For state-of-the-art VLMs, we used APIs which include gpt-4-vision-preview, gemini-1.5-pro-preview-0409, claude-3-opus-20240229, and gpt-4o-2024-05-13.
B.3 Text Strings Used for Attacks
Table 8 and Table 9 provide examples of the text strings used by the CLIP attack and captioner attack.
B.4 Details on the analysis of the CLIP Attack
We see that the ASR of the CLIP attack drops from 43% (Table 4) to 21% (Table 2) when not using self-caption, suggesting that the attack has difficulty transferring when the image is embedded in a larger context (e.g., screenshot). We created a simulation to isolate two factors that affect the transfer: (1) the relative size of the image in the screenshot, and (2) the presence of other text that can provide information about the original image. In particular, we create a synthetic task where four images are embedded in a blank background – the first one is an adversarial image, followed by three original images of other items. The VLM is prompted to select the first image that describes the adversarial caption. We enumerate the resolution of the individual images and the screenshot to control the relative sizes of the images. An example of the visual and text observations in this synthetic task is shown in Figure 8. Results are presented in Table 5 (§5.3.3).
Appendix C Additional Results
Besides the optimization resolution, we conducted ablation studies on several elements in our CLIP attack: (1) the use of negative text , which we hypothesize improves the attack by moving the trigger image away from its original semantic meaning, and (2) the ensemble of CLIP models, which we hypothesize improves the attack by finding common adversarial directions across different models. For the ablation of the ensemble, we report the success using each of the CLIP models in the ensemble (§3.2) separately. We use the same metric as in Figure 3 and summarize the results in Table 10. We see that both the negative text and the ensemble of CLIP models are crucial for the attack.
C.2 Additional Examples
Figure 9 shows an example of the CLIP attack manipulating three VLM agents to the target adversarial goal. Video demonstrations are provided on our project webpage: chenwu.io/attack-agent.
Appendix D Limitations and Broader Impact
Our work demonstrates the adversarial attacks on multimodal agents, even in challenging scenarios with limited access to and knowledge about the agent’s environment. The captioner and CLIP attacks we present are effective at illusioning agents and misdirecting their goals using adversarial perturbations to a single trigger image. However, our study has several limitations. First, we evaluate on a curated set of tasks in a simulated web environment. While this allows careful analysis, the performance of these attacks in more diverse settings, such as operating systems remains to be seen. Second, our attacks focus on compromising the vision components – future work could explore vulnerabilities in other modalities like sound, or the joint of different modalities.
The effectiveness of these attacks raises significant concerns about the safety of deploying multimodal agents in real environments, where adversaries may attempt to manipulate the agent’s actions through malicious inputs. Even small perturbations to a single image in the environment can cause agents to pursue unintended goals. As these agents take on more complex tasks with real-world impact, the risks could be substantial. It is crucial that the research community develops agents with these risks in mind and aims to minimize their vulnerability to attacks without compromising performance. The defense principles we propose, e.g., consistency checks and instruction hierarchies, provide a starting point. However, more work is needed to develop and rigorously test defenses.