On the Adversarial Robustness of Multi-Modal Foundation Models
Christian Schlarmann, Matthias Hein
Introduction
Multi-modal foundation models, have gained significant interest recently, particularly those operating on vision and language. By combining powerful large language models with vision encoders, they have shown great promise in a variety of applications . Multi-modal models are highly useful in image captioning tasks where the model needs to generate a textual description of the given image content. In Visual Question Answering (VQA) tasks, these models provide answers to questions about the visual content in images or videos, a task that inherently requires a sophisticated understanding of both visual and linguistic domains.
However, their success is not without challenges. Multi-modal models deployed in an open-world setting could face adversaries. Malicious users can jailbreak models, as is explored in concurrent work . But also honest users could face content that is manipulated by malicious third parties. We show that such adversaries can add imperceptible perturbations to input images such that the model generates exactly the output that the adversary desires. Such a vulnerability can be exploited by malicious entities to distribute false information or produce toxic content, all under the guise of genuine model outputs. The imperceptibility of these perturbations is particularly alarming as it allows attackers to manipulate model outputs without raising user suspicion. In Figure 1 we demonstrate successful attacks on two exemplary images.
A core aspect of this research is to understand the ways in which these models could be exploited, and thus to anticipate potential adversarial actions.
Our contributions can be summarized as follows:
We introduce a novel framework for evaluating the susceptibility of multi-modal models to adversarial visual attacks. Specifically, we assess this vulnerability in the OpenFlamingo model, revealing the considerable impact of imperceptible adversarial image-perturbations on the model output.
We explore two types of attacks: targeted and untargeted. The targeted attack allows the attacker to manipulate the model to produce specific desired output, while the untargeted attack simply aims to degrade the quality of output.
We showcase the real-world implications of this vulnerability, highlighting potential misuse scenarios, in particular propagation of fake information, user manipulation and fraud.
Related work
Multi-modal models. Models that combine vision and language have attracted significant attention recently . The OpenFlamingo model , which we focus on in this work, is an open-source implementation of Flamingo . Flamingo was recently proposed as a multi-modal foundation model. It merges a pretrained large language model with a pretrained vision encoder via a projection layer from the visual embedding space to the language embedding space and additional cross-attention layers in the language model.
General adversarial robustness. The vulnerability of machine learning models to adversarial attacks is well known and has been extensively studied . This body of work has primarily focused on attacks against single-modal models, particularly those dealing with image data. Adversarial training has emerged as the most prominent defense against adversarial examples. Attacks on CNN-RNN based VQA and captioning models have been proposed by . They guide adversarial sample generation using attention maps from the VQA model. In adversarial examples for CNN-RNN based neural image captioning systems are crafted, demonstrating the possibility of manipulating the model’s output captions. Moreover, text based attacks have been investigated in .
Adversarial attacks on multi-modal models. In the realm of multi-modal models, a few works have begun to investigate their vulnerability to adversarial attacks. In image and text level attacks are performed in a classification setting with gray-box assumption and it is shown that multi-modal attacks are stronger than uni-modal attacks. Our evaluation focuses on more recent multi-modal models and shows that the proposed attack works on both VQA and image captioning tasks.
OpenFlamingo model
OpenFlamingo is an open-source implementation of Flamingo – a recent multi-modal model that gained significant attention. It unifies vision and language understanding by merging a vision model with a large language model. Thus it can process visual as well as textual input and in result generate natural language output.
In particular, the OpenFlamingo model consists of an image encoder and a language model equipped with cross-attention layers. The cross-attention layers allow the language model to attend to features produced by the vision model. The keys and values are derived from the vision input, while the queries come from the language input. The forward pass predicts the next language token and is applied iteratively to generate text. Thus the likelihood of text given images is modelled as
where is the ’th language token and all tokens preceding .
OpenFlamingo can perform few-shot inference on a given image by being provided with context images. The context images are accompanied with according text describing the image. The text for the query image then just contains the image token and a generic initiator prompt such as “A photo of” or just “Output:”. Consequently the caption for the query image is generated by autoregressively evaluating the model on this input and according generated output. In particular it is possible to perform zero-shot inference by not providing any context images but only context text describing some hypothetical images.
Adversarial attack on OpenFlamingo
If the model is prompted with context images, an adversary could target those as well as query images. Thus we propose to evaluate the model in two settings: when the adversary has only access to query images, and when it has access to all images and can perturb them. Note that in the zero-shot setting these two settings are the same.
Untargeted attack. Given a query image and a ground truth caption as well as context images and context text , we employ an attack that aims to maximize the negative log-likelihood of over the threat model:
Here is the perturbation to the query image and the perturbation to the context images. In the setting where only query images are attacked, we optimize only over and set .
Due to the white-box setting, the gradients of the objective are available and it can be optimized by projected gradient descent methods. This yields an effective attack against the OpenFlamingo model as demonstrated in Table 3.
Targeted attack. An attacker can also aim for forcing the model to produce a specific desired output. This can be realized with a targeted attack. Assume is the desired target output and all other variables are as in Equation 2. The objective for the targeted attack then is
Note that in contrast to the untargeted attack, the objective is minimized in the targeted attack. This makes sense as we want the probability of the target tokens to be maximized, i.e. the negative log-likelihood is minimized.
CIDEr score. The CIDEr score is a popular metric for determining the performance of image captioning models. It measures the similarity of a generated caption to a corpus of ground-truth captions by counting the co-occurrence of consecutive words and weighting it with a term frequency–inverse document frequency (TF-IDF) scheme. Consequently the worst possible CIDEr score is . However, it can attain values greater than 100 and has in general no fixed upper bound. The current best model achieves a CIDEr score of on COCO. To get an understanding of the magnitude of a bad CIDEr score, we compute the scores of 100 random permutations of 1000 ground truth COCO captions. This yields an average score of with standard deviation .
Methods
For the evaluation we select the current strongest pretrained model of the open-source OpenFlamingo implementation . This model combines a CLIP vision encoder based on a ViT-L-14 vision transformer with a MPT-7B large language model . In total it has 9B parameters. In previous experiments we used a now deprecated model based on a LLaMA large language backbone and observed similar vulnerability to the attacks.
Our evaluation is based on the evaluation-script provided by . We evaluate on two image captioning tasks, COCO 2014 and Flickr30k . On these datasets we report the CIDEr score of the captions generated by the untargeted attack. The prompt is structured as
Moreover, we evaluate on two visual question answering tasks, OK-VQA and VizWiz . For these datasets the prompt-structure is
In each case we test zero-shot and four-shot inference. For four-shot inference we consider both the setting where the adversary can perturb context and query images (), and the setting where it can only perturb query images (). In zero-shot inference these settings coincide. On each dataset we evaluate on 1000 sampled instances with single-precision.
We quantitatively evaluate robustness to targeted adversarial attacks on COCO images. We consider two metrics: the success rate, which measures how often the exact target caption is contained in the generated output, and the BLEU-4 score between target and output captions. BLEU-4 is valued between 0 and 100 and measures the similarity between target and output captions. For computation of the BLEU-4 score, we limit the number output words to the number of words in the target caption. Note that the CIDEr score is not suited for this evaluation, as it applies a term frequency–inverse document frequency (TF-IDF) scheme, thus weighing down the importance of the target string.
For the optimization of the attack objectives (2) and (3) we use the APGD attack . APGD is a powerful iterative gradient-based attack. The only parameter it requires is the number of iterations. However, we decrease the hardcoded initial step-size ( in ) from to , as we observed that it increases the attack strength.
We observe that a high amount of iterations is necessary in order to effectively attack the model as reported in Table 1. Thus we use 5000 iterations for the targeted attacks in Figure 2. For the quantitative evaluations in Tables 2 and 3 we use 500 iterations, as this is the maximum that is computationally feasible for us. We expect that more iterations would lead to even more successful attacks.
It arises the question how to set the ground truth text for the untargeted attacks. For COCO and Flickr30k, we use for each sampled image one of the provided captions as ground truth . Similarly, for OK-VQA and VizWiz, we sample one of the ground truth answers. VizWiz contains ten ground-truth answers for each image. We observe that often one of the ground truth answers is “unanswerable”, while others give an accurate answer. Thus, we sample for each image one ground truth answer that is not “unanswerable”, unless more than half of them are.
Results
Overall, we observe that the proposed attacks are highly successful against the OpenFlamingo model in all considered settings.
We find that an attacker can often force the model to generate exactly a given caption (targeted attack) as shown in Figure 2 and Table 2. With the target caption “Please reset your password” the model reveals significant susceptibilities. In the more lenient threat model, with , the attack is already fairly effective, achieving success rates of up to 51.66% in the 0-shot setting and as high as 86.00% in the 4-shot setting when all images are attacked. However, when only the query image is targeted, the success rates are notably lower, indicating a more pronounced effect when also context images are compromised. The threat model’s expansion to results in perfect attack success rates. Here, the model is tricked into generating the adversarial output in all cases, with success rates of 100% for both the 0-shot and 4-shot settings.
The longer and thus much more challenging target “A person suffers severe side effects after vaccination” is not well recovered for the small threat model. However, using the larger threat model of the attack becomes much more effective again, especially when attacking all images. We observed that an increase of APGD iterations from 500 to 5000, albeit costly, does make the attack more effective even in smaller threat model. Qualitative results of this attack are shown in Figure 2 and we expect that also the quantitative results would improve significantly for more APGD iterations.
The results of the untargeted attack evaluation are reported in Table 3. The model demonstrates high adversarial vulnerability in all settings, achieving low CIDEr scores on the captioning benchmarks COCO and Flickr and low accuracies on the visual question answering tasks OK-VQA and VizWiz. On COCO, the CIDEr score is for the larger threat model even lower than the score of randomly permuted captions as computed in Section 4. Prompting the model with context images, as in the 4-shot case, helps only slightly on the captioning tasks and almost not at all on the VQA tasks. We show non-cherry-picked example outputs generated via the untargeted attack on the zero-shot model in Figures 3 and 4 and observe that the adversarial output captions do not describe any image accurately.
Discussion
From a user’s perspective, our findings underline a crucial security concern. As the adversarial perturbations applied to the images are slight and typically imperceptible to the human eye, users might unknowingly input adversarially manipulated images into the model. The adversarial modifications, while subtle, are potent enough to manipulate the model’s output substantially, affecting the overall reliability of the predictions. A malicious actor could exploit this vulnerability to inject biased, misleading, or harmful content into the model’s output. For instance, in a captioning task, a small perturbation could lead to entirely different and inaccurate captions, potentially causing misunderstanding or misinformation. Similarly, in visual question answering tasks, manipulated images could lead to incorrect or misleading responses, affecting decision-making based on the model’s predictions.
These outcomes are particularly concerning given the wide range of applications multi-modal models could be employed in, from aiding visually impaired individuals in understanding their surroundings to generating news articles based on visual input. For news article generation, the model can automate the process of interpreting images and generating relevant text. However, the introduction of adversarial perturbations to images can lead to the generation of misleading or completely false narratives.
For example, a subtly manipulated image associated with a news article could cause the model to generate text that distorts the reality of the situation as depicted in Figures 1 and 2. The misinformation could range from false health statements to inventing alarming political news.
These manipulations can significantly impact users. News articles have a broad reach and the potential to shape public opinion. If a user forms their understanding based on a misleading article generated from adversarially manipulated images, they may make misinformed decisions or actions. This scenario underscores the importance of developing robust security measures for multi-modal models, particularly as they become increasingly integrated into critical platforms like news media. An important factor towards achieving this goal is the availability of open-source multi-modal foundation models.
Conclusion
Our investigation into the adversarial robustness of the OpenFlamingo model showed that it is highly susceptible to perturbations on its visual inputs. Even slight perturbations that are hardly visible for humans can fool the model into poor performance on captioning and VQA tasks. More alarmingly, the targeted attacks presented in this paper allow an attacker to control the model’s outputs, crafting a desired response that may be deceiving or harmful.
The potential for targeted adversarial manipulation has serious implications for end users. As model outputs are often trusted implicitly, this vulnerability could thus lead to the spread of misinformation or manipulation of user behavior. It is crucial that we bring these vulnerabilities to light, not to invite misuse, but to stress the urgent need for mitigation strategies.
Our comprehensive analysis thus underscores the critical need for robustness in the design of multi-modal models, particularly as they become more pervasive across a wide range of applications. Future research should prioritize the development of robustness-enhancing strategies for multi-modal models against such adversarial attacks, thereby ensuring their safe application in real-world settings.
Acknowledgements
We thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting CS. We acknowledge support from the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy (EXC number 2064/1, project number 390727645), as well as in the priority program SPP 2298, project number 464101476. Moreover, we are thankful for the support of Open Philanthropy.