How Robust is Google's Bard to Adversarial Image Attacks?

Yinpeng Dong, Huanran Chen, Jiawei Chen, Zhengwei Fang, Xiao Yang, Yichi Zhang, Yu Tian, Hang Su, Jun Zhu

Introduction

The recent progress of Large Language Models (LLMs) has demonstrated unprecedented levels of proficiency in language understanding, reasoning, and generation. Leveraging the powerful LLMs, numerous studies have attempted to seamlessly integrate visual inputs into LLMs. They often employ pre-trained vision encoders (e.g., CLIP ) to extract image features and then align image and language embeddings. These Multimodal Large Language Models (MLLMs) have demonstrated impressive abilities in vision-related tasks, such as image description, visual reasoning, etc. Recently, Google’s Bard released its multimodal capability which allows users to submit prompt containing both image and text, demonstrating superior performance over open-source MLLMs .

Despite these commendable achievements, the security and safety problems associated with these large-scale foundation models are still inevitable and remain a significant challenge . These problems can be amplified for MLLMs, that the integration of vision inputs introduces a compelling attack surface due to the continuous and high-dimensional nature of images . It is a well-established fact that vision models are inherently susceptible to small adversarial perturbations . The adversarial vulnerability of vision encoders can be inherited by MLLMs, resulting in security and safety risks in practical applications of large models.

Some recent studies have explored the robustness of MLLMs to adversarial image attacks . However, these works mainly focus on open-source MLLMs (e.g., MiniGPT4 ), leaving the robustness of commercial MLLMs (e.g., Bard) unexplored. It would be more challenging to attack commercial MLLMs because they are black-box models with unknown model configurations and training datasets, they have much more parameters with significantly better performance, and they are equipped with elaborate defense mechanisms. A common way of performing black-box attacks is based on adversarial transferability , i.e., adversarial examples generated for white-box models are likely to mislead black-box models. Although extensive efforts have been devoted to improving the adversarial transferability, they mainly consider image classification models . Due to the large difference between MLLMs and conventional classifiers, it is worth exploring the effective strategies to fool commercial MLLMs, with the purpose of fully understanding the vulnerabilities of these prominent models.

Given the vulnerabilities of Bard identified in our experiments under adversarial image attacks, we further discuss broader impacts to the practical use of MLLMs and suggest some potential solutions to improve their robustness. We hope this work can provide a deeper understanding of the weaknesses of MLLMs in the aspect of adversarial robustness under the completely black-box setting, and facilitate future research to develop more robust and trustworthy multimodal foundation models.

Related work

Multimodal large language models. The breakthrough of Large Language Models (LLMs) in language-oriented tasks and the emergence of GPT-4 motivate researchers to harness the powerful capabilities of LLMs to assist in various tasks across multimodal scenarios, and further lead to the new realm of Multimodal Large Language Models (MLLMs) . There have been different strategies and models to bridge the gap between text and other modalities. Some works leverage learnable queries to extract visual information and generate language using LLMs conditioned on the visual features. Models including MiniGPT-4 , LLaVA and PandaGPT learn simple projection layers to align the visual features from visual encoders with text embeddings for LLMs. Also, parameter-efficient fine-tuning is adopted by introducing lightweight trainable adapters into models . Several benchmarks have verified that MLLMs show satisfying performance on visual perception and comprehension.

Adversarial robustness of MLLMs. Despite achieving impressive performance, MLLMs still face issues of adversarial robustness due to their architecture based on deep neural networks . Multiple primary attempts have been conducted to study the robustness of MLLMs from different aspects. evaluates the adversarial robustness of MLLMs on image captioning under white-box settings, while conducts both transfer-based and query-based attacks on MLLMs assuming black-box access. trigger LLMs to generate toxic content by imposing adversarial perturbations to the input images. studies image hijacks to achieve specific string, leak context, and jailbreak attacks. These exploratory works demonstrate that MLLMs still face stability and security issues under adversarial perturbations. However, they only consider popular open-source models, but do not study commercial MLLMs (e.g., Bard ). Not only are their model and training configurations unknown, but they are also equipped with multiple auxiliary modules to enhance the performance and ensure the safety, making it more challenging to attack.

Black-box adversarial attacks. Black-box adversarial attacks can be generally categorized into query-based and transfer-based methods. Query-based methods require repeatedly invoking the victim model for gradient estimation, incurring higher costs. In contrast, transfer-based methods only need local surrogate models, leveraging the transferability across models of adversarial samples to carry out the attack. Some methods improve the optimization process by correcting gradients similar to the methods in model training that enhance generalization. Besides, incorporating diversities into the optimization could also raise the transferability , which applies various transformations to inputs to boost the generalization. The ensemble-based attack is also effective when generating the adversarial samples on a group of surrogate models or adjusting one model to simulate diverse models .

Attack on image description

Google’s Bard is a representative MLLM that allows users to assess its multimodal capability through API access. This work aims to identify the adversarial vulnerabilities of Bard to highlight the risks associated with it and the importance of designing more robust models in the future. Specifically, we evaluate the performance of Bard to describe image contents perturbed by imperceptible adversarial noises. We choose the image description task since it is one of the fundamental tasks of MLLMs and we can avoid the influence of instruction following ability on our evaluation. As the model will evole over time, we perform all evaluations during September 10th to 15th, 2023 using the latest update of Bard at July 13th, 2023.

MLLMs usually first extract image embeddings using vision encoders and then generate corresponding text based on image embeddings. Thus, we propose two attacks for MLLMs – image embedding attack and text description attack. As their names indicate, image embedding attack makes the embedding of the adversarial image diverge from that of the original image, based on the fact that if adversarial examples can successfully disrupt the image embeddings of Bard, the generated text will inevitably be affected. On the other hand, text description attack targets the entire pipeline directly to make the generated description different from the correct one.

Formally, let xnat\bm{x}_{nat} denote a natural image and {fi}i=1N\{f_{i}\}_{i=1}^{N} denote a set of surrogate image encoders. The image embedding attack can be formulated as solving

For text description attack, we collect a set of surrogate MLLMs as {gi}i=1N\{g_{i}\}_{i=1}^{N}, where gig_{i} can predict a probability distribution of the next word wtw_{t} given the image x\bm{x}, text prompt p\bm{p}, and previously predicted words w<tw_{<t} as pgi(wt∣x,p,w<t)p_{g_{i}}(w_{t}|\bm{x},\bm{p},w_{<t}). The text description attack maximizes the log-likelihood of predicting a target sentence Y:={yt}t=1LY:=\{y_{t}\}_{t=1}^{L} as

Note that we perform a targeted attack in Eq. 2 rather than an untargeted attack that minimizes the log-likelihood of the ground-truth description. This is because there are multiple correct descriptions of an image. If we only minimize the log-likelihood of predicting a single ground-truth description, the model can also output other correct descriptions given the adversarial example, making the attack ineffective.

To solve the optimization problems in Eq. 1 and Eq. 2, we adopt the state-of-the-art transfer-based attack methods in this paper. The spectrum simulation attack (SSA) performs a spectrum transformation to the input to improve the adversarial transferability. The common weakness attack (CWA) proposes to find the common weakness of an ensemble of surrogate models by promoting the flatness of loss landscapes and closeness between local optima of surrogate models. SSA and CWA can be combined as SSA-CWA, which demonstrates superior transferability for black-box models. Therefore, we adopt SSA-CWA as our attack. More details can be found in .

2 Experimental results

Results. Tab. 1 shows the results. The image embedding attack achieves 22% success rate while the text description attack achieves 10% success rate against Bard. The superiority of image embedding attack over text description attack may be due to the similarity between vision encoders but large differences between LLMs, as commercial models like Bard usually adopt much larger LLMs than open-source LLMs used in our experiments. Note that some of the adversarial examples are wrongly rejected by the defenses of Bard. Fig. 2 shows two successful adversarial examples that Bard provides incorrect descriptions, e.g., Bard describes a panda’s face as a painting of a woman’s face as shown in Fig. 2(b). The experiment demonstrates that large vision-language models like Bard are vulnerable to adversarial attacks and can readily misidentify objects in adversarial images.

Ablation study on model ensemble. To prove the effectiveness of the ensemble attack, we conduct an ablation study with different surrogate models. For simplicity, we only choose 20 images in the NIP2017 dataset to perform image embedding attack. As illustrated in Tab. 2, the attack success rate increases with the number of surrogate models. Therefore, in this work, we choose to ensemble three surrogate models to strike a balance between efficacy and time complexity.

Generalization across different prompts. To assess the generalization of the adversarial examples across different prompts, we measure the attack success rate using the prompts in (e.g., "Provide a brief description of the given image.", "Offer a succinct explanation of the picture presented.", "Take a look at this image and describe what you notice", "Summarize the visual content of the image.", etc.). Remarkably, the adversarial examples that are successful given the original prompt "Describe this image", can also mislead Bard using the prompts given above, demonstrating good generalization of adversarial examples across different prompts.

3 Attack on other MLLMs

We then examine the attack performance of our generated adversarial examples against other commercial MLLMs. GPT-4V is very recently accessible at October 2023 after the first version of this paper. We further evaluate its robustness at October 13th, 2023 in the second version of this paper. In the first version, we also consider two other commercial MLLMs, including Bing Chat and ERNIE Bot . We adopt the 100 adversarial examples generated by the image embedding attack method to directly evaluate the performance of these two models.

Tab. 3 shows the results of attacking GPT-4V, Bing Chat, and ERNIE Bot. Our attack achieves 45%, 26%, and 86% attack success rates against GPT-4V, Bing Chat, and ERNIE bot, respectively, while most of the natural images can be correctly described. There are 30% adversarial images being rejected by Bing Chat since it finds noises in them. Based on the results, we find that Bard is the most robust model among the commercial MLLMs we study, and ERNIE Bot is the least robust one under our attack with 86% success rate. We find that the attack success rate is higher for GPT-4V since it will provide vague descriptions for adversarial images rather than rejecting them like Bing Chat. Fig. 3, Fig. 4, and Fig. 5 show the successful examples of attacking GPT-4V, Bing Chat, and ERNIE Bot, respectively. The results indicate that commercial MLLMs have similar robustness issues under adversarial attacks, requiring further improvement of robustness.

Attack on defenses of Bard

In our evaluation of Bard, we found that Bard is equipped with (at least) two defense mechanisms, including face detection and toxicity detection. Bard will directly reject images containing human faces or toxic contents (e.g., violent, bloody, or pornographic images). The defenses may be deployed to protect human privacy and avoid abuse. However, the robustness of the defenses under adversarial attacks is unknown. Therefore, we evaluate their robustness in this section.

Modern face detection models employ deep neural networks to identify human faces with impressive performance. To attack the face detection module of Bard, we select several face detectors as white-box surrogate models for ensemble attacks. Let {Di}i=1K\{D_{i}\}_{i=1}^{K} denote the set of surrogate face detectors. The output of a face detector DiD_{i} contains three elements: the anchor AA, the bounding box BB, and the face confidence score S∈{0,1}S\in\{0,1\}. Therefore, our face attack minimizes the confidence score such that the model cannot detect the face, which can be formulated as

where LL is the binary cross-entropy (BCE) loss and y^=0\hat{y}=0 (i.e., we minimize the confidence score SDi(x)S_{D_{i}}(\bm{x})). xnat\bm{x}_{nat} is the natural image containing human face and we aim to generate an adversarial example x\bm{x} without being detected. We also adopt the SSA-CWA method to solve Eq. 3.

Experimental settings. (1) Dataset: The experiments are conducted on FFHQ and LFW . The FFHQ dateset comprises 70,000 images, each with a resolution of 1024 ×{\times} 1024. The LFW dataset contains 13,233 celebrity images with a resolution of 250 ×{\times} 250. We randomly select 100 images from each dataset for manual testing. (2) Surrogate models: We choose three public face detection models for ensemble attack, including PyramidBox , S3FD and DSFD . (3) Hyper-parameters: We consider perturbation budgets ϵ=16/255\epsilon=16/255 and ϵ=32/255\epsilon=32/255. (4) Evaluation metric: We consider an attack successful if Bard does not reject the image and provides a description.

Experimental results and analyses. In Fig. 6, we present examples of successful attacks on FFHQ dataset. The quantitative results are summarized in Tab. 4. The experimental results suggest that even if the detailed model configurations of Bard are unknown, we still can successfully attack the face detector of Bard under the black-box setting based on the transferability of adversarial examples. In addition, it seems that the attack success rate is positively correlated with the value of the perturbation budget and negatively correlated with the image resolution. In other words, the attack success rate is higher when the ϵ\epsilon is larger and the image resolution is lower.

2 Attack on toxicity detection

To prevent providing descriptions for toxic images, Bard employs a toxicity detector to filter out such images. To attack it, we need to select certain white-box toxicity detectors as surrogate models. We find that some existing toxicity detectors are linear probed versions of pre-trained vision models like CLIP . To target these surrogate models, we only need to perturb the features of these pre-trained models. Therefore, we employ the exact same objective function as given in Eq. 1 and use the same attack method SSA-CWA. Note that this procedure could also affect the description of the image as shown in Sec. 3. But as the attack success rate on image description is not very high, we could find successful examples that not only evade the toxicity detector but also lead to correct description of the image.

Experiment. We manually collect a set of 100 toxic images containing violent, bloody, or pornographic contents. The other experimental settings are the same as Sec. 3.2. We achieve 36% attack success rate against Bard’s toxicity detector. As shown in Fig. 7, the toxicity detector fails to identify the toxic images with adversarial noises. Consequently, Bard provides inappropriate descriptions for these images. This experiment underscores the potential for malicious adversaries to exploit Bard to generate unsuitable descriptions for harmful contents.

Discussion and Conclusion

In this paper, we analyzed the robustness of Google’s Bard to adversarial attacks on images. By using the state-of-the-art transfer-based attacks to optimize the objectives on image embedding or text description, we achieved a 22% attack success rate against Bard on the image description task. The adversarial examples can also mislead other commercial MLLMs, including Bing Chat with a 26% attack success rate and ERNIE Bot with a 86% attack success rate. The results demonstrate the vulnerability of commercial MLLMs under black-box adversarial attacks. We also found that the current defense mechanisms of Bard can also be easily evaded by adversarial examples.

As large-scale foundation models (e.g., ChatGPT, Bard) have been increasingly used by humans for various purposes, their security and safety problems become a big concern to the public. Adversarial robustness is an important aspect of model security. Although we consider adversarial attacks on the typical image description task, which is not very harmful in some sense, some works demonstrate that adversarial attacks can be used to break the alignment of LLMs or MLLMs . For example, by attaching an adversarial suffix to harmful prompts, LLMs would produce objectionable responses. This problem will be more severe for MLLMs since attacks can be conducted on images. And it will be harder to defend against adversarial image perturbations than adversarial text perturbations due to the continuous space of images. Although previous works have studied this problem for MLLMs, they only consider white-box attacks. We will study black-box attacks against the alignment of commercial MLLMs in future work.

Given the problems of AT, we think that preprocessing-based defenses are more suitable for large-scale foundation models as they can be used in a plug-and-play manner. Some recent works leverage advanced generative models (e.g., diffusion models ) to purify adversarial perturbations (e.g., diffusion purification , likelihood maximization ), which could serve as promising strategies to defend against adversarial examples. We hope this work can motivate future research on developing more effective defense strategies for large-scale foundation models.

References