Attacking Visual Language Grounding with Adversarial Examples: A Case Study on Neural Image Captioning
Hongge Chen, Huan Zhang, Pin-Yu Chen, Jinfeng Yi, Cho-Jui Hsieh
Introduction
In recent years, language understanding grounded in machine vision and perception has made remarkable progress in natural language processing (NLP) and artificial intelligence (AI), such as image captioning and visual question answering. Image captioning is a multimodal learning task and has been used to study the interaction between language and vision models Shekhar et al. (2017). It takes an image as an input and generates a language caption that best describes its visual contents, and has many important applications such as developing image search engines with complex natural language queries, building AI agents that can see and talk, and promoting equal web access for people who are blind or visually impaired. Modern image captioning systems typically adopt an encoder-decoder framework composed of two principal modules: a convolutional neural network (CNN) as an encoder for image feature extraction and a recurrent neural network (RNN) as a decoder for caption generation. This CNN+RNN architecture includes popular image captioning models such as Show-and-Tell Vinyals et al. (2015), Show-Attend-and-Tell Xu et al. (2015) and NeuralTalk Karpathy and Li (2015).
Recent studies have highlighted the vulnerability of CNN-based image classifiers to adversarial examples: adversarial perturbations to benign images can be easily crafted to mislead a well-trained classifier, leading to visually indistinguishable adversarial examples to human Szegedy et al. (2014); Goodfellow et al. (2015). In this study, we investigate a more challenging problem in visual language grounding domain that evaluates the robustness of multimodal RNN in the form of a CNN+RNN architecture, and use neural image captioning as a case study. Note that crafting adversarial examples in image captioning tasks is strictly harder than in well-studied image classification tasks, due to the following reasons: (i) class attack v.s. caption attack: unlike classification tasks where the class labels are well defined, the output of image captioning is a set of top-ranked captions. Simply treating different captions as distinct classes will result in an enormous number of classes that can even precede the number of training images. In addition, semantically similar captions can be expressed in different ways and hence should not be viewed as different classes; and (ii) CNN v.s. CNN+RNN: attacking RNN models is significantly less well-studied than attacking CNN models. The CNN+RNN architecture is unique and beyond the scope of adversarial examples in CNN-based image classifiers.
In this paper, we tackle the aforementioned challenges by proposing a novel algorithm called Show-and-Fool. We formulate the process of crafting adversarial examples in neural image captioning systems as optimization problems with novel objective functions designed to adopt the CNN+RNN architecture. Specifically, our objective function is a linear combination of the distortion between benign and adversarial examples as well as some carefully designed loss functions. The proposed Show-and-Fool algorithm provides two approaches to craft adversarial examples in neural image captioning under different scenarios:
Targeted caption method: Given a targeted caption, craft adversarial perturbations to any image such that its generated caption matches the targeted caption.
Targeted keyword method: Given a set of keywords, craft adversarial perturbations to any image such that its generated caption contains the specified keywords. The captioning model has the freedom to make sentences with target keywords in any order.
As an illustration, Figure 1 shows an adversarial example crafted by Show-and-Fool using the targeted caption method. The adversarial perturbations are visually imperceptible while can successfully mislead Show-and-Tell to generate the targeted captions. Interestingly and perhaps surprisingly, our results pinpoint the Achilles heel of the language and vision models used in the tested image captioning systems. Moreover, the adversarial examples in neural image captioning highlight the inconsistency in visual language grounding between humans and machines, suggesting a possible weakness of current machine vision and perception machinery. Below we highlight our major contributions:
We propose Show-and-Fool, a novel optimization based approach to crafting adversarial examples in image captioning. We provide two types of adversarial examples, targeted caption and targeted keyword, to analyze the robustness of neural image captioners. To the best of our knowledge, this is the very first work on crafting adversarial examples for image captioning.
We propose powerful and generic loss functions that can craft adversarial examples and evaluate the robustness of the encoder-decoder pipelines in the form of a CNN+RNN architecture. In particular, our loss designed for targeted keyword attack only requires the adversarial caption to contain a few specified keywords; and we allow the neural network to make meaningful sentences with these keywords on its own.
We conduct extensive experiments on the MSCOCO dataset. Experimental results show that our targeted caption method attains a 95.8% attack success rate when crafting adversarial examples with randomly assigned captions. In addition, our targeted keyword attack yields an even higher success rate. We also show that attacking CNN+RNN models is inherently different and more challenging than only attacking CNN models.
We also show that Show-and-Fool can produce highly transferable adversarial examples: an adversarial image generated for fooling Show-and-Tell can also fool other image captioning models, leading to new robustness implications of neural image captioning systems.
Related Work
In this section, we review the existing work on visual language grounding, with a focus on neural image captioning. We also review related work on adversarial attacks on CNN-based image classifiers. Due to space limitations, we defer the second part to the supplementary material.
Visual language grounding represents a family of multimodal tasks that bridge visual and natural language understanding. Typical examples include image and video captioning Karpathy and Li (2015); Vinyals et al. (2015); Donahue et al. (2015b); Pasunuru and Bansal (2017); Venugopalan et al. (2015), visual dialog Das et al. (2017); De Vries et al. (2017), visual question answering Antol et al. (2015); Fukui et al. (2016); Lu et al. (2016); Zhu et al. (2017), visual storytelling (Huang et al., 2016), natural question generation (Mostafazadeh et al., 2017, 2016), and image generation from captions Mansimov et al. (2016); Reed et al. (2016). In this paper, we focus on studying the robustness of neural image captioning models, and believe that the proposed method also sheds lights on robustness evaluation for other visual language grounding tasks using a similar multimodal RNN architecture.
Many image captioning methods based on deep neural networks (DNNs) adopt a multimodal RNN framework that first uses a CNN model as the encoder to extract a visual feature vector, followed by a RNN model as the decoder for caption generation. Representative works under this framework include Chen and Zitnick (2015); Devlin et al. (2015); Donahue et al. (2015a); Karpathy and Li (2015); Mao et al. (2015); Vinyals et al. (2015); Xu et al. (2015); Yang et al. (2016); Liu et al. (2017a, b), which are mainly differed by the underlying CNN and RNN architectures, and whether or not the attention mechanisms are considered. Other lines of research generate image captions using semantic information or via a compositional approach Fang et al. (2015); Gan et al. (2017); Tran et al. (2016); Jia et al. (2015); Wu et al. (2016); You et al. (2016).
The recent work in Shekhar et al. (2017) touched upon the robustness of neural image captioning for language grounding by showing its insensitivity to one-word (foil word) changes in the language caption, which corresponds to the untargeted attack category in adversarial examples. In this paper, we focus on the more challenging targeted attack setting that requires to fool the captioning models and enforce them to generate pre-specified captions or keywords.
Methodology of Show-and-Fool
We now formally introduce our approaches to crafting adversarial examples for neural image captioning. The problem of finding an adversarial example for a given image can be cast as the following optimization problem:
where denotes the inverse hyperbolic tangent function and is applied element-wisely. Since , the transformation will automatically satisfy the box constraint. Consequently, the constrained optimization problem in (1) is equivalent to
In the following sections, we present our designed loss functions for different attack settings.
2 Targeted Caption Method
Note that a targeted caption is denoted by
where indicates the index of the -th word in the vocabulary list , is a start symbol and indicates the end symbol. is the length of caption , which is not fixed but does not exceed a predefined maximum caption length. To encourage the neural image captioning system to output the targeted caption , one needs to ensure the log probability of the caption conditioned on the image attains the maximum value among all possible captions, that is,
where is the set of all possible captions. It is also common to apply the chain rule to the joint probability and we have
In neural image captioning networks, is usually computed by a RNN/LSTM cell , with its hidden state and input :
Following the definition of softmax function:
Intuitively, to maximize the targeted caption’s probability, we can directly use its negative log probability (5) as a loss function. The inputs of the RNN are the first words of the targeted caption .
Applying (5) to (3.1), the formulation of targeted caption method given a targeted caption is:
Alternatively, using the definition of the softmax function,
Instead of making each as large as possible, it is sufficient to require the target word to attain the largest (top-1) logit (or probability) among all the words in the vocabulary at position . In other words, we aim to minimize the difference between the maximum logit except , denoted by , and the logit of , denoted by . We also propose a ramp function on top of this difference as the final loss function:
where is a confidence level accounting for the gap between and . When , the corresponding term in the summation will be kept at and does not contribute to the gradient of the loss function, encouraging the optimizer to focus on minimizing other terms where is not large enough.
Applying the loss (7) to (1), the final formulation of targeted caption method given a targeted caption is
We note that Carlini and Wagner (2017) has reported that in CNN-based image classification, using logits in the attack loss function can produce better adversarial examples than using probabilities, especially when the target network deploys some gradient masking schemes such as defensive distillation Papernot et al. (2016b). Therefore, we provide both logit-based and probability-based attack loss functions for neural image captioning.
3 Targeted Keyword Method
In addition to generating an exact targeted caption by perturbing the input image, we offer an intermediate option that aims at generating captions with specific keywords, denoted by . Intuitively, finding an adversarial image generating a caption with specific keywords might be easier than generating an exact caption, as we allow more degree of freedom in caption generation. However, as we need to ensure a valid and meaningful inferred caption, finding an adversarial example with specific keywords in its caption is difficult in an optimization perspective. Our target keyword method can be used to investigate the generalization capability of a neural captioning system given only a few keywords.
In our method, we do not require a target keyword to appear at a particular position. Instead, we want a loss function that allows to become the top-1 prediction (plus a confidence margin ) at any position. Therefore, we propose to use the minimum of the hinge-like loss terms over all as an indication of appearing at any position as the top-1 prediction, leading to the following loss function:
We note that the loss functions in (4) and (5) require an input to predict for each . For the targeted caption method, we use the targeted caption as the input of RNN. In contrast, for the targeted keyword method we no longer know the exact targeted sentence, but only require the presence of specified keywords in the final caption. To bridge the gap, we use the originally inferred caption from the benign image as the initial input to RNN. Specifically, after minimizing (8) for iterations, we run inference on and set the RNN’s input as its current top-1 prediction, and continue this process. With this iterative optimization process, the desired keywords are expected to gradually appear in top-1 prediction.
Another challenge arises in targeted keyword method is the problem of “keyword collision”. When the number of keywords , more than one keywords may have large values of at a same position . For example, if dog and cat are top-2 predictions for the second word in a caption, the caption can either start with “A dog …” or “A cat …”. In this case, despite the loss (8) being very small, a caption with both dog and cat can hardly be generated, since only one word is allowed to appear at the same position. To alleviate this problem, we define a gate function which masks off all the other keywords when a keyword becomes top- at position :
where A is a predefined value that is significantly larger than common logits values. Then (8) becomes:
The log-prob loss for targeted keyword method is discussed in the Supplementary Material.
Experiments
We performed extensive experiments to test the effectiveness of our Show-and-Fool algorithm and study the robustness of image captioning systems under different problem settings. In our experimentsOur source code is available at: https://github.com/huanzhang12/ImageCaptioningAttack, we use the pre-trained TensorFlow implementationhttps://github.com/tensorflow/models/tree/master/research/im2txt of Show-and-Tell Vinyals et al. (2015) with Inception-v3 as the CNN for visual feature extraction. Our testbed is Microsoft COCO Lin et al. (2014) (MSCOCO) data set. Although some more recent neural image captioning systems can achieve better performance than Show-and-Tell, they share a similar framework that uses CNN for feature extraction and RNN for caption generation, and Show-and-Tell is the vanilla version of this CNN+RNN architecture. Indeed, we find that the adversarial examples on Show-and-Tell are transferable to other image captioning models such as Show-Attend-and-Tell Xu et al. (2015) and NeuralTalk2https://github.com/karpathy/neuraltalk2, suggesting that the attention mechanism and the choice of CNN and RNN architectures do not significantly affect the robustness. We also note that since Show-and-Fool is the first work on crafting adversarial examples for neural image captioning, to the best of our knowledge, there is no other method for comparison.
We use ADAM to minimize our loss functions and set the learning rate to 0.005. The number of iterations is set to . All the experiments are performed on a single Nvidia GTX 1080 Ti GPU. For targeted caption and targeted keyword methods, we perform a binary search for 5 times to find the best : initially , and will be increased by times until a successful adversarial example is found. Then, we choose a new to be the average of the largest where an adversarial example can be found and the smallest where an adversarial example cannot be found. We fix except for transferability experiments. For each experiment, we randomly select 1,000 images from the MSCOCO validation set. We use BLEU-1 Papineni et al. (2002), BLEU-2, BLEU-3, BLEU-4, ROUGE Lin (2004) and METEOR Lavie and Agarwal (2005) scores to evaluate the correlations between the inferred captions and the targeted captions. These scores are widely used in NLP community and are adopted by image captioning systems for quality assessment. Throughout this section, we use the logits loss (7)(9). The results of using the log-prob loss (5) are similar and are reported in the supplementary material.
2 Targeted Caption Results
Unlike the image classification task where all possible labels are predefined, the space of possible captions in a captioning system is almost infinite. However, the captioning system is only able to output relevant captions learned from the training set. For instance, the captioning model cannot generate a passive-voice sentence if the model was never trained on such sentences. Therefore, we need to ensure that the targeted caption lies in the space where the captioning system can possibly generate. To address this issue, we use the generated caption of a randomly selected image (other than the image under investigation) from MSCOCO validation set as the targeted caption . The use of a generated caption as the targeted caption excludes the effect of out-of-domain captioning, and ensures that the target caption is within the output space of the captioning network.
3 Targeted Keyword Results
In this task, we use (9) as our loss function, and choose the number of keywords . We run an inference step on every iterations, and use the top-1 caption as the input of RNN/LSTMs. Similar to Section 4.2, for each image the targeted keywords are selected from the caption generated by a randomly selected validation set image. To exclude common words like “a”, “the”, “and”, we look up each word in the targeted sentence and only select nouns, verbs, adjectives or adverbs. We say an adversarial image is successful when its caption contains all specified keywords. The overall success rate and average distortion are shown in Table 1. When compared to the targeted caption method, targeted keyword method achieves an even higher success rate (at least 96% for 3-keyword case and at least 97% for 1-keyword and 2-keyword cases). Figure 2 shows an adversarial example crafted from our targeted keyword method with three keywords - “dog”, “cat” and “frisbee”. Using Show-and-Fool, the top-1 caption of a cake image becomes “A dog and a cat are playing with a frisbee” while the adversarial image remains visually indistinguishable to the original one. When and , even if we cannot find an adversarial image yielding all specified keywords, we might end up with a caption that contains some of the keywords (partial success). For example, when , Table 3 shows the number of keywords appeared in the captions () for those failed examples (not all 3 targeted keywords are found). These results clearly show that the failed examples are still partially successful: the generated captions contain about 1.5 targeted keywords on average.
4 Transferability of Adversarial Examples
It has been shown that in image classification tasks, adversarial examples found for one machine learning model may also be effective against another model, even if the two models have different architectures Papernot et al. (2016a); Liu et al. (2017c). However, unlike image classification where correct labels are made explicit, two different image captioning systems may generate quite different, yet semantically similar, captions for the same benign image. In image captioning, we say an adversarial example is transferable when the adversarial image found on model with a target sentence can generate a similar (rather than exact) sentence on model .
In our setting, model is Show-and-Tell, and we choose Show-Attend-and-Tell Xu et al. (2015) as model . The major differences between Show-and-Tell and Show-Attend-and-Tell are the addition of attention units in LSTM network for caption generation, and the use of last convolutional layer (rather than the last fully-connected layer) feature maps for feature extraction. We use Inception-v3 as the CNN architecture for both models and train them on the MSCOCO 2014 data set. However, their CNN parameters are different due to the fine-tuning process.
To investigate the transferability of adversarial examples in image captioning, we first use the targeted caption method to find adversarial examples for 1,000 images in model with different and , and then transfer successful adversarial examples (which generate the exact target captions on model ) to model . The generated captions by model are recorded for transferability analysis. The transferability of adversarial examples depends on two factors: the intrinsic difference between two models even when the same benign image is used as the input, i.e., model mismatch, and the transferability of adversarial perturbations.
5 Attacking Image Captioning v.s. Attacking Image Classification
In this section we show that attacking image captioning models is inherently more challenging than attacking image classification models. In the classification task, a targeted attack usually becomes harder when the number of labels increases, since an attack method needs to change the classification prediction to a specific label over all the possible labels. In the targeted attack on image captioning, if we treat each caption as a label, we need to change the original label to a specific one over an almost infinite number of possible labels, corresponding to a nearly zero volume in the search space. This constraint forces us to develop non-trivial methods that are significantly different from the ones designed for attacking image classification models.
Conclusion
In this paper, we proposed a novel algorithm, Show-and-Fool, for crafting adversarial examples and providing robustness evaluation of neural image captioning. Our extensive experiments show that the proposed targeted caption and keyword methods yield high attack success rates while the adversarial perturbations are still imperceptible to human eyes. We further demonstrate that Show-and-Fool can generate highly transferable adversarial examples. The high-quality and transferable adversarial examples in neural image captioning crafted by Show-and-Fool highlight the inconsistency in visual language grounding between humans and machines, suggesting a possible weakness of current machine vision and perception machinery. We also show that attacking neural image captioning systems are inherently different from attacking CNN-based image classifiers.
Our method stands out from the well-studied adversarial learning on image classifiers and CNN models. To the best of our knowledge, this is the very first work on crafting adversarial examples for neural image captioning systems. Indeed, our Show-and-Fool algorithm1 can be easily extended to other applications with RNN or CNN+RNN architectures. We believe this paper provides potential means to evaluate and possibly improve the robustness (for example, by adversarial training or data augmentation) of a wide range of visual language grounding and other NLP models.
References
Supplementary Material
Related Work on Adversarial Attacks to CNN-based Image Classifiers
Despite the remarkable progress, CNNs have been shown to be vulnerable to adversarial examples Szegedy et al. (2014); Goodfellow et al. (2015); Carlini and Wagner (2017). In image classification, an adversarial example is an image that is visually indistinguishable to the original image but can cause a CNN model to misclassify. With different objectives, adversarial attacks can be divided into two categories, i.e., untargeted attack and targeted attack. In the literature, a successful untargeted attack refers to finding an adversarial example that is close to the original example but yields different class prediction. For targeted attack, a target class is specified and the adversarial example is considered successful when the predicted class matches the target class. Surprisingly, adversarial examples can also be crafted even when the parameters of target CNN model are unknown to an attacker Liu et al. (2017c); Chen et al. (2017). In addition, adversarial examples crafted from one image classification model can be made transferable to other models Liu et al. (2017c); Papernot et al. (2016a), and there exists a universal adversarial perturbation that can lead to misclassification of natural images with high probability Moosavi-Dezfooli et al. (2017).
Without loss of generality, there are two factors contributing to crafting adversarial examples in image classification: (i) a distortion metric between the original and adversarial examples that regularizes visual similarity. Popular choices are the , and distortions Kurakin et al. (2017); Carlini and Wagner (2017); Chen et al. (2018); and (ii) an attack loss function accounting for the success of adversarial examples. For finding adversarial examples in neural image captioning, while the distortion metric can be identical, the attack loss function used in image classification is invalid, since the number of possible captions easily outnumbers the number of image classes, and captions with similar meaning should not be considered as different classes. One of our major contributions is to design novel attacking loss functions to handle the CNN+RNN architectures in neural image captioning tasks.
More Adversarial Examples with Logits Loss
Figure 4 shows another successful example with targeted caption method. Figures 5, 6 and 7 show three adversarial examples generated by the proposed 3-keyword method. The adversarial examples generated by our methods have small distortions and are visually indistinguishable from the original images. One advantage of using logits losses is that it helps to bypass defensive distillation by overcoming the gradient vanishing problem. To see this, the partial derivative of the softmax function
which vanishes as or The defensive distillation method uses a large distillation temperature in the training process and removes it in the inference process. This makes the inference probability close to 0 or 1, thus leads to a vanished gradient problem. However, by using the proposed logits loss (7), before the word at position in target sentence reaches top-1 probability, we have
It is evident that the gradient (with regard to ) becomes a constant now, since it equals to when and 0 otherwise.
Targeted Caption Results with Log Probability Loss
In this experiment, we use the log probability loss (5) plus a distortion term (as in (3.1)) as our objective function. Similar to the previous experiments, a successful adversarial example is found if the inferred caption after adding the adversarial perturbation exactly matches the targeted caption. The overall success rate and average distortion of adversarial perturbation are shown in Table 5. Among all the tested images, our log-prob loss attains 95.4% success rate, which is about the same as using logits loss. Besides, similar to using logits loss, the adversarial examples generated by using log-prob loss also yield small distortions. In Table 6, we summarize the statistics of the failed adversarial examples. It shows that their generated captions, though not entirely identical to the targeted caption, are also highly relevant to the target captions.
In our experiments, log probability loss exhibits a similar performance as the logits loss, as our target model is undefended and the gradient vanishing problem of softmax is not significant. However, when evaluating the robustness of a general image captioning model, it is recommended to use the logits loss as it does not suffer from potentially vanished gradients and can reveal the intrinsic robustness of the model.
Targeted Keyword Results with Log Probability Loss
Similar to the logits loss, the log-prob loss does not require a particular position for the target keywords . Instead, it encourages to become the top-1 prediction at its most probable position:
To tackle the “keyword collision” problem, we also employ a gate function to avoid the keywords appearing at the positions where the most probable word is already a keyword:
In our methods, the initial input is the originally inferred caption from the benign image, and after minimizing (13) for iterations, we run inference on and set the RNN’s input as its current top-1 prediction, and repeat this procedure until all the targeted keywords are found or the maximum number of iterations is met. With this iterative optimization process, the probabilities of the desired keywords gradually increase, and finally become the top-1 predictions.
The overall success rate and average distortion are shown in Table 5. Table 7 summarizes the number of keywords () appeared in the captions for those failed examples when , i.e., the examples that not all the 3 targeted keywords are found. They account only of all the tested images. Table 7 clearly shows that when is properly chosen, more than of the failed examples contain at least 1 targeted keyword, and more than of the failed examples contain 2 targeted keywords. This result verifies that even the failed examples are reasonably good attacks.
Transferability of Adversarial Examples with Log Probability Loss
Similar to the experiments in Section 4.4, to assess the transferability of adversarial examples, we first use the targeted caption method with log-prob loss to find adversarial examples for 1,000 images in Show-and-Tell model (model ) with different . We then transfer successful adversarial examples, i.e., the examples that generate the exact target captions on model , to Show-Attend-and-Tell model (model ). The generated captions by model are recorded for transferability analysis. The results for transferability using log-prob loss is summarized in Table 8. The definitions of tgt, ori and mis are the same as those in Table 4. Comparing with Table 4 (), the log probability loss shows inferior ori and tgt values, indicating that the additional parameter in the logits loss helps improve transferability.
Attention on Original and Transferred Adversarial Images
Figures 11, 12 and 13 show the original and adversarial images’ attentions over time. In the original images, the Show-Attend-and-Tell model’s attentions align well with human perception. However, the transferred adversarial images obtained on Show-and-Tell model yield significantly misaligned attentions.