Evading Defenses to Transferable Adversarial Examples by Translation-Invariant Attacks

Yinpeng Dong, Tianyu Pang, Hang Su, Jun Zhu

Introduction

Despite the great success, deep neural networks have been shown to be highly vulnerable to adversarial examples . These maliciously generated adversarial examples are indistinguishable from legitimate ones by adding small perturbations, but make deep models produce unreasonable predictions. The existence of adversarial examples, even in the physical world , has raised concerns in security-sensitive applications, e.g., self-driving cars, healthcare and finance.

Attacking deep neural networks has drawn an increasing attention since the generated adversarial examples can serve as an important surrogate to evaluate the robustness of different models and improve the robustness . Several methods have been proposed to generate adversarial examples with the knowledge of the gradient information of a given model, such as fast gradient sign method , basic iterative method , and Carlini & Wagner’s method , which are known as white-box attacks. Moreover, it is shown that adversarial examples have cross-model transferability , i.e., the adversarial examples crafted for one model can fool a different model with a high probability. The transferability enables practical black-box attacks to real-world applications and induces serious security issues.

The threat of adversarial examples has motivated extensive research on building robust models or techniques to defend against adversarial attacks. These include training with adversarial examples , image denoising/transformation , theoretically-certified defenses , and others . Although the non-certified defenses have demonstrated robustness against common attacks, they do so by causing obfuscated gradients, which can be easily circumvented by new attacks . However, some of the defenses claim to be resistant to transferable adversarial examples, making it difficult to evade them by black-box attacks.

The resistance of the defense models against transferable adversarial examples is largely due to the phenomenon that the defenses make predictions based on different discriminative regions compared with normally trained models. For example, we show the attention maps of several normally trained models and defense models in Fig. 2, to represent the discriminative regions for their predictions. It can be seen that the normally trained models have similar attention maps while the defenses induce different attention maps. A similar observation is also found in that the gradients of the defenses in the input space align well with human perception, while those of normally trained models appear very noisy. This phenomenon of the defenses is caused by either training under different data distributions or transforming the inputs before classification . For black-box attacks based on the transferability , an adversarial example is usually generated for a single input against a white-box model. So the generated adversarial example is highly correlated with the discriminative region or gradient of the white-box model at the given input point, making it hard to transfer to other defense models that depend on different regions for predictions. Therefore, the transferability of adversarial examples is largely reduced to the defenses.

To mitigate the effect of different discriminative regions between models and evade the defenses by transferable adversarial examples, we propose a translation-invariant attack method. In particular, we generate an adversarial example for an ensemble of images composed of a legitimate one and its translated versions. We expect that the resultant adversarial example is less sensitive to the discriminative region of the white-box model being attacked, and has a higher probability to fool another black-box model with a defense mechanism. However, to generate such a perturbation, we need to calculate the gradients for all images in the ensemble, which brings much more computations. To improve the efficiency of our attacks, we further show that our method can be implemented by convolving the gradient at the untranslated image with a pre-defined kernel under a mild assumption. By combining the proposed method with any gradient-based attack method (e.g., fast gradient sign method , etc.), we obtain more transferable adversarial examples with similar computation complexity.

Extensive experiments on the ImageNet dataset demonstrate that the proposed translation-invariant attack method helps to improve the success rates of black-box attacks against the defense models by a large margin. Our best attack reaches an average success rate of 82%82\% to evade eight state-of-the-art defenses based only on the transferability, thus demonstrating the insecurity of the current defenses.

Related Work

Adversarial examples. Deep neural networks have been shown to be vulnerable to adversarial examples first in the visual domain . Then several methods are proposed to generate adversarial examples for the purpose of high success rates and minimal size of perturbations . They also exist in the physical world . Although adversarial examples are recently crafted for many other domains, we focus on image classification tasks in this paper.

Black-box attacks. Black-box adversaries have no access to the model parameters or gradients. The transferability of adversarial examples can be used to attack a black-box model. Several methods have been proposed to improve the transferability, which enable powerful black-box attacks. Besides the transfer-based black-box attacks, there is another line of work that performs attacks based on adaptive queries. For example, Papernot et al. use queries to distill the knowledge of the target model and train a surrogate model. They therefore turn the black-box attacks to the white-box attacks. Recent methods use queries to estimate the gradient or the decision boundary of the black-box model to generate adversarial examples. However, these methods usually require a large number of queries, which is impractical in real-world applications. In this paper, we resort to transfer-based black-box attacks.

Attacks for an ensemble of examples. An adversarial perturbation can be generated for an ensemble of legitimate examples. In , the universal perturbations are generated for the entire data distribution, which can fool the models on most of natural images. In , the adversarial perturbation is optimized over a distribution of transformations, which is similar to our method. The major difference between the method in and ours is three-fold. First, we want to generate transferable adversarial examples against the defense models, while the authors in propose to synthesize robust adversarial examples in the physical world. Second, we only use the translation operation, while they use a lot of transformations such as rotation, translation, addition of noise, etc. Third, we develop an efficient algorithm for optimization that only needs to calculate the gradient for the untranslated image, while they calculate the gradients for a batch of transformed images by sampling.

Defend against adversarial attacks. A large variety of methods have been proposed to increase the robustness of deep learning models. Besides directly making the models produce correct predictions for adversarial examples, some methods attempt to detect them instead . However, most of the non-certified defenses demonstrate the robustness by causing obfuscated gradients, which can be successfully circumvented by new attacks . Although these defenses are not robust in the white-box setting, some of them empirically show the resistance against transferable adversarial examples in the black-box setting. In this paper, we focus on generating more transferable adversarial examples against these defenses.

Methodology

In this section, we provide the detailed description of our algorithm. Let xreal\bm{x}^{real} denote a real example and yy denote the corresponding ground-truth label. Given a classifier f(x):X→Yf(\bm{x}):\mathcal{X}\rightarrow\mathcal{Y} that outputs a label as the prediction for an input, we want to generate an adversarial example xadv\bm{x}^{adv} which is visually indistinguishable from xreal\bm{x}^{real} but fools the classifier, i.e., f(xadv)≠yf(\bm{x}^{adv})\neq y.This corresponds to untargeted attack. The method in this paper can be simply extended to targeted attack. In most cases, the LpL_{p} norm of the adversarial perturbation is required to be smaller than a threshold ϵ\epsilon as ∣∣xadv−xreal∣∣p≤ϵ||\bm{x}^{adv}-\bm{x}^{real}||_{p}\leq\epsilon. In this paper, we use the L∞L_{\infty} norm as the measurement. For adversarial example generation, the objective is to maximize the loss function J(xadv,y)J(\bm{x}^{adv},y) of the classifier, where JJ is often the cross-entropy loss. So the constrained optimization problem can be written as

To solve this optimization problem, the gradient of the loss function with respect to the input needs to be calculated, termed as white-box attacks. However, in some cases, we cannot get access to the gradients of the classifier, where we need to perform attacks in the black-box manner. We resort to transferable adversarial examples which are generated for a different white-box classifier but have high transferability for black-box attacks.

Several methods have been proposed to solve the optimization problem in Eq. (1). We give a brief introduction of them in this section.

Fast Gradient Sign Method (FGSM) generates an adversarial example xadv\bm{x}^{adv} by linearizing the loss function in the input space and performing one-step update as

Basic Iterative Method (BIM) extends FGSM by iteratively applying gradient updates multiple times with a small step size α\alpha, which can be expressed as

where x0adv=xreal\bm{x}_{0}^{adv}=\bm{x}^{real}. To restrict the generated adversarial examples within the ϵ\epsilon-ball of xreal\bm{x}^{real}, we can clip xtadv\bm{x}_{t}^{adv} after each update, or set α=\nicefracϵT\alpha=\nicefrac{{\epsilon}}{{T}}, with TT being the number of iterations. It has been shown that BIM induces much more powerful white-box attacks than FGSM at the cost of worse transferability .

Momentum Iterative Fast Gradient Sign Method (MI-FGSM) proposes to improve the transferability of adversarial examples by integrating a momentum term into the iterative attack method. The update procedure is

where gt\bm{g}_{t} gathers the gradient information up to the tt-th iteration with a decay factor μ\mu.

Diverse Inputs Method applies random transformations to the inputs and feeds the transformed images into the classifier for gradient calculation. The transformation includes random resizing and padding with a given probability. This method can be combined with the momentum-based method to further improve the transferability.

Carlini & Wagner’s method (C&W) is a powerful optimization-based method, which solves

where the loss function JJ could be different from the cross-entropy loss. This method aims to find adversarial examples with minimal size of perturbations, to measure the robustness of different models. It also lacks the effectiveness for black-box attacks like BIM.

2 Translation-Invariant Attack Method

Although many attack methods can generate adversarial examples with very high transferability across normally trained models, they are less effective to attack defense models in the black-box manner. Some of the defenses are shown to be quite robust against black-box attacks. So we want to answer that: Are these defenses really free from transferable adversarial examples?

We find that the discriminative regions used by the defenses to identify object categories are different from those used by normally trained models, as shown in Fig. 2. When generating an adversarial example by the methods introduced in Sec. 3.1, the adversarial example is only optimized for a single legitimate example. So it may be highly correlated with the discriminative region or gradient of the white-box model being attacked at the input data point. For other black-box defense models that have different discriminative regions or gradients, the adversarial example can hardly remain adversarial. Therefore, the defenses are shown to be robust against transferable adversarial examples.

To generate adversarial examples that are less sensitive to the discriminative regions of the white-box model, we propose a translation-invariant attack method. In particular, rather than optimizing the objective function at a single point as Eq. (1), the proposed method uses a set of translated images to optimize an adversarial example as

where Tij(x)T_{ij}(\bm{x}) is the translation operation that shifts image x\bm{x} by ii and jj pixels along the two-dimensions respectively, i.e., each pixel (a,b)(a,b) of the translated image is Tij(x)a,b=xa−i,b−jT_{ij}(\bm{x})_{a,b}=x_{a-i,b-j}, and wijw_{ij} is the weight for the loss J(Tij(xadv),y)J(T_{ij}(\bm{x}^{adv}),y). We set i,j∈{−k,...,0,...,k}i,j\in\{-k,...,0,...,k\} with kk being the maximal number of pixels to shift. With this method, the generated adversarial examples are less sensitive to the discriminative regions of the white-box model being attacked, which may be transferred to another model with a higher success rate. We choose the translation operation in this paper rather than other transformations (e.g., rotation, scaling, etc.), because we can develop an efficient algorithm to calculate the gradient of the loss function by the assumption of the translation-invariance in convolutional neural networks.

To solve the optimization problem in Eq. (7), we need to calculate the gradients for (2k+1)2(2k+1)^{2} images, which introduces much more computations. Sampling a small number of translated images for gradient calculation is a feasible way . But we show that we can calculate the gradient for only one image under a mild assumption.

Convolutional neural networks are supposed to have the translation-invariant property , that an object in the input can be recognized in spite of its position. In practice, CNNs are not truly translation-invariant . So we make an assumption that the translation-invariant property is nearly held with very small translations (which is empirically validated in Sec. 4.2). In our problem, we shift the image by no more than 10 pixels along each dimension (i.e., k≤10k\leq 10). Therefore, based on this assumption, the translated image Tij(x)T_{ij}(\bm{x}) is almost the same as x\bm{x} as inputs to the models, as well as their gradients

We then calculate the gradient of the loss function defined in Eq. (7) at a point x^\hat{\bm{x}} as

Given Eq. (9), we do not need to calculate the gradients for (2k+1)2(2k+1)^{2} images. Instead, we only need to get the gradient at the untranslated image x^\hat{\bm{x}} and then average all the shifted gradients. This procedure is equivalent to convolving the gradient with a kernel composed of all the weights wijw_{ij} as

where W\bm{W} is the kernel matrix of size (2k+1)×(2k+1)(2k+1)\times(2k+1), with Wi,j=w−i−jW_{i,j}=w_{-i-j}. We will specify W\bm{W} in the next section.

2.2 Kernel Matrix

There are many options to generate the kernel matrix W\bm{W}. A basic design principle is that the images with bigger shifts should have relatively lower weights to make the adversarial perturbation fool the model at the untranslated image effectively. In this paper, we consider three different choices:

A uniform kernel that Wi,j=\nicefrac1(2k+1)2W_{i,j}=\nicefrac{{1}}{{(2k+1)^{2}}};

We will empirically compare the three kernels in Sec. 4.3.

2.3 Attack Algorithms

Note that in Sec. 3.2.1, we only illustrate how to calculate the gradient of the loss function defined in Eq. (7), but do not specify the update algorithm for generating adversarial examples. This indicates that our method can be integrated into any gradient-based attack method, e.g., FGSM, BIM, MI-FGSM, etc. For gradient-based attack methods presented in Sec. 3.1, in each step we calculate the gradient ∇xJ(xtadv,y)\nabla_{\bm{x}}J(\bm{x}^{adv}_{t},y) at the current solution xtadv\bm{x}^{adv}_{t}, then convolve the gradient with the pre-defined kernel W\bm{W}, and finally obtain the new solution xt+1adv\bm{x}^{adv}_{t+1} following the update rule in different attack methods. For example, the combination of our translation-invariant method and the fast gradient sign method (TI-FGSM) has the following update rule

Also, the integration of the translation-invariant method into the basic iterative method yields the TI-BIM algorithm

The translation-invariant method can be similarly integrated into MI-FGSM and DIM as TI-MI-FGSM and TI-DIM, respectively.

Experiments

In this section, we present the experimental results to demonstrate the effectiveness of the proposed method. We first specify the experimental settings in Sec. 4.1. Then we validate the translation-invariant property of convolutional neural networks in Sec. 4.2. We further conduct two experiments to study the effects of different kernels and size of kernels in Sec. 4.3 and Sec. 4.4. We finally compare the results of the proposed method with baseline methods in Sec. 4.5 and Sec. 4.6.

We use an ImageNet-compatible datasethttps://github.com/tensorflow/cleverhans/tree/master/examples/nips17_adversarial_competition/dataset comprised of 1,000 images to conduct experiments. This dataset was used in the NIPS 2017 adversarial competition. We include eight defense models which are shown to be robust against black-box attacks on the ImageNet dataset. These are

high-level representation guided denoiser (HGD, rank-1 submission in the NIPS 2017 defense competition) ;

input transformation through random resizing and padding (R&P, rank-2 submission in the NIPS 2017 defense competition) ;

input transformation through JPEG compression or total variance minimization (TVM) ;

rank-3 submissionhttps://github.com/anlthms/nips-2017/tree/master/mmd in the NIPS 2017 defense competition (NIPS-r3).

To attack these defenses based on the transferability, we also include four normally trained models—Inception v3 (Inc-v3) , Inception v4 (Inc-v4), Inception ResNet v2 (IncRes-v2) , and ResNet v2-152 (Res-v2-152) , as the white-box models to generate adversarial examples.

In our experiments, we integrate our method into the fast gradient sign method (FGSM) , momentum iterative fast gradient sign method (MI-FGSM) , and diverse inputs method (DIM) . We do not include the basic iterative method and C&W’s method since that they are not good at generating transferable adversarial examples . We denote the attacks combined with our translation-invariant method as TI-FGSM, TI-MI-FGSM, and TI-DIM, respectively.

For the settings of hyper-parameters, we set the maximum perturbation to be ϵ=16\epsilon=16 among all experiments with pixel values in $.Fortheiterativeattackmethods,wesetthenumberofiterationas. For the iterative attack methods, we set the number of iteration as10andthestepsizeasand the step size as\alpha=1.6.ForMI−FGSMandTI−MI−FGSM,weadoptthedefaultdecayfactor. For MI-FGSM and TI-MI-FGSM, we adopt the default decay factor\mu=1.0.ForDIMandTI−DIM,thetransformationprobabilityissetto. For DIM and TI-DIM, the transformation probability is set to0.7$. Please note that the settings for each attack method and its translation-invariant version are the same, because our method is not concerned with the specific attack procedure.

2 Translation-Invariant Property of CNNs

We first verify the translation-invariant property of convolutional neural networks in this section. We use the original 1,000 images from the dataset and shift them by −10-10 to 1010 pixels in each dimension. We input the original images as well as the translated images into Inc-v3, Inc-v4, IncRes-v2, and Res-v2-152, respectively. The loss of each input image is given by the models. We average the loss over all translated images at each position, and show the loss surfaces in Fig. 3.

It can be seen that the loss surfaces are generally smooth with the translations going from −10-10 to 1010 in each dimension. So we could make the assumption that the translation-invariant property is almost held within a small range. In our attacks, the images are shifted by no more than 1010 pixels along each dimension. The loss values would be very similar for the original and translated images. Therefore, we regard that a translated image is almost the same as the corresponding original image as inputs to the models.

3 The Results of Different Kernels

In the section, we show the experimental results of the proposed translation-invariant attack method with different choices of kernels. We attack the Inc-v3 model by TI-FGSM, TI-MI-FGSM, and TI-DIM with three types of kernels, i.e., uniform kernel, linear kernel, and Gaussian kernel, as introduced in Sec. 3.2.2. In Table 1, we report the success rates of black-box attacks against the eight defense models we study, where the success rates are the misclassification rates of the corresponding defense models with the generated adversarial images as inputs.

We can see that for TI-FGSM, the linear kernel leads to better results than the uniform kernel and the Gaussian kernel. And for more powerful attacks such as TI-MI-FGSM and TI-DIM, the Gaussian kernel achieves similar or even better results than the linear kernel. However, both of the linear kernel and the Gaussian kernel are more effective than the uniform kernel. It indicates that we should design the kernel that has lower weights for bigger shifts, as discussed in Sec. 3.2.2. We simply adopt the Gaussian kernel in the following experiments.

4 The Effect of Kernel Size

The size of the kernel W\bm{W} also plays a key role for improving the success rates of black-box attacks. If the kernel size equals to 1×11\times 1, the translation-invariant based attacks degenerate to their vanilla versions. Therefore, we conduct an ablation study to examine the effect of kernel sizes.

We attack the Inc-v3 model by TI-FGSM, TI-MI-FGSM, and TI-DIM with the Gaussian kernel, whose length ranges from 11 to 2121 with a granularity 22. In Fig. 4, we show the success rates against five defense models—IncRes-v2ens, HGD, R&P, TVM, and NIPS-r3. The success rate continues increasing at first, and turns to remain stable after the kernel size exceeds 15×1515\times 15. Therefore, the size of the kernel is set to 15×1515\times 15 in the following.

We also show the adversarial images generated for the Inc-v3 model by TI-FGSM with different kernel sizes in Fig. 5. Due to the smooth effect given by the kernel, we can see that the adversarial perturbations are smoother when using a bigger kernel.

5 Single-Model Attacks

In this section, we compare the black-box success rates of the translation-invariant based attacks with baseline attacks. We first perform adversarial attacks for Inc-v3, Inc-v4, IncRes-v2, and Res-v2-152 respectively using FGSM, MI-FGSM, DIM, and their extensions by combining with the translation-invariant attack method as TI-FGSM, TI-MI-FGSM, and TI-DIM. We adopt the 15×1515\times 15 Gaussian kernel in this set of experiments. We then use the generated adversarial examples to attack the eight defense models we consider based only on the transferability. We report the success rates of black-box attacks in Table 2 for FGSM and TI-FGSM, Table 3 for MI-FGSM and TI-MI-FGSM, and Table 4 for DIM and TI-DIM.

From the tables, we observe that the success rates against the defenses are improved by a large margin when using the proposed method regardless of the attack algorithms or the white-box models being attacked. In general, the translation-invariant based attacks consistently outperform the baseline attacks by 5%∼30%5\%\sim 30\%. In particular, when using TI-DIM, the combination of our method and DIM, to attack the IncRes-v2 model, the resultant adversarial examples have about 60%60\% success rates against the defenses (as shown in Table 4). It demonstrates the vulnerability of the current defenses against black-box attacks. The results also validate the effectiveness of the proposed method. Although we only compare the results of our attack method with baseline methods against the defense models, our attacks remain the success rates of baseline attacks in the white-box setting and the black-box setting against normally trained models, which will be shown in the Appendix.

We show two adversarial images generated for the Inc-v3 model by FGSM and TI-FGSM in Fig. 1. It can be seen that by using TI-FGSM, in which the gradients are convolved by a kernel W\bm{W} before applying to the raw images, the adversarial perturbations are much smoother than those generated by FGSM. The smooth effect also exists in other translation-invariant based attacks.

6 Ensemble-based Attacks

In this section, we further present the results when adversarial examples are generated for an ensemble of models. Liu et al. have shown that attacking multiple models at the same time can improve the transferability of the generated adversarial examples. It is due to that if an example remains adversarial for multiple models, it is more likely to transfer to another black-box model.

We adopt the ensemble method proposed in , which fuses the logit activations of different models. We attack the ensemble of Inc-v3, Inc-v4, IncRes-v2, and Res-v2-152 with equal ensemble weights using FGSM, TI-FGSM, MI-FGSM, TI-MI-FGSM, DIM, and TI-DIM respectively. We also use the 15×1515\times 15 Gaussian kernel in the translation-invariant based attacks.

In Table 5, we show the results of black-box attacks against the eight defenses. The proposed method also improves the success rates across all experiments over the baseline attacks. It should be noted that the adversarial examples generated by TI-DIM can fool the state-of-the-art defenses at an 82%82\% success rate on average based on the transferability. And the adversarial examples are generated for normally trained models unaware of the defense strategies. The results in the paper demonstrate that the current defenses are far from real security, and cannot be deployed in real-world applications.

Conclusion

In this paper, we proposed a translation-invariant attack method to generate adversarial examples that are less sensitive to the discriminative regions of the white-box model being attacked, and have higher transferability against the defense models. Our method optimizes an adversarial image by using a set of translated images. Based on an assumption, our method is efficiently implemented by convolving the gradient with a pre-defined kernel, and can be integrated into any gradient-based attack method. We conducted experiments to validate the effectiveness of the proposed method. Our best attack, TI-DIM, the combination of the proposed translation-invariant method and diverse inputs method , can fool eight state-of-the-art defenses at an 82%82\% success rate on average, where the adversarial examples are generated against four normally trained models. The results identify the vulnerability of the current defenses, and thus raise security issues for the development of more robust deep learning models. We make our codes public at https://github.com/dongyp13/Translation-Invariant-Attacks.

Acknowledgements

This work was supported by the National Key Research and Development Program of China (No. 2017YFA0700904), NSFC Projects (Nos. 61620106010, 61621136008, 61571261), Beijing NSF Project (No. L172037), DITD Program JCKY2017204B064, Tiangong Institute for Intelligent Computing, NVIDIA NVAIL Program, and the projects from Siemens and Intel.

References