Mixup Inference: Better Exploiting Mixup to Defend Adversarial Attacks

Tianyu Pang, Kun Xu, Jun Zhu

Introduction

Deep neural networks (DNNs) have achieved state-of-the-art performance on various tasks (Goodfellow et al., 2016). However, counter-intuitive adversarial examples generally exist in different domains, including computer vision (Szegedy et al., 2014), natural language processing (Jin et al., 2019), reinforcement learning (Huang et al., 2017), speech (Carlini & Wagner, 2018) and graph data (Dai et al., 2018). As DNNs are being widely deployed, it is imperative to improve model robustness and defend adversarial attacks, especially in safety-critical cases. Previous work shows that adversarial examples mainly root from the locally unstable behavior of classifiers on the data manifolds (Goodfellow et al., 2015; Fawzi et al., 2016; 2018; Pang et al., 2018b), where a small adversarial perturbation in the input space can lead to an unreasonable shift in the feature space.

On the one hand, many previous methods try to solve this problem in the inference phase, by introducing transformations on the input images. These attempts include performing local linear transformation like adding Gaussian noise (Tabacof & Valle, 2016), where the processed inputs are kept nearby the original ones, such that the classifiers can maintain high performance on the clean inputs. However, as shown in Fig. 1(a), the equivalent perturbation, i.e., the crafted adversarial perturbation, is still δ\delta and this strategy is easy to be adaptively evaded since the randomness of x‾0\overline{x}_{0} w.r.t x0x_{0} is local (Athalye et al., 2018). Another category of these attempts is to apply various non-linear transformations, e.g., different operations of image processing (Guo et al., 2018; Xie et al., 2018; Raff et al., 2019). They are usually off-the-shelf for different classifiers, and generally aim to disturb the adversarial perturbations, as shown in Fig. 1(b). Yet these methods are not quite reliable since there is no illustration or guarantee on to what extent they can work.

On the other hand, many efforts have been devoted to improving adversarial robustness in the training phase. For examples, the adversarial training (AT) methods (Madry et al., 2018; Zhang et al., 2019; Shafahi et al., 2019) induce locally stable behavior via data augmentation on adversarial examples. However, AT methods are usually computationally expensive, and will often degenerate model performance on the clean inputs or under general-purpose transformations like rotation (Engstrom et al., 2019). In contrast, the mixup training method (Zhang et al., 2018) introduces globally linear behavior in-between the data manifolds, which can also improve adversarial robustness (Zhang et al., 2018; Verma et al., 2019a). Although this improvement is usually less significant than it resulted by AT methods, mixup-trained models can keep state-of-the-art performance on the clean inputs; meanwhile, the mixup training is computationally more efficient than AT. The interpolated AT method (Lamb et al., 2019) also shows that the mixup mechanism can further benefit the AT methods.

In experiments, we evaluate MI on CIFAR-10 and CIFAR-100 (Krizhevsky & Hinton, 2009) under the oblivious attacks (Carlini & Wagner, 2017) and the adaptive attacks (Athalye et al., 2018). The results demonstrate that our MI method is efficient in defending adversarial attacks in inference, and is also compatible with other variants of mixup, e.g., the interpolated AT method (Lamb et al., 2019). Note that Shimada et al. (2019) also propose to mixup the input points in the test phase, but they do not consider their method from the aspect of adversarial robustness.

Preliminaries

In this section, we first introduce the notations applied in this paper, then we provide the formula of mixup in training. We introduce the adversarial attacks and threat models in Appendix A.1.

Given an input-label pair (x,y)(x,y), a classifier FF returns the softmax prediction vector F(x)F(x) and the predicted label y^=arg max⁡j∈[L]Fj(x){\hat{y}}=\operatorname*{arg\,max}_{j\in[L]}F_{j}(x), where LL is the number of classes and [L]={1,⋯ ,L}[L]=\{1,\cdots,L\}. The classifier FF makes a correct prediction on xx if y=y^y={\hat{y}}. In the adversarial setting, we augment the data pair (x,y)(x,y) to a triplet (x,y,z)(x,y,z) with an extra binary variable zz, i.e.,

2 Mixup in training

In supervised learning, the most commonly used training mechanism is the empirical risk minimization (ERM) principle (Vapnik, 2013), which minimizes 1n∑i=1nL(F(xi),yi)\frac{1}{n}\sum_{i=1}^{n}\mathcal{L}(F(x_{i}),y_{i}) on the training dataset D={(xi,yi)}i=1n\mathcal{D}=\{(x_{i},y_{i})\}_{i=1}^{n} with the loss function L\mathcal{L}. While computationally efficient, ERM could lead to memorization of data (Zhang et al., 2017) and weak adversarial robustness (Szegedy et al., 2014).

Methodology

Although the mixup mechanism has been widely shown to be effective in different domains (Berthelot et al., 2019; Beckham et al., 2019; Verma et al., 2019a; b), most of the previous work only focuses on embedding the mixup mechanism in the training phase, while in the inference phase the global linearity of the trained model is not well exploited. Compared to passively defending adversarial examples by directly classifying them, it would be more effective to actively utilize the globality of mixup-trained models in the inference phase to break the locality of adversarial perturbations.

The above insight inspires us to propose the mixup inference (MI) method, which is a specialized inference principle for the mixup-trained models. In the following, we apply colored {\color[rgb]{0,0,1}y}, {\color[rgb]{1,0,0}\hat{y}} and \color[rgb]{1,.5,0}y_{s} to visually distinguish different notations. Consider an input triplet (x,y,z)(x,y,z), where zz is unknown in advance. When directly feeding xx into the classifier FF, we can obtain the predicted label {\color[rgb]{1,0,0}{\hat{y}}}. In the adversarial setting, we are only interested in the cases where xx is correctly classified by FF if it is clean, or wrongly classified if it is adversarial (Kurakin et al., 2018). This can be formally denoted as

2 Theoretical analyses

Theoretically, with unlimited capability and sufficient clean samples, a well mixup-trained model FF can be denoted as a linear function HH on the convex combinations of clean examples (Hornik et al., 1989; Guo et al., 2019), i.e., ∀xi,xj∼p(x)\forall x_{i},x_{j}\sim p(x) and λ∈\lambda\in, there is

Specially, we consider the case where the training objective L\mathcal{L} is the cross-entropy loss, then H(xi)H(x_{i}) should predict the one-hot vector of label yiy_{i}, i.e., Hy(xi)=1y=yiH_{y}(x_{i})=\bm{1}_{y=y_{i}}. If the input x=x0+δx=x_{0}+\delta is adversarial, then there should be an extra non-linear part G(δ;x0)G(\delta;x_{0}) of FF, since xx is off the data manifolds. Thus for any input xx, the prediction vector can be compactly denoted as

Furthermore, according to Eq. (2), there is \bm{1}_{{\color[rgb]{0,0,1}y}={\color[rgb]{1,0,0}\hat{y}}}=\bm{1}_{z=0}. We can represent the {\color[rgb]{1,0,0}\hat{y}}-th components as

MI with predicted label (MI-PL): In this case, the sampling label {\color[rgb]{1,.5,0}y_{s}} is the same as the predicted label {\color[rgb]{1,0,0}\hat{y}}, i.e., p_{s}(y)=\bm{1}_{y={\color[rgb]{1,0,0}\hat{y}}} is a Dirac distribution on {\color[rgb]{1,0,0}\hat{y}}.

MI with other labels (MI-OL): In this case, the label {\color[rgb]{1,.5,0}y_{s}} is uniformly sampled from the labels other than {\color[rgb]{1,0,0}\hat{y}}, i.e., p_{s}(y)=\mathcal{U}_{{\color[rgb]{1,0,0}\hat{y}}}(y) is a discrete uniform distribution on the set \{y\in[L]|y\neq{\color[rgb]{1,0,0}\hat{y}}\}.

We list the simplified formulas of Eq. (7) and Eq. (8) under different cases in Table 1 for clear representation. With the above formulas, we can evaluate how the model performance changes with and without MI by focusing on the formula of

Specifically, in the general-purpose setting where we aim to correctly classify adversarial examples (Madry et al., 2018), we claim that the MI method improves the robustness if the prediction value on the true label {\color[rgb]{0,0,1}y} increases while it on the adversarial label {\color[rgb]{1,0,0}\hat{y}} decreases after performing MI when the input is adversarial (z=1z=1). This can be formally denoted as

We refer to this condition in Eq. (10) as robustness improving condition (RIC). Further, in the detection-purpose setting where we want to detect the hidden variable zz and filter out adversarial inputs, we can take the gap of the {\color[rgb]{1,0,0}\hat{y}}-th component of predictions before and after the MI operation, i.e., \Delta F_{{\color[rgb]{1,0,0}\hat{y}}}(x;p_{s}) as the detection metric (Pang et al., 2018a). To formally measure the detection ability on zz, we use the detection gap (DG), denoted as

A higher value of DG indicates that \Delta F_{{\color[rgb]{1,0,0}\hat{y}}}(x;p_{s}) is better as a detection metric. In the following sections, we specifically analyze the properties of different versions of MI according to Table 1, and we will see that the MI methods can be used and benefit in different defense strategies.

General-purpose defense: If MI-PL can improve the general-purpose robustness, it should satisfy RIC in Eq. (10). By simple derivation and the results of Table 1, this means that

The second mechanism is perturbation shrinkage, where the original perturbation δ\delta shrinks by a factor λ\lambda. This equivalently shrinks the perturbation threshold since ∥λδ∥p=λ∥δ∥p≤λϵ\|\lambda\delta\|_{p}=\lambda\|\delta\|_{p}\leq\lambda\epsilon, which means that MI generally imposes a tighter upper bound on the potential attack ability for a crafted perturbation. Besides, empirical results in previous work also show that a smaller perturbation threshold largely weakens the effect of attacks (Kurakin et al., 2018). Therefore, if an adversarial attack defended by these two mechanisms leads to a prediction degradation as in Eq. (12), then applying MI-PL would improve the robustness against this adversarial attack. Similar properties also hold for MI-OL as described in Sec. 3.2.2. In Fig. 2, we empirically demonstrate that most of the existing adversarial attacks, e.g., the PGD attack (Madry et al., 2018) satisfies these properties.

Detection-purpose defense: According to Eq. (11), the formula of DG for MI-PL is

By comparing Eq. (12) and Eq. (14), we can find that they are consistent with each other, which means that for a given adversarial attack, if MI-PL can better defend it in general-purpose, then ideally MI-PL can also better detect the crafted adversarial examples.

2.2 Mixup inference with other labels

Note that the conditions in Eq. (15) is strictly looser than Eq. (12), which means MI-OL can defend broader range of attacks than MI-PL, as verified in Fig. 2.

Detection-purpose defense: According to Eq. (11) and Table 1, the DG for MI-OL is

Experiments

In this section, we provide the experimental results on CIFAR-10 and CIFAR-100 (Krizhevsky & Hinton, 2009) to demonstrate the effectiveness of our MI methods on defending adversarial attacks. Our codes are available at https://github.com/P2333/Mixup-Inference.

In training, we use ResNet-50 (He et al., 2016) and apply the momentum SGD optimizer (Qian, 1999) on both CIFAR-10 and CIFAR-100. We run the training for 200200 epochs with the batch size of 6464. The initial learning rate is 0.010.01 for ERM, mixup and AT; 0.10.1 for interpolated AT (Lamb et al., 2019). The learning rate decays with a factor of 0.10.1 at 100100 and 150150 epochs. The attack method for AT and interpolated AT is untargeted PGD-10 with ϵ=8/255\epsilon=8/255 and step size 2/2552/255 (Madry et al., 2018), and the ratio of the clean examples and the adversarial ones in each mini-batch is 1:11:1 (Lamb et al., 2019). The hyperparameter α\alpha for mixup and interpolated AT is 1.01.0 (Zhang et al., 2018). All defenses with randomness are executed 3030 times to obtain the averaged predictions (Xie et al., 2018).

2 Empirical verification of theoretical analyses

To verify and illustrate our theoretical analyses in Sec. 3, we provide the empirical relationship between the output predictions of MI and the hyperparameter λ\lambda in Fig. 2. The notations and formulas annotated in Fig. 2 correspond to those introduced in Sec. 3. We can see that the results follow our theoretical conclusions under the assumption of ideal global linearity. Besides, both MI-PL and MI-OL empirically satisfy RIC in this case, which indicates that they can improve robustness under the untargeted PGD-10 attack on CIFAR-10, as quantitatively demonstrated in the following sections.

3 Performance under oblivious attacks

In this subsection, we evaluate the performance of our method under the oblivious-box attacks (Carlini & Wagner, 2017). The oblivious threat model assumes that the adversary is not aware of the existence of the defense mechanism, e.g., MI, and generate adversarial examples based on the unsecured classification model. We separately apply the model trained by mixup and interpolated AT as the classification model. The AUC scores for the detection-purpose defense are given in Fig. 3(a). The results show that applying MI-PL in inference can better detect adversarial attacks, while directly detecting by the returned confidence without MI-PL performs even worse than a random guess.

We also compare MI with previous general-purpose defenses applied in the inference phase, e.g., adding Gaussian noise or random rotation (Tabacof & Valle, 2016); performing random padding or resizing after random cropping (Guo et al., 2018; Xie et al., 2018). The performance of our method and baselines on CIFAR-10 and CIFAR-100 are reported in Table 2 and Table 3, respectively. Since for each defense method, there is a trade-off between the accuracy on clean samples and adversarial samples depending on the hyperparameters, e.g., the standard deviation for Gaussian noise, we carefully select the hyperparameters to ensure both our method and baselines keep a similar performance on clean data for fair comparisons. The hyperparameters used in our method and baselines are reported in Table 4 and Table 5. In Fig. 3(b), we further explore this trade-off by grid searching the hyperparameter space for each defense to demonstrate the superiority of our method.

As shown in these results, our MI method can significantly improve the robustness for the trained models with induced global linearity, and is compatible with training-phase defenses like the interpolated AT method. As a practical strategy, we also evaluate a variant of MI, called MI-Combined, which applies MI-OL if the input is detected as adversarial by MI-PL with a default detection threshold; otherwise returns the prediction on the original input. We also perform ablation studies of ERM / AT + MI-OL in Table 2, where no global linearity is induced. The results verify that our MI methods indeed exploit the global linearity of the mixup-trained models, rather than simply introduce randomness.

4 Performance under white-box adaptive attacks

Following Athalye et al. (2018), we test our method under the white-box adaptive attacks (detailed in Appendix B.2). Since we mainly adopt the PGD attack framework, which synthesizes adversarial examples iteratively, the adversarial noise will be clipped to make the input image stay within the valid range. It results in the fact that with mixup on different training examples, the adversarial perturbation will be clipped differently. To address this issue, we average the generated perturbations over the adaptive samples as the final perturbation. The results of the adversarial accuracy w.r.t the number of adaptive samples are shown in Fig. 4. We can see that even under a strong adaptive attack, equipped with MI can still improve the robustness for the classification models.

Conclusion

In this paper, we propose the MI method, which is specialized for the trained models with globally linear behaviors induced by, e.g., mixup or interpolated AT. As analyzed in Sec. 3, MI can exploit this induced global linearity in the inference phase to shrink and transfer the adversarial perturbation, which breaks the locality of adversarial attacks and alleviate their aggressivity. In experiments, we empirically verify that applying MI can return more reliable predictions under different threat models.

Acknowledgements

This work was supported by the National Key Research and Development Program of China (No. 2017YFA0700904), NSFC Projects (Nos. 61620106010, U19B2034, U1811461), Beijing NSF Project (No. L172037), Beijing Academy of Artificial Intelligence (BAAI), Tsinghua-Huawei Joint Research Program, a grant from Tsinghua Institute for Guo Qiang, Tiangong Institute for Intelligent Computing, the JP Morgan Faculty Research Program and the NVIDIA NVAIL Program with GPU/DGX Acceleration.

References

Appendix A More backgrounds

In this section, we provide more backgrounds which are related to our work in the main text.

Adversarial attacks. Although deep learning methods have achieved substantial success in different domains (Goodfellow et al., 2016), human imperceptible adversarial perturbations can be easily crafted to fool high-performance models, e.g., deep neural networks (DNNs) (Nguyen et al., 2015).

One of the most commonly studied adversarial attack is the projected gradient descent (PGD) method (Madry et al., 2018). Let rr be the number of iteration steps, x0x_{0} be the original clean example, then PGD iteratively crafts the adversarial example as

where clip⁡x,ϵ(⋅)\operatorname{clip}_{x,\epsilon}(\cdot) is the clipping function. Here x0∗x_{0}^{*} is a randomly perturbed image in the neighborhood of x0x_{0}, i.e., U˚(x0,ϵ)\mathring{U}(x_{0},\epsilon), and the finally returned adversarial example is x=xr∗=x0+δx=x^{*}_{r}=x_{0}+\delta, following our notations in the main text.

Threat models. Here we introduce different threat models in the adversarial setting. As suggested in Carlini et al. (2019), a threat model includes a set of assumptions about the adversary’s goals, capabilities, and knowledge.

Adversary’s goals could be simply fooling the classifiers to misclassify, which is referred to as untargeted mode. Alternatively, the goals can be more specific to make the model misclassify certain examples from a source class into a target class, which is referred to as targeted mode. In our experiments, we evaluate under both modes, as shown in Table 2 and Table 3.

Adversary’s knowledge describes what knowledge the adversary is assumed to have. Typically, there are three settings when evaluating a defense method:

Oblivious adversaries are not aware of the existence of the defense DD and generate adversarial examples based on the unsecured classification model FF (Carlini & Wagner, 2017).

White-box adversaries know the scheme and parameters of DD, and can design adaptive methods to attack both the model FF and the defense DD simultaneously (Athalye et al., 2018).

Black-box adversaries have no access to the parameters of the defense DD or the model FF with varying degrees of black-box access (Dong et al., 2018).

In our experiments, we mainly test under the oblivious setting (Sec. 4.3) and white-box setting (Sec. 4.4), since previous work has already demonstrated that randomness itself is efficient on defending black-box attacks (Guo et al., 2018; Xie et al., 2018).

A.2 Interpolated adversarial training

To date, the most widely applied framework for adversarial training (AT) methods is the saddle point framework introduced in Madry et al. (2018):

Here θ\theta represents the trainable parameters in the classifier FF, and SS is a set of allowed perturbations. In implementation, the inner maximization problem for each input-label pair (x,y)(x,y) is approximately solved by, e.g., the PGD method with different random initialization (Madry et al., 2018).

As a variant of the AT method, Lamb et al. (2019) propose the interpolated AT method, which combines AT with mixup. Interpolated AT trains on interpolations of adversarial examples along with interpolations of unperturbed examples (cf. Alg. 1 in Lamb et al. (2019)). Previous empirical results demonstrate that interpolated AT can obtain higher accuracy on the clean inputs compared to the AT method without mixup, while keeping the similar performance of robustness.

Appendix B Technical details

We provide more technical details about our method and the implementation of the experiments.

Generality. According to Sec. 3, except for the mixup-trained models, the MI method is generally compatible with any trained model with induced global linearity. These models could be trained by other methods, e.g., manifold mixup (Verma et al., 2019a; Inoue, 2018; Lamb et al., 2019). Besides, to better defend white-box adaptive attacks, the mixup ratio λ\lambda in MI could also be sampled from certain distribution to put in additional randomness.

Empirical gap. As demonstrated in Fig. 2, there is a gap between the empirical results and the theoretical formulas in Table 1. This is because that the mixup mechanism mainly acts as a regularization in training, which means the induced global linearity may not satisfy the expected behaviors. To improve the performance of MI, a stronger regularization can be imposed, e.g., training with mixup for more epochs, or applying matched λ\lambda both in training and inference.

B.2 Adaptive attacks for mixup inference

Following Athalye et al. (2018), we design the adaptive attacks for our MI method. Specifically, according to Eq. (6), the expected model prediction returned by MI is:

Note that generally the λ\lambda in MI comes from certain distribution. For simplicity, we fix λ\lambda as a hyperparameter in our implementation. Therefore, the gradients of the prediction w.r.t. the input xx is:

In the implementation of adaptive PGD attacks, we first sample a series of examples {xs,k}k=1NA\{x_{s,k}\}_{k=1}^{N_{A}}, where NAN_{A} is the number of adaptive samples in Fig. 3. Then according to Eq. (18), the sign of gradients used in adaptive PGD can be approximated by

B.3 Hyperparameter settings

The hyperparameter settings of the experiments shown in Table 2 and Table 3 are provided in Table 4 and Table 5, respectively. Since the original methods in Xie et al. (2018) and Guo et al. (2018) are both designed for the models on ImageNet, we adapt them for CIFAR-10 and CIFAR-100. Most of our experiments are conducted on the NVIDIA DGX-1 server with eight Tesla P100 GPUs.