Mixup Inference: Better Exploiting Mixup to Defend Adversarial Attacks
Tianyu Pang, Kun Xu, Jun Zhu
Introduction
Deep neural networks (DNNs) have achieved state-of-the-art performance on various tasks (Goodfellow et al., 2016). However, counter-intuitive adversarial examples generally exist in different domains, including computer vision (Szegedy et al., 2014), natural language processing (Jin et al., 2019), reinforcement learning (Huang et al., 2017), speech (Carlini & Wagner, 2018) and graph data (Dai et al., 2018). As DNNs are being widely deployed, it is imperative to improve model robustness and defend adversarial attacks, especially in safety-critical cases. Previous work shows that adversarial examples mainly root from the locally unstable behavior of classifiers on the data manifolds (Goodfellow et al., 2015; Fawzi et al., 2016; 2018; Pang et al., 2018b), where a small adversarial perturbation in the input space can lead to an unreasonable shift in the feature space.
On the one hand, many previous methods try to solve this problem in the inference phase, by introducing transformations on the input images. These attempts include performing local linear transformation like adding Gaussian noise (Tabacof & Valle, 2016), where the processed inputs are kept nearby the original ones, such that the classifiers can maintain high performance on the clean inputs. However, as shown in Fig. 1(a), the equivalent perturbation, i.e., the crafted adversarial perturbation, is still and this strategy is easy to be adaptively evaded since the randomness of w.r.t is local (Athalye et al., 2018). Another category of these attempts is to apply various non-linear transformations, e.g., different operations of image processing (Guo et al., 2018; Xie et al., 2018; Raff et al., 2019). They are usually off-the-shelf for different classifiers, and generally aim to disturb the adversarial perturbations, as shown in Fig. 1(b). Yet these methods are not quite reliable since there is no illustration or guarantee on to what extent they can work.
On the other hand, many efforts have been devoted to improving adversarial robustness in the training phase. For examples, the adversarial training (AT) methods (Madry et al., 2018; Zhang et al., 2019; Shafahi et al., 2019) induce locally stable behavior via data augmentation on adversarial examples. However, AT methods are usually computationally expensive, and will often degenerate model performance on the clean inputs or under general-purpose transformations like rotation (Engstrom et al., 2019). In contrast, the mixup training method (Zhang et al., 2018) introduces globally linear behavior in-between the data manifolds, which can also improve adversarial robustness (Zhang et al., 2018; Verma et al., 2019a). Although this improvement is usually less significant than it resulted by AT methods, mixup-trained models can keep state-of-the-art performance on the clean inputs; meanwhile, the mixup training is computationally more efficient than AT. The interpolated AT method (Lamb et al., 2019) also shows that the mixup mechanism can further benefit the AT methods.
In experiments, we evaluate MI on CIFAR-10 and CIFAR-100 (Krizhevsky & Hinton, 2009) under the oblivious attacks (Carlini & Wagner, 2017) and the adaptive attacks (Athalye et al., 2018). The results demonstrate that our MI method is efficient in defending adversarial attacks in inference, and is also compatible with other variants of mixup, e.g., the interpolated AT method (Lamb et al., 2019). Note that Shimada et al. (2019) also propose to mixup the input points in the test phase, but they do not consider their method from the aspect of adversarial robustness.
Preliminaries
In this section, we first introduce the notations applied in this paper, then we provide the formula of mixup in training. We introduce the adversarial attacks and threat models in Appendix A.1.
Given an input-label pair , a classifier returns the softmax prediction vector and the predicted label , where is the number of classes and . The classifier makes a correct prediction on if . In the adversarial setting, we augment the data pair to a triplet with an extra binary variable , i.e.,
2 Mixup in training
In supervised learning, the most commonly used training mechanism is the empirical risk minimization (ERM) principle (Vapnik, 2013), which minimizes on the training dataset with the loss function . While computationally efficient, ERM could lead to memorization of data (Zhang et al., 2017) and weak adversarial robustness (Szegedy et al., 2014).
Methodology
Although the mixup mechanism has been widely shown to be effective in different domains (Berthelot et al., 2019; Beckham et al., 2019; Verma et al., 2019a; b), most of the previous work only focuses on embedding the mixup mechanism in the training phase, while in the inference phase the global linearity of the trained model is not well exploited. Compared to passively defending adversarial examples by directly classifying them, it would be more effective to actively utilize the globality of mixup-trained models in the inference phase to break the locality of adversarial perturbations.
The above insight inspires us to propose the mixup inference (MI) method, which is a specialized inference principle for the mixup-trained models. In the following, we apply colored {\color[rgb]{0,0,1}y}, {\color[rgb]{1,0,0}\hat{y}} and \color[rgb]{1,.5,0}y_{s} to visually distinguish different notations. Consider an input triplet , where is unknown in advance. When directly feeding into the classifier , we can obtain the predicted label {\color[rgb]{1,0,0}{\hat{y}}}. In the adversarial setting, we are only interested in the cases where is correctly classified by if it is clean, or wrongly classified if it is adversarial (Kurakin et al., 2018). This can be formally denoted as
2 Theoretical analyses
Theoretically, with unlimited capability and sufficient clean samples, a well mixup-trained model can be denoted as a linear function on the convex combinations of clean examples (Hornik et al., 1989; Guo et al., 2019), i.e., and , there is
Specially, we consider the case where the training objective is the cross-entropy loss, then should predict the one-hot vector of label , i.e., . If the input is adversarial, then there should be an extra non-linear part of , since is off the data manifolds. Thus for any input , the prediction vector can be compactly denoted as
Furthermore, according to Eq. (2), there is \bm{1}_{{\color[rgb]{0,0,1}y}={\color[rgb]{1,0,0}\hat{y}}}=\bm{1}_{z=0}. We can represent the {\color[rgb]{1,0,0}\hat{y}}-th components as
MI with predicted label (MI-PL): In this case, the sampling label {\color[rgb]{1,.5,0}y_{s}} is the same as the predicted label {\color[rgb]{1,0,0}\hat{y}}, i.e., p_{s}(y)=\bm{1}_{y={\color[rgb]{1,0,0}\hat{y}}} is a Dirac distribution on {\color[rgb]{1,0,0}\hat{y}}.
MI with other labels (MI-OL): In this case, the label {\color[rgb]{1,.5,0}y_{s}} is uniformly sampled from the labels other than {\color[rgb]{1,0,0}\hat{y}}, i.e., p_{s}(y)=\mathcal{U}_{{\color[rgb]{1,0,0}\hat{y}}}(y) is a discrete uniform distribution on the set \{y\in[L]|y\neq{\color[rgb]{1,0,0}\hat{y}}\}.
We list the simplified formulas of Eq. (7) and Eq. (8) under different cases in Table 1 for clear representation. With the above formulas, we can evaluate how the model performance changes with and without MI by focusing on the formula of
Specifically, in the general-purpose setting where we aim to correctly classify adversarial examples (Madry et al., 2018), we claim that the MI method improves the robustness if the prediction value on the true label {\color[rgb]{0,0,1}y} increases while it on the adversarial label {\color[rgb]{1,0,0}\hat{y}} decreases after performing MI when the input is adversarial (). This can be formally denoted as
We refer to this condition in Eq. (10) as robustness improving condition (RIC). Further, in the detection-purpose setting where we want to detect the hidden variable and filter out adversarial inputs, we can take the gap of the {\color[rgb]{1,0,0}\hat{y}}-th component of predictions before and after the MI operation, i.e., \Delta F_{{\color[rgb]{1,0,0}\hat{y}}}(x;p_{s}) as the detection metric (Pang et al., 2018a). To formally measure the detection ability on , we use the detection gap (DG), denoted as
A higher value of DG indicates that \Delta F_{{\color[rgb]{1,0,0}\hat{y}}}(x;p_{s}) is better as a detection metric. In the following sections, we specifically analyze the properties of different versions of MI according to Table 1, and we will see that the MI methods can be used and benefit in different defense strategies.
General-purpose defense: If MI-PL can improve the general-purpose robustness, it should satisfy RIC in Eq. (10). By simple derivation and the results of Table 1, this means that
The second mechanism is perturbation shrinkage, where the original perturbation shrinks by a factor . This equivalently shrinks the perturbation threshold since , which means that MI generally imposes a tighter upper bound on the potential attack ability for a crafted perturbation. Besides, empirical results in previous work also show that a smaller perturbation threshold largely weakens the effect of attacks (Kurakin et al., 2018). Therefore, if an adversarial attack defended by these two mechanisms leads to a prediction degradation as in Eq. (12), then applying MI-PL would improve the robustness against this adversarial attack. Similar properties also hold for MI-OL as described in Sec. 3.2.2. In Fig. 2, we empirically demonstrate that most of the existing adversarial attacks, e.g., the PGD attack (Madry et al., 2018) satisfies these properties.
Detection-purpose defense: According to Eq. (11), the formula of DG for MI-PL is
By comparing Eq. (12) and Eq. (14), we can find that they are consistent with each other, which means that for a given adversarial attack, if MI-PL can better defend it in general-purpose, then ideally MI-PL can also better detect the crafted adversarial examples.
2.2 Mixup inference with other labels
Note that the conditions in Eq. (15) is strictly looser than Eq. (12), which means MI-OL can defend broader range of attacks than MI-PL, as verified in Fig. 2.
Detection-purpose defense: According to Eq. (11) and Table 1, the DG for MI-OL is
Experiments
In this section, we provide the experimental results on CIFAR-10 and CIFAR-100 (Krizhevsky & Hinton, 2009) to demonstrate the effectiveness of our MI methods on defending adversarial attacks. Our codes are available at https://github.com/P2333/Mixup-Inference.
In training, we use ResNet-50 (He et al., 2016) and apply the momentum SGD optimizer (Qian, 1999) on both CIFAR-10 and CIFAR-100. We run the training for epochs with the batch size of . The initial learning rate is for ERM, mixup and AT; for interpolated AT (Lamb et al., 2019). The learning rate decays with a factor of at and epochs. The attack method for AT and interpolated AT is untargeted PGD-10 with and step size (Madry et al., 2018), and the ratio of the clean examples and the adversarial ones in each mini-batch is (Lamb et al., 2019). The hyperparameter for mixup and interpolated AT is (Zhang et al., 2018). All defenses with randomness are executed times to obtain the averaged predictions (Xie et al., 2018).
2 Empirical verification of theoretical analyses
To verify and illustrate our theoretical analyses in Sec. 3, we provide the empirical relationship between the output predictions of MI and the hyperparameter in Fig. 2. The notations and formulas annotated in Fig. 2 correspond to those introduced in Sec. 3. We can see that the results follow our theoretical conclusions under the assumption of ideal global linearity. Besides, both MI-PL and MI-OL empirically satisfy RIC in this case, which indicates that they can improve robustness under the untargeted PGD-10 attack on CIFAR-10, as quantitatively demonstrated in the following sections.
3 Performance under oblivious attacks
In this subsection, we evaluate the performance of our method under the oblivious-box attacks (Carlini & Wagner, 2017). The oblivious threat model assumes that the adversary is not aware of the existence of the defense mechanism, e.g., MI, and generate adversarial examples based on the unsecured classification model. We separately apply the model trained by mixup and interpolated AT as the classification model. The AUC scores for the detection-purpose defense are given in Fig. 3(a). The results show that applying MI-PL in inference can better detect adversarial attacks, while directly detecting by the returned confidence without MI-PL performs even worse than a random guess.
We also compare MI with previous general-purpose defenses applied in the inference phase, e.g., adding Gaussian noise or random rotation (Tabacof & Valle, 2016); performing random padding or resizing after random cropping (Guo et al., 2018; Xie et al., 2018). The performance of our method and baselines on CIFAR-10 and CIFAR-100 are reported in Table 2 and Table 3, respectively. Since for each defense method, there is a trade-off between the accuracy on clean samples and adversarial samples depending on the hyperparameters, e.g., the standard deviation for Gaussian noise, we carefully select the hyperparameters to ensure both our method and baselines keep a similar performance on clean data for fair comparisons. The hyperparameters used in our method and baselines are reported in Table 4 and Table 5. In Fig. 3(b), we further explore this trade-off by grid searching the hyperparameter space for each defense to demonstrate the superiority of our method.
As shown in these results, our MI method can significantly improve the robustness for the trained models with induced global linearity, and is compatible with training-phase defenses like the interpolated AT method. As a practical strategy, we also evaluate a variant of MI, called MI-Combined, which applies MI-OL if the input is detected as adversarial by MI-PL with a default detection threshold; otherwise returns the prediction on the original input. We also perform ablation studies of ERM / AT + MI-OL in Table 2, where no global linearity is induced. The results verify that our MI methods indeed exploit the global linearity of the mixup-trained models, rather than simply introduce randomness.
4 Performance under white-box adaptive attacks
Following Athalye et al. (2018), we test our method under the white-box adaptive attacks (detailed in Appendix B.2). Since we mainly adopt the PGD attack framework, which synthesizes adversarial examples iteratively, the adversarial noise will be clipped to make the input image stay within the valid range. It results in the fact that with mixup on different training examples, the adversarial perturbation will be clipped differently. To address this issue, we average the generated perturbations over the adaptive samples as the final perturbation. The results of the adversarial accuracy w.r.t the number of adaptive samples are shown in Fig. 4. We can see that even under a strong adaptive attack, equipped with MI can still improve the robustness for the classification models.
Conclusion
In this paper, we propose the MI method, which is specialized for the trained models with globally linear behaviors induced by, e.g., mixup or interpolated AT. As analyzed in Sec. 3, MI can exploit this induced global linearity in the inference phase to shrink and transfer the adversarial perturbation, which breaks the locality of adversarial attacks and alleviate their aggressivity. In experiments, we empirically verify that applying MI can return more reliable predictions under different threat models.
Acknowledgements
This work was supported by the National Key Research and Development Program of China (No. 2017YFA0700904), NSFC Projects (Nos. 61620106010, U19B2034, U1811461), Beijing NSF Project (No. L172037), Beijing Academy of Artificial Intelligence (BAAI), Tsinghua-Huawei Joint Research Program, a grant from Tsinghua Institute for Guo Qiang, Tiangong Institute for Intelligent Computing, the JP Morgan Faculty Research Program and the NVIDIA NVAIL Program with GPU/DGX Acceleration.
References
Appendix A More backgrounds
In this section, we provide more backgrounds which are related to our work in the main text.
Adversarial attacks. Although deep learning methods have achieved substantial success in different domains (Goodfellow et al., 2016), human imperceptible adversarial perturbations can be easily crafted to fool high-performance models, e.g., deep neural networks (DNNs) (Nguyen et al., 2015).
One of the most commonly studied adversarial attack is the projected gradient descent (PGD) method (Madry et al., 2018). Let be the number of iteration steps, be the original clean example, then PGD iteratively crafts the adversarial example as
where is the clipping function. Here is a randomly perturbed image in the neighborhood of , i.e., , and the finally returned adversarial example is , following our notations in the main text.
Threat models. Here we introduce different threat models in the adversarial setting. As suggested in Carlini et al. (2019), a threat model includes a set of assumptions about the adversary’s goals, capabilities, and knowledge.
Adversary’s goals could be simply fooling the classifiers to misclassify, which is referred to as untargeted mode. Alternatively, the goals can be more specific to make the model misclassify certain examples from a source class into a target class, which is referred to as targeted mode. In our experiments, we evaluate under both modes, as shown in Table 2 and Table 3.
Adversary’s knowledge describes what knowledge the adversary is assumed to have. Typically, there are three settings when evaluating a defense method:
Oblivious adversaries are not aware of the existence of the defense and generate adversarial examples based on the unsecured classification model (Carlini & Wagner, 2017).
White-box adversaries know the scheme and parameters of , and can design adaptive methods to attack both the model and the defense simultaneously (Athalye et al., 2018).
Black-box adversaries have no access to the parameters of the defense or the model with varying degrees of black-box access (Dong et al., 2018).
In our experiments, we mainly test under the oblivious setting (Sec. 4.3) and white-box setting (Sec. 4.4), since previous work has already demonstrated that randomness itself is efficient on defending black-box attacks (Guo et al., 2018; Xie et al., 2018).
A.2 Interpolated adversarial training
To date, the most widely applied framework for adversarial training (AT) methods is the saddle point framework introduced in Madry et al. (2018):
Here represents the trainable parameters in the classifier , and is a set of allowed perturbations. In implementation, the inner maximization problem for each input-label pair is approximately solved by, e.g., the PGD method with different random initialization (Madry et al., 2018).
As a variant of the AT method, Lamb et al. (2019) propose the interpolated AT method, which combines AT with mixup. Interpolated AT trains on interpolations of adversarial examples along with interpolations of unperturbed examples (cf. Alg. 1 in Lamb et al. (2019)). Previous empirical results demonstrate that interpolated AT can obtain higher accuracy on the clean inputs compared to the AT method without mixup, while keeping the similar performance of robustness.
Appendix B Technical details
We provide more technical details about our method and the implementation of the experiments.
Generality. According to Sec. 3, except for the mixup-trained models, the MI method is generally compatible with any trained model with induced global linearity. These models could be trained by other methods, e.g., manifold mixup (Verma et al., 2019a; Inoue, 2018; Lamb et al., 2019). Besides, to better defend white-box adaptive attacks, the mixup ratio in MI could also be sampled from certain distribution to put in additional randomness.
Empirical gap. As demonstrated in Fig. 2, there is a gap between the empirical results and the theoretical formulas in Table 1. This is because that the mixup mechanism mainly acts as a regularization in training, which means the induced global linearity may not satisfy the expected behaviors. To improve the performance of MI, a stronger regularization can be imposed, e.g., training with mixup for more epochs, or applying matched both in training and inference.
B.2 Adaptive attacks for mixup inference
Following Athalye et al. (2018), we design the adaptive attacks for our MI method. Specifically, according to Eq. (6), the expected model prediction returned by MI is:
Note that generally the in MI comes from certain distribution. For simplicity, we fix as a hyperparameter in our implementation. Therefore, the gradients of the prediction w.r.t. the input is:
In the implementation of adaptive PGD attacks, we first sample a series of examples , where is the number of adaptive samples in Fig. 3. Then according to Eq. (18), the sign of gradients used in adaptive PGD can be approximated by
B.3 Hyperparameter settings
The hyperparameter settings of the experiments shown in Table 2 and Table 3 are provided in Table 4 and Table 5, respectively. Since the original methods in Xie et al. (2018) and Guo et al. (2018) are both designed for the models on ImageNet, we adapt them for CIFAR-10 and CIFAR-100. Most of our experiments are conducted on the NVIDIA DGX-1 server with eight Tesla P100 GPUs.