Attacks Which Do Not Kill Training Make Adversarial Learning Stronger
Jingfeng Zhang, Xilie Xu, Bo Han, Gang Niu, Lizhen Cui, Masashi Sugiyama, Mohan Kankanhalli
Introduction
Existing empirical defense methods formulate the adversarial training as a minimax optimization problem (Section 2.1) (Madry et al., 2018). To conduct this minimax optimization, projected gradient descent (PGD) is a common method to generate the most adversarial data that maximizes the loss, updating the current model. PGD perturbs the natural data for a fixed number of steps with small step size. After each step of perturbation, PGD projects the adversarial data back onto the -norm ball of the natural data.
However, this minimax formulation is conservative (or even pessimistic), such that it sometimes hurts the natural generalization (Tsipras et al., 2019). For example, the top panels in Figure 1 show that at step to in PGD, the adversarial variants of the natural data significantly cross over the decision boundary and are located at their peer’s (natural data) area. Since adversarial training aims to fit natural data and its adversarial variants simultaneously, such the cross-over mixture makes adversarial training extremely difficult. Therefore, the most adversarial data generated by the PGD-10 (i.e., step 10 in top panel of Figure 1) directly “kill” the training, thus rendering the training unsuccessful.
Inspired by philosopher Friedrich Nietzsche’s quote “that which does not kill us makes us stronger,” we propose friendly adversarial training (FAT): rather than employing the most adversarial data for updating the current model, we search for the friendly adversarial data minimizing the loss. The friendly adversarial data are confidently misclassified by the current model. We design the learning objective of FAT and theoretically justify it by deriving an upper bound of the adversarial risk (Section 3). Essentially, FAT updates the current model using friendly adversarial data. FAT trains a DNN using the wrongly-predicted adversarial data minimizing the loss and the correctly-predicted adversarial data maximizing the loss.
FAT is a reasonable strategy due to two reasons: It removes the existing inconsistency between attack and defense, and it adheres to the spirit of curriculum learning. First, the ways of generating adversarial data by adversarial attackers and adversarial defense methods are inconsistent. Adversarial attacks (Szegedy et al., 2014; Carlini & Wagner, 2017; Athalye et al., 2018) aim to find the adversarial data (not maximizing the loss) to confidently fool the model. On the other hand, existing adversarial defense methods generate the most adversarial data maximizing the loss regardless of the model’s predictions. These two should be harmonized. Second, the curriculum learning strategy has been shown to be effective (Bengio et al., 2009). Fitting most adversarial data initially makes the learning extremely difficult, sometimes even killing the training. Instead, FAT learns initially from the least adversarial data and progressively utilizes increasingly adversarial data.
FAT is easy to implement by just early stopping most-adversarial-data searching algorithms such as PGD, which we call early-stopped PGD (Section 4.1). Once adversarial data is misclassified by the current model, we stop the PGD iterations early. Early-stopped PGD has the benefit of alleviating the cross-over mixture problem. For example, as shown in the bottom panels of Figure 1, adversarial data generated by early-stopped PGD will not be located at their peer areas (extensive details in Section 4.2). Thus, it will not hurt the generalization ability much. In addition, FAT based on early-stopped PGD progressively employs stronger and stronger adversarial data (with more PGD steps), engendering increasingly enhanced robustness of the model over the training progression (Section 5.3). This implies that attacks that do not kill the training indeed make the adversarial learning stronger.
A brief overview of our contributions is as follows. We propose a novel formulation for adversarial learning (Section 3.1) and theoretically justify it by an upper bound of the adversarial risk (Section 3.2). Our FAT approximately realizes this formulation by just stopping PGD early. FAT has the following benefits.
Conventional adversarial training methods, e.g., standard adversarial training (Madry et al., 2018), TRADES (Zhang et al., 2019b) and MART (Wang et al., 2020), can be easily modified to become friendly adversarial training counterparts, i.e., FAT, FAT for TRADES, and FAT for MART (Section 5).
Compared with conventional adversarial training, FAT has a better standard accuracy for natural data, while keeping a competitively robust accuracy for adversarial data (Sections 5.1 and 6.2).
FAT is computationally efficient because the early stopped PGD saves a large number of backward propagations for searching adversarial data (Section 5.2).
FAT can enable larger values of the perturbation bound, i.e., (Section 6.1), due to that FAT can alleviate the cross-over mixture problem (Section 4.2).
With these benefits, FAT allows us to answer that adversarial robustness can indeed be achieved without compromising the natural generalization.
Standard Adversarial Training
Let be the input feature space with the infinity distance metric , and
be the closed ball of radius centered at in .
Given a dataset , where and , the objective function of standard adversarial training (Madry et al., 2018) is
For the sake of conceptual consistency, the objective function (i.e., Eq. (2)) can also be re-written as
It implies the optimization of adversariallly robust network, with one step maximizing loss to find adversarial data and one step minimizing loss on the adversarial data w.r.t. the network parameters .
2 Projected Gradient Descent (PGD)
To generate adversarial data, standard adversarial training uses PGD to approximately solve the inner maximization of Eq. (4) (Madry et al., 2018).
PGD formulates the problem of finding adversarial data as a constrained optimization problem. Namely, given a starting point and step size , PGD works as follows:
There are also other ways to generate adversarial data, e.g., the fast gradient signed method (Szegedy et al., 2014; Goodfellow et al., 2015), the CW attack (Carlini & Wagner, 2017), deformation attack (Alaifari et al., 2019; Xiao et al., 2018), and Hamming distance method (Shamir et al., 2019).
Friendly Adversarial Training
In this section, we develop a novel learning objective for friendly adversarial training (FAT). Theoretically, we justify FAT by deriving a tight upper bound of the adversarial risk.
Let be a margin such that our adversarial data would be misclassified with a certain amount of confidence.
Note that the operator in Eq. (4) is replaced with here, and there is a constraint on the margin of loss values (i.e., the misclassification confidence).
2 Upper Bound on Adversarial Risk
Minimizing the adversarial risk based on our upper bound aids in fine-tuning the decision boundary using friendly adversarial data. On one side, the wrongly-predicted adversarial data have a small distance (in term of the loss value) from the decision boundary (e.g., “Step 10 panel” at the bottom series in Figure 1) so that it will not cause the severe issue of cross-over mixture but fine-tunes the decision boundary. On the other side, correctly-predicted adversarial data maintain the largest distance (in term of maximizing the loss value) from their natural data so that the decision boundary is kept far away.
Key Component of FAT
To search friendly adversarial data, we design an efficient early-stopped PGD algorithm called PGD-- (Section 4.1), which could alleviate the cross-over mixture problem (Section 4.2) and therefore helps adversarial training. Note that besides PGD--, there are other ways to search for friendly adversarial data. We show one example in Appendix B.
Algorithm 1 returns the misclassified adversarial data with small loss values or correctly classified adversarial data with large loss values. Step controls the extent of loss minimization when misclassified adversarial data are found. When is larger, the misclassified adversarial data with slightly larger loss values are returned, and vice versa. is an approximation to in our learning objective. Note that when , the conventional PGD- is the special case of our PGD--. As is an important hyper-parameter of PGD-- for FAT (Section 5), we discuss how to select in Sections 5.1 and 5.2 in detail.
2 PGD-K𝐾K-τ𝜏\tau Alleviates Cross-over Mixture
In deep neural networks, the cross-over mixture problem may not trivially appear in the original input space, but occur in the output of the intermediate layer. Our proposed PGD-- is an effective solution to overcome this problem, which leads to successful adversarial training.
In Figure 3, we trained an 8-layer convolutional neural network (6 convolutional layers and 2 fully-connected layers, namely, Small CNN) on images of two selected classes in CIFAR-10. We conducted a warm-up training using natural training data, then included their adversarial variants generated by PGD- (middle panel) and PGD-- (right panel), where means PGD iterations stop immediately once adversarial data are wrongly predicted by the current network.
Figure 3 shows the output distribution of layer by principal component analysis (PCA, (Abdi & Williams, 2010)), which projects high-dimensional output into a two-dimensional subspace. As shown in Figure 3, the output distribution on natural data (left panel) of the intermediate layer is clearly not mixed. By contrast, the conventional PGD- (middle panel) leads to severe mixing between outputs of adversarial data with different classes. It is more difficult to fit these mixed adversarial data, which leads to inaccurate classifiers. By comparing with PGD-, our PGD-- (right panel) could greatly overcome the mixture issue of adversarial data. Thus, it helps the training algorithm to return an accurate classifier while ensuring adversarial robustness.
To further justify the above fact, we plot output distributions of other layers, e.g., layer and layer . Moreover, we train a Wide ResNet (WRN-40-4) (Zagoruyko & Komodakis, 2016) on 10 classes and randomly select 3 classes for illustrating their output distributions of its intermediate layers. Instead of PCA, we also use a non-linear technique for dimensionality reduction, i.e., t-distributed stochastic neighbor embedding (t-SNE) (Maaten & Hinton, 2008) to visualize output distributions of different classes. All these can be found in Appendix C.
Realization of FAT
Based on the proposed PGD--, we have a new algorithm termed FAT (Algorithm 2). FAT treats the standard adversarial training (Madry et al., 2018) as a special case when we set in Algorithm 1. Besides, we also design FAT for TRADES (Appendix D.2) and FAT for MART (Appendix E.2), making two effective adversarial training methods (TRADES (Zhang et al., 2019b) and MART (Wang et al., 2020)) special cases when . Since the essential component of FAT is PGD--, we should discuss the effects of step w.r.t. standard accuracy and adversarial robustness (Section 5.1) and computational efficiency (Section 5.2). Besides, we should discuss the relation between FAT and curriculum learning (Section 5.3), since FAT is a progressive training strategy. It is worth noting that Sitawarin et al. (2020) independently propose adversarial training with early stopping (ATES). Along with our FAT, ATES corroborates the new formulation (Section 3.1) for adversarial training.
As shown in Section 2.2, the conventional PGD- is a special case, when step in PGD--. Thus, standard adversarial training is a special case of FAT. Here, we investigate how step affects the performance of FAT empirically, and summarize that larger may not increase adversarial robustness but hurt the standard test accuracy. Detailed experimental setups of Figure 4 are in Appendix F.1.
Figure 4 shows that, with the increase of , the standard test accuracy for natural data decreases significantly; while the robust test accuracy for adversarial data increases at smaller values of but reaches its plateau at larger values of . For example, when is bigger than 2, the standard test accuracy continues to decrease with larger . However, the robust test accuracy begins to maintain a plateau for both Small CNN and ResNet-18. Such observation manifests that a larger step may not be necessary for adversarial training. Namely, it may not increase adversarial robustness but hurt the standard accuracy. This reflects a trade-off between the standard accuracy and adversarially robust accuracy (Tsipras et al., 2019) and suggests that our step helps manage this trade-off.
can be treated as a hyper-parameter. Based on the observations in Figure 4, it is enough to select from the set . Note that the size of the set is also influenced by step size and maximum PGD step . In Section 6.2, we use to fine-tune the performance of FAT.
2 Smaller τ𝜏\tau is Computationally Efficient
Adversarial training is time-consuming since it needs multiple backward propagations (BPs) to produce adversarial data. The time-consuming factor depends on the number of BPs used for generating adversarial data (Shafahi et al., 2019; Zhang et al., 2019a; Wong et al., 2020; B.S. & Babu, 2020).
Our FAT uses PGD-- to generate adversarial data, and PGD-- is early-stopped. This implies that FAT is computationally efficient, since FAT does not need to compute maximum BPs on each mini-batch. To illustrate this, we count the number of BPs for generating adversarial data during training. The training setup is the same as the one in Section 5.1, but we only choose from and train for 100 epochs with learning rate divided by 10 at epochs 60 and 90. In Figure 5, we compare the standard adversarial training (Madry et al., 2018) (dashed line) with our FAT (solid line) and adversarial training TRADES (Zhang et al., 2019b) (dashed line) with our FAT for TRADES (solid line, detailed in Algorithm 4 in Appendix D.2). For each epoch, we compute average BPs over all training data for generating the adversarial counterpart.
Figure 5 shows that conventional adversarial training uses PGD- which takes BPs for generating adversarial data in each mini-batch. By contrast, our adversarial training FAT that uses PGD-- significantly reduces the number of required BPs. In addition, with the smaller , FAT needs less BPs on average for generating adversarial data. The magnitude of controls the number of extra BPs, once misclassified adversarial data is found.
Moreover, there are some interesting phenomena observed in our FAT (or FAT for TRADES). As the training progresses, the number of BPs gradually increases. This shows that more steps are needed by PGD to find misclassified adversarial data. This signifies that it is increasingly difficult to find adversarial data that misclassifies the model. Thus, DNNs become more and more adversarially robust over training epochs. In addition, there is a slight surge in average BPs at epochs 60 and 90, where we divided the learning rate by 10. This means that the robustness of the model gets substantially improved at epochs 60 and 90. It is a common trick to decrease the learning rate over training DNNs for good standard accuracy (He et al., 2016). Figure 5 confirms that it is similarly meaningful to decrease the learning rate during adversarial training.
3 Relation to Curriculum Learning
Curriculum learning (Bengio et al., 2009) is a machine learning strategy that gradually makes the learning task more difficult. Curriculum learning is shown effective in improving standard generalization and providing faster convergence (Bengio et al., 2009).
In adversarial training, curriculum learning can also be used to improve adversarial robustness. Namely, DNNs learn from milder adversarial data first, and gradually adapt to stronger adversarial data. There are different ways to determine the hardness of adversarial data. For example, curriculum adversarial training (CAT) uses the perturbation step of PGD as the hardness measure (Cai et al., 2018). Dynamic adversarial training (DAT) uses their proposed criterion, the first-order stationary condition (FOSC), as the hardness measure (Wang et al., 2019). However, both methods do not have a principled way to decide when the hardness should be increased during training. To increase the hardness at the right time, both methods need domain knowledge to fine-tune the curriculum training sequence. For example, CAT needs to decide when to increase step in PGD over training epochs; while DAT needs to decide FOSC for generating adversarial data at different training stages.
Our FAT can also be regarded as a type of curriculum training. As shown in Figure 5, as the training progresses, more and more backward propagations (BPs) are needed to generate adversarial data to fool the classifier. Thus, more and more PGD steps are needed to generate adversarial data. Meanwhile, the network gradually and automatically learns from stronger and stronger adversarial data (adversarial data generated by more and more PGD steps). Differently from CAT and DAT, FAT could automatically increase the hardness of friendly adversarial data based on the model’s predictions. As a result, Table 1 in Section 6.2 shows that empirical results of FAT can outperform the best results in CAT (Cai et al., 2018), DAT (Wang et al., 2019) and standard Madry’s adversarial training (not a curriculum learning) (Madry et al., 2018). Thus, attacks which do not kill training indeed make adversarial learning stronger.
Experiments
To evaluate the efficacy of FAT, we firstly use CIFAR-10 (Krizhevsky, 2009) and SVHN (Netzer et al., 2011) datasets to verify that FAT can help achieve a larger perturbation bound . Then, we train Wide ResNet (Zagoruyko & Komodakis, 2016) on the CIFAR-10 dataset to achieve state-of-the-art results.
All images of CIFAR-10 and SVHN are normalized into $\tau=0,1,3\epsilon_{train}\epsilon_{train}\in[0.03,0.15]\epsilon_{train}\in[0.01,0.06]K=10\alpha=\epsilon/10$. DNNs were trained using SGD with 0.9 momentum for 80 epochs with the initial learning rate of 0.01 divided by 10 at epoch 60.
In addition, we have experiments of training a different DNN, e.g., Small CNN. We also set maximum PGD step . Besides, we compare our FAT for TRADES and TRADES (Zhang et al., 2019b) with different values of . These extensive results are presented in Appendix G.
Figures 6 and 7 show the performance of FAT () and standard adversarial training w.r.t. standard test accuracy and adversarially robust test accuracy of the DNNs. We obtain standard test accuracy for natural test data and robust test accuracy for adversarial test data. The adversarial test data are bounded by perturbations with for CIFAR-10 and for SVHN, which are generated by FGSM, PGD-20, PGD-100 and CW∞ ( version of CW optimized by PGD-30 (Carlini & Wagner, 2017)). Moreover, we also evaluate the robust DNNs using stronger adversarial test data generated by PGD-20 with a larger perturbation bound for CIFAR-10 and for SVHN (bottom right panels in Figures 6 and 7). All PGD attacks have random start, i.e, the uniformly random perturbation of added to the natural test data before PGD perturbations. Step size of PGD is fixed to 2/255.
From top-left panels of Figures 6 and 7, DNNs trained by FAT () have higher standard test accuracy compared with those trained by standard adversarial training (i.e., Madry). This gap significantly widens as perturbation bound increases. Larger will allow the generated adversarial data deviate more from natural data. In standard adversarial training, natural generalization is significantly hurt with larger due to the cross-over mixture issue. In contrast, DNNs trained by FAT could have better standard generalization, which is less affected by an increasing perturbation bound .
In addition, with the increase of the perturbation bound , robust test accuracy (e.g., PGD and CW) of DNNs trained by standard adversarial training (Madry) gets a slight increase first but is followed by a sharp drop. For a larger (e.g., in CIFAR-10 and in SVHN), standard adversarial training (Madry) basically fails, and thus its robust test accuracy drops sharply. Without early-stopped PGD, the generated adversarial data has a severe cross-over mixture problem, which makes the adversarial learning extremely difficult and sometimes even “kills” the learning.
However, it is still meaningful to enable a stronger defense over a weaker attack, i.e., in adversarial training should be larger than in adversarial attack. The right bottom panels in Figures 6 and 7 show for the stronger attack ( for CIFAR-10 and for SVHN), it is meaningful to have larger to attain a better robustness. Our FAT is able to achieve larger . Figures 6 and 7 show our FAT () maintain higher robust accuracy with larger .
Note that the performance of FAT with PGD-10-3 (, red lines) in Figure 6 drops with . We believe that FAT with larger could also have the issue of cross-over mixture, which is detrimental to adversarial learning. In addition, both standard adversarial training and friendly adversarial training do not perform well under CW attack with larger (e.g., in Figure 6), we believe it is due to the mismatch between PGD-adversarial training and CW attack. We discuss the reasons in detail in Appendix G.4.
To sum up, deep models by FAT with (green line) have higher standard test accuracy, but lower robust test accuracy. By increasing to 1, deep models have slightly reduced standard test accuracy but have the increased adversarial robustness. This sheds light on the importance of , which handles the trade-off between robustness and standard accuracy. In order not to “kill” the training at the initial stage, we could vary from a smaller value to a larger value over training epochs. In addition, due to benefits that our FAT could enable larger , we could also make larger over training. Those tricks echo our paper’s philosophy of “attacks which do not kill training make adversarial training stronger”. In the next subsection, we unleash the full power of FAT (and FAT for TRADES) and show its superior performance over the state-of-the-art methods.
2 Performance Evaluations on Wide ResNet
To manifest the full power of friendly adversarial training, we adversarially train Wide ResNet (Zagoruyko & Komodakis, 2016) to achieve the state-of-the-art performance on CIFAR-10. Similar to (Wang et al., 2019; Zhang et al., 2019b), we employ WRN-32-10 (Table 1) and WRN-34-10 (Table 2) as our deep models.
In Table 1, we compare FAT with standard adversarial training (Madry) (Madry et al., 2018), CAT (Cai et al., 2018) and DAT (Wang et al., 2019) on WRN-32-10. Training and evaluation details are in Appendix H.1. The performance evaluations are done exactly as in DAT (Wang et al., 2019). In Table 2, we compare FAT for TRADES with TRADES (Zhang et al., 2019b) on WRN-34-10. Training and evaluation details are in Appendix H.2. The performance evaluations are done exactly as in TRADES (Zhang et al., 2019b).
Moreover, Nakkiran (2019) states that robust classification needs more complex classifiers (exponentially more complex). We employ FAT for TRADES on even larger WRN-58-10, the performance gets further improved over WRN-34-10 (Appendix H.2). Moreover, we also apply the early-stopped PGD to MART, namely, FAT for MART (in Appendix E.2). As a result, the performance gets improved (detailed in Appendix H.3).
Tables 1 and 2 and results in Appendix H justify the efficacy of friendly adversarial training - adversarial robustness can indeed be achieved without compromising the natural generalization. In addition, we are even able to attain the state-of-the-art robustness.
Conclusion
This paper has proposed a novel formulation for adversarial training. Friendly adversarial training (FAT) approximately realizes this formulation by stopping the PGD early. FAT is computationally efficient and adheres to the spirit of curriculum training. In addition, FAT helps to relieve the problem of cross-over mixture. As a result, FAT can train deep models with larger perturbation bounds . Finally, FAT can achieve competitive performance on the large capacity networks. Further research includes (a) how to choose optimal step in FAT algorithm, (b) besides PGD-K-, how to search for friendly adversarial data effectively, and (c) theoretically studying adversarially robust generalization (Yin et al., 2019), e.g., through the lens of Rademacher complexity (Bartlett & Mendelson, 2002).
This research is supported by the National Research Foundation, Singapore under its Strategic Capability Research Centres Funding Initiative (MK, JZ), JST AIP Acceleration Research Grant Number JPMJCR20U3, Japan (GN, MS), National Key RD Program No.2017YFB1400100, the NSFC No.91846205, the Shandong Key RD Program No.2018YFJH0506 (LC), the Early Career Scheme (ECS) through the Research Grants Council of Hong Kong under Grant No.22200720 (BH), HKBU Tier-1 Start-up Grant (BH) and HKBU CSD Start-up Grant (BH). Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not reflect the views of National Research Foundation, Singapore.
References
Appendix A Friendly Adversarial Training
For completeness, besides the learning objective by loss value, we also give the learning objective of FAT by class probability. Then, based on the learning objective by loss value, we give the proof of Theorem 1, which theoretically justifies FAT.
Note that the operator in Eq. (4) is replaced with here, and there is a constraint on the margin of loss values (i.e., the mis-classification confidence).
Case 2 (by class probability).
A.2 Proofs
Lemma 1. For any classifier , any probability distribution on , we have
Our Theorem 1 informs our strategy to fine tune the decision boundary. To fine-tune the decision boundary, the data “near” the classifier plays an important role. Those data are easily wrongly predicted with small perturbations. As we show in the Figure 2, when adversarial data are wrongly predicted, our adversarial data (purple triangle) increases to minimize the loss by a violation of a small constant . Thus, our adversarial data can help fine-tune the decision boundary “bit by bit” over the training.
Appendix B Alternative Adversarial Data Searching Algorithm
where is a natural data and is step size.
We employ Small CNN and ResNet-18 to show the performance of deep models against FGSM, PGD-20, PGD-100 and CW∞ in Figure 9 and Figure 10. Deep models are trained using SGD with 0.9 momentum for 80 epochs with the initial learning rate 0.01 divided by 10 at 60 epoch. We compare FAT combined with the alternative adversarial data searching algorithm () and standard adversarial training (Madry) with different step size , i.e., . The maximum PGD step is fixed to 10. All the testing settings are the same as those are stated in Section 6.1.
Appendix C Mixture Alleviation
In Figure 3 in Section 4.2, we only visualize layer ’s output distribution by Small CNN (8-layer convolutional neural network with 6 convolutional layers and 2 fully connected layers). For completeness, we visualize the output distributions of layers and .
We conduct warm-up training using natural training data of two randomly selected classes (bird and deer) in CIFAR-10, then involve its adversarial variants generated by PGD-20 with step size and maximum perturbation . We show output distributions of layer , and by PCA in Figure 11(a) and t-SNE in Figure 11(b).
C.2 Output distributions of WRN-40-4’s intermediate layers
We train a Wide ResNet (WRN-40-4, totally 41 layers) using natural data on 10 classes in CIFAR-10 and then include adversarial variants. We randomly select 3 classes (deer, horse and truck) for illustrating output distributions of WRN-40-4’s intermediate layers. Adversarial data are generated by PGD-20 with step size and maximum perturbation on WRN-40-4. We show output distributions by layer , and by PCA in Figure 12(a) and t-SNE in Figure 12(b).
Appendix D FAT for TRADES
Similar to virtual adversarial training (VAT) adding a regularization term to the loss function (Miyato et al., 2016), which regularizes the output distribution by its local sensitivity of the output w.r.t. input, the objective of TRADES is
D.2 FAT for TRADES - Realization
Given a dataset , where and , adversarial training (TRADES) with early stopped PGD-- returns a classifier :
Based on our early stopped PGD-- for TRADES in Algorithm 3, our friendly adversarial training for TRADES (FAT for TRADES) is
Appendix E FAT for MART
MART (Wang et al., 2020) emphasizes the importance of misclassified natural data on the adversarial robustness. Wang et al. (2020) propose a regularized adversarial learning objective which contains an explicit differentiation of misclassified data as the regularizer. The learning objective of MART is
E.2 FAT for MART - Realization
Given a dataset , where and , FAT for MART returns a classifier :
Appendix F Experimental Setup
Figure 4 presents empirical results on CIFAR-10 via our FAT algorithm,where we train 8-layer convolutional neural network (Small CNN, blue line) and 18-layer residual neural network (ResNet-18, red line) (He et al., 2016). The maximum step , , step size , and step . We train deep networks for 80 epochs using SGD with 0.9 momentum, where learning rate starts at 0.1 and divided by 10 at 60 epoch.
For each , we take five trials, where each trial will obtain standard test accuracy evaluated on natural test data and robust test accuracy evaluated on adversarial test data that are generated by attacks FGSM (Goodfellow et al., 2015), PGD-10 and PGD-20, PGD-100 (Madry et al., 2018) and CW attack (Carlini & Wagner, 2017) respectively. All those attacks are white box attacks, which are constrained by the same perturbation bound . Following Zhang et al. (2019b), all attacks have the random start, and the step size in PGD-10, PGD-20, PGD-100 and CW is fixed to .
In this section, we provide extensive experimental results. The test settings are the same as those are stated in Section 6.1. In Section G.1, instead of using ResNet-18, we conduct adversarial training on the deep model of Small CNN. In Section G.2, instead of applying FAT, we compare our FAT for TRADES and TRADE (Zhang et al., 2019b) under different values of perturbation bound on the deep models ResNet-18 and Small CNN. In Section G.3, we set maximum PGD steps and report results of FAT and FAT for TRADES over existing methods with larger perturbation bound . To sum up, all those extensive results verify that FAT and FAT for TRADES can enable deep models trained under larger values of perturbation bound .
We train Small CNN on CIFAR-10 and SVHN using the same settings as those stated in Section 6.1. We show standard and robust test accuracy of deep model (Small CNN) on CIFAR-10 dataset (Figure 13) and SVHN dataset (Figure 14).
G.2 FAT for TRADES
We apply FAT for TRADES(Algorithm 4) to Small CNN and ResNet-18 on CIFAR-10 dataset. All training settings are the same as those are stated in Section 6.1. Regularization parameter . We present standard and robust test results of Small CNN (Figure 15) and ResNet-18 (Figure 16).
G.3 Maximum PGD Step K=20𝐾20K=20
By setting maximum PGD step , we conduct more experiments on Small CNN and ResNet-18 using FAT and FAT for TRADES. Except maximum PGD steps , training settings are the same as those are stated in Section 6.1. Test results of robust deep models are shown in Figures 17, 18, 19 and 20.
G.4 C&\&W Attack Analysis
As is shown in Figure 6 along with Figure 13 in Section G.1, both standard adversarial training and friendly adversarial training do not perform well under CW (Carlini & Wagner, 2017) attack with larger (e.g., ). We discuss the reasons for these phenomena.
Analysis.
In Figure 6, with larger , the performance evaluated by PGD attacks increases, while performance evaluated by CW attack decreases. The reason is that CW and PGD have different ways of generating adversarial data according to Eq. (12) and Eq. (4) respectively. The two interactive methods search adversarial data in different directions due to gradients w.r.t. different loss. Therefore, the distributions of CW and PGD adversarial data are inconsistent. As perturbation bound increases, there are more PGD adversarial data generated within -ball. A DNN learned from more PGD adversarial data becomes more defensive to PGD attacks, but this deep model may not effectively defend CW adversarial data.
Appendix H Extensive State-of-the-art Results on Wide ResNet
In Table 1, we compare our FAT with standard adversarial training (Madry), CAT (Cai et al., 2018) and DAT (Wang et al., 2019).
We use FAT ( and respectively) to train WRN-32-10 for 120 epochs using SGD with 0.9 momentum, and weight decay is 0.0002. Maximum PGD step is 10 and step size is fixed to 0.007. The initial learning rate is 0.1 reduced to 0.01, 0.001 and 0.0005 at epoch 60, 90 and 110. We set step initially and increase by one at epoch 50 and 90 respectively. The maximum step . We report performance of the deep model at the last epoch. For fair comparison, in Table 1 we use the same test settings as those in DAT (Wang et al., 2019). Performance of robust deep model is evaluated standard test accuracy for natural data and robust test accuracy for adversarial data, that are generated by FGSM, PGD-20 (20-steps PGD with random start), PGD-100 and CW∞(L∞ version of CW optimized by PGD-30).
All attacks have the same perturbation bound and step size in PGD is . The same as DAT (Wang et al., 2019), there is random start in PGD attack, i.e., uniformly random perturbations () added to natural data before PGD perturbations. We report the median test accuracy and its standard deviation over 5 repeated trails of adversarial training in Table 1.
H.2 Training details of FAT for TRADES on Wide ResNet
In Table 2, we use FAT for TRADES ( and respectively) train WRN-34-10 by FAT for TRADES for 85 epochs using SGD with 0.9 momentum and 0.0002 weight decay. Maximum PGD step and step size . The initial learning rate is 0.1 and divided 10 at epoch 75. We set step initially and increased by one at epoch 30, 50 and 70. Since TRADES has a trade-off parameter , for fair comparison, our FAT for TRADES use the same . In Table 2, we set and separately, which are endorsed by (Zhang et al., 2019b).
For fair comparison, we use the same test settings as those are stated in TRADES (Zhang et al., 2019b). All attacks have the same perturbation bound ( without random start), and step size , which is the same as stated in the paper (Zhang et al., 2019b). Performance of robust deep model is evaluated standard test accuracy for natural data and robust test accuracy for adversarial data, that are generated by FGSM, PGD-20, PGD-100 and CW∞(L∞ version of CW optimized by PGD-30). We report the median test accuracy and its standard deviation of the deep model at the last epoch over 3 repeated trials of adversarial training in Table 2.
However, in TRADES’s experimental testingTRADES GitHub, they use random start before PGD perturbation that is deviated from the statements in the paper (Zhang et al., 2019b). For fair comparison, we also retest the robust deep models under PGD attacks with random start. We evaluate their publicly released robust deep modelTRADES’s pre-trained model WRN-34-10 and compare it with ours trained by FAT for TRADES. The test results are reported in Table 3.
FAT for TRADES on larger WRN-58-10.
We employ Wide ResNet with larger capacity, i.e., WRN-58-10 to show our superior performance achieved by FAT for TRADES in Table 3. All the training settings are the same as details on WRN-34-10 in this section. The regularization parameter is fixed to 6.0. All attacks have the same perturbation bound and step size , which is the same as TRADES’s experimental setting. Robustness against FGSM, PGD-20(20-steps PGD with random start) and CW∞ is reported in Table 3.
H.3 FAT for MART on Wide ResNet
We train WRN-34-10 by FAT for MART ( and respectively) using SGD with 0.9 momentum and 0.0002 weight decay. Maximum PGD step and step size . The initial learning rate is 0.1 and divided 10 at epoch 60 and 90 respectively. We set step initially and increase by one at epoch 20, 40, 60 and 80. The regularization parameter is fixed to 6.0. The maximum step size .
For fair comparison, all attacks have the same perturbation bound and step size , which is the same setting in MART (Wang et al., 2020). White-box robustness of the deep model against attacks such as FGSM, PGD-20 (20-steps PGD with random start) and CW∞ (L∞ version of CW optimized by PGD-30) is reported. We evaluate Wang et al. (2020) publicly released robust deep modelMART’s pre-trained model WRN-34-10 and compare it with ours trained by FAT for MART. In Table 4, we report the median test accuracy and its standard deviation over 3 repeated trails of FAT for MART on WRN-34-10.
FAT for MART on larger WRN-58-10.
In Table 4, we also employ WRN-58-10 to show the performance achieved by FAT for MART. All the training and testing settings are the same as those on WRN-34-10.