Uncovering the Limits of Adversarial Training against Norm-Bounded Adversarial Examples

Sven Gowal, Chongli Qin, Jonathan Uesato, Timothy Mann, Pushmeet Kohli

Introduction

Neural networks are being deployed in a wide variety of applications with great success (Goodfellow et al., 2016; Krizhevsky et al., 2012; Hinton et al., 2012). As neural networks tackle challenges ranging from ranking content on the web (Covington et al., 2016) to autonomous driving (Bojarski et al., 2016) via medical diagnostics (De Fauw et al., 2018), it has become increasingly important to ensure that deployed models are robust and generalize to various input perturbations. Despite their success, neural networks are not intrinsically robust. In particular, the addition of small but carefully chosen deviations to the input, called adversarial perturbations, can cause the neural network to make incorrect predictions with high confidence (Carlini & Wagner, 2017a; Goodfellow et al., 2015; Kurakin et al., 2016; Szegedy et al., 2014). Starting with Szegedy et al. (2014), there has been a lot of work on understanding and generating adversarial perturbations (Carlini & Wagner, 2017b; Athalye & Sutskever, 2018), and on building models that are robust to such perturbations (Goodfellow et al., 2015; Papernot et al., 2016; Madry et al., 2018; Kannan et al., 2018). Robust optimization techniques, like the one developed by Madry et al. (2018), learn robust models by finding worst-case adversarial examples (by running an inner optimization procedure) at each training step and adding them to the training data. This technique has proven to be effective and is now widely adopted.

Since Madry et al. (2018), various modifications to their algorithm have been proposed (Xie et al., 2019; Pang et al., 2020b; Huang et al., 2020; Qin et al., 2019; Zoran et al., 2019; Andriushchenko & Flammarion, 2020). We highlight the work by Zhang et al. (2019) who proposed TRADES which balances the trade-off between standard and robust accuracy, and the work by Carmon et al. (2019); Uesato et al. (2019); Najafi et al. (2019); Zhai et al. (2019) who simultaneously proposed the use of additional unlabeled data in this context. As shown in Fig. 1, despite this flurry of activity, progress over the past two years has been slow. In a similar spirit to works exploring transfer learning (Raffel et al., 2020) or recurrent architectures (Jozefowicz et al., 2015), we perform a systematic study around the landscape of adversarial training in a bid to discover its limits. Concretely, we study the effects of (i) the objective used in the inner and outer optimization procedures, (ii) the quantity and quality of additional unlabeled data, (iii) the model size, as well as (iv) other factors (such as the use of model weight averaging). In total we trained more than 150 adversarially robust models and dissected each of them to uncover new ideas that could improve adversarial training and our understanding of robustness. Here is a non-exhaustive list highlighting our findings (in no specific order):

TRADES (Zhang et al., 2019) combined with early stopping outperforms regular adversarial training (as proposed by Madry et al., 2018). This is in contrast to the observations made by Rice et al. (2020) (Sec. 4.1).

As observed by Madry et al. (2018); Xie & Yuille (2019); Uesato et al. (2019), increasing the capacity of models improves robustness (Sec. 4.4).

The choice of activation function matters. Similar to observations made by Xie et al. (2020b), we found that Swish/SiLU (Hendrycks & Gimpel, 2016; Ramachandran et al., 2017; Elfwing et al., 2018) performs best. However, in contrast to Xie et al., we found that other “smooth” activation functions do not necessarily improve robustness (Sec. 4.5.2).

The way in which additional unlabeled data is extracted from the 80 Million Tiny Images dataset (80M-Ti) (Torralba et al., 2008) and used during training can have a significant impact (Sec. 4.3).

Model weight averaging (WA) (Izmailov et al., 2018) consistently provides a boost in robustness. In the setting without additional unlabeled data, WA provides improvements as large as those provided by TRADES on top of classical adversarial training (Sec. 4.5.1).

In this study, we aim to find the current limits of adversarial robustness. What we found was that an accumulation of small factors can significantly improve upon the state-of-the-art when combined. As fundamentally new techniques might be needed, it is important to understand the limitations of current approaches. We hope that this new set of baselines can further new understanding about adversarial robustness.

Background

Since Szegedy et al. (2014) observed that neural networks which achieve high accuracy on test data are highly vulnerable to adversarial examples, the art of crafting increasingly sophisticated adversarial examples has received a lot of attention. Goodfellow et al. (2015) proposed the Fast Gradient Sign Method (FGSM) which generates adversarial examples with a single normalized gradient step. It was followed by R+FGSM (Tramèr et al., 2017), which adds a randomization step, and the Basic Iterative Method (BIM) (Kurakin et al., 2016), which takes multiple smaller gradient steps. These are often grouped under the term Projected Gradient Descent (PGD) which usually refers to the optimization procedure used to search norm-bounded perturbations.

Many other alternative defenses are not covered in the scope of this paper. They range from preprocessing techniques (Guo et al., 2018; Buckman et al., 2018) to detection algorithms (Metzen et al., 2017; Feinman et al., 2017), and also include the definition of new regularizers (Moosavi-Dezfooli et al., 2019; Xiao et al., 2019; Qin et al., 2019). The difficulty of adversarial evaluation also drove the need for certified defenses (Wong et al., 2018; Mirman et al., 2018; Gowal et al., 2019a; Zhang et al., 2020; Cohen et al., 2019; Salman et al., 2019), but the guarantees that these techniques provide do not yet match the empirical robustness obtained through adversarial training.

It is worth noting that many of the defense strategies proposed in the literature (Papernot et al., 2016; Lu et al., 2017; Kannan et al., 2018; Tao et al., 2018; Zhang & Wang, 2019) were broken by stronger adversaries (Carlini & Wagner, 2016, 2017b; Athalye & Sutskever, 2018; Engstrom et al., 2018; Carlini, 2019; Uesato et al., 2018; Athalye et al., 2018). Hence, the robust accuracy obtained under different evaluation protocols cannot be easily compared and care has to be taken to make sure the evaluation is as strong as is possible (Carlini et al., 2019). In this manuscript, we evaluate each model against two of the strongest adversarial attacks, AutoAttack and MultiTargeted, developed by Croce & Hein (2020) and Gowal et al. (2019b), respectively.

2 Adversarial training

Madry et al. (2018) formulate a saddle point problem whose goal is to find model parameters θ{\bm{\theta}} that minimize the adversarial risk:

For each example x{\bm{x}} with label yy, adversarial training minimizes the loss given by

where δ^\hat{{\bm{\delta}}} is given by Eq. 2 and lxentl_{\textrm{xent}} is the softmax cross-entropy loss.

Setup and implementation details

For consistency with prior work on adversarial robustness (Madry et al., 2018; Rice et al., 2020; Zhang et al., 2019; Uesato et al., 2019), we use Wide ResNets (Wrns) (He et al., 2016; Zagoruyko & Komodakis, 2016). Our baseline model (on which most experiments are performed) is 28 layers deep with a width multiplier of 10, and is denoted by Wrn-28-10. We also use deeper (up to 70 layers) and wider (up to 20) models. Our largest model is a Wrn-70-16 containing 267M parameters, our smallest model is a Wrn-28-10 containing 36M parameters. A popular option in the literature is the Wrn-34-20, which contains 186M parameters.

We use Stochastic Gradient Descent (SGD) with Nesterov momentum (Polyak, 1964; Nesterov, 1983). The initial learning rate of 0.1 is decayed by 10×10\times half-way and three-quarters-of-the-way through training (we refer to this schedule as the multistep schedule). We use a global weight decay parameter of 5×10−45\times 10^{-4}. In the basic setting without additional unlabeled data, we use a batch size of 128128 and train for 200200 epochs (i.e., 78K steps). With additional unlabeled data, we use a batch size of 10241024 with 512512 samples from Cifar-10 and 512512 samples from a subset of 500K images extracted from the tiny images dataset 80M-Ti (Torralba et al., 2008) We use the dataset from Carmon et al. (2019) available at https://github.com/yaircarmon/semisup-adv. and train for 19.5K steps (i.e., 400400 Cifar-10-equivalent epochs). To use the unlabeled data with adversarial training, we use the pseudo-labeling mechanism presented by Carmon et al. (2019): a separate non-robust classifier is trained on clean data from Cifar-10 to provide labels to the unlabeled samples (the dataset created by Carmon et al., 2019 is already annotated with such labels). Our batches are split over 3232 Google Cloud TPU v33 cores. As is common on Cifar-10, we augment our samples with random crops (i.e., pad by 4 pixels and crop back to 32×3232\times 32) and random horizontal flips.

As adversarial training is a bit more noisy than regular training, for each hyperparameter setting, we train two models. Throughout training we measure the robust accuracy using Pgd4040 on 1024 samples from a separate validation set (disjoint from the training and test set). Similarly to Rice et al. (2020), we perform early stopping by keeping track of model parameters that achieve the highest robust accuracy (i.e., lowest adversarial risk as shown in Eq. 1) on the validation set. From both models, we pick the one with highest robust accuracy on the validation set and use this model for a further more thorough evaluation. Finally, the robust accuracy is reported on the full test set against a mixture of AutoAttack (Croce & Hein, 2020) and MultiTargeted (Gowal et al., 2019b). We execute the following sequence of attacks: AutoPgd on the cross-entropy loss with 5 restarts and 100 steps, AutoPgd on the difference of logits ratio loss with 5 restarts and 100 steps, MultiTargeted on the margin loss with 10 restarts and 200 steps.For reference, the AutoAttack leaderboard at https://github.com/fra31/auto-attack evaluates the model from Rice et al. (2020) to 53.42%, while this evaluation computes a robust accuracy of 53.38% for the same model. The AutoAttack leaderboard also evaluates the model from Carmon et al. (2019) to 59.53%, while this evaluation computes a robust accuracy of 59.47% for the same model. We also report the clean accuracy which is the top-1 accuracy without adversarial perturbations.

2 Baseline.

As a comparison point, we train a Wrn-28-10 ten times with the default settings given above for both the Cifar-10-only and the additional data settings. The resulting robust accuracy on the test set is 50.80±\pm0.23% (for Cifar-10-only) and 58.41±\pm0.25% (with additional unlabeled data). When using a Wrn-34-20, we obtain 52.91% which is in line with the model obtained by Rice et al. (2020) for the same settings (i.e., 53.38%). When using TRADES (instead of regular adversarial training) in the additional data setting, we obtain 59.45% which is in line with the model obtain by Carmon et al. (2019) for the same settings (i.e., 59.47%). As our pipeline is implemented in JAX (Bradbury et al., 2018) and Haiku (Hennigan et al., 2020), we do not exclude some slight differences in data preprocessing and network initialization.

Experiments and analysis

The following sections detail different independent experiments (more experiments are available in the appendix). Each section is self-contained to allow the reader to jump to any section of interest. The outline is as follows: \etocsettocstyle \localtableofcontents

As explained in Sec. 2.2, adversarial training as proposed by Madry et al. (2018) aims to minimize the loss given in Eq. 3 and is usually implemented as

Here, lxentl_{\textrm{xent}} denotes the cross-entropy loss and δ^\hat{{\bm{\delta}}} is treated as a constant (i.e., there is no back-propagation through the inner optimization procedure). One of most successful variant of adversarial training is TRADES (Zhang et al., 2019) which derives a theoretically grounded regularizer that balances the trade-off between standard and robust accuracy. The overall loss used by TRADES is given by

where DKLD_{\textrm{KL}} denotes the Kullback-Leibler divergence. TRADES is one of the core components used by Carmon et al. (2019) within RST and by Uesato et al. (2018) within UAT-OT. A follow-up work (Wang et al., 2020), known as Misclassification Aware Adversarial Training (MART), introduced a boosted loss that differentiates between the misclassified and correctly classified examples in a bid to improve this trade-off further. We invite the reader to refer to Wang et al. (2020) for more details.

Table 1 shows the performance of TRADES and MART compared to adversarial training (AT). We observe that while TRADES systematically improves upon AT in both data settings, this is not the case for MART. We find that MART has the propensity to create moderate forms of gradient masking against weak attacks like Pgd2020. In one extreme case, a Wrn-70-16 trained with additional unlabeled data using MART obtained an accuracy of 71.08% against Pgd2020 which dropped to 60.63% against our stronger suite of attacks (this finding is consistent with the AutoAttack leaderboard at https://github.com/fra31/auto-attack).

Contrary to the suggestion of Rice et al. (2020) (i.e., “the original PGD-based adversarial training method can actually achieve the same robust performance as state-of-the-art method”, see Sec. 2.1), TRADES (when combined with early-stopping – as our setup dictates) is more competitive than classical adversarial training. The results also highlight the importance of strong evaluations beyond Pgd2020 (including evaluations of the validation set used for early stopping).

2 Inner maximization

Most works use the same loss (e.g., cross-entropy loss or Kullback-Leibler divergence) for solving the inner maximization problem (i.e., finding an adversarial example) and the outer minimization problem (i.e., training the neural network). However, there are many plausible approximations for the inner maximization. For example, instead of using the cross-entropy loss or Kullback-Leibler divergence for the inner maximization, Uesato et al. (2018) used the margin loss max⁡i≠yf(x;θ)i−f(x;θ)y\max_{i\neq y}f({\bm{x}};{\bm{\theta}})_{i}-f({\bm{x}};{\bm{\theta}})_{y} which improved attack convergence speed. Similarly, we could mix the cross-entropy loss with TRADES via the following:

Or mix the Kullback-Leibler divergence with classical adversarial training:

Below are our observations on the effects of robustness and clean accuracy when we combine different inner and outer optimisation losses.

Table 2 shows the performance of different inner and outer loss combinations. Each combination is evaluated against two adversaries: a weaker Pgd4040 attack using the margin loss (and denoted by Pgdmargin40{}^{40}_{\textrm{margin}}) and our stronger combination of attacks (as detailed in Sec. 3.1 and denoted AA+MT). The first observation is that TRADES obtains higher robust accuracy compared to classical adversarial training (AT) across most combinations of inner losses. We note that the clean accuracy is higher for AT compared to TRADES and address this trade-off in the next section (Sec. 4.2.2). In the low-data setting, the combination of TRADES with cross-entropy (TRADES-XENT) yields the best robust accuracy. In the high-data setting, we obtain higher robust accuracy using TRADES-KL.

Our second observation is that using margin loss during training can show signs of gradient masking: models trained using margin loss show higher degradation in robust accuracy when evaluated against the stronger adversary (AA+MT) as opposed to the weaker Pgdmargin40{}^{40}_{\textrm{margin}}. Table 2 shows that the drop in robust accuracy can be as high as -6.35%.Using margin loss for evaluation is not uncommon, since it has been suggested to yield a stronger adversary (see Carlini & Wagner, 2017b; Liu et al., 2017). This level of gradient masking is most prominent in AT-MARGIN and is mitigated when we use TRADES for the outer minimization. This degradation in robust accuracy is significantly reduced when we use cross-entropy as the inner maximization loss.

Similarly to the observation made in Sec. 4.1, TRADES obtains higher adversarial accuracy compared to classical adversarial training. We found that using margin loss during training (for the inner maximization procedure) creates noticeable gradient masking.

2.2 Inner maximization perturbation radius

Several works explored the use of larger (Gowal et al., 2019a) or adaptive (Balaji et al., 2019) perturbation radii. Like TRADES, using different perturbation radii is an attempt to bias the clean to robust accuracy trade-off: as we increase the training perturbation radius we expect increased robustness and lower clean accuracy.

Table 3 shows the effect of increasing the perturbation radius ϵ\epsilon by a factor 1.1×1.1\times and 1.2×1.2\times the original value of 8/2558/255 (during training, not during evaluation). We notice that adversarial training (AT) can close the gap to TRADES as we increase the perturbation radius (especially in the low-data regime). Although not reported here, we noticed that TRADES does not get similar improvements in performance with larger radii (possibly because TRADES is already actively managing the trade-off with clean accuracy).

Tuning the training perturbation radius can marginally improve robustness (when using classical adversarial training). We posit that understanding when and how to use larger perturbation radii might be critical towards our understanding of robust generalization. Work from Balaji et al. (2019) do provide additional insights, but the topic remains largely under-explored.

3 Additional unlabeled data

After the work from Schmidt et al. (2018) which posits that robust generalization requires more data, Hendrycks et al. (2019) demonstrated that one could leverage additional labeled data from ImageNet to improve the robust accuracy of models on Cifar-10. Uesato et al. (2019) and Carmon et al. (2019) were among the first to introduce additional unlabeled data to Cifar-10 by extracting images from 80M-Ti (i.e., subset of high scoring images from a Cifar-10 classifier). Uesato et al. (2019) used 200K additional images to train their best model as they observed that using more data (i.e., 500K images) worsen their results. This suggests that additional data needs to be close enough to the original Cifar-10 images to be useful. This is something already suggested by Oliver et al. (2018) in the context of semi-supervised learning.

In this experiment, we use four different subsets of additional unlabeled images extracted from 80M-Ti. The first set with 500K images is the additional data used by Carmon et al. (2019).https://github.com/yaircarmon/semisup-adv The second, third and fourth sets consist of 200K, 500K and 1M images regenerated using a process identical to Carmon et al. (2019) with a another pre-trained classifier (achieving 95.86% accuracy on the test set). To generate a dataset of size NN, we remove duplicates from the Cifar-10 test set, score the remaining images using a standard Cifar-10 classifier, and pick the top-N/10N/10 scoring images from each class. Hence, the dataset with 500K images contains all the images from the dataset with 200K images and, similarly, the dataset with 1M images contains all the images from the dataset with 500K images. As the datasets increase in size, the images they contain may become less relevant to the Cifar-10 classification task (as their standard classifier score becomes smaller).

Table 4 shows the performance of adversarial training with pseudo-labeling as the quantity of additional unlabeled data increases. This training scheme is identical to UAT-FT (as introduced by Uesato et al., 2019) and is slightly different to the one proposed in Carmon et al. (2019) (which used TRADES). Firstly, we note that our regenerated set of 500K images improves robustness (+0.71%) compared to Carmon et al. (2019). Secondly, similar to the observations made by Uesato et al. (2019), there is a sweet spot where additional data is maximally useful. Going from 200K to 500K additional images improves robust accuracy by +1.83%. However, increasing the amount of additional images further to 1M is detrimental (-0.23%). This suggests that more data improves robustness as long as the additional images relate to the original Cifar-10 dataset (e.g., the more images we extract, the less likely it is that these images correspond to classes within Cifar-10).

Small differences (e.g., different classifiers used for pseudo-labeling) in the process that extracts additional unlabeled data can have significant impact on robustness (i.e., models trained on our regenerated dataset obtain higher robust accuracy than those trained with the data from Carmon et al., 2019). There is also a trade-off between the quantity and the quality of the extra unlabeled data (i.e., increasing the amount of unlabeled data to 1M did not increase robustness).

3.2 Ratio of labeled-to-unlabelled data per batch

In this section, all the experiments use the unlabeled data from Carmon et al. (2019). In the baseline setting, as done in Carmon et al. (2019), each batch during training uses 50% labeled data and the rest for unlabeled data. This effectively downweighs unlabeled images by a factor 10×\times, as we have 50K labeled images and 500K unlabeled images. Increasing this ratio reduces the weight given to additional data and puts more emphasis of the original Cifar-10 images.

Fig. 2 shows the robust accuracy as we vary that ratio. We observe that giving slightly more importance to unlabeled data helps. More concretely, we find an optimal ratio of labeled-to-unlabeled data of 3:7, which provides a boost of +0.95% over the 1:1 ratio. This suggest that the additional data extracted by Carmon et al. (2019) is well aligned with the original Cifar-10 data and that we can improve robust generalization by allowing the model to see this additional data more frequently. Increasing this ratio (i.e., reducing the importance of the unlabeled data) gradually degrades robustness and, eventually, the robust accuracy matches the one obtained by models trained without additional data.

We also experimented with label smoothing (for both labeled and unlabeled data independently). Label smoothing should counteract the effect of the noisy labels resulting from the classifier used in the pseudo-labeling process. However, we did not observe improvements in performance for any of the settings tried.

Tuning the weight given to unlabeled examples (by varying the labeled-to-unlabeled data ratio per batch) can provide improvements in robustness. This experiment highlight that a careful treatment of the extra unlabeled data can provide improvements in adversarial robustness.

4 Effects of scale

Rice et al. (2020) observed that increasing the model width improves robust accuracy despite the phenomenon of robust overfitting (which favors the use of early stopping in adversarial training). Uesato et al. (2019) also trialed deeper models with a Wrn-106-8 and observed improved robustness. A most systematic study of the effect of network depth was also conducted by Xie & Yuille (2020) on ImageNet (scaling a ResNet to 638 layers). However, there have not been any controlled experiments on Cifar-10 that varied both depth and width of Wrns. We note that most work on adversarial robustness on Cifar-10 use either a Wrn-34-10 or a Wrn-28-10 network, with Wrn-34-20 being another popular option.

Fig. 3 shows the effect of increasing the depth and width of our baseline network. It is possible to observe that, while both depth and width increase the number of effective model parameters, they do not always provide the same effect on robustness. For example, a Wrn-46-15 which trains in roughly the same time as a Wrn-28-20 (about 5 hours in our setup) reaches higher robust accuracy: +0.96% and +0.66% on the settings without and with additional data, respectively. Table 14 in the appendix also shows that the clean accuracy improves as networks become larger.

Larger models provide improved robustness (Madry et al., 2018; Xie & Yuille, 2019; Uesato et al., 2019) and, for identical parameter count, deeper models can perform better.

5 Other tricks

Model Weight Averaging (WA) (Izmailov et al., 2018) is widely used in classical training (Tan & Le, 2019), and leads to better generalization. The WA procedure finds much flatter solutions than SGD, and approximates ensembling with a single model. To the best of our knowledge, the effect of WA on robustness have not been studied in the literature. Ensembling has received some attention (Pang et al., 2019; Strauss et al., 2017), but requires training multiple models. We implement WA using an exponential moving average θ′{\bm{\theta}}^{\prime} of the model parameters θ{\bm{\theta}} with a decay rate τ\tau: we execute θ′←τ⋅θ′+(1−τ)⋅θ{\bm{\theta}}^{\prime}\leftarrow\tau\cdot{\bm{\theta}}^{\prime}+(1-\tau)\cdot{\bm{\theta}} at each training step. During evaluation, the weighted parameters θ′{\bm{\theta}}^{\prime} are used instead of the trained parameters θ{\bm{\theta}}.

Fig. 4 summarizes the performance of model weight averaging (detailed results are in Table 15 in the appendix). Panel 4(a) demonstrates significant improvements in robustness: +1.41% and +0.73% with respect to the baseline without WA for settings without and with additional data, respectively. In fact, in the low-data regime, WA provides an improvement similar to that of TRADES (+1.11%). A possible explanation for this phenomena is possibly given by panel 4(b). We observe that not only that WA achieves higher robust accuracy, but also that it maintains this higher accuracy over a few training epochs (about 25 epochs). This reduces the sensitivity to early stopping which may miss the most robust checkpoint (happening shortly after the second learning rate decay). Another important benefit of WA is its rapid convergence to about 50% robust accuracy (against Pgd4040) in the early stages of training, which suggests that it could be combined with efficient adversarial training techniques such as the one presented by Wong et al. (2020).

Although widely used for standard training, WA has not been explored within adversarial training. We discover that WA provides sizeable improvements in robustness and hope that future work can explore this phenomenon.

5.2 Activation functions

With the exception from work by Xie et al. (2020a) which focused on ImageNet, there have been little to no investigations into the effect of different activation functions on adversarial training. Xie et al. (2020a) discovered that “smooth” activation functions yielded higher robustness when using weak adversaries during training (PGD with a low number of steps). In particular, they posit that they allow adversarial training to find harder adversarial examples and compute better gradient updates. Qin et al. (2019) also experimented with softplus activations with success.

While we observe in Fig. 5 that activation functions other than ReLU (Nair & Hinton, 2010) can positively affect clean and robust accuracy, the trend is not as clear as the one observed by Xie et al. (2020a) on ImageNet. In fact, apart from Swish/SiLU (Hendrycks & Gimpel, 2016; Ramachandran et al., 2017; Elfwing et al., 2018), which provides improvements of +0.8% and +1.13% in the settings without and with additional unlabeled data, the order of the best performing activation functions changes when we go from the Cifar-10-only setting to the setting with additional unlabeled data. Overall, ReLU remains a good choice.

The choice of activation function matters: we found that while Swish/SiLU performs best in our experiments. Other “smooth” activation functions (Xie et al., 2020b) do not necessarily correlate positively with robustness in our experiments.

A new state-of-the-art

Conclusion

References

Appendix A Additional experiments

In order to keep the main text concise, we relegated additional experiments to this section. Similarly to Sec. 4, each section is self-contained to allow the reader to jump to any section of interest. The outline is as follows: \etocsettocstyle \localtableofcontents

In this section, we test different learning rate schedules. In particular, we compare the multistep schedule introduced in Sec. 3.1, where the initial learning rate of 0.10.1 is decayed by 10×10\times half-way and three-quarters-of-the-way through training, with the cosine and exponential schedules. For the cosine schedule, we set the initial learning rate to 0.10.1 and decay it to by the end of training. For the exponential schedule, we set the initial learning rate to 0.10.1 and decay it every 55 epochs such that by the end of training the final learning rate is 0.0010.001. Learning rates are all scaled according to the batch size (i.e., effective learning=max⁡(learning rate×batch size/256,learning rate)\textrm{effective learning}=\max(\textrm{learning rate}\times\textrm{batch size}/256,\textrm{learning rate})).

Table 6 shows that, as they are currently implemented, the multistep schedule is superior to the cosine and exponential schedules. While, we did our best to tune all schedules, we do not exclude the possibility that better schedules exist. In particular, the optimal schedule may depend on the model architecture and method used to find adversarial examples (e.g., AT or TRADES). Smoother schedules (like cosine or exponential) are also less sensitive to the early stopping criterion and may result in less noisy results. When using additional unlabeled data, model weight averaging (Sec. 4.5.1) and a Wrn-70-16, we found that the cosine schedule performed slightly better than the multistep schedule (+0.24%).

The multistep schedule developed over the years, which has been tuned to Wide-ResNets and adversarial training, works well for settings with and without additional unlabeled data.

A.1.2 Number of optimization steps

Training for longer is not always beneficial, especially when is comes to adversarial training. The robust overfitting phenomenon studied by Rice et al. (2020) attests to the difficulty of finding the right number of steps to optimize for. In this experiment, we use the multistep learning rate schedule and change the number of training epochs. We expect to find an optimal schedule that balances overfitting with robustness.

In Table 7, we vary the number of training epochs between {50,100,200,400}\{50,100,200,400\} for the setting without additional data and between {100,200,400,800}\{100,200,400,800\} for the setting with additional data. For both settings, training for longer is not beneficial. Without additional data, using 400 epochs instead of 200 leads to degradation of the robust accuracy by -0.64%. With additional data, using 800 epochs instead of 400 leads to degradation of -1.43%. Fig. 6 shows the robust accuracy as training progresses for the setting without additional data. We observe that letting the model train for longer leads to robust overfitting.

A.2 Inner maximization

It is important that the attack used during training be strong enough. Weak attacks tend to provide a false sense of security by allowing the trained network to use obfuscation as a defense mechanism (Qin et al., 2019). In this experiment, we study the effect of the attack strength on robustness. Although recent work demonstrated that it is possible to train robust model with single-step attacks (Wong et al., 2020), it is generally accepted that the number of steps used for the inner optimization correlates with the strength of that optimization procedure (i.e., its ability to the find a minima close to the global minima). Note, however, that more inner steps leads to increased training time.

In Table 8, we vary the number of attack steps KK between 11 and 1616. We adapt the step-size by setting α\alpha to max⁡(1.25ϵ/K,0.007)\max(1.25\epsilon/K,0.007). We observe that stronger attacks yield more robust models (with diminishing returns). For example, increasing KK from 44 to 88, 88 to 1010 and 1010 to 1616 improves robust accuracy by +2.59%, +0.75% and +0.51% respectively (in the setting without additional data).

Strong inner optimizations improve robustness (at the cost of increased training time).

A.2.2 Inner maximization perturbation radius (continued)

This section continues the evaluation made in Sec. 4.2.2. In particular, we evaluate the effect of using a larger perturbation radius within TRADES (Zhang et al., 2019) (Sec. 4.2.2 explored this effect within classical adversarial training).

Table 9 shows the effect of increasing the perturbation radius ϵ\epsilon by a factor 1.1×1.1\times and 1.2×1.2\times the original value of 8/2558/255 (during training, not during evaluation). We notice that to the contrary of adversarial training, which benefits from increases perturbation radii, TRADES’ performance is inconsistent (possibly because TRADES is already actively managing the trade-off with clean accuracy).

When using TRADES, tuning the perturbation radius is not always beneficial.

A.3 Additional unlabeled data

This section continues the evaluation made in Sec. 4.3.2. We evaluate the effect of varying the labeled-to-unlabeled data ratio on our largest unlabeled dataset consisting of 1M images.

Fig. 8 shows the robust accuracy as we vary that ratio. Similarly to Fig. 2 (which used the dataset from Carmon et al., 2019), we observe that giving slightly more importance to unlabeled data helps. We find an identical optimal ratio of labeled-to-unlabeled data of 3:7, which provides a boost of +1.32% over the 1:1 ratio. When we use the dataset from Carmon et al. (2019), the optimal ratio provides a boost of +0.95% only. This could indicate that larger gains in robustness are possible when using larger unlabeled datasets.

Larger unlabeled datasets can provide larger improvements in robustness.

A.3.2 Label smoothing

As hinted in Sec. 4.3.2, we experiment with label smoothing (for the additional unlabeled data). Label smoothing should counteract the effect of the noisy labels resulting from the classifier used in the pseudo-labeling process. Label smoothing modifies one-hot labels yy by creating smoother targets y^=(1−γ)y+γ1\hat{y}=(1-\gamma)y+\gamma{\bm{1}}. More specifically, we minimize the following loss

In Table 10, we apply label smoothing to the examples originating from the additional unlabeled dataset (from Carmon et al., 2019). We vary γ\gamma for these examples only (i.e., the labeled data from Cifar-10 continues to use hard labels with γ=0\gamma=0). We observe that label smoothing is detrimental and that the resulting robust accuracy is inconsistent, without a clear correlation with γ\gamma.

Applying label smoothing to the additional unlabeled data is detrimental to robustness.

A.4 Other tricks

Using larger batch size not only influences resource utilization but also affects the optimization process. Larger batches provide less noisy gradients (e.g., when using Stochastic Gradient Descent) and more precise batch statistics. In classical adversarial training, it is common to let the batch statistics “float” as the inner optimization process is run.Xie & Yuille (2019) experimented with alternative batch normalization strategies. As such, batch size also influences the quality of the adversarial examples generated during training. In this experiment, we vary the batch size and compensate for the loss of gradient noise by scaling the outer learning rate using the linear scaling rule introduced by Goyal et al. (2017) (i.e., effective learning=max⁡(learning rate×batch size/256,learning rate)\textrm{effective learning}=\max(\textrm{learning rate}\times\textrm{batch size}/256,\textrm{learning rate})).

Table 11 shows the robust accuracy resulting from using different batch sizes. We observe that our default batch size of 128128 is sub-optimal for the setting without additional unlabeled data: a batch size of 512512 improves robust accuracy by +0.88%. For the setting with additional unlabeled data, the largest batch size (i.e., 10241024) remains the best, as it improves on smaller batch sizes by at least +0.66%.

Training batch size has an effect on robustness.

A.4.2 Data augmentation

Data augmentation can reduce generalization error. For image classification tasks, random flips, rotations and crops are commonly used (He et al., 2016). As is common for Cifar-10, in our baseline settings, we apply random translations by up to 4 pixels and random horizontal flips. More sophisticated techniques such as Cutout (DeVries & Taylor, 2017) (which produces random occlusions) and mixup (Zhang et al., 2018) (which linearly interpolates between two images) demonstrate compelling results on standard classification tasks. However, both techniques are not very effective when used in conjunction with adversarial training (Rice et al., 2020). In this experiment, we evaluate AutoAugment (Cubuk et al., 2019), RandAugment (Cubuk et al., 2020) and the AugMix augmentation (Hendrycks et al., 2020) (with their default settings). We also evaluate the color augmentation scheme proposed by Chen et al. (2020) (as part of the SimCLR pipeline) by varying the color jittering strength. Irrespective of the strength, this color augmentation scheme always includes a random color drop (i.e., conversion to gray-scale) with a probability of 20%.

Table 12 summarizes the results. We observe that AutoAugment, RandAugment and AugMix, which have mainly been tuned for ImageNet, reduce robust accuracy. These techniques would require further fine-tuning to be competitive with the simplest augmentation scheme (i.e., random translation by 4 pixels). Furthermore, we observe that increasing the strength of color jittering correlates negatively with robustness, as robust accuracy drops by -1.98% and -1.50% in the settings without and with additional data, respectively. Finally, we note that randomly dropping color does improve robustness in both settings (i.e., +0.69% and +0.29% with respect to the baselines).

Data augmentation schemes that perform well for standard classification tasks do not necessarily improve robust generalization.

A.4.3 Label smoothing

In this section, we complement the label smoothing experiment done in Sec. A.3.2. To the contrary of Sec. A.3.2, which explored label smoothing for unlabeled data only, this section explores label smoothing for the labeled data.

Table 13 shows the results. We do not observe any clear correlation between label smoothing and robust accuracy. In particular, setting the label smoothing factor γ\gamma to 0.02 and 0.2 seems helpful (in both data settings), while setting it to 0.05 or 0.1 seems detrimental in at least one of the two data settings.

Applying label smoothing to the labeled data has minor effects on robustness.

Appendix B Comparison with concurrent works

Independently of our work, Pang et al. (2020a)Accepted to the 2021 Conference on Learning Representations and available on ArXiv on October 1st, 2020. also investigate the limits of current approaches to adversarial training. We also highlight concurrent works by Chen et al. (2021)Accepted to the 2021 Conference on Learning Representations and available on OpenReview on September 28th, 2020. and Wu et al. (2020)Wu et al. (2020)Accepted at the 2020 Conference on Neural Information Processing Systems, but only available in its final form (with their latest results) from ArXiv on October 13th, 2020.. We briefly summarize both these works here for clarity and completeness.

This work is the closest in essence to ours. Pang et al. (2020a) make a thorough literature review and observe that different papers on adversarial robustness differ in their hyper-parameters (despite using similar techniques). While they analyze some properties that we also analyze in this manuscript (such as training batch size, label smoothing, weight decay, activation functions), they also complement our analyses with experiments on early stopping and perturbation radius warm-up, optimizers, model architectures beyond Wide-ResNets and batch normalization. The combination of their findings applied to a Wrn-34-20 reaches 54.39% robust accuracy without additional unlabeled data (in comparison, our Wrn-34-20 reaches a robust accuracy of 56.86%). Their study hints at further improvements by more finely tuning the weight decay.

In their manuscript, Wu et al. (2020) explore adversarial weight perturbation as way to improve robust generalization. Their technique interleaves model weight perturbations with example perturbations (within a single training step). They observe that the resulting loss landscapes become flatter and, as a result, robustness improves. Using this approach, Wu et al. (2020) improve robust accuracy by significant margins reaching 56.17% without additional data. By combining our findings with their technique, we expect that further improvements are possible (although we posit that model weight averaging may already have a similar effect).

Simultaneously to us, Chen et al. (2021) discovered that model weight averaging can significantly improve robustness on a wide range of models and datasets. They argue (similarly to Wu et al., 2020) that WA leads to a flatter adversarial loss landscape, and thus a smaller robust generalization gap. In addition to matching our experimental results on WA, they provide a deeper, noteworthy analysis.

Appendix C Additional detailed results

Appendix D Loss landscape analysis

Fig. 9, Fig. 10 and Fig. 11 shows the loss landscapes around the first 5 images of the Cifar-10 test set for our largest model trained with additional unlabeled data, our largest model trained without additional unlabeled data and Carmon et al.’s model trained with additional unlabeled data, respectively. Most landscapes are smooth and do not exhibit patterns of gradient obfuscation. Interestingly, the landscapes corresponding to the fifth image (of a dog) are quite similar across all models, with the larger models’ landscapes being less smooth (with a cliff). Overall, it is difficult to interpret these figures further and we rely on the fact that AutoAttack and MultiTargeted are accurate. We note that the black-box Square attack Andriushchenko et al. (2020) does not find any misclassified attack that was not already found by AutoPgd and MultiTargeted.