Adversarial Training Can Hurt Generalization
Aditi Raghunathan, Sang Michael Xie, Fanny Yang, John C. Duchi, Percy Liang
Introduction
Neural networks trained using standard training have very low accuracies on perturbed inputs commonly referred to as adversarial examples . Even though adversarial training can be effective at improving the accuracy on such examples (robust accuracy), these modified training methods decrease accuracy on natural unperturbed inputs (standard accuracy) . Table 1 shows the discrepancy between standard and adversarial training on Cifar-10. While adversarial training improves robust accuracy from 3.5% to 45.8%, standard accuracy drops from 95.2% to 87.3%.
Another explanation could be that the hypothesis class is not rich enough to contain predictors that have optimal standard and high robust accuracy, even if they exist . However, Table 1 shows that adversarial training achieves 100% standard and robust accuracy on the training set, suggesting that the hypothesis class is expressive enough in practice.
Having ruled out a conflict in the objectives and expressivity issues, Table 1 suggests that the tradeoff stems from the worse generalization of adversarial training either due to (i) the statistical properties of the robust objective or (ii) the dynamics of optimizing the robust objective on neural networks. In an attempt to disentangle optimization and statistics, we ask does the tradeoff indeed disappear if we rule out optimization issues? After all, from a statistical perspective, the robust objective adds information (constraints on the outputs of perturbations) which should intuitively aid generalization, similar to Lasso regression which enforces sparsity .
We answer the above question negatively by constructing a learning problem with a convex loss where adversarial training hurts generalization even when the optimal predictor has both optimal standard and robust accuracy. Convexity rules out optimization issues, revealing a fundamental statistical explanation for why adversarial training requires more samples to obtain high standard accuracy. Furthermore, we show that we can eliminate the tradeoff in our constructed problem using the recently-proposed robust self-training on additional unlabeled data.
In an attempt to understand how predictive this example is of practice, we subsample Cifar-10 and visualize trends in the performance of standard and adversarially trained models with varying training sample sizes. We observe that the gap between the accuracies of standard and adversarial training decreases with larger sample size, mirroring the trends observed in our constructed problem. Recent results from show that, similarly to our constructed setting, robust self-training also helps to mitigate the trade-off in Cifar-10.
Standard vs. robust generalization.
Recent work has focused on the sample complexity of learning a predictor that has high robust accuracy (robust generalization), a different objective. In contrast, we study the finite sample behavior of adversarially trained predictors on the standard learning objective (standard generalization), and show that adversarial training as a particular training procedure could require more samples to attain high standard accuracy.
Convex learning problem: the staircase
We construct a learning problem with the following properties. First, fitting the majority of the distribution is statistically easy—it can be done with a simple predictor. Second, perturbations of these majority points are low in probability and require complex predictors to be fit. These two ingredients cause standard estimators to perform better than their adversarially trained robust counterparts with a few samples. Standard training only fits the training points, which can be done with a simple estimator that generalizes well; adversarial training encourages fitting perturbations of the training points making the estimator complex and generalize poorly.
The central premise of this work is that the optimal predictor is robust. In our construction, we let be robust by enforcing the invariance property (see Appendix A)
2 Construction
In our construction, we consider linear predictors as “simple” predictors that generalize well and staircase predictors as “complex” predictors that generalize poorly (Figure 1(a)).
In order to satisfy the property that a simple predictor fits most of the distribution, we define to be linear on the set , where
for parameters and a positive integer . Any predictor that fits points in will have low (but not optimal) standard error when is small.
Perturbations.
Output distribution.
For any point in the support ,
for some parameter . Setting the slope as makes resemble a staircase. Such an satisfies the invariance property (1) that ensures that the optimal predictor for standard error is also robust. Note that (a simple linear function) when restricted to in . Note also that the invariance sets are disjoint. This is in contrast to the example in , where any invariant function is also globally constant. Our construction allows a non-trivial robust and accurate estimator.
We generate the output by adding Gaussian noise to the optimal predictor , i.e., where .
3 Simulations
We empirically validate the intuition that the staircase problem is sensitive to robust training by simulating training with various sample sizes and comparing the test MSE of the standard and robust estimators (2) and (3). We report final test errors here; trends in generalization gap (difference between train and test error) are nearly identical. See Appendix D for more details.
Figure 2 shows the difference in test errors of the two estimators. For each sample size , we compare the standard and robust estimators by performing a grid search over regularization parameters that individually minimize the test MSE of each estimator. With few samples, most training samples are from and standard training learns a simple linear predictor that fits all of . On the other hand, robust estimators fit the low probability perturbations , leading to staircases that generalize poorly. Figure 1(b) visualizes the two estimators for small samples. However, as we increase the size of the training set, the training set contains all points from , and robust estimators also generalize well despite being more complex. Furthermore, in this regime, robust estimators indeed see the expected “regularization” benefit where the robust objective helps fit points in the low probability regions , even when they are not yet sampled in the training points. In general, we see that robust training has higher test error with a small sample size, but the difference in the test error of standard and robust estimators decreases as sample size increases, and robust training eventually obtains lower test error.
Another common approach to encoding invariances is data augmentation, where perturbations are sampled from and added to the dataset. Data augmentation is less demanding than adversarial training which minimizes loss on the worst-case point within the invariance set. We find that for our staircase example, an estimator trained even with the less demanding data augmentation sees a similar tradeoff with small training sets, due to increased complexity of the augmented estimator.
4 Robust self-training mostly eliminates the tradeoff
Section 2.3 shows that the gap between the standard errors of robust and standard estimators decreases as training sample size increases. Moreover, if we obtained training points spanning , then the robust estimator (staircase) would also generalize well and have lower error than the standard estimator. Thus, a natural strategy to eliminate the tradeoff is to sample more training points. In fact, we do not need additional labels for the points on —a standard trained estimator fits points on with just a few labels, and can be used to generate labels on additional unlabeled points. Recent works have proposed robust self-training (RST) to leverage unlabeled data for robustness . RST is a robust variant of the popular self-training algorithm for semi-supervised learning , which uses a standard estimator trained on a few labels to generate psuedo-labels for unlabeled data as described above. See Appendix C for details on RST.
For the staircase problem (), RST mostly eliminates the tradeoff and achieves similar test error to standard training (while also being robust, see Appendix C.2) as shown in Figure 2.
Experiments on Cifar-10
In our staircase problem from Section 2, robust estimators perform worse on the standard objective because these predictors are more complex, thereby generalizing poorly. Does this also explain the drop in standard accuracy we see for adversarially trained models on real datasets like Cifar-10?
Adversarial training can also help
One of the key ingredients that causes the tradeoff in the staircase problem is the complexity of robust predictors. If we change our construction such that robust predictors are also simple, we see that adversarial training instead offers a regularization benefit. When , the optimal predictor (which is robust) is linear (Figure 1(d)). We find that adversarial training has lower standard error by enforcing invariance on making the robust estimator less sensitive to target noise (Figure 4(a)).
Similarly, on Mnist , the adversarially trained model has lower test error than standard trained model. As we increase the sample size, both standard and adversarially trained models converge to obtain same small test error. We remark that our observation on Mnist is contrary to that reported in , due to a different initialization that led to better optimization (see Appendix Section D.2).
Conclusion
In this work, we shed some light on the counter-intuitive phenomenon where enforcing invariance respected by the optimal function could actually degrade performance. Being invariant could require complex predictors and consequently more samples to generalize well. Our experiments support that the tradeoff between robustness and accuracy observed in practice is indeed due to insufficient samples and additional unlabeled data is sufficient to mitigate this tradeoff.
Acknowledgements
We are grateful to Tengyu Ma for several helpful discussions. This work was funded by an Open Philanthropy Project Award and NSF Frontier Award Grant no. 1805310. AR was supported by Google Fellowship and Open Philanthropy AI Fellowship. FY was supported by the Institute for Theoretical Studies ETH Zurich and the Dr. Max Rössler and the Walter Haefner Foundation. FY and JCD were supported by the Office of Naval Research Young Investigator Award N00014-19-1-2288.
References
Appendix A Consistency of robust and standard estimators
If both (2) and (3) converge to the same Bayes optimal as , we say that the two estimators and are consistent. In this section, we show that the invariance condition (7) implies consistency of and .
Intuitively, from (7), since is invariant for all in , the maximum over in the robust objective is achieved by the unperturbed input (and also achieved by any other element of ). Hence the standard and robust loss of are equal. For any other predictor, the robust loss upper bounds the standard loss, which in turn is an upper bound on the standard loss of (since is Bayes optimal). Therefore also obtains optimal robust loss and and are consistent and converge to with infinite data.
For the classification case, consistency requires label invariance, which is that
such that the adversary cannot change the label that achieves the maximum but can perturb the distribution.
The optimal standard classifier here is the Bayes optimal classifier . Assuming that is in , then consistency follows by essentially the same argument as in the regression case.
In our staircase problem, from (1), we assume that the target is generated as follows: where , we see that the points within an invariance sets have the same target distribution (target distribution invariance).
The target invariance condition above implies consistency in both the regression and classification case.
Appendix B Convex staircase example
We focus on a 1-dimensional regression case. Let be the total number of “stairs” in the staircase problem. Let be the number of stairs that have a large weight in the data distribution. Define to be the probability of sampling a perturbation point, i.e. , which we will choose to be close to zero. The size of the perturbations is , which is bounded by so that , for any . The standard deviation of the noise in the targets is . Finally, is a parameter controlling the slope of the points in .
Let be a distribution over where is the probability simplex of dimension . We define the data distribution with the following generative process for one sample . First, sample a point from according to the categorical distribution described by , such that . Second, sample by perturbing with probability such that
Note that this is just a formalization of the distribution described in Section 2. The sampled is in with probability and with probability , where we choose to be small.
In addition, in order to exaggerate the difference between robust and standard estimators for small sample sizes, we set such that the first stairs have the majority of probability mass. To achieve this, we set the unnormalized probabilities of as
and define by normalizing . For our examples, we fix . In general, even though we can increase to create versions of our example with more stairs, is fixed to highlight the bad extrapolation behavior of the robust estimator.
Distribution of 𝒴𝒴\mathcal{Y}.
We define the target distribution as , where rounds to the nearest integer. The invariance sets are . We define the distribution such that for any , all points in have the same mean target value . See Figure 1 for an illustration.
B.2 Model
For some regularization parameter we optimize with the penalized smoothing spline loss function over parameters ,
where measures smoothness in terms of the second derivative.With respect to the regularized objectives (2) and (3), the norm regularizer is .
B.3 Role of different parameters
To construct an example where robustness hurts generalization, the main parameters needed are that the slope is large and that the probability of drawing samples from perturbation points is small. When slope is large, the complexity of the true function increases such that good generalization requires more samples. A small ensures that a low-norm linear solution has low test error. This example is insensitive to whether there is label noise, meaning that is sufficient to observe that robustness hurts generalization.
If , then the complexity of the true function is low and we observe that robustness helps generalization. In contrast, this example relies on the fact that there is label noise () so that the noise-cancelling effect of robust training improves generalization. In the absence of noise, robustness neither hurts nor helps generalization since both the robust and standard estimators converge to the true function () with only one sample.
B.4 Plots of other values
We show plots for a variety of quantities against number of samples . For each , we pick the best regularization parameter with respect to standard test MSE individually for robust and standard training. in the (robustness hurts) and (robustness helps) cases, with all the same parameters as before. In both cases, the test MSE and generalization gap (difference between training MSE and test MSE) are almost identical due to robust and standard training having similar training errors. In the case where robustness hurts (Figure 6), robust training finds higher norm estimators for all sample sizes. With enough samples, standard training begins to increase the norm of its solution as it starts to converge to the true function (which is complex) and the robust train MSE starts to drop accordingly.
In the case where robustness helps (Figure 7), the optimal predictor is the line , which has 0 norm. The robust estimator has consistently low norm. With small sample size, the standard estimator has low norm but has high test MSE. This happens when the standard estimator is close to linear (has low norm), but the estimator has the wrong slope, causing high test MSE. However, in the infinite data limit, both standard and robust estimators converge to the optimal solution.
Appendix C Robust self-training algorithm
We describe the robust self-training procedure, which performs robust training on a dataset augmented with unlabeled data. The targets for the unlabeled data are generated from a standard estimator trained on the labeled training data. Since the standard estimator has good standard generalization, the generated targets for the unlabeled data have low error on expectation. Robust training on the augmented dataset seeks to improve both the standard and robust test error of robust training (over just the labeled training data). Intuitively, robust self-training achieves these gains by mimicking the standard estimator on more of the data distribution (by using unlabeled data) while also optimizing the robust objective.
Compute the standard estimator (2) on the labeled data with regularization parameter .
Generate pseudo-targets by evaluating the standard estimator obtained above on the unlabeled data .
Construct an augmented dataset , .
Return a robust estimator (3) with the augmented dataset as training data.
We present relevant results from the recent work of on robust self-training applied on Cifar-10 augmented with unlabeled data in Table 2. The procedure employed in is identical to the procedure describe above, using a modified version of adversarial training (TRADES) as the robust estimator.
C.2 Robust self-training doesn’t sacrifice robustness
In Section 2.4, we show that if we have access to additional unlabeled samples from the data distribution, robust self-training (RST) can mitigate the tradeoff in standard error between robust and standard estimators. It is important that we do not sacrifice robustness in order to have better standard error. Figure 5 shows that in the case where robustness hurts generalization in our convex construction (), RST improves over robust training not only in standard test error (Section 2.4), but also in robust test error. Therefore, by leveraging some unlabeled data, we can recover the standard generalization performance of standard training using RST while simultaneously improving robustness.
Appendix D Experimental details
D.2 Mnist
We note here that the tradeoff for adversarial training reported in is because the adversarially trained model hasn’t converged (even after a large number of epochs). Using the Xavier initialization, we get faster convergence with adversarial training and see no drop in clean accuracy at the same level of robust accuracy. Interestingly, standard training is not affected by initialization, while adversarial training is dramatically affected.