The Effects of Regularization and Data Augmentation are Class Dependent

Randall Balestriero, Leon Bottou, Yann LeCun

Introduction

Machine learning and deep learning aim at learning systems to solve as accurately as possible a given task at hand (LeCun et al.,, 1998; Bishop and Nasrabadi,, 2006; Jordan and Mitchell,, 2015). This process often takes the form of being given a finite training set and a performance measure, optimizing the system’s parameters e.g. from gradient updates, and assessing the system’s performance on test set samples, i.e. samples that were not used during the system optimization. As the training set is finite, and the optimal design of the system is unknown, it is common to employ regularization during the optimization phase to reduce over-fitting (Tikhonov,, 1943; Tihonov,, 1963) i.e. to decrease the system’s performance gap between train set and test set samples (Simard et al.,, 1991; Chapelle et al.,, 2000; Bottou,, 2012; Neyshabur et al.,, 2014).

Data-Augmentation (DA) is a data-driven and informed regularization strategy that artificially increase the number of training samples (Shorten and Khoshgoftaar,, 2019). As opposed to most explicit regularizers e.g. Tikhonov regularization (Krogh and Hertz,, 1991), also denoted as weight decay, DA’s regularization is implicit as it is not a function of a model’s parameter, but a function of the training samples (Neyshabur et al.,, 2014; Hernández-García and König,, 2018; LeJeune et al.,, 2019); although some DA strategies can be turned into explicit regularizers Balestriero et al., (2022). Nevertheless, a key distinction between DA and weight decay is that DA requires more domain knowledge to be successful than weight decay. Most —if not all— of current state-of-the-art employ such regularizers (Huang et al.,, 2018; Chen et al., 2020b, ; Liu et al.,, 2021; Tan and Le,, 2021; Liu et al.,, 2022).

In this paper, we will demonstrate that when employing regularization such as DA or weight decay, a significant bias is introduced into the trained model. In particular, the regularized model exhibits strong per-class favoritism i.e. while the average test accuracy over all classes improves when employing regularization, it is at the cost of the model becoming arbitrarily inaccurate on some specific classes as illustrated in fig. 1. After a brief theoretical justification on why and when DA can be the cause of bias (section 2.1), we propose a dedicated sensitivity analysis of the bias produced by different amounts of DA in sections 2.2 and 2.3, which is followed by a similar study dedicated to weight decay in section 2.4 and transfer learning in section 2.5. We shall highlight that although we perform a class-level study, it is possible to refine this entire analysis at the sample-level.

For readers familiar with statistical estimation results e.g. the bias-variance trade-off (Kohavi et al.,, 1996; Von Luxburg and Schölkopf,, 2011) or bayesian estimation e.g. Tikhonov regularization (Box and Tiao,, 2011; Gruber,, 2017), it should not be surprising that regularization produces bias. In fact, it is often beneficial to introduce bias through regularization if it results in a significant reduction of the estimator variance —when one aims to minimize the average empirical risk. This is one of the main reason behind the success of techniques such as ridge regression. However, what is potentially dangerous is that the bias introduced by regularization treats classes differently, including on transfer learning tasks as we will demonstrate in section 2.5. Those observations also support recent theoretical results tying a model’s performance to its robustness and to DA, as we discuss in appendix A.

Regularization Creates Class-Dependent Model Bias that can be Harmful even for Transfer Learning Tasks

The first part of our study focuses on DA, a technique that regularizes a model by introducing new training samples, derived from the observed ones. DA samples have been known to sometimes disregard the semantic information of the original samples (Krizhevsky et al.,, 2012). Nevertheless, DA remains applied universally, and fearlessly across tasks and datasets (Shorten and Khoshgoftaar,, 2019) as it provides significant performance improvements, even in semi-supervised and unsupervised settings Guo et al., (2018); Xie et al., (2020); Misra and Maaten, (2020). We first provide in sections 2.1 and 2.2 some intuition on why DA can be a source of bias regardless of the task, dataset and model at hand. We then quantify the amount of bias caused by DA in various realistic scenarios in LABEL: and 2.3; and extend our analysis to weight decay in section 2.4. Finally, we conclude by demonstrating how the bias introduced by regularization transfers to downstream tasks e.g. when deploying an Imagenet (source) trained model on the INaturalist (target) dataset in section 2.5; that scenario is key as it demonstrates the potential harm of selecting the best performing model —on average— on the source dataset which could turn out to also be the most biased model on the target dataset class of interest. In fact, it is crucial to remember that regularization, or any other form of structural risk minimization, improves generalization performances by increasing the bias of the estimator so that the estimator’s variance is decreased by a greater amount. However, nothing guarantees the fairness of this bias i.e. for it to be equally distributed amongst the dataset classes.

Whenever the transformations produced by Tα,∀α\mathcal{T}_{\alpha},\forall\alpha do not respect the level-set of f∗f^{*}, and whenever the model has enough capacity to minimize the training loss, the DA will create irreducible bias in fθf_{\theta} as in

The main idea of the proof, provided in appendix B, is to show that if a transformation does not move samples on the level-set of the true function (left-hand-side of eq. 1), then fθf_{\theta} will learn a different level (since it has training error), and thus ∥f∗−fθ∥>0\|f^{*}-f_{\theta}\|>0 i.e. fθf_{\theta} is biased regardless of the training set.

Whenever the left-hand-side of eq. 1 is , the DA is denoted as label-preserving (Cui et al.,, 2015; Taylor and Nitschke,, 2018). From the above, we see that unless the target y{\bm{y}} associated to Tα(x)\mathcal{T}_{\alpha}({\bm{x}}) is modified accordingly to encode the shift in the target function level-set produced by Tα\mathcal{T}_{\alpha}, any DA that is not label-preserving will introduce a bias. Some DAs propose to incorporate label transformation i.e. not only x{\bm{x}} but also y{\bm{y}} is augmented to better inform on the uncertainty that has been added into Tθ(x)\mathcal{T}_{\theta}({\bm{x}}). This is for example the case for MixUp (Zhang et al.,, 2017), ManifoldMixUp (Verma et al.,, 2019), CutMix (Yun et al.,, 2019) and their extensions.

Our goal in the next section 2.2 is to demonstrate how DAs such as random crop, color jittering, or CutOut are only label preserving for some values of α\alpha that vary with the sample class. As a consequence, while the use of the DA improves the average test performance, it is at the cost of a significant reduction in performance for some of the classes.

2 The Same Data-Augmentation can be Label-Preserving or Not Between Different Classes

In the previous section 2.1 we provided a general argument on the sufficient conditions for DA to produce a biased model. We hope in this section to provide a more concrete example that applies to current DN training. To that end, we will demonstrate that a DA can be label-preserving or not depending on the sample’s class, hence, since the same DA policy is employed for all classes, the augmented dataset will exhibit a class-imbalance in favor of the classes for which the DA is most label-preserving.

To measure by how much a given DA, Tα\mathcal{T}_{\alpha}, is label-preserving, we propose to take 66 popular architectures that are pre-trained on Imagenet (Deng et al.,, 2009) from the official PyTorch (Paszke et al.,, 2019) repository, and to evaluate their accuracy performances for varying DA settings (top of fig. 2). We observe that when considering the dataset as a whole, it is possible to identify a DA regime as a function of α\alpha for which the amount of information present in Tα(x)\mathcal{T}_{\alpha}({\bm{x}}) becomes insufficient to predict the correct label on average. But more interestingly, we also take the per-class accuracy performance (bottom of fig. 2) and observe that for some classes, any level of transformation α\alpha can produce augmented samples with enough information to be correctly classified, while other classes see their samples become unpredictable as soon as Tα\mathcal{T}_{\alpha} moves away from the identity mapping. Note that we report test set performances ensuring that the observed performances are not due to ad-hoc memorization.

To further ensure that the observed relation between label-preservation, sample class, and amount of transformation α\alpha is sound, we provide in fig. 3 the per-class test accuracy on different models, all exhibit the same trends. In short, we identify that when creating an augmented dataset by applying the same DA across classes, the number of per-class samples that actually contain enough information about their true labels will become largely imbalance between classes, even if the original dataset was balanced. Any model trained on the augmented dataset will thus focus on the classes for which the DA is the most label-preserving. We propose in the next section to precisely quantify the impact of DA on each dataset class.

3 Measuring the Average Treatment Effect of Data-Augmentation on Models’ Class-Dependent Bias

This section aims at quantifying precisely the amount of downward or upward per-class performance shift that came as a result from using DA. We thus propose a sensitivity analysis by training a large collection of models with varying DA policies to precisely assess the relation between DA and class-dependent model bias.

First, we propose in fig. 4 a sensitivity analysis by training the same architecture on Imagenet with varying DA policies. In particular, we consider a given DA (random crop in this case) and we vary the support of the parameter α\alpha which represents how much of the original image is kept in the crop. We train DNs using α∈[100,τ]\alpha\in[100,\tau] with τ\tau varying from 100100 to 88 and for each case, we report our metrics averaged over 2020 trained models. We observe a clear relation between increase in the strength of the DA, increase in the average test accuracy overall classes, and decrease in some per-class test accuracies. For example, on a resnet50 Imagenet setting, the accuracy on the “academic gown” class goes from 62% to 40% steadily as τ\tau decreases. We defer the same experiment but using weight decay in the next section 2.4.

4 Uninformed Weight Decay Also Creates Class-Dependent Model Bias

We quantified in section 2.3 how much per-class bias was produced by DA. As per the arguments given in sections 2.1 and 2.2, it would be natural to assume that what makes DA responsible for creating class-dependent bias in DNs is our misfortune in defining correct augmentation policies, and thus, that other regularization techniques that are uninformed e.g. weight decay would behave differently. The goal of this section is to demonstrate that weight decay also suffers from the same class-dependent behavior indicating that designing a regularizer for DNs that is fair across classes might require novel innovative solutions.

The basic formulation of weight-decay consists in having any loss to be minimized L\mathcal{L} and add to it an additional term denoted as γ∥θ∥22\gamma\|\theta\|_{2}^{2} where θ\theta collects the model’s parameters and γ\gamma is the strength of the regularization (proportional to the restriction on the trained model complexity). In general, we do not incorporate the bias term(s) within θ\theta as this would directly remove the ability of the model to learn the natural class prior (Hastie et al.,, 2009). As a result, and throughout this study, we consider θ\theta to incorporate all the DN parameters except for the ones of batch-normalization layers, as commonly done in deep learning (Leclerc et al.,, 2022). We report in fig. 5 the per-class performance of a resnet50 trained on Imagenet with varying weight decay coefficient γ\gamma (as was done for DA in fig. 4) and we observe that different classes have different test accuracy sensitivities to variations in γ\gamma. Some will see their generalization performance increase, while others will have decreasing generalization performances. That is, even for uninformative regularizers such as weight decay a per-class bias is introduced, reducing performances for some of the classes. The next section 2.5 proposes to quantifying how much bias transfers to different downstream tasks.

5 The Class-Dependent Bias Transfers to Other Downstream Tasks

The last experiment we propose is to quantify the amount of class-dependent bias that transfers to other downstream tasks, a common situation in transfer learning and in system deployment to the real world (Pan and Yang,, 2009). We thus want to measure how regularization applied during the pre-training phase on a source dataset impacts the per-class accuracy of that model on the target dataset.

In order to keep the setting similar to section 2.3, we adopt a resnet50 model with random crop DA. That model is pre-trained on Imagenet dataset (source) with varying value of τ\tau (random crop lower bound) and then, the trained model is transferred to the INaturalist dataset (Van Horn et al.,, 2018) (target) that consists of 10,000 classes. When transferring the model to INaturalist, the parameters are kept frozen, and only a linear classifier is trained on top of it. We report in fig. 6 the performance of the trained models with varying τ\tau on different INaturalist classes. We observe once again that the best resnet50 —on average— is not necessarily the one that should be deployed as there exists a strong per-class bias that varies with τ\tau. As a result, picking the best performing model from a source dataset to a target dataset, might leave the pipeline to perform poorly since that model might also be the one that is the most biased against the class of interest in the target dataset.

This result should motivate the design of novel regularizers that do not reduce performances between classes at different regimes. Additionally, due to the cost of training multiple models with varying regularization settings, one might wonder on the possible alternative solutions to detect trends such as shown in fig. 6 only when given a single pre-trained model.

Conclusion

We proposed in this study to understand the impact of regularization, in particular data-augmentation and weight decay, into the final performances of a deep network. We obtained that the use of regularization increases the average test performances at the cost of significant performance drops on some specific classes. By focusing on maximizing aggregate performance statistics we have produced learning mechanisms that can be potentially harmful, especially in transfer learning tasks. In fact, we have also observed that varying the amount of regularization employed during pre-training of a specific dataset impacts the per-class performances of that pre-trained model on different downstream tasks e.g. going from Imagenet to INaturalist.

References

Appendix A Theoretical Relations Between Data-Augmentation, Generalization, and a Model’s Bias and Robustness

Bias-Variance Decomposition. The expected error of a trained model can always be decomposed into two terms that can be tune by altering the considered model class and regularization, and a third term which is the inherent measurement noise, that can not be reduced. In fact, the true labels are obtained from y=f(x)+ϵ{\bm{y}}=f({\bm{x}})+\epsilon with ff the true, unknown, data model, and ϵ\epsilon some irreducible error i.e. coming from measurements or data compression. For the Mean-Squared Error (MSE) we obtain this decomposition as

Although we only derive here the MSE case, the same can be obtained in the multivariate case and in classification settings (Valentini and Dietterich,, 2004; Hastie et al.,, 2009). When the model is with low complexity e.g. the model is under-parametrized, or the regularization is applied aggressively, the variance term becomes small and the bias term increases. Conversely, when the model is not regularized and is overparametrized, the variance term increases and the bias term reduces. One strategy to find the best model is offered by structural risk minimization.

Structural Risk Minimization (SRM). As introduced by Vapnik and Chervonenkis, (1974), SRM proposes a search strategy to obtain the best model. One first chooses a class of function that f^\hat{f} must live in, hopefully informed from a priori knowledge e.g. polynomials of degree kk, resnet50 DN architecture. Then, one finds a hierarchical construction of nested functional spaces that relate to the function complexity. For example, it could be considering polynomials of increasing degree, up to kk, or considering resnet50s with weights being bounded by an increasing constant. Finally, one minimizes the empirical risk (training error) for each model and picks the one with best valid set performances. The valid set performance can be estimated empirically from a subset of the training set that has been set apart before fitting, or from other measures such as the VC-dimension of the trained model (Vapnik and Chervonenkis,, 2015). For example, cross-validating the weight decay parameter of a model corresponds to SRM. We make this connection precise below.

From Data-Augmentation and Weight Decay to Nested Functional Spaces.Regularization is commonly employed during training to prevent overfitting i.e. reduce the model complexity. One might argue that regularization does not strictly restrict the model functional space, it simply favors simpler models to be used. Yet, it turns out that regularization can be cast as explicitly restricting the parameter space of the model i.e. restricting the functional space of the model.

setting λ←γ\lambda\leftarrow\gamma, θ←θ^(γ)\theta\leftarrow\hat{\theta}(\gamma), and β←∥θ^(γ)∥22,s←0\beta\leftarrow\|\hat{\theta}(\gamma)\|_{2}^{2},s\leftarrow 0 solves the system. As a result, solving the constrained optimization problem with β=∥θ^(γ)∥22\beta=\|\hat{\theta}(\gamma)\|_{2}^{2} and solving the Tikhonov regularized problem with γ\gamma coefficient is equivalent. Since we also have that ∥θ^(γ1)∥22≤∥θ^(γ2)∥22\|\hat{\theta}(\gamma_{1})\|_{2}^{2}\leq\|\hat{\theta}(\gamma_{2})\|_{2}^{2}, we obtain that starting from a high value of Tikhonov regularization and gradually reducing it produces nested functional spaces where the model live in. More general results can be obtained in different regularization settings in Kloft et al., (2009); Hastie et al., (2009). Moving to the case of DA, the same procedure applies. In fact, one can use the results of Phaisangittisagul, (2016); LeJeune et al., (2019); Balestriero et al., (2022) to cast DAs such as dropout, and image perturbations as explicit regularizers à la Tikhonov and repeat the above procedure. In fact, DA can be seen as producing a model with lower variance and a bias towards the augmented dataset. Hence, if the augmented dataset is not aligned with the true underlying data distribution, DA will increase the model’s bias. For results studying DA in the context of structural risk minimization, we refer the reader to Chapelle et al., (2000); Arjovsky et al., (2019); Chen et al., 2020a ; Ilse et al., (2021). In short, it is clear from the above that DA is yet another form of regularization that can be applied for structural risk minimization. Yet, even if DA effectively restrict the functional space of the model, it can be at the cost of producing strong biases, as we empirically observed in section 2.

Provable Bias Induced by Data-Augmentation. Especially relevant to our results is a recent result of Xu et al., (2020). In this work, it was theorized what when the underlying dataset contains inherent biases, training on the original data is more effective than employing an i.i.d. DA policy, i.e. applying the same random augmentation to all samples/classes, to produce an unbiased model. In short, the DA policy exacerbates the already present biases and makes the trained model further away from the unbiased optimum. As per our experiments from section 2, we observe that bias can take many form. For example, having classes with different statistics e.g. objects that appear at very specific scale in the image. One strategy proposed by Xu et al., (2020), assuming that the bias of a model can be measured, was to adapt the DA policy to the dataset at hand to actively correct the present bias, as done for example in McLaughlin et al., (2015); Iosifidis and Ntoutsi, (2018); Jaipuria et al., (2020). In fact, when the bias is measurable, many recent studies have shown that DAs could be designed to reduce the bias present in a dataset (McLaughlin et al.,, 2015; Jaipuria et al.,, 2020).

A.2 Known Tradeoff Between Regularization, Generalization and Robustness

There exist a few previous work in that direction, especially in the field of robust and overparametrized machine learning.

In Raghunathan et al., (2020) a surprising result demonstrated that the minimum norm interpolant of the original + DA could have a larger standard error than that of the original data’s minimum norm interpolant. Even more surprising, this result was obtained using consistent DA i.e. transformations that do not alter the label information of the samples as in py∣x=py∣Tα(x)p_{{\bm{y}}|{\bm{x}}}=p_{{\bm{y}}|\mathcal{T}_{\alpha}({\bm{x}})}, or f∗(x)=f∗(Tα(x))f^{*}({\bm{x}})=f^{*}(\mathcal{T}_{\alpha}({\bm{x}})) as discussed in eq. 1. This possible degradation of performance occurs as long as the model remains over-parametrized, even when considering the original+DA dataset.

In addition to the implication of DA into bias, there exists an intertwined relationship between DA and model robustness. In fact, even assuming the use of perfectly adapted DAs, there exists an inherent tradeoff between accuracy and robustness that holds even in the infinite data limit (Tsipras et al.,, 2018; Fawzi et al.,, 2018; Zhang et al.,, 2019). For example, Min et al., (2021) proved in the robust linear classification regime that (i) more data improves generalization in a weak adversary regime, (ii) more data can improve generalization up to a point where additional data starts to hurt generalization in a medium adversary regime, and that (iii) more data immediately decreases generalization error in a strong adversary regime.

Lastly, a few studies have started to study the bias the could be cause from some DAs, although those studies focused on learned DAs e.g. from Generative Adversarial Networks (Hu and Li,, 2019). In that scenario, the predicament is that the GAN in itself is biased, and thus any GAN-generated DA will inherit those biases.

Appendix B Proof of 1

To streamline our derivations and without loss of generality, we first impose that T−α\mathcal{T}_{-\alpha} inverts the action of Tα\mathcal{T}_{\alpha}, that T0\mathcal{T}_{0} acts as the identity mapping and that Tα∘Tβ=Tα+β\mathcal{T}_{\alpha}\circ\mathcal{T}_{\beta}=\mathcal{T}_{\alpha+\beta}. In short, we have a group structure that allows us to define the following equivalence class ∼\sim that is positive iff two images are related by a transformation, formally defined as

which is reflexive (u∼u{\bm{u}}\sim{\bm{u}} with α=0\alpha=0), symmetric (u∼v{\bm{u}}\sim{\bm{v}} with α\alpha implies that v∼u{\bm{v}}\sim{\bm{u}} with −α-\alpha), and transitive (u∼v{\bm{u}}\sim{\bm{v}} with α\alpha and v∼w{\bm{v}}\sim{\bm{w}} with β\beta implies that u∼w{\bm{u}}\sim{\bm{w}} with α+β\alpha+\beta).

where the last equality assumes that the model has training error, and will thus predict the same constant for all DA version of the same sample. And since that last equation is always positive if the DA does not respect the true level-sets of the true function f∗f^{*} we obtain the desired result. The same derivation can be carried out with any desired loss e.g. for classification tasks. ∎

Appendix C Additional Figures