Improving Robustness Without Sacrificing Accuracy with Patch Gaussian Augmentation

Raphael Gontijo Lopes, Dong Yin, Ben Poole, Justin Gilmer, Ekin D. Cubuk

Introduction

Modern deep neural networks can achieve impressive performance at classifying images in curated datasets . Yet their performance is not robust to corruptions that typically occur in real-world settings. For example, neural networks are sensitive to small translations and changes in scale , blurring and additive noise , small objects placed in images , and even different images from a similar distribution of the training set . For models to be useful in the real world, they need to be both accurate on a high-quality held-out set of images , which we refer to as “clean accuracy,” and robust on corrupted images, which we refer to as “robustness.” Most of the literature in machine learning has focused on architectural changes to improve clean accuracy but has recently become interested in robustness as well.

Research in neural network robustness has tried to quantify the problem by establishing benchmarks that directly measure it and comparing the performance of humans and neural networks . Others have tried to understand robustness by highlighting systemic failure modes of current learning methods. For instance, networks exhibit excessive invariance to visual features , texture bias , sensitivity to worst-case (adversarial) perturbations , and a propensity to rely solely on non-robust, but highly predictive features for classification . Of particular relevance to our work, Ford et al show that in order to get adversarial robustness, one needs robustness to noise-corrupted data.

Another line of work has attempted to increase model robustness performance, either by directly projecting out superficial statistics , via architectural improvements , pre-training schemes , or through the use of data augmentations. Data augmentation increases the size and diversity of the training set, and provides a simple method for learning invariances that are challenging to encode architecturally . Recent work in this area includes learning better transformations , inferring combinations of transformations , unsupervised methods , theory of data augmentation , and applications for one-shot learning .

Despite these advances, individual data augmentation methods that improve robustness do so at the expense of reduced clean accuracy. Some have even claimed that there exists a fundamental trade-off between the two . Because of this, many recent works focus on improving either one or the other . In this work we propose a data augmentation that overcomes this trade-off, achieving both improved robustness and clean accuracy. Our contributions are as follows:

We characterize a trade-off between robustness and accuracy among two standard data augmentations: Cutout and Gaussian (Section 2.1).

We devise a simple data augmentation method (which we term Patch Gaussian) that allows us to interpolate between the two augmentations above. (Section 3.1)

We find that Patch Gaussian allows us to overcome the observed trade-off (Figure 1, Section 4.1), and achieves a new state of the art in the Common Corruptions benchmark on CIFAR-C and ImageNet-C. (Section 4.2)

We demonstrate that Patch Gaussian can be combined with other regularization strategies (Section 4.3) and data augmentation policies (Section 4.4) , and can improve COCO object detection performance as well (Section 4.5).

We perform a frequency-based analysis of models trained with Patch Gaussian and find that they can better leverage high-frequency information in lower layers, while not being too sensitive to them at later ones (Section 5.1).

Preliminaries

We start by considering two data augmentations: Cutout and Gaussian . The former sets a random patch of the input images to a constant (the mean pixel in the dataset) and is successful at improving clean accuracy. The latter works by adding independent Gaussian noise to each pixel of the input image, which can increase robustness to Gaussian noise directly.

To apply Gaussian, we uniformly sample a standard deviation σ\sigma from 0 up to some maximum value σmax\sigma_{max}, and add i.i.d. noise sampled from N(0,σ2)\mathcal{N}(0,\sigma^{2}) to each pixel. To apply Cutout, we use a fixed patch size WW, and randomly set a square region with size W×WW\times W to the constant mean of each RGB channel in the dataset. As in , the patch location is randomly sampled and can lie outside of the 32×3232\times 32 CIFAR-10 (or 224×224224\times 224 ImageNet) image but its center is constrained to lie within it. Patch sizes and σmax\sigma_{max} are selected based on the method described in Section 3.2.

We compare the effectiveness of Gaussian and Cutout data augmentation for accuracy and robustness by measuring the performance of models trained with each on clean data, as well as data corrupted by various standard deviations of Gaussian noise. Figure 2 highlights an apparent trade-off in using these methods. In accordance to previous work , Cutout improves accuracy on clean test data. Despite this, we find it does not lead to increased robustness. Conversely, training with higher σ\sigma of Gaussian can lead to increased robustness to Gaussian noise, but it also leads to decreased accuracy on clean data. Therefore, any robustness gains are offset by poor overall performance.

At first glance, these results seem to reinforce the findings of previous work , indicating that robustness comes at the cost of generalization. In the following sections, we will explore whether there exists augmentation strategies that do not exhibit this limitation.

Method

Each of the two methods seen so far achieves one half of our stated goal: either improving robustness or improving clean test accuracy, but never both. To overcome the limitations of existing data augmentation techniques, we introduce Patch Gaussian, a new technique that combines the noise robustness of Gaussian with the improved clean accuracy of Cutout.

Patch Gaussian works by adding a WWxWW patch of Gaussian noise to the image (Figure 3)A TensorFlow implementation of Patch Gaussian can be found in Appendix (Figure 10).. As with Cutout, the center of the patch is sampled to be within the image. By varying the size of this patch and the maximum standard deviation of noise sampled σmax\sigma_{max}, we can interpolate between Gaussian (which applies additive Gaussian noise to the whole image) and an approximation of Cutout (which removes all information inside the patch). See Figure 9 for more examples.

All image transformations, including Patch Gaussian are performed on images with unnormalized pixel values in range.Forallimages,standardrandomflippingandcroppingisappliedimmediatelyafteranyaugmentationsmentionedonCIFAR−10(before,onImagenet).Afternoise−basedaugmentations,imagesareclippedtotherange. For all images, standard random flipping and cropping is applied immediately after any augmentations mentioned on CIFAR-10 (before, on Imagenet). After noise-based augmentations, images are clipped to the range.

2 Hyper-parameter selection

Our goal is to learn models that achieve both good clean accuracy and improved robustness to corruptions. When selecting hyper-parameters we need to decide how to weight these two metrics. Here, we focused on identifying the models that were most robust while still achieving a minimum accuracy (Z) on the clean test data. Values of Z vary per dataset and model, and can be found in the Appendix (Table 6). If no model has clean accuracy ≥\geqZ, we report the model with highest clean accuracy, unless otherwise specified. We find that patch sizes around 2525 on CIFAR (≤\leq250 on ImageNet, i.e.: uniformly sampled with maximum value 250) with σ≤1.0\sigma\leq 1.0 generally perform the best. A complete list of selected hyper-parameters for all augmentations can be found in Table 6.

Here, robustness is defined as average accuracy of the model, when tested on data corrupted by various σ\sigma (0.10.1, 0.20.2, 0.30.3, 0.50.5, 0.80.8, 1.01.0) of Gaussian noise, relative to the clean accuracy. This metric is correlated with mCE , so it ensures model rosbustness is generally useful beyond Gaussian corruptions. By picking models based on their Gaussian noise robustness, we ensure that our selection process does not overfit to the Common Corruptions benchmark .

3 Models, Datasets, & Implementation Details

We run our experiments on CIFAR-10 and ImageNet datasets. On CIFAR-10, we use the Wide-ResNet-28-10 model , as well as the Shake-shake-112 model , trained for 200 epochs and 600 epochs respectively. The Wide-ResNet model uses a initial learning rate of 0.1 with a cosine decay schedule. Weight decay is set to be 55e-44 and batch size is 128. We train all models, including the baseline, with standard data augmentation of horizontal flips and pad-and-crop. Our code uses the same hyper parameters as Available at https://github.com/tensorflow/models/tree/master/research/autoaugment.

On ImageNet, we use the ResNet-50 and Resnet-200 models , trained for 90 epochs. We use a weight decay rate of 11e-44, global batch size of 512 and learning rate of 0.2. The learning rate is decayed by 10 at epochs 30, 60, and 80. We use standard data augmentation of horizontal flips and crops. All CIFAR-10 and ImageNet experiments use the listed hyper-parameters above, unless specified otherwise. Our code uses the same hyper parameters as open-sourced implementationsAvailable at https://github.com/tensorflow/tpu/tree/master/models/official/resnet.

Results

We show that models trained with Patch Gaussian can overcome the trade-off observed in Fig. 2 and learn models that are robust while maintaining their generalization accuracy (Section 4.1). In doing so, we establish a new state of the art in CIFAR-C and ImageNet-C Common Corruptions benchmark (Section 4.2). We then show that Patch Gaussian can be used in complement to other common regularization strategies (Section 4.3), data augmentation policies (Section 4.4), and that it can also improve training of object detection models (Section 4.5)

We train models on various hyper-parameters of Patch Gaussian and find that the model selected by the method in Section 3.2 leads to improved robustness to Gaussian noise, like Gaussian, while also improving clean accuracy, much like Cutout. In Figure 4, we visualize these results in an ablation study, either varying patch sizes and fixing σ\sigma to the selected value (1.01.0) or varying noise level σ\sigma and fixing the patch size to the selected value (350350).

2 Training with Patch Gaussian leads to improved Common Corruption robustness

In this section, we look at how our augmentations impact robustness in a more realistic setting beyond robustness to additive Gaussian noise. Rather than focusing on adversarial examples that are worst-case bounded perturbations, we focus on a more general set of corruptions that models are likely to encounter in real-world settings: the Common Corruptions benchmark . This benchmark, also referred to as CIFAR-C and ImageNet-C, is composed of images transformed with 15 corruptions, at 5 severities each. The corruptions are designed to model those commonly found in real-world settings, such as brightness, different weather conditions, and different kinds of noise.

Tables 1 and 2 show that Patch Gaussian achieves state of the art on both of these benchmarks in terms of mean Corruption Error (mCE). However, ImageNet-C was released in compressed JPEG format , which alters the corruptions applied to the raw pixels. Therefore, we report results on the benchmark as-released (“Original mCE”) as well as a version of 12 corruptions without the extra compression (“mCE”) Available at https://github.com/tensorflow/datasets under imagenet2012_corrupted.

To compute mCE, we normalize the corruption error for each model and dataset to the baseline with only flip and crop data augmentation. The one exception is Original mCE ImageNet, where we use the AlexNet baseline to be directly comparable with previous work .

Because Patch Gaussian is a noise-based augmentation, we wanted to verify whether its gains on this benchmark were solely due to improved performance on noise-based corruptions (Gaussian Noise, Shot Noise, and Impulse Noise). To do this, we also measure the models’ average performance on all other corruptions, reported as “Original mCE (-noise)”, and “mCE (-noise)”. We observe that Patch Gaussian outperforms all other models, even on corruptions like fog where Gaussian hurts performance . Scores for each corruption can be found in the Appendix (Tables 7 and 8).

Comparing the lower capacity ResNet-50 and Wide ResNet models to higher-capacity ResNet-200 and Shake 112 models, we find diminished gains in clean accuracy and robustness. Still, Patch Gaussian achieves a substantial increase in mCE relative to other augmentation strategies.

3 Patch Gaussian can be combined with other regularization strategies

Since Patch Gaussian has a regularization effect on the models trained above, we compare it with other regularization methods: larger weight decay, label smoothing, and dropblock (Table 3). We find that while label smoothing improves clean accuracy, it weakens the robustness in all corruption metrics we have considered. This agrees with the theoretical prediction from , which argued that increasing the confidence of models would improve robustness, whereas label smoothing reduces the confidence of predictions. We find that increasing the weight decay from the default value used in all models does not improve clean accuracy or robustness.

Here, we focus on analyzing the interaction of different regularization methods with Patch Gaussian. Previous work indicates that improvements on the clean accuracy appear after training with Dropblock for 270 epochs , but we did not find that training for 270 epochs changed our analysis. Thus, we present models trained at 90 epochs for direct comparison with other results. Due to the shorter training time, Dropblock does not improve clean accuracy, yet it does make the model more robust (relative to baseline) according to all corruption metrics we consider.

We find that using label smoothing in addition to Patch Gaussian has a mixed effect, it improves clean accuracy while slightly improving robustness metrics except for the Original mCE. Combining Dropblock with Patch Gaussian reduces the clean accuracy relative to the Patch Gaussian-only model, as Dropblock seems to be a strong regularizer when used for 90 epochs. However, using Dropblock and Patch Gaussian together leads to the best robustness performance. These results indicate that Patch Gaussian can be used in conjunction with existing regularization strategies.

4 Patch Gaussian can be combined with AutoAugment policies for improved results

Knowing that Patch Gaussian can be combined with other regularizers, it’s natural to ask whether it can also be combined with other data augmentation policies. Table 4 highlights models trained with AutoAugment . For fair comparison of mCE scores, we train all models with the best AutoAugment policy, but without contrast and Inception color pre-processing, as those are present in the Common Corruptions benchmark. This process is imperfect since AutoAugment has many operations, some of which could still be correlated with corruptions. Regardless, we find that Patch Gaussian improves accuracy and robustness over simply using AutoAugment.

Because AutoAugment leads to state of the art accuracies, we are interested in seeing how far it can be combined with Patch Gaussian to improve results. Therefore, and unlike previous experiments, models are trained for 180 epochs to yield best results possible.

5 Patch Gaussian improves performance in object detection

Since Patch Gaussian can be combined with both regularization strategies as well as data augmentation policies, we want to see if it is generally useful beyond classification tasks. We train a RetinaNet detector with ResNet-50 backbone on the COCO dataset . Images for both baseline and Patch Gaussian models are horizontally flipped half of the time, after being resized to 640×640640\times 640. We train both models for 150 epochs using a learning rate of 0.08 and a weight decay of 1e−41e-4. The focal loss parameters are set to be α=0.25\alpha=0.25 and γ=1.5\gamma=1.5.

Despite being designed for classification, Patch Gaussian improves detection performance according to all metrics when tested on the clean COCO validation set (Table 5). On the primary COCO metric mean average precision (mAP), the model trained with Patch Gaussian achieves a 1% higher accuracy over the baseline, whereas the model trained with Gaussian suffers a 2.9% loss.

Next, we evaluate these models on the validation set corrupted by i.i.d. Gaussian noise, with σ=0.25\sigma=0.25. We find that model trained with Gaussian and Patch Gaussian achieve the highest mAP of 26.1% on the corrupted data, whereas the baseline achieves 11.6%. It is interesting to note that Patch Gaussian model achieves a better result on the harder metrics of small object detection and stricter intersection over union (IOU) thresholds, whereas the Gaussian model achieves a better result on the easier tasks of large object detection and less strict IOU threshold metric.

Overall, as was observed for the classification tasks, training object detection models with Patch Gaussian leads to significantly more robust models without sacrificing clean accuracy.

Discussion

In an attempt to understand Patch Gaussian’s performance, we perform a frequency-based analysis of models trained with various augmentations using the method introduced in .

For CIFAR-10 models, we present this analysis for the entire Fourier domain, with noise sampled with norm 44. For ImageNet, we focus our analysis on lower frequencies that are more visually salient add noise with norm 15.715.7.

Note that for Cutout and Gaussian, we chose larger patch sizes and σ\sigmas than those selected with the method in Section 3.2 in order to highlight the effect of these augmentations on sensitivity. Heatmaps of other models can be found in the Appendix (Figure 12).

We confirm findings by that Gaussian encourages the model to learn a low-pass filter of the inputs. Models trained with this augmentation, then, have low test error sensitivity at high frequencies, which could help robustness. However, valuable high-frequency information is being thrown out at low layers, which could explain the lower test accuracy.

We further find that Cutout encourages the use of high-frequency information, which could help explain its improved generalization performance. Yet, it does not encourage lower test error sensitivity, which explains why it doesn’t improve robustness either.

Patch Gaussian, on the other hand, seems to allow high-frequency information through at lower layers, but still encourages relatively lower test error sensitivity at high frequencies. Indeed, when we measure accuracy on images filtered with a high-pass filter, we see that Patch Gaussian models can maintain accuracy in a similar way to the baseline and to Cutout, where Gaussian fails to. See Figure 5 for full results.

Understanding the impact of data distributions and noise on representations has been well-studied in neuroscience . The data augmentations that we propose here alter the distribution of inputs that the network sees, and thus are expected to alter the kinds of representations that are learned. Prior work on efficient coding and autoencoders has shown how filter properties change with noise in the unsupervised setting, resulting in lower-frequency filters with Gaussian, as we observe in Fig. 5. Consistent with prior work on natural image statistics , we find that networks are least sensitive to low frequency noise where the spectral density is largest. Performance drops at higher frequencies when the amount of noise we add grows relative to the typical spectral density observed at these frequencies. In future work, we hope to better understand the relationship between naturally occuring properties of images and sensitivity, and investigate whether training with more naturalistic noise can yield similar gains in corruption robustness.

Conclusion

In this work, we introduced a single data augmentation operation, Patch Gaussian, which improves robustness to common corruptions without incurring a drop in clean accuracy. For models that are large relative to the dataset size (like ResNet-200 on ImageNet and all models on CIFAR-10), Patch Gaussian improves clean accuracy and robustness concurrently. We showed that Patch Gaussian achieves this by interpolating between two standard data augmentation operations Cutout and Gaussian. We also demonstrate that Patch Gaussian can be used in conjunction with other regularization and data augmentation strategies, and can also improve the performance of object detection models, indicating it is generally useful. Finally, we analyzed the sensitivity to noise in different frequencies of models trained with Cutout and Gaussian, and showed that Patch Gaussian combines their strengths without inheriting their weaknesses.

We would like to thank Benjamin Caine, Trevor Gale, Keren Gu, Jonathon Shlens, Brandon Yang, Nic Ford, Samy Bengio, Alex Alemi, Matthew Streeter, Robert Geirhos, Alex Berardino, Jon Barron, Barret Zoph, and the Google Brain and Google AI Residency teams.

References

Appendix