Exploring the Landscape of Spatial Robustness
Logan Engstrom, Brandon Tran, Dimitris Tsipras, Ludwig Schmidt, Aleksander Madry
Introduction
Neural networks are now widely embraced as dominant solutions in computer vision (Krizhevsky et al., 2012; He et al., 2016), speech recognition (Graves et al., 2013), and natural language processing (Collobert & Weston, 2008). However, while their accuracy scores often match (and sometimes go beyond) human-level performance on key benchmarks (He et al., 2015; Taigman et al., 2014), models experience severe performance degradation in the worst case.
In this case, models are vulnerable to so-called adversarial examples, or slightly perturbed inputs that are almost indistinguishable from natural data to a human but cause state-of-the-art classifiers to make incorrect predictions (Szegedy et al., 2014; Goodfellow et al., 2015). This raises concerns about the use of neural networks in contexts where reliability, dependability, and security are important.
As such, one may suspect that adversarial examples are a problem only in the presence of a truly malicious attacker, and are unlikely to arise in more benign environments. However, recent work has shown that neural network–based vision classifiers are vulnerable to input images that have been spatially transformed through small rotations, translations, shearing, scaling, and other natural transformations (Fawzi & Frossard, 2015; Kanbak et al., 2018; Xiao et al., 2018; Tramèr & Boneh, 2017). Such transformations are pervasive in vision applications and hence quite likely to naturally occur in practice. The vulnerability of neural networks to such transformations raises a natural question:
How can we build spatially robust classifiers?
We address this question by first performing an in-depth study of neural network–based classifier robustness to two basic image transformations: translations and rotations. While these transformations appear natural to a human, we show that small rotations and translations alone can significantly degrade accuracy. These transformations are particularly relevant for computer vision applications since real-world objects do not always appear perfectly centered.
Our goal is to obtain a fine-grained understanding of the spatial robustness of standard, near state-of-the-art image classifiers for the MNIST (LeCun, 1998), CIFAR10 (Krizhevsky, 2009), and ImageNet (Russakovsky et al., 2015) datasets.
We find that small rotations and translations consistently and significantly degrade accuracy of image classifiers on a number of tasks, as illustrated in Figure 1. Our results suggest that classifiers are highly brittle: even small random transformations can degrade accuracy by up to 30%. Such brittleness to random transformations suggests that these models might be unreliable even in benign settings.
Relative adversary strength.
Spatial loss landscape.
Improving spatial robustness.
We then develop methods for alleviating these vulnerabilities using insights from our study. As a natural baseline, we augment the training procedure with rotations and translations. While this does largely mitigate the problem on MNIST, additional data augmentation only marginally increases robustness on CIFAR10 and ImageNet. We thus propose two natural methods for further increasing the robustness of these models, based on robust optimization and aggregation of random input transformations. These methods offer significant improvements in classification accuracy against both adaptive and random attackers when compared to both standard models and those trained with additional data augmentation. In particular, on ImageNet, our best model attains a top1 accuracy of 56% against the strongest adversary, versus 34% for a standard network with additional data augmentation.
Related Work
The fact that small rotations and translation can fool neural networks on MNIST and CIFAR10 was, to the best of our knowledge, first observed in (Fawzi & Frossard, 2015). They compute the minimum transformation required to fool the model and use it as a measure for a quantitative comparison of different architectures and training procedures. The main difference to our work is that we focus on the optimization aspect of the problem. We show that a few random queries usually suffice for a successful attack, while first-order methods are ineffective. Moreover, we go beyond standard data augmentation and evaluate the effectiveness of natural baseline defenses.
The concurrent work of Kanbak et al. (2018) proposes a different first-order method to evaluate the robustness of classifiers based on geodesic distances on a manifold. This metric is harder to interpret than our parametrized attack space. Moreover, given our findings on the non-concavity of the optimization landscape, it is unclear how close their method is to the ground truth (exhaustive enumeration). While they perform a limited study of defenses (adversarial fine-tuning) using their method, it appears to be less effective than our baseline worst-of-10 training. We attribute this difference to the inherent obstacles first-order methods face in this optimization landscape.
Adversarial Rotations and Translations
Recall that in the context of image classification, an adversarial example for a given input image and a classifier is an image that satisfies two properties: (i) on the one hand, the adversarial example causes the classifier to output a different label on than on , i.e., we have . (ii) On the other hand, the adversarial example is “visually similar” to .
First, we define the exact range of attacks we want to optimize over. For the case of rotation and translation attacks, we wish to find parameters such that rotating the original image by degrees around the center and then translating it by pixels causes the classifier to make a wrong prediction. Formally, the pixel at position is moved to the following position (assuming the point is the center of the image):
We implement this transformation in a differentiable manner using the spatial transformer blocks of (Jaderberg et al., 2015) We used the open source implementation found here: https://github.com/tensorflow/models/tree/master/research/transformer.. In order to handle pixels that are mapped to non-integer coordinates, the transformer units include a differentiable bilinear interpolation routine. Since our loss function is differentiable with respect to the input and the transformation is in turn differentiable with respect to its parameters, we can obtain gradients of the model’s loss function w.r.t. the perturbation parameters. This enables us to apply a first-order optimization method to our problem.
By defining the spatial transformation for some as , we construct an adversarial perturbation for by solving the problem
where is the loss function of the neural networkThe loss of the classifier is a function from images to real numbers that expresses the performance of the network on the particular example (e.g., the cross-entropy between predicted and correct distributions)., and is the correct label for .
We compute the perturbation from Equation 1 in three distinct ways:
Worst-of-: We randomly sample different choices of attack parameters and choose the one on which the model performs worst. As we increase , this attack interpolates between a random choice and grid search.
We remark that while a first-order attack requires full knowledge of the model to compute the gradient of the loss with respect to the input, the other two attacks do not. They only require the outputs corresponding to chosen inputs, which can be done with only query access to the target model.
Improving Invariance to Spatial Transformations
As we will see in Section 5, augmenting the training set with random rotations and translations does improve the robustness of the model against such random transformations. However, data augmentation does not significantly improve the robustness against worst-case attacks and sometimes leads to a drop in accuracy on unperturbed images. To address these issues, we explore two simple baselines that turn out to be surprisingly effective.
Instead of performing standard empirical risk minimization to train the classification model, we utilize ideas from robust optimization. Robust optimization has a rich history (Ben-Tal et al., 2009) and has recently been applied successfully in the context of defending neural networks against adversarial examples (Madry et al., 2018; Sinha et al., 2018; Raghunathan et al., 2018; Wong & Kolter, 2018). The main barrier to applying robust optimization for spatial transformations is the lack of an efficient procedure to compute the worst-case perturbation of a given example. Performing a grid search (as described in Section 3) is prohibitive as this would increase the training time by a factor close to the grid size, which can easily be a factor 100 or 1,000. Moreover, the non-convexity of the loss landscape prevents potentially more efficient first-order methods from discovering (approximately) worst-case transformations (see Section 5 for details).
Given that we cannot fully optimize over the space of translations and rotations, we instead use a coarse approximation provided by the worst-of-10 adversary (as described in Section 3). So each time we use an example during training, we first sample 10 transformations of the example uniformly at random from the space of allowed transformations. We then evaluate the model on each of these transformations and train on the one perturbation with the highest loss. This corresponds to approximately minimizing a min-max formulation of robust accuracy similar to (Madry et al., 2018). Training against such an adversary increases the overall time by a factor of roughly six.We need to perform 10 forward passes and one backwards pass instead of one forward and one backward pass required for standard training.
Aggregating Random Transformations.
As Section 5 shows, the accuracy against a random transformation is significantly higher than the accuracy against the worst transformation in the allowed attack space. This motivates the following inference procedure: compute a (typically small) number of random transformations of the input image and output the label that occurs the most in the resulting set of predictions. We constrain these random transformations to be within of the input image size in each translation direction and up to of rotation. Note that if an adversary rotates an image by (a valid attack in our threat model), we may end up evaluating the image on rotations of up to . The training procedure and model can remain unchanged while the inference time is increased by a small factor (equal to the number of transformations we evaluate on).
Combining Both Methods.
The two methods outlined above are orthogonal and in some sense complementary. We can therefore combine robust training (using a worst-of-k adversary) and majority inference to further increase the robustness of our models.
Experiments
We evaluate standard image classifiers for the MNIST (LeCun, 1998), CIFAR10 (Krizhevsky, 2009) and ImageNet (Russakovsky et al., 2015) datasets. In order to determine the extent to which misclassification is caused by insufficient data augmentation during training, we examine various data augmentation methods. We begin with a description of our experimental setup.
Attack Space.
In order to maintain the visual similarity of images to the natural ones we restrict the space of allowed perturbations to be relatively small. We consider rotations of at most and translations of at most (roughly) 10% percent of the image size in each direction. This corresponds to 3 pixels for MNIST (image size ) and CIFAR10 (image size ), and 24 pixels for ImageNet (image size ). For grid search attacks, we consider 5 values per translation direction and 31 values for rotations, equally spaced. For first-order attacks, we use steps of projected gradient descent of step size times the parameter range. When rotating and translating the images, we fill the empty space with zeros (black pixels).
Data Augmentation.
We consider five variants of training for our models.
Standard training: The standard training procedure for the respective model architecture.
No random cropping: Standard training for CIFAR-10 and ImageNet includes data augmentation via random crops. We investigate the effect of this data augmentation scheme by also training a model without random crops.
Random rotations and translations: At each training step, we perform a uniformly random perturbation from the attack space on each training example.
Random rotations and translations from larger intervals: As before, we perform uniformly random perturbations, but now from a superset of the attack space (, 13% pixels).
1 Evaluating Model Robustness
We evaluate all models against random and grid search adversaries with rotations and translations considered both separately and together. We report the results in Table 1. We visualize a random subset of successful attacks in Figures 5, 6, and 7 of Appendix A.
Despite the high accuracy of standard models on unperturbed examples and their reasonable performance on random perturbations, a grid search can significantly lower the classifiers’ accuracy on the test set. For the standard models, accuracy drops from 99% to 26% on MNIST, 93% to 3% on CIFAR10, and 76% to 31% on ImageNet (Top 1 accuracy).
The addition of random rotations and translations during training greatly improves both the random and adversarial accuracy of the classifier for MNIST and CIFAR10, but less so for ImageNet. For the first two datasets, data augmentation increases the accuracy against a grid adversary by 60% to 70%, while the same data augmentation technique adds less than 3% accuracy on ImageNet.
We perform a fine-grained investigation of our findings:
In Figure 3 we examine how many examples can be fooled by (i) rotations only, (ii) translations only, (iii) neither transformation, or (iv) both.
We visualize the set of fooling angles for a random sample of the rotations-only grid in Figure 4 on ImageNet, and provide more examples in the appendix in Figure 10. We observe that the set of fooling angles is nonconvex and not contiguous.
To investigate how many transformations are adversarial per image, we analyze the percentage of misclassified grid points for each example in Figure 11. While the majority of images has only a small number of adversarial transformations, a significant fraction of images is fooled by 20% or more of the transformations.
A natural question is whether the reduced accuracy of the models is due to the cropping applied during the transformation. We verify that this is not the case by applying zero and reflection padding to the image datasets. We note that the zero padding creates a “black canvas” version of the dataset, ensuring that no information from the original image is lost after a transformation. We show a random set of adversarial examples in this setting in Figure 8 and a full evaluation in Table 4. We also provide more details regarding reflection padding in Section B and provide an evaluation in Table 6. All of these are in Appendix A.
2 Comparing Attack Methods
In Table 2 we compare different attack methods on various classifiers and datasets. We observe that worst-of-10 is a powerful adversary despite its limited interaction with the target classifier. The first-order adversary performs significantly worse. It fails to approximate the ground-truth accuracy of the models and performs significantly worse than the grid adversary and even the worst-of-10 adversary.
Relation to Black-Box Attacks.
Given its limited interaction with the model, the worst-of-10 adversary achieves a significant reduction in classification accuracy. It performs only 10 random, non-adaptive queries to the model and is still able to find adversarial examples for a large fraction of the inputs (see Table 2). The low query complexity is an important baseline for black-box attacks on neural networks, which recently gained significant interest (Papernot et al., 2017; Chen et al., 2017; Bhagoji et al., 2017; Ilyas et al., 2018). Black-box attacks rely only function evaluations of the target classifier, without additional information such as gradients. The main challenge is to construct an adversarial example from a small number of queries. Our results show that it is possible to find adversarial rotations and translations for a significant fraction of inputs with very few queries.
3 Evaluating Our Defense Methods.
As we see in Table 1, training with a worst-of-10 adversary significantly increases the spatial robustness of the model, also compared to data augmentation with random transformations. We conjecture that using more reliable methods to compute the worst-case transformations will further improve these results. Unfortunately, increasing the number of random transformations per training example quickly becomes computationally expensive. And as pointed out above, current first-order methods also appear to be insufficient for finding worst-case transformations efficiently.
Our results for majority-based inference are presented in Table 5 of Appendix A. By combining these two defenses, we improve the worst-case performance of the models from 26% to 98% on MNIST, from 3% to 82% on CIFAR10, and from 31% to 56% on ImageNet (Top 1).
Conclusions
We examined the robustness of state-of-the-art image classifiers to translations and rotations. We observed that even a small number of randomly chosen perturbations of the input are sufficient to considerably degrade the classifier’s performance.
The fact that common neural networks are vulnerable to simple and naturally occurring spatial transformations (and that these transformations can be found easily from just a few random tries) indicates that adversarial robustness should be a concern not only in a fully worst-case security setting. We conjecture that additional techniques need to be incorporated in the architecture and training procedures of modern classifiers to achieve worst-case spatial robustness. Also, our results underline the need to consider broader notions of similarity than only pixel-wise distances when studying adversarial misclassification attacks. In particular, we view combining the pixel-wise distances with rotations and translations as a next step towards the “right” notion of similarity in the context of images.
Acknowledgements
Dimitris Tsipras was supported in part by the NSF grant CCF-1553428 and the NSF Frontier grant CNS-1413920. Aleksander Mądry was supported in part by an Alfred P. Sloan Research Fellowship, a Google Research Award, and the NSF grants CCF-1553428 and CNS-1815221.
References
Appendix A Omitted Tables and Figures
Appendix B Mirror Padding
In the experiments of Section 5, we filled the remaining pixels of rotated and translated images with black (also known as zero or constant padding). This is the standard approach used when performing random cropping for data augmentation purposes. We briefly examined the effect of mirror padding, that is replacing empty pixels by reflecting the image around the borderhttps://www.tensorflow.org/api_docs/python/tf/pad. The results are shown in Table 6. We observed that training with one padding method and evaluating using the other resulted in a significant drop in accuracy. Training using one of these methods randomly for each example resulted in a model which roughly matched the best-case of the two individual cases.