Model-Based Robust Deep Learning: Generalizing to Natural, Out-of-Distribution Data
Alexander Robey, Hamed Hassani, George J. Pappas
Introduction
Over the last decade, we have witnessed unprecedented breakthroughs in deep learning . Rapidly growing bodies of work continue to improve the state-of-the-art in generative modeling , computer vision , and natural language processing . Indeed, the significant progress made in these fields has prompted large-scale integration of deep learning techniques into a myriad of application domains, including autonomous vehicles, medical diagnostics, and robotics . Importantly, many of these domains are safety-critical, meaning that the detections, recommendations, or decisions made by deep learning systems can directly impact the well-being of humans . For this reason, it is essential that the deep learning systems used in safety-critical applications are robust and trustworthy .
Despite the remarkable progress made toward improving the state-of-the-art in deep learning, it is well-known that many deep learning frameworks including neural networks are fragile to seemingly innocuous and imperceptible changes to their input data . Well-documented examples of fragility to carefully-designed noise can be found in the context of image detection , video analysis , traffic sign misclassification , machine translation , clinical trials , and robotics . In response to this vulnerability, a growing body of work has focused on improving the robustness of deep learning. More specifically, the literature concerning adversarial robustness has sought to improve robustness against small, imperceptible perturbations of data, which have been shown to cause misclassification . Over the last five years, this literature has included the development of robust training algorithms and certifiable defenses . In particular, these robust training approaches, i.e. the method of adversarial training , typically perturb input data via adversarially-chosen, norm-bounded noise in a robust optimization formulation , and have been shown to be effective at improving the robustness of deep learning against norm-bounded perturbations .
While the adversarial training paradigm has provided a rigorous framework for analyzing and improving the robustness of deep learning, the algorithms used in this paradigm have notable limitations. Specifically, most adversarial training algorithms are only applicable for robustness applications in which data is perturbed by norm-bounded, artificially-generated, imperceptible noise. Thus while adversarial training algorithms can resolve security threats arising from artificial tampering of the data, these schemes cannot provide similar levels of robustness to changes that may arise due to other kinds of perturbations or variation . And to this end, numerous recent papers have unanimously shown that deep learning is extremely fragile to unbounded shifts in the data-distribution which commonly occur due to a wide range of natural phenomena and which cannot be modeled by additive, norm-bounded perturbations. Such phenomena include unseen distributional shifts such as changes in image lighting, variable weather conditions, or blurring . And while such unseen distributional shifts are arguably more common in safety-critical domains than norm-bounded perturbations, there are remarkably few general, principled techniques that provide robustness against these forms of out-of-distribution, naturally-occurring variation . Thus, it is of critical importance for the deep learning community to design novel algorithms that are robust against natural, out-of-distribution shifts in data.
In this paper, we formulate the first general-purpose algorithms that (1) use unlabeled data to learn models that describe arbitrary forms of natural variation and (2) exploit these models to provide significant robustness against natural, out-of-distribution shifts in data. To this end, we propose a paradigm shift from perturbation-based adversarial robustness to model-based robust deep learning. In this paradigm, following the observation that data can vary in highly nonlinear and unbounded ways in real-world, safety-critical environments, we first obtain models of natural variation which describe how data varies in natural environments. Such models of natural variation may be known a priori, as is the case for geometric transformations such as rotation or scaling. Alternatively, in some settings a model of natural variation may not be known beforehand and therefore must be learned from data; for example, there are no analytic models that describe how to change the weather conditions in images. Once such models of natural variation have been obtained, in this paradigm, we formulate a novel robust optimization problem that exploits models of natural variation to produce neural networks that are robust to the source of natural variation captured by the model. In this way, the goal of the model-based robust paradigm is to develop general-purpose algorithms that can be used to train neural networks to be robust against natural, out-of-distribution shifts in data.
Our experiments show that across a variety of naturally-occurring and challenging conditions, such as changes in lighting, background color, haze, decolorization, snow, rain, frost, fog, and contrast, in twelve distinct datasets including MNIST, SVHN, GTSRB, CURE-TSR, ImageNet, and ImageNet-c, neural networks trained with our model-based algorithms significantly outperform classifiers trained via empirical risk minimization, norm-bounded robust deep learning algorithms, data augmentation methods, and, when applicable, domain adaptation techniques. In particular, we show that classifiers trained on ImageNet using our model-based algorithms and tested on various subsets of ImageNet-c improve over state-of-the-art classifiers by up to 30 percentage points. Furthermore, we show that the model-based robust deep learning paradigm is model-agnostic and adaptable, meaning that it can be used to provide robustness against arbitrary forms of natural variation in data and regardless of whether models of natural variation are known a priori or learned from data. To demonstrate the broad applicability of our approach, we present apply our paradigm to four novel settings. (1) First, we show that our algorithms are the first to provide out-of-distribution robustness on a range of challenging settings. (2) Next, we show that models of natural variation can be composed to provide robustness against multiple simultaneous distribution shifts. To evaluate this property, we curate several new datasets, each of which has two simultaneous natural shifts. (3) Thirdly, we show that models of natural variation trained on one dataset can be used to provide robustness on datasets that are entirely unseen while training the model. (4) Lastly, we show that in the setting of unsupervised domain adaptation, our algorithms outperform traditional domain adaptation techniques by significant margins.
While the experiments in this paper focus on image classification tasks subject to challenging natural conditions, our model-based robust deep learning paradigm is much broader and can, in principle, be applied to many other application domains as long as one can obtain accurate models describing how data naturally varies. In that sense, we believe that this approach will open up numerous directions for future research.
Contributions. The contributions of our paper can be summarized as follows:
Paradigm shift. We propose a paradigm shift from perturbation-based adversarial robustness to model-based robust deep learning, in which models of natural variation express changes due to natural conditions that frequently appear in data.
Learning models of natural variation. For many forms of natural variation that are commonly encountered in safety-critical applications, we use deep generative models to learn models of natural variation from unlabelled data that are consistent with realistic conditions.
Robust-optimization-based formulation. We formulate a novel training procedure by constructing a general robust optimization problem that searches for challenging out-of-distribution shifts in data to train classifiers to be robust against natural variation.
Novel model-based training algorithms. We propose a family of novel training algorithms that exploit models of natural variation to improve the robustness of deep learning against challenging natural conditions captured my models of natural variation.
ImageNet-c robustness. We show that our algorithms improve the robustness of classifiers trained on ImageNet and tested on ImageNet-c by as much as 30 percentage points on a variety of challenging settings, including changes in snow, contrast, brightness, and frost.
Out-of-distribution robustness. We show that our algorithms are the first to consistently provide robustness against natural, out-of-distribution shifts in data, including changes in snow, rain, fog, and brightness on SVHN, GTSRB, CURE-TSR, and ImageNet, that frequently occur in real-world environments.
Robustness to simultaneous distributional shifts. We show that our framework is composable and thus can be used to improve robustness against multiple simultaneous distributional shifts in data. To evaluate this feature, we curate four new datasets, each of which has two simultaneous distributional shifts.
Robustness to unseen domains. We show that models of natural variation can be reused on datasets that are entirely unseen during training to improve out-of-distribution generalization. This property demonstrates that model-based robustness is transferrable to unseen domains.
Robustness in the setting of unsupervised domain adaptation. We show that in the setting of unsupervised domain adaptation, our algorithms provide higher levels of robustness than traditional domain adaptation techniques.
Perturbation-based robustness in deep learning
In a litany of past works, it has been shown empirically that first-order methods (e.g. SGD or Adam ) can be used to approximately solve (2.1) to obtain weights that engender neural networks which achieve high classification accuracy on a variety of image classification tasks
As observed in previous works , solving the optimization problem stated in (2.1) does not result in robust neural networks. More specifically, neural networks trained by solving (2.1) are known to be susceptible to adversarial attacks. This means that given a datum with a corresponding label , one can find another datum such that (1) is close to with respect to a given Euclidean norm and (2) is predicted by the learned classifier as belonging to a different class where . If such a datum exists, it is called an adversarial example.
To address this striking vulnerability, researchers have sought to improve the robustness of deep learning by developing adversarial training algorithms, which inure neural networks against small, norm-bounded perturbations . The dominant paradigm toward training neural networks to be robust against adversarial examples relies on a robust optimization perspective . Indeed, the approach used in to provide robustness to adversarial examples is formalized by considering a distinct yet related optimization problem to (2.1). In particular, the idea is to train neural networks to be robust against a worst-case perturbation of each instance . This worst-case perspective can be formulated in the following way:
We can think of the optimization problem in (2.2) as comprising two coupled optimization problems: an inner maximization problem and an outer minimization problem:
Limitations of perturbation-based robustness. While there has been significant progress toward developing algorithms that train neural networks to be robust against norm-bounded perturbations , there are significant limitations to adversarial training. Notably, it has been unanimously shown in a spate of recent work that deep learning is also fragile to various forms of natural variation . In the context of image classification, such natural variation includes changes in lighting, weather, or background color , spatial transformations such as rotation or scaling , and sensor-based attacks . These realistic forms of variation in data, which are known as nuisances in the computer vision community, cannot be modeled by the norm-bounded perturbations used in the standard adversarial training paradigm of (2.2) . And while these natural distributional shifts are ubiquitous in numerous application domains, there are remarkably few general, principled techniques that provide robustness against these forms of out-of-distribution, naturally-occurring variation . Therefore, an important open problem in the deep learning community is to develop algorithms that can train neural networks to be robust against natural and realistic forms of out-of-distribution data that are often inherent in safety-critical applications.
Challenges in designing a more general robustness paradigm. Given the efficacy of works that seek to improve the robustness of deep learning against adversarial perturbations, it is of fundamental interest to determine whether the adversarial robustness literature can be leveraged toward developing more general notions of robustness. To this end, in this paper we identify two fundamental challenges toward achieving this objective.
Firstly, unlike in the setting of perturbation-based robustness, in real-world environments, data can vary in unknown and highly nonlinear ways. Thus, the first step toward building a more general robust training procedure must be to design mechanisms that accurately describe how data varies in such environments. Indeed, in many scenarios, known geometric or physical structure can be used to describe how data naturally varies, as is the case for spatial transformations such as rotations or scalings. On the other hand, many transformations, such as changes in weather conditions in images, cannot be described by analytical mathematical expressions. For this reason, it is essential that a more general robustness paradigm be able to leverage known structure and to learn this structure from data when no analytic expression describing how data varies is available.
The second challenge underlying the task of developing a more general robustness paradigm is to formulate a principled training procedure that leverages suitable models that describe how data varies toward generalizing to out-of-distribution shifts in the data distribution. Indeed, assuming one has access a suitable model of natural variation, such a procedure should be agnostic to the specific parameterization of the model and adaptable to both models that are known a priori as well as models that are learned from data.
A unifying solution: Model-Based Robust Deep Learning. In this paper, we present a new training paradigm for deep learning that improves robustness against natural, out-of-distribution shifts in data by addressing both of these unique challenges. Rather than perturbing data in a norm-bounded manner, our robust training approach exploits models of natural variation that describe how data changes with respect to particular shifts in the data distribution. However, we emphasize that our approach is model-agnostic in the sense that it provides a paradigm that is applicable across arbitrary classes of naturally-occurring variation. Indeed, in this paper we will show that even if a model of natural variation is not explicitly known a priori, one can train neural networks to be robust against natural variation by learning a model of this variation in an offline and data-driven fashion. More broadly, we claim that the framework described in this paper represents a new and more general paradigm for robust deep learning as it provides a methodology for improving the robustness of deep learning against arbitrary sources of natural variation.
Model-based robust deep learning
In the following section, we introduce the model-based robust deep learning paradigm. Motivated by past work concerning robustness against adversarially-chosen, norm-bounded perturbations, we formulate a robust optimization problem that characterizes a new notion of robustness with respect to natural variation. To concertize this formulation, we also offer a geometric interpretation of this novel notion of robustness.
While adversarial training provides robustness against the imperceptible perturbations described in Figure 1(a), in natural environments data varies in ways that cannot be captured by additive, norm-bounded perturbations. For example, consider the two traffic signs shown in Figure 1(b). Note that the images on the left and on the right show the same traffic sign; however, the image on the left shows the sign on a sunny day, whereas the image on the right shows the sign in the middle of a snow storm. This example prompts several relevant questions. How do we ensure that neural networks are robust to such natural variation? How can we rethink adversarial training algorithms to provide robustness against natural-varying and challenging data?
For the time being, we assume the existence of a suitable model of natural variation ; later, in Section 4, we will detail our approach for obtaining models of natural variation that correspond to a wide variety of natural shifts in the data distribution. In this way, given a model of natural variation , our immediate goal to develop novel model-based robust training algorithms that train neural networks to be robust against natural variation by exploiting the model . For instance, if models variation in the lighting conditions in an image, our model-based training algorithm will provide robustness against lighting discrepancies. On the other hand, if models changes in weather conditions such as in Figure 1(b), then our model-based algorithms will improve the robustness of trained classifiers against varying weather conditions. More generally, our model-based robust training formulation is agnostic to the source of natural variation, meaning that our paradigm is broadly applicable to any source of natural variation that a model of natural variation can capture.
2 Formulating the model-based robust optimization problem
Our point of departure from the classical adversarial training formulation of (2.2) is in the choice of the so-called adversarial perturbation. In this paper, we assume that the adversary has access to a model of natural variation , which allows it to transform into a distinct yet related instance by choosing different values of from a given nuisance space . The goal in this setting is to train a classifier that achieves high accuracy both on a test set drawn i.i.d. from and on more-challenging test data that has been subjected to the source of natural variation that models. In this sense, we are proposing a new training paradigm for deep learning that provides robustness against models of natural variation .
In order to defend a neural network against such an adversary, we propose the following model-based robust optimization problem, which will be the central object of study in this paper:
Here the nuisance space may be problem-dependent; indeed, different parameterizations of the model of natural variation may influence the choice of . We defer a discussion of the choice of the nuisance space until Section 7.2.
Conceptually, the intuition for this formulation is similar to the intuition for (2.2) given in Section 2. Indeed, the optimization problem in (3.1) also comprises an inner maximization problem and an outer minimization problem:
3 Geometry of model-based robust training
Figure 3(b) shows the geometry of the model-based robust training paradigm. Let us consider a task in which our goal is to correctly classify images of street signs in varying weather conditions. In the model-based robust training paradigm, we assume that we are equipped with a model of natural variation which, by varying the nuisance parameter , changes the output image according to the natural phenomena captured by the model. For example, if our data contains images in sunny weather, the model may be designed to continuously vary the weather conditions in these images without changing the scene, other vehicles on the road, or the size and shape of the street signs in these images.
Models of natural variation
Our model-based robustness paradigm of (3.1) critically relies on the existence of a model of natural variation that maps and consequently describes how a datum can be deformed into via the choice of a nuisance parameter . In this section, we consider cases in which (1) a model is known a priori, and (2) a model is unknown and therefore must be learned offline from data. In this second case in which models of natural variation must be learned from data, we propose a formulation for obtaining such models.
For many problems, a model is known a priori due to underlying physical or geometric laws and can be immediately exploited in our model-based robust training formulation. One straightforward example in which a model of natural variation is known is the classical adversarial training paradigm described by equation (2.2). Indeed, by inspecting equations (2.2) and (3.1), we can immediately extract the well-known norm-bounded adversarial model:
The above example of a known model shows that in some sense the perturbation-based adversarial training paradigm of equation (2.2) is a special case of the model-based robust deep learning paradigm (3.1) when . Of course, for this choice of adversarial perturbations there is a plethora of robust training algorithms .
Another example of a known model of natural variation is shown in Figure 4. Consider a scenario in which we would like to be invariant to changes in the background color for the MNIST dataset . This would require having a model that takes an MNIST digit as input and reproduces the same digit but with various colorized RGB backgrounds which correspond to different values of . This model is relatively simple to describe; pseudocode is provided in Algorithm 1.
More broadly, there are many settings in which naturally-occuring variation in data has geometric structure that is known a priori. For example, in image classification tasks, there are usually intrinsic geometric structures that identify how data can be rotated, translated, or scaled. Indeed, geometric models for rotating an image along a particular axis can be characterized by a one-dimensional angular parameter . In this case, a known model of natural variation for rotation can be described by
where is a rotation matrix. Such geometric models can facilitate adversarial distortions of images using a low-dimensional parameter . In prior work, this idea has been exploited to train neural networks to be robust against rotations of the data around a given axis .
Altogether, these examples show that for a variety of problems, known models can be used to analytically describe how data changes. In the context of known models, our model-based approach offers a more general framework that is model-agnostic in the sense that it is applicable to all such models of how data varies. Before describing our approach for learning models of natural variation from data when a known model is not available, we briefly explore the connection between known models of natural variation and equivariant neural networks.
In contrast to the literature that concerns equivariance, much of the adversarial robustness community has focused on what is often called invariance. A function is said to be invariant to if for any , meaning that transforming an input by has no impact on the output. While previous approaches exploit such transformations for designing architectures that respect this structure, our goal is to exploit this structure toward developing robust training algorithms.
2 Learning unknown models of natural variation G(x,δ)𝐺𝑥𝛿G(x,\delta) from data
While geometry and physics may provide analytical models of natural variation that can be exploited in our model-based robust training procedure, in many situations such models are not known or are too costly to obtain. For example, consider Figure 3(b) in which a model of natural variation describes the impact of adding snowy weather to an image . In this case, the transformation takes an image of a street sign in sunny weather and maps it to an image in snowy weather. Even though there is a relationship between the snowy and the sunny images, obtaining a model relating the two images is extremely challenging if we resort to physics or geometric structure. For such problems, we advocate for learning the model from data prior to model-based robust training. An example of a learned model of natural variation is shown in Figure 5.
In what follows, we introduce a statistical framework for learning models of natural variation from data. In particular, we advocate for a procedure in which a model of natural variation is learned offline using unlabeled and unpaired data prior to performing model-based robust training on a new and possibly different dataset. We note that while the procedure we describe is quite general, there are likely other formulations that may also result in suitable models of natural variation. Indeed, one interesting future direction is to explore approaches in which one learns a model of natural variation and trains a classifier via the model-based robust training paradigm simultaneously.
A statistical framework for learning models of natural variation. In order to learn a model of natural variation , we assume that we have access to two unpaired image domains and that are drawn from a common dataset or distribution. Generally speaking, in our setting domain will contain the original data without any natural variation, and domain will contain data that has been transformed by an underlying natural phenomenon. Thus, in the example of Figure 3(b), domain would contain images of traffic signs in sunny weather, and domain would contain images of street signs in snowy weather. We emphasize that the domains and are unpaired, meaning that it may not be possible to select an image of a traffic sign in sunny weather from domain and find a corresponding image of that same street sign in the same scene with snowy weather in domain .
Here is an appropriately-chosen distance metric that measures the distance between two probability distributions (e.g. the KL-divergence or Wasserstein distance).
This problem has received broad interest in the machine learning community thanks to the recent advances in generative modeling. In particular, in the fields of image-to-image translation and style-transfer, learning mappings between unpaired image domains is a well-studied problem . In the next subsection, we will show how the breakthroughs in these fields can be used to learn a model of natural variation that closely approximates underlying natural phenomena.
3 Using deep generative models to learn models of natural variation
Among the methods mentioned in the previous paragraph, all seek to learn multimodal mappings without relying on class-conditioning, meaning that they seek to disentangle the semantic content of a datum (i.e. its label or the characterizing component of an input image) from the nuisance content (e.g. background color, weather conditions, etc.) to produce a multimodal distribution over varying output images. We highlight these methods because learning a multimodal mapping is a concomitant property toward learning models of natural variation that can produce images subject to a range of natural conditions. Indeed, any of these architectures are suitable for solving (4.4). However, throughout the experiments that are presented in Section 6, for simplicity we adhere to a particular choice for the architecture for .
An architecture for models of natural variation. In this paper we will predominantly use the Multimodal Unsupervised Image-to-Image Translation (MUNIT) framework to learn models of natural variation. At its core, MUNIT combines two autoencoding networks and two generative adversarial networks (GANs) to learn two mappings: one that maps images from domain to corresponding images in and one that maps in the other direction from to . For completeness, we provide a complete characterization of the MUNIT framework and the hyperparamters we used to train models of natural variation using MUNIT in Appendix A. For the purposes of this paper, we will only exploit the mapping from to , although one direction for future work is to incorporate both mappings. Therefore, the map learned in the MUNIT framework can be thought of as taking as input an image and a nuisance parameter and outputting an image that has the same semantic content as the input image but that has a different level of natural variation. This architecture is illustrated in Figure 6. Notice that the encoding network encodes the input image into a semantic component and a nuisance parameter; by varying this nuisance parameter and then decoding, the MUNIT framework can be used to vary the natural conditions in the input image.
4 A gallery of learned models of natural variation
To demonstrate the efficacy of the MUNIT framework toward learning models of natural variation , in Table 1 we show a gallery of learned models of natural variation learned using MUNIT for various datatsets and sources of natural variation. Importantly, this table shows that our framework in conjunction with the MUNIT architecture can be used to learn perceptually realistic models of natural variation for both low-dimension data (e.g. SVHN) and high-dimensional data (e.g. ImageNet). To this end, in Section 7.2, we will more quantitatively evaluate the ability of learned models of natural variation to produce realistic output image distributions.
Model-based robust training algorithms
In the previous section, we described a procedure that can be used to train models of natural variation . In some cases, such models may be known a priori while in other cases such models may be learned offline from data. Regardless of their origin, we will now assume that we have access to a suitable model and shift our attention toward exploiting in the development of novel robust training algorithms.
To begin, recall the optimization-based formulation of (3.1). Given a model of natural variation , (3.1) is a nonconvex-nonconcave min-max problem, and is therefore difficult to solve exactly. We will therefore resort to approximate methods for solving this challenging optimization problem. To elucidate our approach for solving (3.1), we first characterize the problem in the finite-sample setting. That is, rather than assuming access to the full joint distribution , we assume that we are given given a finite number of samples distributed i.i.d. according to the true data distribution . The empirical version of (3.1) in the finite-sample setting can be expressed in the following way:
An overview of the model-based robust training algorithms. Each of the three algorithms we propose in this paper – MRT-, MAT-, and MDA- – seeks a solution to (5.1) by alternating between solving the outer minimization problem and solving the inner maximization problem. Indeed, one similarity amongst these three algorithms is that each procedure seeks a solution to the outer problem by using a standard first-order optimization technique (e.g. SGD or Adam). However, the algorithms differ in how they search for a solution to the inner problem; at a high level, each of these methods seeks such a solution by augmenting the original training dataset with new data generated by a given model of natural variation . In particular, MRT randomly queries to generate several new data points and then selects those generated data that induce the highest loss in the inner maximization problem. On the other hand, MAT employs a gradient-based search in the nuisance space to find loss-maximizing generated data. Finally, MDA augments the training dataset with generated data by sampling randomly in to produce a wide range of natural conditions. We note that past approaches have used similar adversarial and statistical augmentation techniques. However, the main difference between these past works and our algorithms are that our algorithms exploit models of natural variation to generate new data.
In the remainder of this section, we will describe each algorithm in detail and provide psuedocode for each algorithm. Python implementations of each algorithm are available at the following link: https://github.com/arobey1/mbrdl.
In general, solving the inner maximization problem in (5.1) is difficult and motivates the need for methods that yield approximate solutions. In this vein, one simple scheme is to sample different nuisance parameters for each instance-label pair and among those sampled values, find the nuisance parameter that gives the highest empirical loss under . Indeed, this approach is not designed to find an exact solution to the inner maximization problem; rather it aims to find a difficult example by sampling in the nuisance space of the model of natural variation.
Once we obtain this difficult example by sampling in , the next objective is to solve the outer minimization problem. The procedure we propose in this paper for solving this problem amounts to using the worst-case nuisance parameter obtained via the inner maximization problem to perform data-augmentation. That is, for each instance-label pair , we treat as a new instance-label pair that can be used to supplement the original dataset . These training data can be used together with first-order optimization methods to solve the outer minimization problem to a locally optimal solution .
Throughout the experiments in the forthcoming sections, we will train classifiers via MRT with different values of . In this algorithm, controls the number of data points we consider when searching for a loss-maximing datum. To make clear the role of in this algorithm, we will refer to Algorithm 2 as MRT- when appropriate.
2 Model-based Adversarial Training (MAT)
An empirical analysis of the performance of MAT will be given in Section 6. To emphasize the role of the number of gradient steps used to find a loss maximizing nuisance parameter , we will often refer to Algorithm 3 as MAT-.
3 Model-based Data Augmentation (MDA)
Another interpretation of (3.1) is as follows. Rather than taking an adversarial point of view in which we expose neural networks to the most challenging model-generated examples, an alternative is to expose these networks to a diversity of model-generated data during training. In this approach, by augmenting with model-generated data corresponding to a wide range of natural variations , one might hope to achieve higher levels of robustness with respect to a given model of natural variation .
In MDA, the parameter controls the number of model-based data points per data point in that we append to the training set. To make this explicit, we will frequently refer to Algorithm 4 as MDA-.
Experiments
We present experiments in five different and challenging settings over twelve distinct datasets to demonstrate the broad applicability of the model-based robust deep learning paradigm.
Experimental overview. First, in Sections 6.1-6.2, we show that our algorithms are the first to consistently provide out-of-distribution robustness across a range of challenging corruptions, including shifts in brightness, contrast, snow, fog, frost, and haze on CURE-TSR, ImageNet, and ImageNet-c. Next, in Section 6.3, we show that models of natural variation can be composed to provide robustness against simultaneous shifts. To evaluate this feature, we curate several new datasets containing simultaneous sources of natural variation. Following this, in Section 6.4, we show that models of natural variation trained on a fixed dataset can be reused to provide robustness on datasets entirely unseen while training . This demonstrates that model-based robustness is highly transferable, meaning that once a model of natural variation has been learned for a particular source of natural variation, it can be reused to provide robustness against the same source of natural variation in future applications. Finally, in Section 6.5, we show that in the setting of unsupervised domain adaptation, our algorithm outperforms a well-known domain adaptation baseline.
Notation for image domains. Throughout this section, we consider a wide range of datasets, including MNIST , SVHN , GTSRB , CURE-TSR , MNIST-m , Fashion-MNIST , EMNIST , KMNIST , QMNIST , USPS , ImageNet , and ImageNet-c . For many of these datasets, we extract subsets corresponding to different sources of natural variation; henceforth, we will call these subsets domains. To indicate the source of natural variation under consideration in a given experiment, we use the notation “source (AB)” to denote a distributional shift from domain to domain . For example, “contrast (lowhigh)” will denote a shift from low-contrast to high-contrast within a particular dataset. Images from domains and for each of the shifts used in this paper are available in Appendix B. We note that our experiments contain domains with both natural and artificially-generated variation; details concerning how we extracted non-artificial variation can be found in Appendix C.
We emphasize that each image domain used in this paper contains a training set and a test set. While both the training and test set in each domain come from the same distribution, neither the models of natural variation nor the classifiers have access to test data from any domain during the training phase. More explicitly, when learning models of natural variation and training classifiers, we use data from the training set of the relevant domains. Conversely, when testing the classifiers, we use data from the test set of the relevant domains.
Baseline algorithms and evaluation metrics. In the experiments, we consider a variety of baseline algorithms. In particular, where approapriate, we compare our model-based algorithms to empirical risk mimization (ERM), the adversarial training algorithm PGD , a recently-proposed data-augmentation technique called AugMix , and the domain adaptation technique known as Adversarial Discriminative Domain Adaptation (ADDA) . To evaluate the performance of each of these algorithms, we will report the top-1 accuracy (i.e. the standard classification accuracy) and, where appropriate, the top-5 accuracy (i.e. the frequency with which the ground-truth label is one of the top five predicted classes). Further details concerning architecture selection and the hyperparameters used are given in Appendix A.
In many applications, one might have data corresponding to low levels of natural variation, such as a dusting of snow in images of street signs. However, it is often difficult to collect data corresponding to high levels of natural variation, such as images taken during a blizzard. In such cases, we show that our algorithms can be used to provide significant out-of-distribution robustness against data with high levels of natural variation by training on data with relatively low levels of the same source of natural variation. To do so, we use data from the CURE-TSR dataset , which contains images of street signs divided into subsets according to various sources of natural variation and corresponding severity levels. For example, for images in the “snow” subset, challenge-level 0 corresponds to no snow, whereas challenge-level 5 corresponds to a full blizzard.
For each row of Table 2, we use unlabeled data from challenge-levels 0 and 2 to learn a model of natural variation corresponding to a given subset of CURE-TSR. We then train classifiers using MRT-10 with labeled challenge-level 0 data. We also train classifiers using ERM and PGD using the labeled data from challenge-levels 0 and 2. To denote the fact that ERM and PGD are trained using an augmented dataset containing labeled challenge-level 2 data in addition to challenge-level 0 data, we denote these algorithms in Table 2 by ERM+Aug and PGD+Aug. We then test all classifiers on data from challenge-levels 3, 4, and 5. Note that in some sense this is an unfair comparison for our methods, given that the model-based algorithms are not given access to labeled challenge-level 2 data. In spite of this unfair comparison, our algorithms still outperform the baseline algorithms.
By considering each row in 2, a general pattern emerges. When tested on challenge-level 3 and 4 data, MRT improves over ERM+Aug and PGD+Aug, although in some cases the improvements are somewhat modest. These modest improvements are due in part to the fact that challenge-level 3 data has only slightly more natural variation than challenge-level 2 data, which all of the classifiers have seen in either labeled form (in the case of the baselines) or unlabeled form (in the case of MRT). However, by comparing the performance of the trained classifiers on challenge-level 5 data, it is clear that MRT improves significantly over the baselines by as much as 20 percentage points. This shows that as the challenge-level becomes more severe, our model-based algorithms outperform the baselines by relatively larger margins.
2 Model-based robustness on the shift from ImageNet to ImageNet-c
To demonstrate the scalability of our approach, we perform experiments on ImageNet and the recently-curated ImageNet-c dataset . ImageNet-c contains images from the ImageNet test set that are corrupted according to artificial transformations, such as snow, rain, and fog, and are labeled from 1-5 depending on the severity of the corruption. For brevity, in Table 3, we abbreviate ImageNet-c to “IN-c.”
For numerous challenging corruptions, we train models to map from the classes 0-9 of ImageNet to the corresponding classes of ImageNet-c. We then train all classifiers, each of which uses the ResNet50 architecture , on classes 10-59 of ImageNet, and test on the corresponding classes for various subsets of ImageNet-c. Note that in this setting, the ImageNet classes used to train the model of natural variation are disjoint from those that are used to train the classifier, so many techniques, including most domain adaptation methods, do not apply. To offer a point of comparison, we include the accuracies of classifiers trained using AugMix, which is a recently proposed method that adds known transformations to the data .
By considering Table 3, the first notable observation is that the classifiers trained using ERM perform very poorly on the shift from ImageNet to ImageNet-c. That is, when evaluating ResNet50 classifiers trained on ImageNet on the original ImageNet test set, one would expect to attain a peak top-1 accuracy of more than 80 percentage points ; therefore, the drop in classification accuracy for the classifiers trained using ERM is remarkable given that the corruptions included in this table are quite common. This demonstrates the degree to which classifiers trained on ImageNet lack robustness to common corruptions and transformations of data . To this end, the authors of sought to address this fragility by introducing the AugMix algorithm, which adds randomly corrupted data to augment the training dataset. As shown in Table 3, while this method does improve accuracy on two challenges, for other corruptions, the top-1 and top-5 accuracies plummet. On the other hand, in almost all cases that we considered, MDA-3 significantly outperformed both methods, in some cases improving by nearly 30 percentage points in top-1 accuracy.
3 Robustness to simultaneous distributional shifts
In practice, it is common to encounter multiple simultaneous distributional shifts. For example, in image classification, there may be shifts in both brightness and contrast; yet while there may be examples corresponding to shifts in either brightness or contrast in the training data, there may not be any examples of both shifts occurring simultaneously. To address this robustness challenge, for each row of Table 4, we learn two models of natural variation and using unlabeled training data corresponding to two separate shifts, which map domains (e.g. low- to high-brightness) and (e.g. low- to high-contrast). We then compose these models to form a new model
which can be used to vary the natural conditions in a given input image according to both sources of natural variation modeled by and . We then train classifiers on labeled data from (e.g. data with either low-brightness or low-contrast) and test on data from (e.g. data with both high-brightness and high-contrast).
To gather data from containing images with high-brightness and high-contrast from SVHN, we threshold the SVHN training set according to both of these sources of natural variation. On the other hand, to create the data from for the ImageNet experiments, we curate four new datasets by applying pairs of transformations that were originally used to create the ImageNet-c datasets. Images corresponding to these dataset, which contain shifts in brightness and contrast, brightness and snow, brightness and fog, and contrast and fog, as well as more details concerning how they were curated are available in Appendix C.
The results in Table 4 reveal that across each of the settings we considered, MDA-3 significantly outperformed ERM with respect to top-1 classification accuracy. Despite the fact that neither algorithm has access to data from domain during training, MDA is able to improve by between 5 and 35 percentage points over ERM across these five settings. The composable nature of our methods is unique to our framework; we plan to explore this feature further in future work.
4 Transferability of model-based robustness
Because we learn models of natural variation offline before training a classifier, our paradigm can be applied to domains that are entirely unseen while training the model . In particular, we show that models can be reused on similar yet unseen datasets to provide robustness against a common source of natural variation. For example, one might have access to two domains corresponding to the shift from images of European street signs taken during the day to images taken at night. However, one might wish to provide robustness against the same shift from daytime to nighttime on a new dataset of American street signs without access to any nighttime images in this new dataset. Whereas many techniques, including most domain adaptation methods, do not apply in this scenario, in the model-based robust deep learning paradigm, we can simply learn a model of natural variation corresponding to the changes in lighting for the European street signs and then apply this model to the dataset of the American signs.
Table 5 shows several experiments of this stripe in which a model of natural variation is learned on one dataset and then applied on another dataset . Notably, classifiers trained using our model-based algorithms significantly outperform the ERM and PGD baselines in each case, despite the fact that all classifiers have access to the same training data from dataset . In particular, by using models of natural variation trained on , and then training in the model-based paradigm on , we are able to improve the test accuracy on the shift from domain to domain on by as much as 40 percentage points. This demonstrates that models of natural variation can be reused to provide robustness against domains that are entirely unseen during the training of the model of natural variation.
5 Model-based robust deep learning for unsupervised domain adaptation
The results of Sections 6.1, 6.2, 6.3, and 6.4 show that our approach does not require labelled or unlabelled data from domain to improve robustness on the shift from domain to domain . In particular, in Section 6.1, we trained classifiers on low levels of natural variation and then evaluated these classifiers on higher levels of natural variation. Next, in Section 6.2 we used a different distribution of classes to train a model of natural variation than we used to train and evaluate classifiers. Further, in Section 6.3, we assumed access to data corresponding to two fixed distributional shifts, but not to the data corresponding to both shifts occurring simultaneously. Finally, in Section 6.4, we trained models of natural variation on a different dataset from the dataset used to train and evaluate classifiers.
However, when unlabelled data corresponding to a fixed domain shift from domain to domain is available, it is of interest to evaluate how our approach compares to relevant methods such as domain adaptation. We emphasize that while this is one of the most commonly studied settings in domain adaptation, it represents only one particular setting to which the model-based robust deep learning paradigm can be applied.
In Table 6, for each shift from domain to , we assume access to labeled data from domain and unlabeled data from domain . In each row, we use unlabeled data from both domains to train a model of natural variation . We then train classifiers using our algorithms, ERM, and PGD using data from domain ; we then evaluate these classifiers on data from the test set of domain . Furthermore, we compare to ADDA, which is a well-known domain adaptation method . In every scenario, our model-based algorithms significantly outperform the baselines, often by 10-20 percentage points. Notably, while ADDA offers strong performance on SVHN, it fairs significantly worse on GTSRB and CURE-TSR.
Discussion
Following the experiments of Section 6, we now discuss further aspects of model-based training. In particular, we provide an additional experiment which shows that more accurate models of natural variation engender classifiers that are more robust against natural variation in our paradigm. We also discuss choices for the integer parameter in each of the model-based training algorithms. Furthermore, we provide details concerning how we chose to define the nuisance spaces for learned models of natural variation used in the experiments of Section 6.
An essential yet so far undiscussed piece of the efficacy of our model-based paradigm is the impact of the quality of learned models of natural variation on the robustness we are ultimately able to provide. In scenarios where we do not have access to a known model of natural variation, the ability to provide any sort of meaningful robustness relies on learned models that can accurately render realistic looking data consistent with various sources of natural variation. To this end, it is reasonable to expect that models that can more effectively render realistic yet challenging data should result in classifiers that are more robust to shifts in natural variation.
To examine the impact of the quality of learned models of natural variation in our paradigm, we consider a task performed in Section 6.5, wherein we learned a model of natural variation that mapped low-contrast samples from SVHN, which comprised domain , to high-contrast samples from SVHN, which comprised domain . While learning this model, we saved snapshots at various points during the training procedure. In particular, we collected a family of intermediate models , where the index refers to the number of MUNIT training iterations that were completed before saving the snapshot:
In Figure 9(j), we show the result of training classifiers with MRT-10 using each model of natural variation . Note that the models that are trained for more training steps engender classifiers that provide higher levels of robustness against the shift in natural variation. Indeed, as the model produces random noise, the performance of this classifier performs at effectively the level as the baseline classifier (shown in Table 6). On the other hand, the model is able to accurately preserve the semantic content of the input data while varying the nuisance content, and is therefore able to provide higher levels of robustness. In other words, better models induce improved robustness for classifiers trained in the model-based paradigm.
While this experiment demonstrates that models of natural variation which generate more realistic data ultimately lead to classifiers that are more robust against out-of-distribution shifts, further questions remain. In particular, in our setting, as learned models of natural variation critically rely on the ability of deep generative models to render realistic images, there is no guarantee that learned models will always accurately reflect natural conditions. Indeed, it remains an open problem as to how to ensure that deep generative models generalize effectively . To this end, one promising direction for future work is to study the conditions under which reliable models of natural variation can be learned.
2 Algorithm and hyperparameter selection criteria
Throughout the experiments, in general we see that the sampling-based algorithms presented in this paper achieve higher levels of robustness against almost all sources of natural variation. This finding stands in contrast to field of perturbation-based robustness, in which adversarial methods have been shown to be the most effective in improving the robustness against small, norm-bounded perturbations . Going forward, an interesting research direction is not only to consider new algorithms but also to understand whether sampling-based or adversarial techniques provide more robustness with respect to a given model of natural variation.
The impact of in the model-based training algorithms. The discussion surrounding the difference between the sampling-based and adversarial mindsets that characterize our model-based algorithms is intimately related to the parameter in each of the model-based training algorithms. Notably, in each of these algorithms, the parameter serves distinct yet related purposes. In MRT, controls the number of nuisance parameters that are sampled at each iteration. Similarly, in MAT, determines the number of steps of gradient ascent performed at each iteration. On the other hand, in MDA, controls the number of data points that are added to the training set at each iteration.
As we show in our experiments, each of the three model-based algorithms can be used to provide significant out-of-distribution robustness against various sources of natural variation. In this subsection, we focus on the impact of varying the parameter in each of these algorithms. In particular, in Table 7, we see that varying has a different impact for each of the three algorithms. For MAT, we see that increasing decreases the accuracy of the trained classifier; one interpretation of this phenomenon is that larger values of allow MAT to find more challenging forms of natural variation. On the other hand, the test accuracy of MRT improves slightly as increases. Recall that while both MRT and MAT seek to find “worst-case” natural variation, MRT employs a sampling-based approach to solving the inner maximization problem as opposed to the more precise, gradient-based procedure used by MAT. Thus the differences in the impact of varying between MAT and MRT may be due to the fact that MRT only approximately solves the inner problem at each iteration. Finally, we see that increasing slightly decreases the test accuracy of classifiers trained with MDA.
This study can also be used as an algorithm selection criteria. Indeed, when data presents many modes corresponding to different levels of natural variation, it may be more efficacious to use MRT or MDA, which will observe a more diverse set of natural conditions due to their sampling-based approaches. On the other hand, when facing a single challenging source of natural variation, it may be more useful to use MAT, which seeks to find “worst-case,” natural, out-of-distribution data.
To this end, throughout our experiments, we generally scale with . For low-dimensional data (e.g. MNIST, SVHN, etc.), we found that sufficed toward capturing the underlying source of natural variation effectively. However, on datasets such as GTSRB, for which we rescaled instances into arrays, we found that was more appropriate for capturing the full range of natural variation. Indeed, on ImageNet, which contains instances of size , we found that still produced images that captured the essence of the underlying source of natural variation.
Related works
In this section, we attempt to characterize works from a variety of fields, including adversarial robustness, generative modeling, equivariant neural networks, and domain adaptation, which have been influential in the development of the model-based robust deep learning paradigm.
A rapidly growing body of work has addressed adversarial robustness of neural networks with respect to small norm-bounded perturbations. This problem has motivated an arms-race-like amalgamation of adversarial attacks and defenses within the scope of norm-bounded adversaries and has prompted researchers to closely study the theoretical properties of adversarial robustness . And while some defenses have withstood a variety of strong adversaries , it remains to be seen as to whether such progress will ultimately lead to deep learning models that are reliably robust against adversarial, perturbation-based attacks.
Several notable works that propose methods for defending against adversarial attacks formulate so-called adversarial training algorithms, the goal of which is to defend neural networks against worst-case perturbations . Some of the most successful works take a robust optimization perspective, in which the goal is to find the worst-case adversarial perturbation of data by solving a min-max problem . However, it has also been shown that randomized smoothing based defenses are able to withstand strong attacks . In a different yet related line of work, optimization-based methods have been proposed to provide certifiable guarantees on the robustness of neural networks against small perturbations . Others have also studied how different architectural choices can result in classifiers that are more robust against adversarial examples .
As adversarial training methods have become more sophisticated, a range of adaptive adversarial attacks, or attacks specifically targeting a particular defense, have been proposed . Prominent among the attacks on robustly-trained classifiers have been algorithms that circumvent so-called obfuscated gradients . Such attacks generally focus on generating adversarial examples that are perceptually similar to a given input image . Very recently, the authors of developed the notion of a “perceptual adversarial attack,” in which an adversary can corrupt a given input image subject to a learned “neural perceptual distance.” As far as the authors are aware, this is the strongest imperceptible adversarial attack for trained classifiers.
In summary, the commonality among all the approaches mentioned above is that they generally consider norm-bounded adversarial perturbations which are perceptually indistinguishable from true examples. Contrary to these approaches, in this work we propose a paradigm shift from norm-bounded perturbation-based robustness to model-based robustness. To this end, we have provided training algorithms that improve robustness against natural shifts in the data distribution; such changes are often perceptually distinct from input data, as is the case for varying lighting or weather conditions.
2 Generative models in the context of robustness
Another line of work has sought to leverage deep generative models in the loop of training to facilitate strategies for generating and defending against adversarial examples. In , , and , the authors propose attack strategies that use the generator from a generative adversarial network (GAN) to generate additive perturbations that can be used to attack a classifier. On the other hand, a framework called DefenseGAN, which uses a Wasserstein GAN to “denoise” adversarial examples , has been proposed to defend against perturbation-based attacks. This defense method was later broken by the Robust Manifold Defense , which searches over the parameterized manifold induced by a generative model to find worst-case perturbations of data. The min-max formulation used in this work is analogous to the PGD defense , and was foundational in our development of the model-based robust deep learning paradigm.
Closer to the approach we describe in this paper are works that use deep generative models to generate adversarial inputs themselves, rather than generating small perturbations. The authors of and use the generator from a GAN to generate adversarial examples that obey Euclidean norm-based constraints. Alternatively, the authors of use GANs to generate adversarial patterns that can be used to construct adversarial examples in multiple domains. Similarly, and generate unrestricted adversarial examples, or instances that are not subjected to a norm-based constraint , via a deep generative model. Finally, two recent papers use a GAN to perform data-augmentation by generating perceptually realistic samples.
In this work, we use generative networks to learn and model the natural variability within data. This is indeed different than generating norm-based adversarial perturbations or perceptually realistic adversarial examples, as have already been considered in the literature. Our generative models are designed to capture natural shifts in data, whereas the relevant literature has sought to create synthetic adversarial perturbations that fool neural networks.
3 A broader view of robustness in deep learning
Aside from the algorithms we introduced in Section 5, we are not aware of any other algorithms that can be used to address out-of-distribution robustness across the diverse array of tasks presented in Section 6. However, several lines of research have sought to address this problem in constrained settings or under highly restrictive assumptions. Of note are works that have sought to provide robustness against specific transformations of data that are more likely to be encountered in applications than norm-bounded perturbations. Transformations that have recently received attention from the adversarial robustness community include adversarial quilting , adversarial patches and clothing , geometric transformations , distortions , deformations and occlusions , and nuisances encountered by unmanned aerial vehicles . In response to these works and motivated by myriad safety-critical applications, first steps toward robust defenses against specific distributional shifts have recently been proposed . The resulting methodologies generally leverage properties specific to the transformation of interest.
Another line of work has sought to use various mechanisms to create semantically-realistic adversarial examples via more general frameworks. For example, in , the authors seek to create “semantic adversarial examples” via color-based shifts, which relates to the perspective advocated for in , in which the authors argue that system semantics and specifications should be considered when generating meaningful disturbances in the data. Similarly, in , the authors use differentiable renderers to generate semantically meaningful changes. In the same spirit, the idea in is to leverage an information theoretic approach to edit the nuisance content of images to create perceptually realistic data that causes misclassification.
While this progress has helped to motivate new notions of robustness, defenses against specific threat models are limited in the sense that they often cannot be generalized to develop a learning paradigm that is broadly applicable across different forms of natural variation. This contrasts with the motivation behind this paper, which is to provide general robust training algorithms that can improve the robustness of trained neural networks across a variety of scenarios and applications.
More related to the current work are two concurrent papers that formulate robust training procedures under the assumption that data is corrupted according to a fixed generative architecture. The authors of exploit properties specific to the StyleGAN architecture to formulate a training algorithm that provides robustness against color-based shifts on MNIST and CelebA . In our work, we propose a more general framework and three novel robust training algorithms that can exploit any suitable generative model, and we show improvements on more challenging, naturally-occurring shifts across twelve distinct datasets. The authors of use conditional VAEs to learn perturbation sets corresponding to simple corruptions from pairs of images. In our framework we improve robustness against more challenging, natural shifts by learning from unpaired datasets and we do not rely on class-conditioning to generate realistic images.
4 Domain adaptation and domain generalization
In the domain adaptation literature, various methods have been proposed which rely on the restrictive assumption that unlabeled data corresponding to a fixed distributional shift is available during training . Several works in this vain use an adversarial min-max formulation to adapt a classifier trained on a source domain to perform well on a related target domain, for which labels are unavailable at training time . We note that these “adversarial” methods differ from the model-based paradigm we introduce in that rather than formulating the problem of adapting a classifier to perform well on the target domain, we employ a min-max procedure to search for worst-case shifts in data for a fixed generative model. Indeed, the main difference between domain adaptation techniques and our paradigm is that our solution does not assume access to unlabeled data from a fixed shift and can be applied to datasets that are entirely unseen during training.
Also related is the field of domain generalization , in which one assumes access to a variety of training domains, all of which are related to an unseen target domain on which the trained classifier is ultimately evaluated. Such works often rely on transfer learning and distribution matching to improve classification accuracy on the unseen domain. While the experiments in Sections 6.4 tackle a similar problem, in which knowledge from one domain is used to learn a classifier that performs well on an unseen test domain, the experiments in the other subsections of Section 6 show that our model-based paradigm is much more broadly applicable that techniques domain generalization techniques.
Conclusion and future directions
In this paper, we formulated a novel problem that concerns the robustness of deep learning with respect to natural, out-of-distribution shifts in data. Motivated by perceptible nuisances in computer vision, such as lighting or weather changes, we propose the novel model-based robust deep learning paradigm, the goal of which is to improve the robustness of deep learning systems with respect to naturally occurring conditions. This notion of robustness offers a departure from that of norm-bound, perturbation-based adversarial robustness, and indeed our optimization-based formulation results in a new family of training algorithms that can be used to train neural networks to be robust against arbitrary forms of natural variation. Across a range of diverse experiments and across twelve distinct datasets, we empirically show that the model-based paradigm is broadly applicable to a variety of challenging domains.
Our model-based robust deep learning paradigm open numerous directions for future work. In what follows, we briefly highlight several of these broad directions.
Learning a library of nuisance models. One natural future direction is to determine superior ways of learning models of natural variations to perform model-based training. In this paper, we used the MUNIT framework , but other existing architectures may be better suited for specific forms of natural variation or for different datasets. Indeed, a more rigorous statistical analysis of problem (4.4) may lead to the discovery of new architectures designed specifically for model-based training. To this end, recent work that concerns learning invariances in neural networks may provide insight into learning physically meaningful models . Beyond computer vision, learning such models in other domains such as robotics would enable new applications.
Model-based algorithms and architectures. Another important direction involves the development of new algorithms for solving the min-max formulation of (3.1). In this paper, we presented three algorithms – MRT, MDA, and MAT – that can be used to approximately solve this problem, but other algorithms are possible and may result in higher levels of robustness. In particular, adapting first-order methods to search globally over the learned image manifold may provide more efficient, scalable, or robust results. Indeed, one open question is whether it is necessary to decouple the procedures used to learn classifiers and models of natural variation. Another interesting direction is to follow the line of work that has developed equivariant neural network architectures toward designing new architectures that are invariant to models of natural variation.
Applications beyond image classification. Throughout the paper, we have empirically demonstrated the utility of our approach across many image classification tasks. However, our model-based paradigm could also be broadly applied to further applications, both in computer vision as well as outside computer vision. Within computer vision, one can consider other tasks such as semantic segmentation in the presence of challenging natural conditions. Outside computer vision, one exciting area is to exploit physical models of robot dynamics with deep reinforcement learning for applications such as walking in unknown terrains. In any domain where one has access to suitable models of natural variation, our approach allows domain experts to leverage these models in order to make deep learning far more robust.
Theoretical foundations. Finally, we believe that there are many exciting open questions with respect to the theoretical aspects of model-based robust training. What types of models provide significant robustness gains in our paradigm? How accurate does a model need to be to engender neural networks that are robust against natural variation and out-of-distribution shifts? We would like to address such theoretical questions from geometric, physical, and statistical perspectives with an eye toward developing faster algorithms that are both more sample-efficient and that provide higher levels of robustness. A deeper theoretical understanding of our model-based robust deep learning paradigm could result in new approaches that blend model-based and perturbation-based methods and algorithms.
Acknowledgements
This work has been partially supported by the Defense Advanced Research Projects Agency (DARPA) Assured Autonomy under Contract No. FA8750-18-C-0090, AFOSR under grant FA9550-19-1-0265 (Assured Autonomy in Contested Environments), NSF CPS 1837210, and ARL CRA DCIST W911NF-17-2-0181 program.
Authors
All authors are with the Department of Electrical and Systems Engineering, University of Pennsylvania, Philadelphia, PA 19104. The authors can be reached at the following email addresses:
References
Appendix A Training details
In this appendix, we provide details concerning our implementation of both the model-based robust deep learning algorithms as well as models of natural variation. We note that all experiments described in Sections 6 and 7 were run on four NVIDIA RTX 5000 GPUs.
PGD. When training classifiers with PGD , we use a perturbation budget of , a step size of , and twenty iterations of gradient ascent per batch. Our implementation of PGD is included along with our implementation of the model-based robust deep learning algorithms; this implementation can be found at https://github.com/arobey1/mbrdl.
ADDA. We used the following implementation of ADDA : https://github.com/Carl0520/ADDA-pytorch. We trained each classifier for 100 epochs on the source domain, and we then trained for a further 100 epochs when adaptating the weights of the target encoder. We used the LeNet for the encoder networks; to ensure a fair comparison, we used two convolutional layers and two feed-forward layers, which is the same as the classification networks used for model-based training.
AugMix. We used the following implementation of AugMix : https://github.com/google-research/augmix. We used the default parameters for training described in this repository.
Model-based algorithms. All three model-based algorithms (MRT, MAT, and MDA) are implemented in our repository: https://github.com/arobey1/mbrdl. Throughout the experiments, unless stated otherwise, we ran MRT with , MAT with , and MDA with . We also used a trade-off parameter of .
A.2 Classifier training details
When training classifiers for MNIST, Q-MNIST, E-MNIST, K-MNIST, Fashion-MNIST, USPS, SVHN, GTSRB, and CURE-TSR we use a simple CNN architecture with two convolutional layers and two feed forward layers. More specifically, the architecture we used the following architecture:
For each of these experiments, we use the Adadelta optimizer with a learning rate of 1.0. We also use a batch size of 64. Images from MNIST, Q-MNIST, E-MNIST, K-MNIST, Fashion-MNIST, USPS, SVHN, and CURE-TSR are resized to ; for grayscale datasets such as MNIST, we repeat the channels three times. Images from GTSRB are resized to . We train each classifier for 100 epochs.
When training on ImageNet, we use the ResNet-50 architecture. We note that architectural choices are possible and will be explored in future work. For each of the experiments on ImageNet, we use SGD with an initial learning rate of 0.05; we decay the learning rate linearly to 0.001 over 100 epochs. We use a batch size of 64 for ERM and PGD. For the model-based algorithms, we use a batch size of 32, as these algorithms require the GPU to store multiple copies of each image during each training iteration.
A.3 MUNIT framework overview
The MUNIT model consists of an encoder-decoder pair and for each image domain and . These encoder-decoder pairs are trained to learn a mapping that reconstructs its input. That is, and . More specifically, is trained to encode into a content code and a style code . Similarly, is trained to encode into and . Then the decoding networks and are trained to reconstruct the encoded pairs and into the respective images and .
Inter-domain image translation is performed by swapping the decoders. In this way, to map an image from to , is first encoded into . Then, a new style vector is sampled from from a prior distribution on the set and the translated image is equal to . The translation of from to can be described via a similar procedure with , , and a prior supported on . In this paper, we follow the convention used in as use a Gaussian distribution for both and with zero mean and an identity covariance matrix.
Training an MUNIT model involves considering four loss terms. First, the encoder-decoder pairs and are trained to reconstruct their inputs my minimizing the following loss:
Further, when translating an image from one domain to another, the authors of argue that we should be able to reconstruct the style and content codes. By rewriting the encoding networks as and , the constraint on the content codes can be expressed in the following way:
Finally, two GANs corresponding to the two domains and are used to form an adversarial loss term. The GANs use the decoders and as the respective generators for domains and . By denoting the discriminators for these domains by and , we can write the GANs as and . In this way, the final loss term takes the following form:
Using the four loss terms we have described, the MUNIT framework uses first-order methods to solve the following nonconvex optimization problem:
A.4 Hyperparameters and implementation of MUNIT
In Table 8, we record the hyperparameters we used for training models of natural variation via the MUNIT framework. The hyperparameters we selected are generally in line with those suggested in . We use the same architectures for the encoder, decoder, and discriminative networks as are described in Appendix B.2 of .
Appendix B A gallery of learned models of natural variation
We conclude this section by showing images corresponding to the many distributional shifts used in the experiments section. Furthermore, we show images generated by passing domain images through learned models of natural variation.
Appendix C Details concerning datasets and domains
As mentioned in Section 6, we used twelve different datasets in this work to fully evaluate the efficacy of the model-based algorithms we introduced in Section 5. For several of these datasets, we curated subsets corresponding to different factors of natural variation, which we refer to as domains.
Throughout our experiments, we demonstrate that our methods are able to provide robustness against many challenging sources of natural variation. Furthermore, our experiments contain domains with both naturally-occurring and artificially-generated variation. Notably, every experiment involving data from SVHN or GTSRB used naturally-occurring variation. In what follows, we discuss both of these categories.
Throughout the experiments, we used data from SVHN and GTSRB to train neural networks to be robust against contrast and brightness variation. To extract naturally-occurring variation from these datasets, we used simple metrics to threshold the data into subsets corresponding to different levels of natural variation. Specifically, we defined the brightness of an RGB image to be the mean pixel value of , and we define the contrast to be the difference between the largest and smallest pixel values. Table 10 show the thresholds we chose for contrast and brightness on SVHN and GTSRB. Note that these thresholds were chosen somewhat subjectively to reflect our perception of low, medium and high values of brightness and contrast. We intend to experiment with different thresholds in future work.
Figure 23 shows a summary of the subsets of SVHN that we compiled corresponding to brightness. In particular, Figure 23(a) shows a histogram of the brightnesses of images in SVHN. We used this histogram to set thresholds for low, medium, and high brightness, which are given in Table 10. The images below the histogram correspond to the bins of the histogram; that is, images further to the left in Figure 23(a) have lower brightness, whereas images further to the right have high brightness. In Figures 23(b), 23(c), and 23(d), we show samples from the subsets of low, medium and high contrast subsets of SVHN that we compiled. Figure 24 tells the same story as 23 for the contrast nuisances in SVHN. Again, Figure 24(a) shows a histogram and accompanying images corresponding to different values of contrast. Figures 24(b), 24(c), and 24(d) show samples from the subsets of low, medium, and high contrast images we compiled.
We repeat this analysis for the brightness and contrast thresholding operations for GTSRB in Figures 25 and 26. Again, the difference between high- and low-brightness samples is remarkable, as is the difference in the samples corresponding to high- and low-contrast. However, an interesting difference between the distributions of brightness and contrast on GTSRB vis-a-vis SVHN is that the distributions for GTSRB are skewed, whereas the distributions for SVHN are close to being symmetric.
C.1.2 Artificially-generated variation
The remainder of the experiments, including those on MNIST, CURE-TSR, and ImageNet, use artifically-generated variation. Indeed, one challenge in addressing deep learning’s lack of robustness to natural variation is that relatively few datasets contain labeled forms of naturally-occurring sources of variation. To this end, an important research challenge is to curate datasets with naturally-occurring variation; we plan to pursue this goal in future work.
When data with naturally-occurring variation is not available, artificially-generated variation can be used as an effective proxy for testing the robustness of deep learning against different forms of variation . Indeed, the recently curated CURE-TSR and ImageNet-c were created using pre-defined, artifical transformations of data. While these transformations are synthetic, the images in Appendix B show that they are indeed quite realistic.
C.2 Datasets introduced in this paper
In this paper, we introduced several new datasets which contain multiple simultaneous corruptions, including show, brightness, contrast, and fog. In particular, we used the transforms used to create ImageNet-c to add multiple corruptions to the ImageNet test set. To do so, we used the open-source code from , which can be found at https://github.com/hendrycks/robustness. Images of these datasets are shown in Figures 26-29.