Learning perturbation sets for robust machine learning
Eric Wong, J. Zico Kolter
Introduction
Within the last decade, adversarial learning has become a core research area for studying robustness and machine learning. Adversarial attacks have expanded well beyond the original setting of imperceptible noise to more general notions of robustness, and can broadly be described as capturing sets of perturbations that humans are naturally invariant to. These invariants, such as facial recognition should be robust to adversarial glasses (Sharif et al., 2019) or traffic sign classification should be robust to adversarial graffiti (Eykholt et al., 2018), form the motivation behind many real world adversarial attacks. However, human invariants can also include notions which are not inherently adversarial, for example image classifiers should be robust to common image corruptions (Hendrycks & Dietterich, 2019) as well as changes in weather patterns (Michaelis et al., 2019).
How can we learn models that are robust to perturbations without a predefined perturbation set?
In the absence of a mathematical definition, in this work we present a general framework for learning perturbation sets from perturbed data. More concretely, given pairs of examples where one is a perturbed version of the other, we propose learning generative models that can “perturb” an example by varying a fixed region of the underlying latent space. The resulting perturbation sets are well-defined and can naturally be used in robust training and evaluation tasks. The approach is widely applicable to a range of robustness settings, as we make no assumptions on the type of perturbation being learned: the only requirement is to collect pairs of perturbed examples.
Given the susceptibility of deep learning to adversarial examples, such a perturbation set will undoubtedly come under intense scrutiny, especially if it is to be used as a threat model for adversarial attacks. In this paper, we begin our theoretical contributions with a broad discussion of perturbation sets and formulate deterministic and probabilistic properties that a learned perturbation set should have in order to be a meaningful proxy for the true underlying perturbation set. The necessary subset property ensures that the set captures real perturbations, properly motivating its usage as an adversarial threat model. The sufficient likelihood property ensures that real perturbations have high probability, which motivates sampling from a perturbation set as a form of data augmentation. We then prove the main theoretical result, that a learned perturbation set defined by the decoder and prior of a conditional variational autoencoder (CVAE) (Sohn et al., 2015) implies both of these properties, providing a theoretically grounded framework for learning perturbation sets. The resulting CVAE perturbation sets are well motivated, can leverage standard architectures, and are computationally efficient with little tuning required.
Background and related work
Other work has studied perturbation sets that are not necessarily mathematically formulated but well-defined from a human perspective such as spatial transformations (Xiao et al., 2018b). Real-world adversarial attacks tend to try to remain either inconspicuous to the viewer or meddle with features that humans would naturally ignore, such as textures on 3D printed objects (Athalye et al., 2017), graffiti on traffic signs (Eykholt et al., 2018), shapes of objects to avoid LiDAR detection (Cao et al., 2019), irrelevant background noise for audio (Li et al., 2019a), or barely noticeable films on cameras (Li et al., 2019b). Although not necessarily adversarial, Hendrycks & Dietterich (2019) propose the set of common image corruptions as a measure of robustness to informal shifts in distribution.
Generative modeling and adversarial robustness
Adversarial defenses and data augmentation
Successful approaches for learning adversarially robust networks include methods which are both empirically robust via adversarial training (Goodfellow et al., 2014; Kurakin et al., 2016; Madry et al., 2017) and also certifiably robust via provable bounds (Wong & Kolter, 2017; Wong et al., 2018; Raghunathan et al., 2018; Gowal et al., 2018; Zhang et al., 2019) and randomized smoothing (Cohen et al., 2019; Yang et al., 2020). Critically, these defenses require mathematically-defined perturbation sets, which has limited these approaches from learning robustness to more general, real-world perturbations. We directly build upon these approaches by learning perturbation sets that can be naturally and directly incorporated into robust training, greatly expanding the scope of adversarial defenses to new contexts. Our work also relates to using non-adversarial perturbations via data augmentation to reduce generalization error (Zhang et al., 2017; DeVries & Taylor, 2017; Cubuk et al., 2019), which can occasionally also improve robustness to unrelated image corruptions (Geirhos et al., 2018; Hendrycks et al., 2019; Rusak et al., 2020). Our work differs in that rather than aggregating or proposing generic data augmentations, our perturbation sets can provide data augmentation that is targeted for a particular robustness setting.
Perturbation sets learned from data
where has support and is a distribution parameterized by .
To be a reasonable threat model for adversarial examples, one desirable expectation is that a perturbation set should at least contain close approximations of the perturbed data. In other words, the set of perturbed data should be (approximately) a necessary subset of the perturbation set. This notion of containment can be described more formally as follows:
Our second desirable property is specific to the probabilistic view from Equation (3), where we would expect perturbed data to have a high probability of occurring under a probabilistic perturbation set. In other words, a perturbation set should assign sufficient likelihood to perturbed data, described more formally in the following definition:
A model that assigns high likelihood to perturbed observations is likely to generate meaningful samples, which can then be used as a form of data augmentation in settings that care more about average-case over worst-case robustness. To measure this property, the likelihood can be approximated with a standard Monte Carlo estimate by sampling from the prior .
Variational autoencoders for learning perturbations sets
where , , , , and are arbitrary functions representing the respective encoder, prior, and decoder networks. CVAEs are trained by maximizing a likelihood lower bound
Our theoretical results prove that optimizing the CVAE objective naturally results in both the necessary subset and sufficient likelihood properties outlined in Section 3.1, which motivates why the CVAE is a reasonable framework for learning perturbation sets. Note that these results are not immediately obvious, since the likelihood of the CVAE objective is taken over the full posterior while the perturbation set is defined over a constrained latent subspace determined by the prior. The proofs rely heavily on the multivariate normal parameterizations, with requiring several supporting results which relate the posterior and prior distributions. We give a concise, informal presentation of the main theoretical results in this section, deferring the full details, proofs, and supporting results to Appendix B. Our results are based on the minimal assumption that the CVAE objective has been trained to some threshold as described in Assumption 1.
The CVAE objective has been trained to some thresholds as follows
where each bounds the KL-divergence of the th dimension.
Our first theorem, Theorem 1, states that the approximation error of a perturbed example is bounded by the components of the CVAE objective. The implication here is that with enough representational capacity to optimize the objective, one can satisfy the necessary subset property by training a CVAE, effectively capturing perturbed data at low approximation error in the resulting perturbation set.
Let be the Mahalanobis distance which captures of the probability mass for a -dimensional standard multivariate normal for some . Then, there exists a such that and for
where is a constant dependent on . Moreover, as and (the theoretical limits of these boundsIn practice, VAE architectures in general have a non-trivial gap from the approximating posterior which may make these theoretical limits unattainable.), then and .
Our second theorem, Theorem 2, states that the expected approximation error over the truncated prior can also be bounded by components of the CVAE objective. Since the generator parameterizes a multivariate normal with identity covariance, an upper bound on the expected reconstruction error implies a lower bound on the likelihood. This implies that one can also satisfy the sufficient likelihood property by training a CVAE, effectively learning a probabilistic perturbation set that assigns high likelihood to perturbed data.
Let be the Mahalanobis distance which captures of the probability mass for a -dimensional standard multivariate normal for some . Then, the truncated expected approximation error can be bounded with
where is a multivariate normal that has been truncated to radius and is a constant that depends exponentially on and .
The main takeaway from these two theorems is that optimizing the CVAE objective naturally results in a learned perturbation set which satisfies the necessary subset and sufficient likelihood properties. The learned perturbation set is consequently useful for adversarial robustness since the necessary subset property implies that the perturbation set does not “miss” perturbed data. It is also useful for data augmentation since the sufficient likelihood property ensures that perturbed data occurs with high probability. We leave further discussion of these two theorems to Appendix B.3.
Experiments
Finally, we present a variety of experiments to showcase the generality and effectiveness of our perturbation sets learned with a CVAE. Our experiments for each dataset can be broken into two distinct types: the generative modeling problem of learning and evaluating a perturbation set, and the robust optimization problem of learning an adversarially robust classifier to this perturbation. We note that our approach is broadly applicable, has no specific requirements for the encoder, decoder, and prior networks, and avoids the unstable training dynamics found in GANs. Furthermore, we do not have the blurriness typically associated with VAEs since we are modeling perturbations and not the underlying image. All code, configuration files, and pretrained model weights for reproducing our experiments are at https://github.com/locuslab/perturbation_learning.
In all settings, we first train perturbation sets and evaluate them with a number of metrics averaged over the test set. We present a condensed version of these results in Table 1, which establishes a quantitative baseline for learning real-world perturbation sets in three benchmark settings that future work can improve upon. Specifically, the approximation error measures the necessary subset property, the expected approximation error measures the sufficient likelihood property, and the reconstruction error and KL divergence are standard CVAE metrics. The full evaluation is described in Appendix C, and a complete tabulation of our results on evaluating perturbation sets can be found in Tables 4, 7, and 10 in the appendix for each setting.
We qualitatively evaluate our perturbation set in Figure 1, which depicts interpolations, random samples, and adversarial examples from the perturbation set. For additional analysis of the perturbation set, we refer the reader to Appendix F.2 for a study on using different pairing strategies during training and Appendix F.3 which finds semantic latent structure and visualizes additional examples.
We next employ the perturbation set in adversarial training and randomized smoothing to learn models which are robust against worst-case CIFAR10 common corruptions. We report results at three radius thresholds which correspond to the th, th, and th percentiles of latent encodings as described in Appendix F.3. We compare to two data augmentation baselines of training on perturbed data or samples drawn from the learned perturbation set, and also evaluate performance on three extra out-of-distribution corruptions (one for each weather, blur, and digital category denoted OOD) that are not present during training.
2 Multi-illumination
Our last set of experiments looks at learning a perturbation set of multiple lighting conditions using the Multi-Illumination (MI) dataset (Murmann et al., 2019). These consist of a thousand scenes captured in the wild under 25 different lighting variations, and our goal is to learn a perturbation set which captures real-world lighting conditions. Since the process of learning lighting variations is largely similar to the image corruptions setting, we defer most of the discussion on learning and evaluating the CVAE-based lighting perturbation set to Appendix G.1. Note that our perturbation set accurately captures real-world changes in lighting with a low approximation error of 0.006 as seen in Table 1. We qualitatively evaluate our perturbation set by depicting adversarial examples in Figure 2, with more visualizations (including samples and interpolations) in Appendix G.2.
We devote the remainder of this section to studying the task of generating material segmentation maps which are robust to lighting perturbations, using our CVAE perturbation set. We highlight that adversarial training improves robustness to worst-case lighting perturbations over directly training on the perturbed examples, increasing robust accuracy from to at the maximum radius . Additional results on certifiably robust models with randomized smoothing can be found in Appendix G.3.
Conclusion
In this paper, we presented a general framework for learning perturbation sets from data when the perturbation cannot be mathematically-defined. We outlined deterministic and probabilistic properties that measure how well a perturbation set fits perturbed data, and formally proved that a perturbation set based upon the CVAE framework satisfies these properties. This work establishes a principled baseline for learning perturbation sets with quantitative metrics which future work can potentially improve upon, e.g. by using a different generative modeling frameworks. The resulting perturbation sets open up new downstream robustness tasks such as adversarial and certifiable robustness to common image corruptions and lighting perturbations, while also potentially improving non-adversarial robust performance to natural perturbations. Our work opens a pathway for practitioners to learn machine learning models that are robust to targeted, real-world perturbations that can be collected as data.
References
Appendix A Appendix
Appendix B Theoretical results
In this section, we present the theoretical results in their full detail and exposition. Both of the main theorems presented in this work require a number of preceding results in order to formally link the prior and the posterior distribution based on their KL divergence. We will present and prove these supporting results before proving each main theorem.
Theorem 1 connects the CVAE objective to the necessary subset property. In order to prove this, we first prove three supporting lemmas. Lemma 1 states that if the expected value of a function over a normal distribution is low, then for any fixed radius there must exist a point within the radius with proportionally low function value. This leverages the fact that the majority of the probability mass is concentrated around the mean and can be characterized by the Mahalanobis distance.
We will prove this by contradiction. Assume for sake of contradiction that for all such that , we have . We divide the expectation into two integrals over the inside and outside of the Mahalanobis ball:
Using the assumption on the first integral and non-negativity of in the second integrand, we can conclude
Our second lemma, Lemma 2, is an important result that comes from the algebraic form of the KL divergence. It is needed to connect the bound on the KL divergence to the actual variances of the prior and the posterior distribution, and uses the LambertW function to do so.
Let be the LambertW function with branch . Let , and suppose . Then, where and . Additionally, these bounds coincide at when .
The LambertW function is defined as the inverse function of , where the path to multiple solutions is determined by the branch . We can then write the inverse of as one of the following two solutions:
Since is convex with minimum at , and since the two solutions surround , the set of points which satisfy is precisely the interval . Evaluating the bound at completes the proof. ∎
Lemma 3 is the last lemma needed to prove Theorem 1, which explicitly bounds terms involving the mean and variance of the prior and posterior distributions by their KL distance, leveraging Lemma 2 to bound the ratio of the variances. The two quantities bounded in this lemma will be used in the main theorem to bound the distance of a point with low reconstruction error from the prior distribution.
Suppose the KL distance between two normals is bounded, so for some constant . Then,
and also where
Since for , we apply this for to prove the first bound on the squared distance as
Next, since , we can bound the remainder as the following
and using Lemma 2, we can bound where
With these three results, we can now prove the first main theorem, which we is presented below in its complete form, allowing us to formally tie the CVAE objective to the existence of a nearby point with low reconstruction error.
resulting in the following likelihood lower bound:
Suppose we have trained the lower bound to some thresholds
where bounds the KL-divergence of the th dimension. Let be the Mahalanobis distance which captures of the probability mass for a -dimensional standard multivariate normal for some . Then, there exists a such that and for
where is a constant dependent on . Moreover, as and (the theoretical limits of these bounds, e.g. by training), then and .
The high level strategy for this proof will consist of two main steps. First, we will show that there exists a point near the encoding distribution which has low reconstruction error, where we leverage the Mahalanobis distance to capture nearby points with Lemma 1. Then, we apply the triangle inequality to bound its distance under the prior encoding. Finally, we will use Lemma 3 to bound remaining quantities relating the distances between the prior and encoding distributions to complete the proof.
Let be the parameterization trick of the encoding distribution. Rearranging the log probability in the assumption and using the reparameterization trick, we get
We can use Lemma 3 on the KL assumption to bound the following quantities
where are as defined from Lemma 3.
Let , so for all . Plugging this in along with the previous bounds we get the following bound on the norm of before the prior reparameterization:
Thus, the norm (before the prior reparameterization) and reconstruction error of can be bounded by and .
To conclude the proof, we note that from Lemma 2, for all implies , and so . Similarly by inspection, implies that , which concludes the proof. ∎
B.2 Proof of Theorem 2
Let and . Then,
Furthermore if and , then where
The proof here is almost purely algebraic in nature. By definition of and , we have
We will focus on bounding the exponent, . We can bound this by considering two cases. First, suppose that and so the exponent is a concave quadratic. Then, the maximum value of the quadratic is at its root:
Further assume that this is within the interval of the indicator function, so . Then, plugging in the maximum value into the quadratic results in the following bound for this case:
Consider the other case, so either , or the optimal value for in the previous case when is not within the interval . Then, the maximum value of this quadratic must occur at , so . Plugging this into the quadratic for positive , this results in
and so we can bound this case with the following function :
Plugging in the maximum over both cases forms our final bound on the ratio of distributions.
To finish the proof, assume we have the corresponding bounds on and . Then, the first case can be bounded with defined as
The second case can be bounded with defined as
And thus we can bound as
Lemma 4 can be directly applied to the setting of Theorem 2 for the case of a non-conditional VAE with a standard normal distribution for the prior. However, since we are using conditional VAEs instead, the prior distribution has its own mean and variance and so the posterior distribution needs to be unparameterized from the posterior and reparameterized with the prior. Corollary 1 establishes this formally, extending Lemma 4 to the conditional setting.
Let and . Then,
where and . Furthermore, if for some constant , then where is a constant which depends on .
First, note that Lemma 3 implies that and . Observe that and , and so the proof reduces to an application of Lemma 4 to this particular and . ∎
These results allow us to bound the ratio of the prior and posterior distributions within a fixed radius, which will allow us to bound the expectation over the prior with the expectation over the posterior. We can now finally prove Theorem 2, presented in its complete form below.
resulting in the following likelihood lower bound:
Suppose we have trained the lower bound to some thresholds
where bounds the KL-divergence of the th dimension. Let be the Mahalanobis distance which captures of the probability mass for a -dimensional standard multivariate normal for some . Then, the truncated expected reconstruction error can be bounded with
where is a multivariate normal that has been truncated to radius and is a constant that depends exponentially on .
The overall strategy for this proof will be as follows. First, we will rewrite the expectation over a truncated normal into an expectation over a standard normal, using an indicator function to control the radius. Second, we will do an “unparameterization” of the standard normal to match the parameterized objective in the assumption. Finally, we will bound the ratio of the unparameterized density over the normal prior, which allows us to bound the expectation over the prior with the assumption. This last step to bound the ratio of densities is made possible by the truncation, which would otherwise grow exponentially with the tails of the distribution.
For notational simplicity, let . The quantity we wish to bound can be rewritten using the Mahalanobis distance as
where we used the fact that , which simply rewrites the density of a truncated normal using a scaled standard normal density and an indicator function. We can do a parameterization trick to rewrite this as
followed by a reverse parameterization trick with to get the following equivalent expression
where . For convenience, we can let
which can be interpreted as a truncated version of , and so the expectation can be represented more succinctly as
where we also used the fact that is a diagonal normal, so . Each term in this product can be bounded by Lemma 4 to get
where is as defined in Corollary 1 for each using the corresponding , and we let . Thus we can now bound the truncated expected value by plugging in our bound for into Equation (31) to get
Using our lower bound on the expected log likelihood from the assumption, we can bound the remaining integral with
and so combining this into our previous bound from Equation (34) results in the final bound
B.3 Discussion of theoretical results
We first note that the bound on the expected reconstruction error in Theorem 2 is larger than the corresponding bound from Theorem 1 by an exponential factor. This gap is to some extent desired, since Theorem 1 characterizes the existence of a highly accurate approximation whereas Theorem 2 characterizes the average case, and so if this gap were too small, the average case of allowable perturbations would be constrained in reconstruction error, which is not necessarily desired in all settings.
This relates to the overapproximation error, or the maximum reconstruction error within the perturbation set. Initially one may think that having low overapproximation error to be a desirable property of perturbation sets, in order to constrain the perturbation set from deviating by “too much.” However, perturbation sets such as the rotation-translation-skew perturbations for MNIST will always have high overapproximation error, and so reducing this is not necessary desired. Furthermore, it not theoretically guaranteed for the CVAE to minimize the overapproximation error without further assumptions. This is because it is possible for a function to be arbitrarily small in expectation but be arbitrarily high at some nearby point, so optimizing the CVAE objective does not imply low overapproximation error. A simple concrete example demonstrating this is the following piecewise linear function
Appendix C Evaluating a perturbation set
and the probabilistic perturbation set is defined by the truncated normal distribution before the parameterization trick as follows,
The benefits of such a selection is that by taking the maximum, we are selecting a learned perturbation set that includes every point with approximation error as low as the posterior encoder. To some extent, this will also capture additional types of perturbations beyond the perturbed data, which can be both beneficial and unwanted depending on what is captured. However, in the context of adversarial attacks, using a perturbation set which is too large is generally more desirable than one which is too small, in order to not underspecify the threat model.
Evaluation metrics
Encoder approximation error (Enc. AE) We can get a fast upper bound of approximation error by taking the posterior mean, unparameterizing it with respect to the prior, and projecting it to the ball as follows:
PGD approximation error (PGD AE) We can refine the upper bound by solving the following problem with projected gradient descent:
In practice, we implement this by warm starting the procedure with the solution from the encoder approximation error, and run iterations of PGD at step size .
Expected approximation error (EAE) We can compute this by drawing samples and calculating the following Monte Carlo estimate:
In practice, we find that is sufficient for reporting means over the dataset with near-zero standard deviation.
In practice, we implement this by doing a random initialization and run iterations of PGD at step size .
Reconstruction error (Recon. err) This is the typical reconstruction error of a variational autoencoder, which is a Monte Carlo estimate over the full posterior with one sample:
Note that we report the average over all pixels to be consistent with the other metrics in this paper, however it is typical to implement this during training as a sum of squared error instead of a mean.
KL divergence (KL) This is the standard KL divergence between the posterior and the prior distributions
Appendix D Adversarial training, randomized smoothing, and data augmentation with learned perturbation sets
In this section, we describe how various standard, successful techniques in robust training can applied to learned perturbation sets. We note that each method is virtually unchanged, with the only difference being the application of the method to the latent space of the generator instead of directly in the input space.
D.2 Randomized smoothing
Randomized smoothing is a method for learning models robust to adversarial examples which come with certificates that can prove (with high probability) the non-existance of adversarial examples. We follow closely the randomized smoothing procedure from (Cohen et al., 2019), which trains and certifies a network by augmented the inputs with large Gaussian noise. In order to do randomized smoothing on the learned perturbation sets, it suffices to simply train, predict, and certify with augmented noise in the latent space of the generator, as shown in Algorithm 2 for prediction and certification.
The algorithms are almost identical to that propose by Cohen et al. (2019), with the exception that the noise is passed to the generator before going through the classifier. Note that LowerConfidenceBound returns a one-sided confidence interval for the Binomial parameter given a sample , and BinomPValue returns the value of a two-sided hypothesis test that .
Appendix E MNIST
We use the standard MNIST dataset consisting of 60,000 training examples with 1,000 examples randomly set aside for validation purposes, and 10,000 examples in the test set. These experiments were run on a single GeForce RTX 2080 Ti graphics card, with the longest perturbation set taking 1 hour to train.
Both networks are trained for 20 epochs, with step size following a piece-wise linear schedule of over epochs $\beta[0,0.001,0.01]$.
RTS details
For the RTS setting, we follow the formulation studied by Jaderberg et al. (2015), which consists of a random rotation between $[0.7,1.3]42\times 4212842\times 42$ feature map, followed by five parallel spatial transformers which use the latent vector to transform the conditioned input. Multiple spatial transformers here are used simply because training just one can be inconsistent. The outputs of the spatial transformers are concatenated via the channel dimension, followed by a convolutional block and one final convolution to reduce it back to one channel.
All networks are trained for 100 epochs with cyclic learning rate peaking at on the th epoch with the Adam optimizer using batch size 128. The KL divergence is weighted by which follows a piece-wise linear schedule from over epochs $$.
E.2 Samples and visualizations
E.3 Evaluating the perturbation set
The RTS-1 results demonstrate how fixing the number of perturbations seen during training to one per datapoint is still enough to learn a perturbation set with as much approximation error as a perturbation set with an infinite number of samples, denoted RTS. Increasing the number of perturbations to five per datapoint allows the remaining metrics to match the RTS model, and so this suggests that not many samples are needed to learn a perturbation set in this setting.
Appendix F CIFAR10 common corruptions
Blurs: defocus blur, glass blur, motion blur, zoom blur
Digital: brightness, contrast, elastic, pixelate, jpeg
The dataset comes with different corruption levels, so we focus on the highest severity, corruption level five, so the total training set has 600,000 perturbations (12 for each of 50,000 examples) and the test set has 120,000 perturbations. The dataset also comes with several additional corruptions meant to be used as a validation set, however for our purposes will serve as a way to measure performance on out-of-distribution corruptions. These are Gaussian blur, spatter, and saturate which correspond to the blur, weather, and digital categories respectively. We generate a validation set from the training set by randomly setting aside of the CIFAR10 training set and all of their corresponding corrupted variants. These experiments were run on a single GeForce RTX 2080 Ti graphics card, taking 22 hours to train the CVAE and 28 hours to run adversarial training.
We use standard preactivation residual blocks (He et al., 2016), with a convolutional bottleneck which reduces the number of channels rather than downsampling the feature space. This aids in learning to produce CIFAR10 common corruptions since the corruptions are reflective of local rather than global changes, and so there is not so much benefit from compressing the perturbation information into a smaller feature map. Specifically, our residual blocks uses convolutions that go from channels denoted as . Then, our encoder and prior networks are as shown in Table 6, where for the prior network and for the encoder network, and the decoder network is as shown in Table 6.
Note that the CVAE encoders output the log variance of the prior and posterior distributions, which need to be exponentiated in order to calculate the KL distance. This runs the risk of of numerical overflow and exploding gradients, which can suddenly cause normal training to fail. To stabilize the training procedure, it is sufficient to use a scaled Tanh activation function before predicting the log variance, which is a Tanh activation which has been scaled to output a log variance betwen . This has the effect of preventing the variance from being outside the range of , and stops overflow from happening. In practice, the prior and posterior converge to a variance within this range, and so this doesn’t seem to adversely effect the resulting representative power of the CVAE while stabilizing the exponential calculation.
Training is done for 1000 epochs using a cyclic learning rate (Smith, 2017), peaking at on the 400th epoch using the Adam optimizer (Kingma & Ba, 2014) with momentum and batch size . The hyperparameter is also scheduled to increase linearly from to over the first 400 epochs. We use standard CIFAR10 data augmentation with random cropping and flipping.
F.2 Pairing strategies
Here, we quantitatively evaluate the effect of the different pairing strategies for the CIFAR10 common corruptions setting when training a perturbation set with a CVAE in Table 7. We first note that the approach using random pairs of both original and perturbed data has a larger radius than the others, likely since the generator needs to produce not just the perturbed data but also the original as well. We next see that all metrics across the board are substantially lower for the approach which centers the prior by always conditioning on the unperturbed example, and so if such an unperturbed example is known, we recommend conditioning on said example to learn a better perturbation set.
F.3 Properties of the CVAE latent space
Additional samples and interpolations from the CVAE
To further illustrate what corruptions are represented in the CVAE latent space, we plot additional samples and interpolations from the CVAE. In Figure 7 we draw four random samples for eight different examples, which show a range of corruptions and provide qualitative evidence that the perturbation set generates reasonable samples. In Figure 8 we see a number of examples being interpolated between weather, blur, and digital corruptions, which also demonstrate how the perturbation set not only includes the corrupted datapoints, but also variations and interpolations in between corruptions.
F.4 Adversarial training
For learning a robust CIFAR10 classifier, we use a standard wide residual network (Zagoruyko & Komodakis, 2016) with depth 28 and width factor 10. All models are trained with the Adam optimizer with momentum 0.9 for 100 epochs with batch size 128 and cyclic learning rate schedule which peaks at 0.2 at epoch 40. This cyclic learning rate was tuned to optimize the validation performance for the data augmentation baseline, and kept fixed as-is for all other approaches. We use validation-based early stopping for all approaches to combat robust overfitting (Rice et al., 2020). However, we find that in this setting and also likely due to the cyclic learning rate, we do not observe much overfitting. Consequently, the final model at the end of training ends up being selected for all methods regardless. Adversarial training is done with a PGD adversary at the full radius of with step size and 7 iterations. At evaluation time, for a given radius we use a PGD adversary with 50 steps of size . Additional examples of adversarial attacks under this adversary are in Figure 9, which still appear to be reasonable corruptions despite being adversarial.
F.5 Comparisons to other baselines
One may be curious as to how these results compare to other methods for learning models which are robust to different threat models. In this section, we elaborate more on the comparison to other baselines. We preface this discussion with the following disclaimers:
It not necessarily expected for robustness to one threat model to generalize to another threat model.
It is expected that training against a given threat model gives the best results to that threat model, in comparison to training against a different threat model.
Consequently, the fact that these baselines perform universally worse than the CVAE approaches is unsurprising and confirms our expectations. Nonetheless, we provide this comparison to satiate the readers curiosity.
F.6 Randomized smoothing
The results of randomized smoothing at these thresholds are presented in Table 9. Note that by construction, each noise level cannot certify a radius larger than what is calculated for. We find that clean accuracy isn’t significantly harmed when smoothing at the lower noise levels, and actually improves perturbed accuracy over the adversarial training approach from Table 2 but with lower out-of-distribution accuracy than adversarial training. For example, smoothing with a noise level of can achieve 92.5% perturbed accuracy, which is higher than the best empirical approach from Table 2. Furthermore, all noise levels are able to achieve non-trivial degrees of certified robustness, with the highest noise level being the most robust but also taking a hit to standard test set accuracy metrics. Most notably, due to the latent structure present in the perturbation set as discussed in Appendix F.3, even the ability to certify small radii translates to meaningful provable guarantees against certain types of corruptions, like defocus blur and jpeg compression which are largely captured at radius .
Appendix G Multi-illumination
In this final section we present the multi-illumination experiments in greater detail with an expanded discussion. We use the test set provided by Murmann et al. (2019) which consists of 30 held out scenes and hold out 25 additional “drylab” scenes for validation. Unlike in the CIFAR10 common corruptions setting, there is no such thing as an “unperturbed” example in this dataset so we train on random pairs selected from the 25 different lighting variations for each scene. These experiments were run on a single Quadro RTX 8000 graphics card, taking 16 hours to train the CVAE and 12 hours to run adversarial training.
We convert a generic UNet architecture (Ronneberger et al., 2015) to use as a CVAE, by inserting the variational latent space in between the skip connections of similar resolutions. At a high level, the encoder and prior networks will be based on the downsampling half of a UNet architecture, the conditional generator will be based on the full UNet architecture, and the components will be linked via the latent space of the VAE. Specifically, our networks have $1\times 1[(127,187),(62,93),(31,46),(15,23),(7,11)]256$ dimensional latent space vector for the CVAE. Similar to the CIFAR10 setting, we use a scaled Tanh activation to stabilize the log variance calculation.
The generator of the CVAE UNet is implemented as a typical UNet that takes as input the conditioned example, where intermediate feature maps are concatenated with extra feature maps from the latent space of the CVAE. Specifically, each latent space sub-vector is mapped with a fully connected layer back to a feature map with size with one channel. It is then interpolated to the actual feature map size of the UNet (which is a no-op for the MIP5 resolution) and concatenated to the standard UNet feature map before upsampling.
We train the model for 1000 epochs with batch size 64, using the Adam optimizer with momentum, weight decay , and a cyclic learning rate from over $\beta10\beta10$. Random cropping augmentation is proportionately increased with the size of the image, using padding of 20 and 40 for MIP4 and MIP3 respectively.
The results of this fine tuning procedure to learn a perturbation at higher resolutions are summarized in Tables 10 and 11. As expected, the quality metrics of the perturbation set get slightly worse at higher resolutions which is counteracted to some degree by the increase in weighting for the KL divergence at MIP3. Despite using the same architecture size, the perturbation set is able to reasonably scale to images with 16 times more pixels and generate reasonable samples while keeping relatively similar quality metrics. However, to keep computation requirements at a reasonable threshold, we focus our experiments at the MIP5 resolution, which is sufficient for our robustness tasks.
G.2 Additional samples and interpolations from the CVAE
We present additional samples and interpolations of lighting changes learned by the CVAE perturbation set. Figure 11 shows interpolations between three randomly chosen lighting perturbations for four different scenes, while Figure 12 shows 24 additional random samples from the perturbation set, showing a variety of lighting conditions. These demonstrate qualitatively that the perturbation set contains a reasonable set of lighting changes.
G.3 Adversarial training and randomized smoothing
In this final task, we leverage our learned perturbation set to learn a model which is robust to lighting perturbations. Specifically the multi-illumination dataset comes with annotated material maps that label each pixel with the type of material (e.g. fabric, plastic, and paper), so a natural task to perform in this setting is to learn material segmentation maps. We report robustness results at which correspond to the the 50th, 75th, and 100th percentiles. We select a fixed lighting angle which appears to be neutral and train a material segmentation model to generate a baseline, which achieves accuracy on the test set but only 14.9% robust accuracy under the full perturbation model.
We evaluate several methods for improving robustness of the segmentation maps to lighting perturbations, namely data augmentation with the perturbed data, data augmentation with the CVAE, and adversarial training with the CVAE. The results are tabulated in Table 3. The CVAE data augmentation approach is not as effective in this setting at improving robust accuracy, as the pure data augmentation approach does reasonably well. However, the adversarial training approach unsurprisingly has the most robust accuracy, maintaining robust accuracy under the full perturbation set at and outperforming the data augmentation approaches. We plot adversarial perturbations at and the resulting changed segmentation maps in Figure 13 for a model trained with pure data augmentation and in Figure 14 for a model trained to be adversarially robust, the latter of which has, on average, less pixels that are effected by an adversarial lighting perturbation. Adversarial examples at the full radius of are shown in Figure 15, where we see that the perturbation set is beginning to cast dark shadows over regions of the image to force the model to fail.
Finally, we train a certifiably robust material segmentation model using the perturbation set. We train using a noise level of which can certify a radius of at most , or the limit of the perturbation set. The resulting robustness curve is plotted in Figure 16. The model achieves perturbed accuracy and is able to get certified accuracy at the th percentile of radius . The key takeaway is that we can now certify real-world perturbations to some degree, in this case certifying robustness to 50% of lighting perturbations with non-trivial guarantees.