Adversarial examples for generative models
Jernej Kos, Ian Fischer, Dawn Song
Introduction
Adversarial examples have been shown to exist for a variety of deep learning architectures. Adversarial examples are even easier to produce against most other machine learning architectures, as shown in Papernot et al. (2016), but we are focused on deep networks. They are small perturbations of the original inputs, often barely visible to a human observer, but carefully crafted to misguide the network into producing incorrect outputs. Seminal work by Szegedy et al. (2013) and Goodfellow et al. (2014), as well as much recent work, has shown that adversarial examples are abundant and finding them is easy.
Most previous work focuses on the application of adversarial examples to the task of classification, where the deep network assigns classes to input images. The attack adds small adversarial perturbations to the original input image. These perturbations cause the network to change its classification of the input, from the correct class to some other incorrect class (possibly chosen by the attacker). Critically, the perturbed input must still be recognizable to a human observer as belonging to the original input class. Random noise images and “fooling” images (Nguyen et al., 2014) do not belong to this strict definition of an adversarial input, although they do highlight other limitations of current classifiers.
Deep generative models, such as Kingma & Welling (2013), learn to generate a variety of outputs, ranging from handwritten digits to faces (Kulkarni et al., 2015), realistic scenes (Oord et al., 2016), videos (Kalchbrenner et al., 2016), 3D objects (Dosovitskiy et al., 2016), and audio (van den Oord et al., 2016). These models learn an approximation of the input data distribution in different ways, and then sample from this distribution to generate previously unseen but plausible outputs.
To the best of our knowledge, no prior work has explored using adversarial inputs to attack generative models. There are two main requirements for such work: describing a plausible scenario in which an attacker might want to attack a generative model; and designing and demonstrating an attack that succeeds against generative models. We address both of these requirements in this work.
One of the most basic applications of generative models is input reconstruction. Given an input image, the model first encodes it into a lower-dimensional latent representation, and then uses that representation to generate a reconstruction of the original input image. Since the latent representation usually has much fewer dimensions than the original input, it can be used as a form of compression. The latent representation can also be used to remove some types of noise from inputs, even when the network has not been explicitly trained for denoising, due to the lower dimensionality of the latent representation restricting what information the trained network is able to represent. Many generative models also allow manipulation of the generated output by sampling different latent values or modifying individual dimensions of the latent vectors without needing to pass through the encoding step.
These properties of input reconstruction generative networks suggest a variety of different attacks that would be enabled by effective adversaries against generative networks. Any attack that targets the compression bottleneck of the latent representation can exploit natural security vulnerabilities in applications built to use that latent representation. Specifically, if the person doing the encoding step is separated from the person doing the decoding step, the attacker may be able to cause the encoding party to believe they have encoded a particular message for the decoding party, but in reality they have encoded a different message of the attacker’s choosing. We explore this idea in more detail as it applies to the application of compressing images using a VAE or VAE-GAN architecture.
Related work and background
This work focuses on adversaries for variational autoencoders (VAEs, proposed in Kingma & Welling (2013)) and VAE-GANs (VAEs composed with a generative adversarial network, proposed in Larsen et al. (2015)).
Many adversarial attacks on classification models have been described in existing literature (Goodfellow et al., 2014; Szegedy et al., 2013). These attacks can be untargeted, where the adversary’s goal is to cause any misclassification, or the least likely misclassification (Goodfellow et al., 2014; Kurakin et al., 2016); or they can be targeted, where the attacker desires a specific misclassification. Moosavi-Dezfooli et al. (2016) gives a recent example of a strong targeted adversarial attack. Some adversarial attacks allow for a threat model where the adversary does not have access to the target model (Szegedy et al., 2013; Papernot et al., 2016), but commonly it is assumed that the attacker does have that access, in an online or offline setting (Goodfellow et al., 2014; Kurakin et al., 2016). See Papernot et al. (2015) for an overview of different adversarial threat models.
Given a classifier and original inputs , the problem of generating untargeted adversarial examples can be expressed as the following optimization: , where is a chosen distance measure between examples from the input space (e.g., the norm). Similarly, generating a targeted adversarial attack on a classifier can be expressed as , where is some target label chosen by the attacker.
These optimization problems can often be solved with optimizers like L-BFGS or Adam (Kingma & Ba, 2015), as done in Szegedy et al. (2013) and Carlini & Wagner (2016). They can also be approximated with single-step gradient-based techniques like fast gradient sign (Goodfellow et al., 2014), fast gradient (Huang et al., 2015), or fast least likely class (Kurakin et al., 2016); or they can be approximated with iterative variants of those and other gradient-based techniques (Kurakin et al., 2016; Moosavi-Dezfooli et al., 2016).
An interesting variation of this type of attack can be found in Sabour et al. (2015). In that work, they attack the hidden state of the target network directly by taking an input image and a target image and searching for a perturbed variant of that generates similar hidden state at layer of the target network to the hidden state at the same layer generated by . This approach can also be applied directly to attacking the latent vector of a generative model.
2 Background on VAEs and VAE-GANs
Problem definition
We provide a motivating attack scenario for adversaries against generative models, as well as a formal definition of an adversary in the generative setting.
To motivate the attacks presented below, we describe the attack scenario depicted in Figure 1. In this scenario, there are two parties, the sender and the receiver, who wish to share images with each other over a computer network. In order to conserve bandwidth, they share a VAE trained on the input distribution of interest, which will allow them to send only latent vectors .
There are other attacks of this general form, where the sender and the receiver may be separated by distance, as in this example, or by time, in the case of storing compressed images to disk for later retrieval. In the time-separated attack, the sender and the receiver may be the same person or multiple different people. In either case, if they are using the insecure channel of the VAE’s latent space, the messages they share may be under the control of an attacker. For example, an attacker may be able to fool an automatic surveillance system if the system uses this type of compression to store the video signal before it is processed by other systems. In this case, the subsequent analysis of the video signal could be on compromised data showing what the attacker wants to show.
While we do not specifically attack their models, viable compression schemes based on deep neural networks have already been proposed in the literature, showing promising results Toderici et al. (2015; 2016).
2 Defining adversarial examples against generative models
Attack methodology
With the trained classifier, the attacker finds adversarial examples using the methods described in Section 4.4.
Our second approach generates adversarial perturbations using the VAE loss function. The attacker chooses two inputs, (the source) and (the target), and uses one of the standard adversarial methods to perturb into such that its reconstruction matches the reconstruction of , using the methods described in Section 4.4.
3 Latent attack
Our third approach attacks the latent space of the generative model.
This attack is similar to the work of Sabour et al. (2015), in which they use a pair of source image and target image to generate that induces the target network to produce similar activations at some hidden layer as are produced by , while maintaining similarity between and .
is a distance measure between two vectors. We use the norm, under the assumption that the latent space is approximately euclidean.
We also explored a variation on the single latent vector target attack, which we describe in Section A.1 in the Appendix.
4 Methods for solving the adversarial optimization problem
We can use a number of different methods to generate the adversarial examples. We initially evaluated both the fast gradient sign Goodfellow et al. (2014) method and an optimization method. As the latter produces much better results we focus on the optimization method, while we include some FGS results in the Appendix. The attack can be used either in targeted mode (where we want a specific class, , to be reconstructed) or untargeted mode (where we just want an incorrect class to be reconstructed). In this paper, we focus on the targeted mode of the attacks.
The optimization-based approach, explored in Szegedy et al. (2013) and Carlini & Wagner (2016), poses the adversarial generation problem as the following optimization problem:
5 Measuring attack effectiveness
The architecture is the same as shown in Figure 3. We use the generative model to reconstruct the attempted adversarial inputs by computing:
We derive two metrics from classifier predictions after one reconstruction feedback loop. The first metric is , the attack success rate ignoring targeting, i.e., without requiring the output class of the adversarial example to match the target class:
is the total number of reconstructed adversarial examples; is when , the classification of the reconstruction for image , does not equal , the ground truth classification of the original image, and otherwise. The second metric is , the attack success rate including targeting (i.e., requiring the output class of the adversarial example to match the target class), which we define similarly as:
Both metrics are expected to be higher for more successful attacks. Note that . When computing these metrics, we exclude input examples that have the same ground truth class as the target class.
Evaluation
We evaluate the three attacks on MNIST (LeCun et al., 1998), SVHN (Netzer et al., 2011) and CelebA (Liu et al., 2015), using the standard training and validation set splits. The VAE and VAE-GAN architectures are implemented in TensorFlow (Abadi & et al., 2015). We optimized using Adam with learning rate and other parameters set to default values for both the generative model and the classifier. For the VAE, we use two architectures: a simple architecture with a single fully-connected hidden layer with 512 units and ReLU activation function; and a convolutional architecture taken from the original VAE-GAN paper Larsen et al. (2015) (but trained with only the VAE loss). We use the same architecture trained with the additional GAN loss for the VAE-GAN model, as described in that work. For both VAE and VAE-GAN we use a 50-dimensional latent representation on MNIST, a 1024-dimensional latent representation on SVHN and 2048-dimensional latent representation on CelebA.
In this section we only show results where no sampling from latent space has been performed. Instead we use the mean vector as the latent representation . As sampling can have an effect on the resulting reconstructions, we evaluated it separately. We show the results with different number of samples in Figure 22 in the Appendix. On most examples, the visible change is small and in general the attack is still successful.
Both VAE and VAE-GAN by themselves reconstruct the original inputs well as show in Figure 9, although the quality from the VAE-GAN is noticeably better. As a control, we also generate random noise of the same magnitude as used for the adversarial examples (see Figure 13), to show that random noise does not cause the reconstructed noisy images to change in any significant way. Although we ran experiments on both VAEs and VAE-GANs, we only show results for the VAE-GAN as it generates much higher quality reconstructions than the corresponding VAE.
We use a simple classifier architecture to help generate attacks on the VAE and VAE-GAN models. The classifier consists of two fully-connected hidden layers with 512 units each, using the ReLU activation function. The output layer is a 10 dimensional softmax. The input to the classifier is the 50 dimensional latent representation produced by the VAE/VAE-GAN encoder. The classifier achieves accuracy on the validation set after training for 100 epochs.
To see if there are differences between classes, we generate targeted adversarial examples for each MNIST class and present the results per-class. For the targeted attacks we used the optimization method with lambda , where Adam-based optimization was performed for epochs with a learning rate of . The mean norm of the difference between original images and generated adversarial examples using the classifier attack is , while the mean RMSD is .
Numerical results in Table 2 show that the targeted classifier attack successfully fools the classifier. Classifier accuracy is reduced to , while the matching rate (the ratio between the number of predictions matching the target class and the number of incorrectly classified images) is , which means that all incorrect predictions match the target class. However, what we are interested in (as per the attack definition from Section 3.2) is how the generative model reconstructs the adversarial examples. If we look at the images generated by the VAE-GAN for class , shown in Figure 4, the targeted attack is successful on some reconstructed images (e.g. one, four, five, six and nine are reconstructed as zeroes). But even when the classifier accuracy is and matching rate is , an incorrect classification does not always result in a reconstruction to the target class, which shows that the classifier is fooled by an adversarial example more easily than the generative model.
The reconstruction feedback loop described in Section 4.5 can be used to measure how well a targeted attack succeeds in making the generative model change the reconstructed classes. Table 4 in the Appendix shows and for all source and target class pairs. A higher value signifies a more successful attack for that pair of classes. It is interesting to observe that attacking some source/target pairs is much easier than others (e.g. pair vs. ) and that the results are not symmetric over source/target pairs. Also, some pairs do well in , but do poorly in (e.g., all source digits when targeting ). As can be seen in Figure 11, the classifier adversarial examples targeting consistently fail to reconstruct into something easily recognizable as a . Most of the reconstructions look like , but the adversarial example reconstructions of source s instead look like or .
1.3 Latent attack
To generate adversarial examples using the latent attack, we used the optimization method with , where Adam-based optimization was performed for epochs with a learning rate of . The mean norm of the difference between original images and generated adversarial examples using this approach is , while the mean RMSD is .
Table 3 shows and for all source and target class pairs. Comparing with the numerical evaluation results of the classifier attack we can see that the latent attack performs much better. This result remains true when visually comparing the reconstructed images, shown in Figure 5.
We also tried an untargeted version of the latent attack, where we change Equation 2 to maximize the distance in latent space between the encoding of the original image and the encoding of the adversarial example. In this case the loss we are trying to minimize is unbounded, since the distance can always grow larger, so the attack normally fails to generate a reasonable adversarial example.
2 SVHN
The SVHN dataset consists of cropped street number images and is much less clean than MNIST. Due to the way the images have been processed, each image may contain more than one digit; the target digit is roughly in the center. VAE-GAN produces high-quality reconstructions of the original images as shown in Figure 17 in the Appendix.
3 CelebA
The CelebA dataset consists of more than 200,000 cropped faces of celebrities, each annotated with 40 different attributes. For our experiments, we further scale the images to 64x64 and ignore the attribute annotations. VAE-GAN reconstructions of original images after training are shown in Figure 19 in the Appendix.
4 Summary of different attack methods
Table 1 shows a comparison of the mean distances between original images and generated adversarial examples for the three different attack methods. The larger the distance between the original image and the adversarial perturbation, the more noticeable the perturbation will tend to be, and the more likely a human observer will no longer recognize the original input, so effective attacks keep these distances small while still achieving their goal. The latent attack consistently gives the best results in our experiments, and the classifier attack performs the worst.
Conclusion
We explored generating adversarial examples against generative models such as VAEs and VAE-GANs. These models are also vulnerable to adversaries that convince them to turn inputs into surprisingly different outputs. We have also motivated why an attacker might want to attack generative models. Our work adds further support to the hypothesis that adversarial examples are a general phenomenon for current neural network architectures, given our successful application of adversarial attacks to popular generative models. In this work, we are helping to lay the foundations for understanding how to build more robust networks. Future work will explore defense and robustification in greater depth as well as attacks on generative models trained using natural image datasets such as CIFAR-10 and ImageNet.
This material is in part based upon work supported by the National Science Foundation under Grant No. TWC-1409915. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation.
References
Appendix A Appendix
A variant of the single latent vector targeted attack described in Section 4.3, that was not explored in previous work to our knowledge is to take the mean latent vector of many target images and use that vector as . This variant is more flexible, in that the attacker can choose different latent properties to target without needing to find the ideal input. For example, in MNIST, the attacker may wish to have a particular line thickness or slant in the reconstructed digit, but may not have such an image available. In that case, by choosing some images of the target class with thinner lines or less slant, and some with thicker lines or more slant, the attacker can find a target latent vector that closely matches the desired properties.
In this work, we choose to reconstruct “ideal” MNIST digits by taking the mean latent vector of all of the training digits of each class, and using those vectors as . Given a target class , a set of examples and their corresponding ground truth labels , we create a subset as follows:
Both variants of this attack appear to be similarly effective, as shown in Figure 15 and Figure 5. The trade-off between the two in these experiments is between the simplicity of the first attack and the flexibility of the second attack.