Adversarial Transformation Networks: Learning to Generate Adversarial Examples

Shumeet Baluja, Ian Fischer

Introduction and Background

With the resurgence of deep neural networks for many real-world classification tasks, there is an increased interest in methods to generate training data, as well as to find weaknesses in trained models. An effective strategy to achieve both goals is to create adversarial examples that trained models will misclassify. Adversarial examples are small perturbations of the inputs that are carefully crafted to fool the network into producing incorrect outputs. These small perturbations can be used both offensively, to fool models into giving the “wrong” answer, and defensively, by providing training data at weak points in the model. Seminal work by Szegedy et al. 2013 and Goodfellow et al. 2014b, as well as much recent work, has shown that adversarial examples are abundant, and that there are many ways to discover them.

Given a classifier f(x):x∈X→y∈Yf(\mathbf{x}):\mathbf{x}\in\mathcal{X}\rightarrow y\in\mathcal{Y} and original inputs x∈X\mathbf{x}\in\mathcal{X}, the problem of generating untargeted adversarial examples can be expressed as the optimization: arg min⁡x∗L(x,x∗) s.t. f(x∗)≠f(x)\argmin_{\mathbf{x}^{\mathbf{\ast}}}L(\mathbf{x},\mathbf{x}^{\mathbf{\ast}})\ s.t.\ f(\mathbf{x}^{\mathbf{\ast}})\neq f(\mathbf{x}), where L(⋅)L(\cdot) is a distance metric between examples from the input space (e.g., the L2L_{2} norm). Similarly, generating a targeted adversarial attack on a classifier can be expressed as arg min⁡x∗L(x,x∗) s.t. f(x∗)=yt\argmin_{\mathbf{x}^{\mathbf{\ast}}}L(\mathbf{x},\mathbf{x}^{\mathbf{\ast}})\ s.t.\ f(\mathbf{x}^{\mathbf{\ast}})=y_{t}, where yt∈Yy_{t}\in\mathcal{Y} is some target label chosen by the attacker. Another axis to compare when considering adversarial attacks is whether the adversary has access to the internals of the target model. Attacks without internal access are possible by transferring successful attacks on one model to another model, as in Szegedy et al. 2013; Papernot et al. 2016a, and others. A more challenging class of blackbox attacks involves having no access to any relevant model, and only getting online access to the target model’s output, as explored in Papernot et al. 2016b; Baluja et al. 2015; Tramèr et al. 2016. See Papernot et al. 2015 for a detailed discussion of threat models.

Until now, these optimization problems have been solved using three broad approaches: (1) By directly using optimizers like L-BFGS or Adam (Kingma & Ba 2015), as proposed in Szegedy et al. 2013 and Carlini & Wagner 2016. Such optimizer-based approaches tend to be much slower and more powerful than the other approaches. (2) By approximation with single-step gradient-based techniques like fast gradient sign (Goodfellow et al. 2014b) or fast least likely class (Kurakin et al. 2016a). These approaches are fast, requiring only a single forward and backward pass through the target classifier to compute the perturbation. (3) By approximation with iterative variants of gradient-based techniques (Kurakin et al. 2016a; Moosavi-Dezfooli et al. 2016a; Moosavi-Dezfooli et al. 2016b). These approaches use multiple forward and backward passes through the target network to more carefully move an input towards an adversarial classification.

Adversarial Transformation Networks

In this work, we propose Adversarial Transformation Networks (ATNs). An ATN is a neural network that transforms an input into an adversarial example against a target network or set of networks. ATNs may be untargeted or targeted, and trained in a black-box E.g., using Williams 1992 to generate training gradients for the ATN based on a reward signal computed on the result of sending the generated adversarial examples to the target network. or white-box manner. In this work, we will focus on targeted, white-box ATNs.

Formally, an ATN can be defined as a neural network:

where θ\boldsymbol{\theta} is the parameter vector of gg, ff is the target network which outputs a probability distribution across class labels, and x′∼x\mathbf{x}\mathbf{{}^{\prime}}\sim\mathbf{x}, but arg max⁡f(x)≠arg max⁡f(x′)\argmax f(\mathbf{x})\neq\argmax f(\mathbf{x}\mathbf{{}^{\prime}}).

To find gf,θg_{f,\boldsymbol{\theta}}, we solve the following optimization:

where LXL_{\mathcal{X}} is a loss function in the input space (e.g., L2L_{2} loss or a perceptual similarity loss like Johnson et al. 2016), LYL_{\mathcal{Y}} is a specially-formed loss on the output space of ff (described below) to avoid learning the identity function, and β\beta is a weight to balance the two loss functions. We will omit θ\boldsymbol{\theta} from gfg_{f} when there is no ambiguity.

Inference.

At inference time, gfg_{f} can be run on any input x\mathbf{x} without requiring further access to ff or more gradient computations. This means that after being trained, gfg_{f} can generate adversarial examples against the target network ff even faster than the single-step gradient-based approaches, such as fast gradient sign, so long as ∣∣gf∣∣⪅∣∣f∣∣||g_{f}||\lessapprox||f||.

Loss Functions.

The input-space loss function, LXL_{\mathcal{X}}, would ideally correspond closely to human perception. However, for simplicity, L2L_{2} is sufficient. LYL_{\mathcal{Y}} determines whether or not the ATN is targeted; the target refers to the class for which the adversary will cause the classifier to output the maximum value. In this work, we focus on the more challenging case of creating targeted ATNs, which can be defined similarly to Equation 1:

where tt is the target class, so that arg max⁡f(x′)=t\argmax f(\mathbf{x}\mathbf{{}^{\prime}})=t. This allows us to target the exact class the classifier should mistakenly believe the input is.

In this work, we define LY,t(y′,y)=L2(y′,r(y,t))L_{\mathcal{Y},t}(\mathbf{y}\mathbf{{}^{\prime}},\mathbf{y})=L_{2}(\mathbf{y}\mathbf{{}^{\prime}},r(\mathbf{y},t)), where y=f(x)\mathbf{y}=f(\mathbf{x}), y′=f(gf(x))\mathbf{y}\mathbf{{}^{\prime}}=f(g_{f}(\mathbf{x})), and r(⋅)r(\cdot) is a reranking function that modifies y\mathbf{y} such that yk<yt,∀ k≠ty_{k}<y_{t},\forall~k\neq t.

Note that training labels for the target network are not required at any point in this process. All that is required is the target network’s outputs y\mathbf{y} and y′\mathbf{y}\mathbf{{}^{\prime}}. It is therefore possible to train ATNs in a self-supervised manner, where they use unlabeled data as the input and make arg max⁡ f(gf,t(x))=t\argmax~f(g_{f,t}(\mathbf{x}))=t.

Reranking function.

There are a variety of options for the reranking function. The simplest is to set r(y,t)=\onehot(t)r(\mathbf{y},t)=\onehot(t), but other formulations can make better use of the signal already present in y\mathbf{y} to encourage better reconstructions. In this work, we look at reranking functions that attempt to keep r(y,t)∼yr(\mathbf{y},t)\sim\mathbf{y}. In particular, we use r(⋅)r(\cdot) that maintains the rank order of all but the targeted class in order to minimize distortions when computing x′=gf,t(x)\mathbf{x}\mathbf{{}^{\prime}}=g_{f,t}(\mathbf{x}).

The specific r(⋅)r(\cdot) used in our experiments has the following form:

α>1\alpha>1 is an additional parameter specifying how much larger yty_{t} should be than the current max classification. norm(⋅)norm(\cdot) is a normalization function that rescales its input to be a valid probability distribution.

1 Adversarial Example Generation

There are two approaches to generating adversarial examples with an ATN. The ATN can be trained to generate just the perturbation to x\mathbf{x}, or it can be trained to generate an adversarial autoencoding of x\mathbf{x}.

Perturbation ATN (P-ATN): To just generate a perturbation, it is sufficient to structure the ATN as a variation on the residual block (He et al. 2015): gf(x)=tanh⁡(x+G(x))g_{f}(\mathbf{x})=\tanh(\mathbf{x}+\mathcal{G}(\mathbf{x})), where G(⋅)\mathcal{G}(\cdot) represents the core function of gfg_{f}. With small initial weight vectors, this structure makes it easy for the network to learn to generate small, but effective, perturbations.

Adversarial Autoencoding (AAE): AAE ATNs are similar to standard autoencoders, in that they attempt to accurately reconstruct the original input, subject to regularization, such as weight decay or an added noise signal. For AAE ATNs, the regularizer is LYL_{\mathcal{Y}}. This imposes an additional requirement on the AAE to add some perturbation p\mathbf{p} to x\mathbf{x} such that r(f(x′))=y′r(f(\mathbf{x}\mathbf{{}^{\prime}}))=\mathbf{y}\mathbf{{}^{\prime}}.

For both ATN approaches, in order to enforce that x′\mathbf{x}\mathbf{{}^{\prime}} is a plausible member of X\mathcal{X}, the ATN should only generate values in the valid input range of ff. For images, it suffices to set the activation function of the last layer to be the tanhtanh function; this constrains each output channel to $$.

2 Related Network Architectures

This training objective resembles standard Generative Adversarial Network training (Goodfellow et al. 2014a) in that the goal is to find weaknesses in the classifier. It is interesting to note the similarity to work outside the adversarial training paradigm — the recent use of feed-forward neural networks for artistic style transfer in images (Gatys et al. 2015)(Ulyanov et al. 2016). Gatys et al. 2015 originally proposed a gradient descent procedure based on “back-driving networks” (Linden & Kindermann 1989) to modify the inputs of a fully-trained network to find a set of inputs that maximize a desired set of outputs and hidden unit activations. Unlike standard network training in which the gradients are used to modify the weights of the network, here, the network weights are frozen and the input itself is changed. In subsequent work, Ulyanov et al. 2016 created a method to approximate the results of the gradient descent procedure through the use of an off-line trained neural network. Ulyanov et al. 2016 removed the need for a gradient descent procedure to operate on every source image to which a new artistic style was to be applied, and replaced it with a single forward pass through a separate network. Analagously, we do the same for generating adverarial examples: a separately trained network approximates the usual gradient descent procedure done on the target network to find adversarial examples.

MNIST Experiments

To begin our empirical exploration, we train five networks on the standard MNIST digit classification task (LeCun et al. 1998). The networks are trained and tested on the same data; they vary only in the weight initialization and architecture, as shown in Table 1. Each network has a mix of convolution (Conv) and Fully Connected (FC) layers. The input to the networks is a 28x28 grayscale image and the output is 10 logit units. Classifierp and Classifiera0 use the same architecture, and only differ in the initialization of the weights. We will primarily use Classifierp for the experiments in this section. The other networks will be used later to analyze the generalization capabilities of the adversaries. Table 1 shows that all of the networks perform well on the digit recognition task. It is easy to get better performance than this on MNIST, but for these experiments, it was more important to have a variety of architectures that achieved similar accuracy, than to have state-of-the-art performance.

We attempt to create an Adversarial Autoencoding ATN that can target a specific class given any input image. The ATN is trained against a particular classifier as illustrated in Figure 1. The ATN takes the original input image, x\mathbf{x}, as input, and outputs a new image, x′\mathbf{x}\mathbf{{}^{\prime}}, that the target classifier should erroneously classify as tt. We also add the constraint that the ATN should maintain the ordering of all the other classes as initially output by the classifier. We train ten ATNs against Classifierp – one for each target digit, tt.

An example is provided to make this concrete. If a classifier is given an image, x3\mathbf{x}_{3}, of the digit 3, a successful ordering of the outputs (from largest to smallest) may be as follows: Classifierp(x3)→{}_{p}(\mathbf{x}_{3})\rightarrow. If ATN7 is applied to x3\mathbf{x}_{3}, when the resulting image, x′3\mathbf{x}\mathbf{{}^{\prime}}_{3}, is fed into the same classifier, the following ordering of outputs is desired (note that the 7 has moved to the highest output): Classifierp({}_{p}(ATN7(x3))→{}_{7}(\mathbf{x}_{3}))\rightarrow.

Training for a single ATNt proceeds as follows. The weights of Classifierp are frozen and never change during ATN training. Every training image, x\mathbf{x}, is passed through Classifierp to obtain output y\mathbf{y}. As described in Equation 4, we then compute rα(y,t)r_{\alpha}(\mathbf{y},t) by copying y\mathbf{y} to a new value, y′\mathbf{y}\mathbf{{}^{\prime}}, setting yt′=α∗max⁡(y)y_{t}^{\prime}=\alpha*\max(\mathbf{y}), and then renormalizing y′\mathbf{y}\mathbf{{}^{\prime}} to be a valid probability distribution. This sets the target class, tt, to have the highest value in y′\mathbf{y}\mathbf{{}^{\prime}} while maintaining the relative order of the other original classifications. In the MNIST experiments, we empirically set α=1.5\alpha=1.5.

Given y′\mathbf{y}\mathbf{{}^{\prime}}, we can now train ATNt to generate x′\mathbf{x}\mathbf{{}^{\prime}} by minimizing β∗LX=β∗L2(x,x′)\beta*L_{\mathcal{X}}=\beta*L_{2}(\mathbf{x},\mathbf{x}\mathbf{{}^{\prime}}) and LY=L2(y,y′)L_{\mathcal{Y}}=L_{2}(\mathbf{y},\mathbf{y}\mathbf{{}^{\prime}}) using Equation 2. Though the weights of Classifierp are frozen, error derivatives are still passed through them to train the ATN. We explore several values of β\beta to balance the two loss functions. The results are shown in Table 2.

We tried three ATN architectures for the AAE task, and each was trained with three values of β\beta against all ten targets, tt. The full 3×33\times 3 set of experiments are shown in Table 2. The accuracies shown are the ability of ATNt to transform an input image x\mathbf{x} into x′\mathbf{x}\mathbf{{}^{\prime}} such that Classifierp mistakenly classifies x′\mathbf{x}\mathbf{{}^{\prime}} as tt. Images that were originally classified as tt were not counted in the test as no transformation on them was required. Each measurement in Table 2 is the average of the 10 networks, ATN0-9.

Results.

In Figure 2(top), each row represents the transformation that ATNt makes to digits that were initially correctly classified as 0-9 (columns). For example, in the top row, the digits 1-9 are now all classified as 0. In all cases, their second highest classification is the original correct classification (0-9).

The reconstructions shown in Figure 2(top) have the largest β\beta; smaller β\beta values are shown in the bottom row. The fidelity to the underlying digit diminishes as β\beta is reduced. However, by loosening the constraints to stay similar to the original input, the number of trials in which the transformer network is able to successfully “fool” the classification network increases dramatically, as seen in Table 2. Interestingly, with β=0.010\beta=0.010, in Figure 2(second row), where there should be a ‘0’ that is transformed into a ‘1’, no digit appears. With this high β\beta, no example was found that could be transformed to successfully fool Classifierp. With the two smaller β\beta values, this anomaly does not occur.

In Figure 3, we provide a closer look at examples of x\mathbf{x} and x′\mathbf{x}\mathbf{{}^{\prime}} for ATNc with β=0.005\beta=0.005. A few points should be noted:

The transformations maintain the large, empty regions of the image. Unlike many previous studies in attacking classifiers, the addition of salt-and-pepper type noise did not appear (Nguyen et al. 2014; Moosavi-Dezfooli et al. 2016b).

In the majority of the generated examples, the shape of the digit does not dramatically change. This is the desired behavior: by training the networks to maintain the order beyond the top-output, only minimal changes should be made to the image. The changes that are often introduced are patches where the light strokes have become darker.

Vertical-linear components of the original images are emphasized in several digits; it is especially noticeable in the digits transformed to 1. With other digits (e.g., 8), it is more difficult to find a consistent pattern of what is being (de)emphasized to cause the classification network to be fooled.

A novel aspect of ATNs is that though they cause the target classifier to output an erroneous top-class, they are also trained to ensure that the transformation preserves the existing output ordering of the target-classifier (other than the top-class). For the examples that were successfully transformed, Table 3 gives the average rank-difference of the outputs with the pre-and-post transformed images (excluding the intentional targeted misclassification).

A Deeper Look into ATNs

This section explores three extensions to the basic ATNs: increasing the number of networks the ATNs can attack, using hidden state from the target network, and using ATNs in serial and parallel.

So far, we have examined ATNs in the context of attacking a single classifier. Can ATNs create adversarial examples that generalize to other classifiers? Much research has studied adversarial transfer for traditional adversaries, including the recent work of Moosavi-Dezfooli et al. 2016a; Liu et al. 2016.

To test transfer, we take the adversarial examples from the previously trained ATNs and test them against Classifiera0,a1,a2,a3 (described in Table 1).

The results in Table 4 clearly show that the transformations made by the ATN are not general; they are tied to the network it is trained to attack. Even Classifiera0, which has the same architecture as Classifierp, is not more susceptible to the attacks than those with different architectures. Looking at the second place correctness scores (in the same Table 4), it may, at first, seem counter-intuitive that the conditional probability of a correct second-place classification remains high despite a low first-place classification. The reason for this is that in the few cases in which the ATN was able to successfully change the classifier’s top choice, the second choice (the real classification) remained a close second (i.e., the image was not transformed in a large manner), thereby maintaining the high performance in the conditional second rank measurement.

Training against multiple networks.

Is it possible to create a network that will be able to create a single transform that can attack multiple networks? Will such an ATN generalize better to unseen networks? To test this, we created an ATN that receives training signals from multiple networks, as shown in Figure 4. As with the earlier training, the LXL_{\mathcal{X}} reconstruction error remains.

The new ATN was trained with classification signals from three networks: Classifierp, and Classifiera1,2. The training proceeds in exactly the same manner as described earlier, except the ATN attempts to minimize LYL_{\mathcal{Y}} for all three target networks at the same time. The results are shown in Table 5. First, examine the columns corresponding to the networks that were used in the training (marked with an *). Note that the success rates of attacking these three classifiers are consistently high, comparable with those when the ATN was trained with a single network. Therefore, it is possible to learn a transformation network that modifies images such that perturbation defeats multiple networks.

Next, we turn to the remaining two networks to which the adversary was not given access during training. There is a large increase in success rates over those when the ATN was trained with a single target network (Table 4). However, the results do not match those of the networks used in training. It is possible that training against larger numbers of target networks at the same time could further increase the transferability of the adversarial examples.

Finally, we look at the success rates of image transformations. Do the same images consistenly fool the networks, or are the failure cases of the networks different? As shown in Figure 5, for the 3 networks the ATN was trained to defeat, the majority of transformations attacked all three networks successfully. For the unseen networks, the results were mixed; the majority of transformations successfully attacked only a single network.

2 “Insider” Information

In the experiments thus far, the classifier, CC, was treated as a white box. From this box, two pieces of information were needed to train the ATN. First, the actual outputs of CC were used to create the new target vector. Second, the error derivatives from the new target vector were passed through CC and propagated into the ATN.

In this section, we examine the possibility of “opening” the classifier, and accessing more of its internal state. From CC, the actual hidden unit activations for each example are used as additional inputs to the ATN. Intuitively, because the goal is to maintain as much similarity as possible to the original image and to maintain the same order of the non-top-most classifications as the original image, access to these activations may convey usable signals.

Because of the very large number of hidden units that accompany convolution layers, in practice, we only use the penultimate fully-connected layer from CC. The results of training the ATNs with this extra information are shown in Table 6. Interestingly, the most salient difference does not come from the ability of the ATN to attack the networks in the first-position. Rather, when looking at the conditional-successes of the second-position, the numbers are improved (compare to Table 2). We speculate that this is because the extra hints provided by the classifier’s internal activations (with the unmodified image) could be used to also ensure that the second-place classification, after input modification, was also correctly maintained.

3 Serial and Parallel ATNs

Separate ATNs are created for each digit (0-9). In this section, we examine whether the ATNs can be used in parallel (can the same original image be transformed by each of the ATNs successfully?) and in serial (can the same image be transformed by one ATN then that resulting image be transformed by another, successfully?).

In the first test, we started with 1000 images of digits from the test set. Each was passed through all 10 ATNs (ATNc, β=0.005\beta=0.005); the resulting images were then classified with Classifierp. For each image, we measured how many ATNs were able to successfully transform the image (success is defined for ATNt as causing the classifier to output tt as the top-class). Out of the 1000 trials, 283 were successfully transformed by all 10 of the ATNs. Samples results and a histogram of the results are shown in Figure 6.

A second experiment is constructed in which the 10 ATNs are applied serially, one-after-the-other. In this scenario, first ATN0 is applied to image x\mathbf{x}, yielding x′\mathbf{x}\mathbf{{}^{\prime}}. Then ATN1 is applied to x′\mathbf{x}\mathbf{{}^{\prime}} yielding x′′\mathbf{x}\mathbf{{}^{\prime}}^{\prime} … to ATN9. The goal is to see whether the transformations work on previously transformed images. The results of chaining the ATNs together in this manner are shown in Figure 6(right). The more transformations that are applied, the larger the image degradation. As expected, by the ninth transformation (rightmost column in Figure 6) the majority of images are severely degraded and usually not recognizable. Though we expected the degradation in images, there were two additional, surprising, findings. First, in the parallel application of ATNs (the first experiment described above), out of 1000 images, 283 of them were successfully transformed by 10 of the ATNs. In this experiment, 741 images were successfully transformed by 10 ATNs. The improvement in the number of all-10 successes over applying the ATNs in parallel occurs because each transformation effectively diminishes the underlying original image (to remove the real classification from the top-spot). Meanwhile, only a few new pixels are added by the ATN to cause the misclassification as it is also trained to minimize the reconstruction error. The overarching effect is a fading of the image through chaining ATNs together.

Second, it is interesting to examine what happens to the second-highest classifications that the networks were also trained to preserve. Order preservation did not occur in this test. Had the test worked perfectly, then for an input-image, x\mathbf{x} (e.g., of the digit 8), after ATN0 was applied, the first and second top classifications of x′\mathbf{x}\mathbf{{}^{\prime}} should be 0,8, respectively. Subsequently, after ATN1 is then applied to x′\mathbf{x}\mathbf{{}^{\prime}}, the classifications of x′′\mathbf{x}\mathbf{{}^{\prime}}^{\prime} should be 1,0,8, etc. The reason this does not hold in practice is that though the networks were trained to maintain the high classification (8) of the original digit, x\mathbf{x}, they were not trained to maintain the potentially small perturbations that ATN0 made to x\mathbf{x} to achieve a top-classification of 0. Therefore, when ATN1 is applied, the changes that ATN0 made may not survive the transformation. Nonetheless, if chaining adversaries becomes important, then training the ATNs with images that have been previously modified by other ATNs may be a sufficient method to address the difference in training and testing distributions. This is left for future work.

ImageNet Experiments

We explore the effectiveness of ATNs on the ImageNet dataset (Deng et al. 2009), which consists of 1.2 million natural images categorized into 1 of 1000 classes. The target classifier, ff, used in these experiments is a pre-trained state-of-the-art classifier, Inception ResNet v2 (IR2), that has a top-1 single-crop error rate of 19.9% on the 50,000 image validation set, and a top-5 error rate of 4.9%. It is described fully in Szegedy et al. 2016.

We trained AAE ATNs and P-ATNs as described in Section 2 to attack IR2. Training an ATN against IR2 follows the process described in Section 3.

IR2 takes as input images scaled to 299×299299\times 299 pixels of 3 channels each. To autoencode images of this size for the AAE task, we use three different fully convolutional architectures (Table 7):

IR2-Base-Deconv, a small architecture that uses the first few layers of IR2 and loads the pre-trained parameter values at the start of training the ATN, followed by deconvolutional layers;

IR2-Resize-Conv, a small architecture that avoids checkerboard artifacts common in deconvolutional layers by using bilinear resize layers to downsample and upsample between stride 1 convolutions; and

IR2-Conv-Deconv, a medium architecture that is a tower of convolutions followed by deconvolutions.

For the perturbation approach, we use IR2-Base-Deconv and IR2-Conv-FC, which has many more parameters than the other architectures due to two large fully-connected layers. The use of fully-connected layers cause the network to learn too slowly for the autoencoding approach (AAE ATN), but can be used to learn perturbations quickly (P-ATN).

All five architectures across both tasks are trained with the same hyperparameters. For each architecture and task, we trained four networks, one for each target class: binoculars, soccer ball, volcano, and zebra. In total, we trained 20 different ATNs to attack IR2.

To find a good set of hyperparameters for these networks, we did a series of grid searches through reasonable parameter values for learning rate, α\alpha, and β\beta, using only Volcano as the target class. Those training runs were terminated after 0.025 epochs, which is only 1600 training steps with a batch size of 20. Based on the parameter search, for the results reported here, we set the learning rate to 0.0001, α=1.5\alpha=1.5, and β=0.01\beta=0.01. All runs were trained for 0.1 epochs (6400 steps) on shuffled training set images, using the Adam optimizer and the TensorFlow default settings.

In order to avoid cherrypicking the best results after the networks were trained, we selected four images from the unperturbed validation set to use for the figures in this paper prior to training. Once training finished, we evaluated the ATNs by passing 1000 images from the validation set through the ATN and measuring IR2’s accuracy on those adversarial examples.

2 Results Overview

Table 8 shows the top-1 adversarial accuracy for each of the 20 model/target combinations. The AAE approach is superior to the perturbation approach, both in terms of top-1 adversarial accuracy, and in terms of training success. Nonetheless, the results in Figures 9 and 7 show that using an architecture like IR2-Conv-FC can provide a qualitatively different type of adversary from the AAE approach.The examples generated using the perturbation approach preserve more pixels in the original image, at the expense of a small region of large perturbations.

In contrast to the perturbation approaches, the AAE architectures distribute the differences across wider regions of the image. However, IR2-Base-Deconv and IR2-Conv-Deconv tend to exhibit checkerboard patterns, which is a common problem in image generation with deconvolutions (Odena et al. 2016). The checkerboarding led us to try IR2-Resize-Conv, which avoids the checkerboard pattern, but gives smooth outputs (Figure 9). Interestingly, in all three AAE networks, many of the original high-frequency patterns are replaced with high frequencies that encode the adversarial signal.

The results from IR2-Base-Deconv show that the same network architectures perform substantially differently when trained as P-ATNs and AAE ATNs. Since P-ATNs are only learning to perturb the input, these networks are much better at preserving the original image, but the perturbations end up being focused along the edges or in the corners of the image. The form of the perturbations often manifests itself as “DeepDream”-like images, as in Figure 8. Approximately the same perturbation, in the same place, is used across all input examples. Placing the perturbations in that manner is less likely to disrupt the other top classifications, thereby keeping LYL_{\mathcal{Y}} lower. This is in stark contrast to the AAE ATNs, which creatively modify the input, as seen in Figures 9 and 7.

3 Detailed Discussion

Figure 7 shows that ATNs are capable of generating a wide variety of adversarial perturbations targeting a single network. Previous approaches to generating adversarial examples often produced qualitatively uniform results – they add various amounts of “noise” to the image, generally concentrating the noise at pixels with large gradient magnitude for the particular adversarial loss function. Indeed, Hendrik Metzen et al. 2017 recently showed that it may be possible to train a detector for previous adversarial attacks. From the perspective of an attacker, then, adversarial examples produced by ATNs may provide a new way past defenses in the cat-and-mouse game of security, since this somewhat unpredictable diversity will likely challenge such approaches to defense. Perhaps a much more interesting consequence of this diversity is its potential application for more comprehensive adversarial training, as described below.

Adversarial Training with ATNs.

In Kurakin et al. 2016b, the authors show the current state-of-the-art in using adveraries for improving training. With single step and iterative gradient methods, they find that it is possible to increase a network’s robustness to adversarial examples, while suffering a small loss of accuracy on clean inputs. However, it works only for the adversary the network was trained against. It appears that ATNs could be used in their adversarial training architecture, and could provide substantially more diversity to the trained model than current adversaries. This adversarial diversity might improve model test-set generalization and adversarial robustness.

Because ATNs are quick to train relative to the target network (in the case of IR2, hours instead of weeks), reliably produce diverse adversarial examples, and can be automatically checked for quality (by checking their success rate against the target network and the LXL_{\mathcal{X}} magnitude of the adversarial examples), they could be used as follows: Train a set of ATNs targeting a random subset of the output classes on a checkpoint of the target network. Once the ATNs are trained, replace a fraction of each training batch with corresponding adversarial examples, subject to two constraints: the current classifier incorrectly classifies the adversarial example as the target class, and the LXL_{\mathcal{X}} loss of the adversarial example is below a threshold that indicates it is similar to the original image. If a given ATN stops producing successful adversarial examples, replace it with a newly trained ATN targeting another randomly selected class. In this manner, throughout training, the target network would be exposed to a shifting set of diverse adversaries from ATNs that can be trained in a fully-automated manner. This procedure conceptually resembles GAN training (Goodfellow et al. 2014a) in many ways, but the goal is different: for GANs, the focus is on using an easy-to-train discriminator to learn a hard-to-train generator; for this adversarial training system, the focus is on using easy-to-train generators to learn a hard-to-train multi-class classifier. Note also that we can run the adversarial example generation in this algorithm on unlabeled data, as described in Section 2. Miyato et al. 2016 also describe a method for using unlabeled data in a manner conceptually similar to adversarial training.

DeepDream perturbations.

IR2-Conv-FC exhibits interesting behavior not seen in any of the other architectures. The network builds a perturbation that generally contains spatially coherent, recognizable regions of the target class. For example, in Figure 8, a consistent soccer-ball “ghost” image appears in all of the transformed images. While the methods and goals of these perturbations are quite different from those generated by DeepDream (Mordvintsev et al. 2015), the qualitative results appear similar. IR2-Conv-FC seems to learn to distill the target network’s representation of the target class in a manner that can be drawn across a large fraction of the image. This is likely due to the final fully-connected layer, which has one weight for each pixel and channel, allowing the network to specify a particular output at each pixel. This result hints at a direct relationship between DeepDream-style techniques and adversarial examples that may improve our ability to find and correct weaknesses in our models.

High frequency data.

The AAE ATNs all remove high frequency data from the images when building their reconstructions. This is likely to be due to limitations of the underlying architectures. In particular, all three convolutional architectures have difficulty exactly recreating edges from the input image, due to spatial data loss introduced when downsampling and padding. Consequently, the LXL_{\mathcal{X}} loss penalizes high confidence predictions of edge locations, leading the networks to learn to smooth out boundaries in the reconstruction. This strategy minimizes the overall loss, but it also places a lower bound on the error imposed by pixels in regions with high frequency information.

This lower bound on the loss in some regions provides the network with an interesting strategy when generating an AAE output: it can focus the adversarial perturbations in regions of the input image that have high-frequency noise. This strategy is visible in many of the more interesting images in Figure 7. For example, many of the networks make minimal modification to the sky in the dog image, but add substantial changes around the edges of the dog’s face, exactly where the LXL_{\mathcal{X}} error would be high in a non-adversarial reconstruction.

IR2-Base-Deconv (3.4M parameters) IR2 MaxPool 5a (35x35x192) →\rightarrow Pad (37x37x192) →\rightarrow Deconv (4x4x512, stride=2) →\rightarrow Deconv (3x3x256, stride=2) →\rightarrow Deconv (4x4x128, stride=2) →\rightarrow Pad (299x299x128) →\rightarrow Deconv (4x4x3) →\rightarrow Image (299x299x3) IR2-Resize-Conv (3.8M parameters) Conv (5x5x128) →\rightarrow Bilinear Resize (0.5) →\rightarrow Conv (4x4x256) →\rightarrow Bilinear Resize (0.5) →\rightarrow Conv (3x3x512) →\rightarrow Bilinear Resize (0.5) →\rightarrow Conv (1x1x512) →\rightarrow Bilinear Resize (2) →\rightarrow Conv (3x3x256) →\rightarrow Bilinear Resize (2) →\rightarrow Conv (4x4x128) →\rightarrow Pad (299x299x128) →\rightarrow Conv (3x3x3) →\rightarrow Image (299x299x3) IR2-Conv-Deconv (12.8M parameters) Conv (3x3x256, stride=2) →\rightarrow Conv (3x3x512, stride=2) →\rightarrow Conv (3x3x768, stride=2) →\rightarrow Deconv (4x4x512, stride=2) →\rightarrow Deconv (3x3x256, stride=2) →\rightarrow Deconv (4x4x128, stride=2) →\rightarrow Pad (299x299x128) →\rightarrow Deconv (4x4x3) →\rightarrow Image (299x299x3) IR2-Conv-FC (233.7M parameters) Conv (3x3x512, stride=2) →\rightarrow Conv (3x3x256, stride=2) →\rightarrow Conv (3x3x128, stride=2) →\rightarrow FC (512) →\rightarrow FC (268203) →\rightarrow Image (299x299x3)

Conclusions and Future Work

Current methods for generating adversarial samples involve a gradient descent procedure on individual input examples. We have presented a fundamentally different approach to finding examples by training neural networks to convert inputs into adversarial examples. Our method is efficient to train, fast to execute, and produces remarkably diverse, successful adversarial examples.

Future work should explore the possibility of using ATNs in adversarial training. A successful ATN-based system may pave the way towards models with better generalization and robustness.

Hendrik Metzen et al. 2017 recently showed that it is possible to detect when an input is adversarial, for current types of adversaries. It may be possible to train such detectors on ATN output. If so, using that signal as an additional loss for the ATN may improve the outputs. Similarly, exploring the use of a GAN discriminator during training may improve the realism of the ATN outputs. It would be interesting to explore the impact of ATNs on generative models, rather than just classifiers, similar to work in Kos et al. 2017. Finally, it may also be possible to train ATNs in a black-box manner, similar to recent work in Tramèr et al. 2016; Baluja et al. 2015, or using REINFORCE (Williams 1992) to compute gradients for the ATN using the target network simply as a reward signal.

References