EAD: Elastic-Net Attacks to Deep Neural Networks via Adversarial Examples

Pin-Yu Chen, Yash Sharma, Huan Zhang, Jinfeng Yi, Cho-Jui Hsieh

Introduction

Deep neural networks (DNNs) achieve state-of-the-art performance in various tasks in machine learning and artificial intelligence, such as image classification, speech recognition, machine translation and game-playing. Despite their effectiveness, recent studies have illustrated the vulnerability of DNNs to adversarial examples (?; ?). For instance, a carefully designed perturbation to an image can lead a well-trained DNN to misclassify. Even worse, effective adversarial examples can also be made virtually indistinguishable to human perception. For example, Figure 1 shows three adversarial examples of an ostrich image crafted by our algorithm, which are classified as “safe”, “shoe shop” and “vacuum” by the Inception-v3 model (?), a state-of-the-art image classification model.

The lack of robustness exhibited by DNNs to adversarial examples has raised serious concerns for security-critical applications, including traffic sign identification and malware detection, among others. Moreover, moving beyond the digital space, researchers have shown that these adversarial examples are still effective in the physical world at fooling DNNs (?; ?). Due to the robustness and security implications, the means of crafting adversarial examples are called attacks to DNNs. In particular, targeted attacks aim to craft adversarial examples that are misclassified as specific target classes, and untargeted attacks aim to craft adversarial examples that are not classified as the original class. Transfer attacks aim to craft adversarial examples that are transferable from one DNN model to another. In addition to evaluating the robustness of DNNs, adversarial examples can be used to train a robust model that is resilient to adversarial perturbations, known as adversarial training (?). They have also been used in interpreting DNNs (?; ?).

Throughout this paper, we use adversarial examples to attack image classifiers based on deep convolutional neural networks. The rationale behind crafting effective adversarial examples lies in manipulating the prediction results while ensuring similarity to the original image. Specifically, in the literature the similarity between original and adversarial examples has been measured by different distortion metrics. One commonly used distortion metric is the LqL_{q} norm, where ∥x∥q=(∑i=1p∣xi∣q)1/q\|\mathbf{x}\|_{q}=(\sum_{i=1}^{p}|\mathbf{x}_{i}|^{q})^{1/q} denotes the LqL_{q} norm of a pp-dimensional vector x=[x1,…,xp]\mathbf{x}=[\mathbf{x}_{1},\ldots,\mathbf{x}_{p}] for any q≥1q\geq 1. In particular, when crafting adversarial examples, the L∞L_{\infty} distortion metric is used to evaluate the maximum variation in pixel value changes (?), while the L2L_{2} distortion metric is used to improve the visual quality (?). However, despite the fact that the L1L_{1} norm is widely used in problems related to image denoising and restoration (?), as well as sparse recovery (?), L1L_{1}-based adversarial examples have not been rigorously explored. In the context of adversarial examples, L1L_{1} distortion accounts for the total variation in the perturbation and serves as a popular convex surrogate function of the L0L_{0} metric, which measures the number of modified pixels (i.e., sparsity) by the perturbation. To bridge this gap, we propose an attack algorithm based on elastic-net regularization, which we call elastic-net attacks to DNNs (EAD). Elastic-net regularization is a linear mixture of L1L_{1} and L2L_{2} penalty functions, and it has been a standard tool for high-dimensional feature selection problems (?). In the context of attacking DNNs, EAD opens up new research directions since it generalizes the state-of-the-art attack proposed in (?) based on L2L_{2} distortion, and is able to craft L1L_{1}-oriented adversarial examples that are more effective and fundamentally different from existing attack methods.

To explore the utility of L1L_{1}-based adversarial examples crafted by EAD, we conduct extensive experiments on MNIST, CIFAR10 and ImageNet in different attack scenarios. Compared to the state-of-the-art L2L_{2} and L∞L_{\infty} attacks (?; ?), EAD can attain similar attack success rate when breaking undefended and defensively distilled DNNs (?). More importantly, we find that L1L_{1} attacks attain superior performance over L2L_{2} and L∞L_{\infty} attacks in transfer attacks and complement adversarial training. For the most difficult dataset (MNIST), EAD results in improved attack transferability from an undefended DNN to a defensively distilled DNN, achieving nearly 99% attack success rate. In addition, joint adversarial training with L1L_{1} and L2L_{2} based examples can further enhance the resilience of DNNs to adversarial perturbations. These results suggest that EAD yields a distinct, yet more effective, set of adversarial examples. Moreover, evaluating attacks based on L1L_{1} distortion provides novel insights on adversarial machine learning and security implications of DNNs, suggesting that L1L_{1} may complement L2L_{2} and L∞L_{\infty} based examples toward furthering a thorough adversarial machine learning framework.

Related Work

Here we summarize related works on attacking and defending DNNs against adversarial examples.

FGM and I-FGM: Let x0\mathbf{x}_{0} and x\mathbf{x} denote the original and adversarial examples, respectively, and let tt denote the target class to attack. Fast gradient methods (FGM) use the gradient ∇J\nabla J of the training loss JJ with respect to x0\mathbf{x}_{0} for crafting adversarial examples (?). For L∞L_{\infty} attacks, x\mathbf{x} is crafted by

where ϵ\epsilon specifies the L∞L_{\infty} distortion between x\mathbf{x} and x0\mathbf{x}_{0}, and sign(∇J)\textnormal{sign}(\nabla J) takes the sign of the gradient. For L1L_{1} and L2L_{2} attacks, x\mathbf{x} is crafted by

for q=1,2q=1,2, where ϵ\epsilon specifies the corresponding distortion. Iterative fast gradient methods (I-FGM) were proposed in (?), which iteratively use FGM with a finer distortion, followed by an ϵ\epsilon-ball clipping. Untargeted attacks using FGM and I-FGM can be implemented in a similar fashion.

C&W attack: Instead of leveraging the training loss, Carlini and Wagner designed an L2L_{2}-regularized loss function based on the logit layer representation in DNNs for crafting adversarial examples (?). Its formulation turns out to be a special case of our EAD formulation, which will be discussed in the following section. The C&W attack is considered to be one of the strongest attacks to DNNs, as it can successfully break undefended and defensively distilled DNNs and can attain remarkable attack transferability.

JSMA: Papernot et al. proposed a Jacobian-based saliency map algorithm (JSMA) for characterizing the input-output relation of DNNs (?). It can be viewed as a greedy attack algorithm that iteratively modifies the most influential pixel for crafting adversarial examples.

DeepFool: DeepFool is an untargeted L2L_{2} attack algorithm (?) based on the theory of projection to the closest separating hyperplane in classification. It is also used to craft a universal perturbation to mislead DNNs trained on natural images (?).

Black-box attacks: Crafting adversarial examples in the black-box case is plausible if one allows querying of the target DNN. In (?), JSMA is used to train a substitute model for transfer attacks. In (?), an effective black-box C&W attack is made possible using zeroth order optimization (ZOO). In the more stringent attack scenario where querying is prohibited, ensemble methods can be used for transfer attacks (?).

Defenses in DNNs

Defensive distillation: Defensive distillation (?) defends against adversarial perturbations by using the distillation technique in (?) to retrain the same network with class probabilities predicted by the original network. It also introduces the temperature parameter TT in the softmax layer to enhance the robustness to adversarial perturbations.

Adversarial training: Adversarial training can be implemented in a few different ways. A standard approach is augmenting the original training dataset with the label-corrected adversarial examples to retrain the network. Modifying the training loss or the network architecture to increase the robustness of DNNs to adversarial examples has been proposed in (?; ?; ?; ?).

Detection methods: Detection methods utilize statistical tests to differentiate adversarial from benign examples (?; ?; ?; ?). However, 10 different detection methods were unable to detect the C&W attack (?).

EAD: Elastic-Net Attacks to DNNs

Elastic-net regularization is a widely used technique in solving high-dimensional feature selection problems (?). It can be viewed as a regularizer that linearly combines L1L_{1} and L2L_{2} penalty functions. In general, elastic-net regularization is used in the following minimization problem:

EAD Formulation and Generalization

Inspired by the C&W attack (?), we adopt the same loss function ff for crafting adversarial examples. Specifically, given an image x0\mathbf{x}_{0} and its correct label denoted by t0t_{0}, let x\mathbf{x} denote the adversarial example of x0\mathbf{x}_{0} with a target class t≠t0t\neq t_{0}. The loss function f(x)f(\mathbf{x}) for targeted attacks is defined as

It is worth noting that the term [Logit(x)]t[\textbf{Logit}(\mathbf{x})]_{t} is proportional to the probability of predicting x\mathbf{x} as label tt, since by the softmax classification rule,

Consequently, the loss function in (4) aims to render the label tt the most probable class for x\mathbf{x}, and the parameter κ\kappa controls the separation between tt and the next most likely prediction among all classes other than tt. For untargeted attacks, the loss function in (4) can be modified as

In this paper, we focus on targeted attacks since they are more challenging than untargeted attacks. Our EAD algorithm (Algorithm 1) can directly be applied to untargeted attacks by replacing f(x,t)f(\mathbf{x},t) in (4) with f(x)f(\mathbf{x}) in (6).

In addition to manipulating the prediction via the loss function in (4), introducing elastic-net regularization further encourages similarity to the original image when crafting adversarial examples. Our formulation of elastic-net attacks to DNNs (EAD) for crafting an adversarial example (x,t)(\mathbf{x},t) with respect to a labeled natural image (x0,t0)(\mathbf{x}_{0},t_{0}) is as follows:

where f(x,t)f(\mathbf{x},t) is as defined in (4), c,β≥0c,\beta\geq 0 are the regularization parameters of the loss function ff and the L1L_{1} penalty, respectively. The box constraint x∈p\mathbf{x}\in^{p} restricts x\mathbf{x} to a properly scaled image space, which can be easily satisfied by dividing each pixel value by the maximum attainable value (e.g., 255). Upon defining the perturbation of x\mathbf{x} relative to x0\mathbf{x}_{0} as δ=x−x0\boldsymbol{\delta}=\mathbf{x}-\mathbf{x}_{0}, the EAD formulation in (EAD Formulation and Generalization) aims to find an adversarial example x\mathbf{x} that will be classified as the target class tt while minimizing the distortion in δ\boldsymbol{\delta} in terms of the elastic-net loss β∥δ∥1+∥δ∥22\beta\|\boldsymbol{\delta}\|_{1}+\|\boldsymbol{\delta}\|_{2}^{2}, which is a linear combination of L1L_{1} and L2L_{2} distortion metrics between x\mathbf{x} and x0\mathbf{x}_{0}. Notably, the formulation of the C&W attack (?) becomes a special case of the EAD formulation in (EAD Formulation and Generalization) when β=0\beta=0, which disregards the L1L_{1} penalty on δ\boldsymbol{\delta}. However, the L1L_{1} penalty is an intuitive regularizer for crafting adversarial examples, as ∥δ∥1=∑i=1p∣δi∣\|\boldsymbol{\delta}\|_{1}=\sum_{i=1}^{p}|\boldsymbol{\delta}_{i}| represents the total variation of the perturbation, and is also a widely used surrogate function for promoting sparsity in the perturbation. As will be evident in the performance evaluation section, including the L1L_{1} penalty for the perturbation indeed yields a distinct set of adversarial examples, and it leads to improved attack transferability and complements adversarial learning.

EAD Algorithm

When solving the EAD formulation in (EAD Formulation and Generalization) without the L1L_{1} penalty (i.e., β=0\beta=0), Carlini and Wagner used a change-of-variable (COV) approach via the tanh⁡\tanh transformation on x\mathbf{x} in order to remove the box constraint x∈p\mathbf{x}\in^{p} (?). When β>0\beta>0, we find that the same COV approach is not effective in solving (EAD Formulation and Generalization), since the corresponding adversarial examples are insensitive to the changes in β\beta (see the performance evaluation section for details). Since the L1L_{1} penalty is a non-differentiable, yet piece-wise linear, function, the failure of the COV approach in solving (EAD Formulation and Generalization) can be explained by its inefficiency in subgradient-based optimization problems (?).

To efficiently solve the EAD formulation in (EAD Formulation and Generalization) for crafting adversarial examples, we propose to use the iterative shrinkage-thresholding algorithm (ISTA) (?). ISTA can be viewed as a regular first-order optimization algorithm with an additional shrinkage-thresholding step on each iteration. In particular, let g(x)=c⋅f(x)+∥x−x0∥22g(\mathbf{x})=c\cdot f(\mathbf{x})+\|\mathbf{x}-\mathbf{x}_{0}\|_{2}^{2} and let ∇g(x)\nabla g(\mathbf{x}) be the numerical gradient of g(x)g(\mathbf{x}) computed by the DNN. At the k+1k+1-th iteration, the adversarial example x(k+1)\mathbf{x}^{(k+1)} of x0\mathbf{x}_{0} is computed by

for any i∈{1,…,p}i\in\{1,\ldots,p\}. If ∣zi−x0i∣>β|\mathbf{z}_{i}-{\mathbf{x}_{0}}_{i}|>\beta, it shrinks the element zi\mathbf{z}_{i} by β\beta and projects the resulting element to the feasible box constraint between 0 and 1. On the other hand, if ∣zi−x0i∣≤β|\mathbf{z}_{i}-{\mathbf{x}_{0}}_{i}|\leq\beta, it thresholds zi\mathbf{z}_{i} by setting [Sβ(z)]i=x0i[S_{\beta}(\mathbf{z})]_{i}={\mathbf{x}_{0}}_{i}. The proof of optimality of using (8) for solving the EAD formulation in (EAD Formulation and Generalization) is given in the supplementary materialhttps://arxiv.org/abs/1709.04114. Notably, since g(x)g(\mathbf{x}) is the attack objective function of the C&W method (?), the ISTA operation in (8) can be viewed as a robust version of the C&W method that shrinks a pixel value of the adversarial example if the deviation to the original image is greater than β\beta, and keeps a pixel value unchanged if the deviation is less than β\beta.

Our EAD algorithm for crafting adversarial examples is summarized in Algorithm 1. For computational efficiency, a fast ISTA (FISTA) for EAD is implemented, which yields the optimal convergence rate for first-order optimization methods (?). The slack vector y(k)\mathbf{y}^{(k)} in Algorithm 1 incorporates the momentum in x(k)\mathbf{x}^{(k)} for acceleration. In the experiments, we set the initial learning rate α0=0.01\alpha_{0}=0.01 with a square-root decay factor in kk. During the EAD iterations, the iterate x(k)\mathbf{x}^{(k)} is considered as a successful adversarial example of x0\mathbf{x}_{0} if the model predicts its most likely class to be the target class tt. The final adversarial example x\mathbf{x} is selected from all successful examples based on distortion metrics. In this paper we consider two decision rules for selecting x\mathbf{x}: the least elastic-net (EN) and L1L_{1} distortions relative to x0\mathbf{x}_{0}. The influence of β\beta, κ\kappa and the decision rules on EAD will be investigated in the following section.

Performance Evaluation

In this section, we compare the proposed EAD with the state-of-the-art attacks to DNNs on three image classification datasets - MNIST, CIFAR10 and ImageNet. We would like to show that (i) EAD can attain attack performance similar to the C&W attack in breaking undefended and defensively distilled DNNs, since the C&W attack is a special case of EAD when β=0\beta=0; (ii) Comparing to existing L1L_{1}-based FGM and I-FGM methods, the adversarial examples using EAD can lead to significantly lower L1L_{1} distortion and better attack success rate; (iii) The L1L_{1}-based adversarial examples crafted by EAD can achieve improved attack transferability and complement adversarial training.

We compare EAD with the following targeted attacks, which are the most effective methods for crafting adversarial examples in different distortion metrics.

C&W attack: The state-of-the-art L2L_{2} targeted attack proposed by Carlini and Wagner (?), which is a special case of EAD when β=0\beta=0.

FGM: The fast gradient method proposed in (?). The FGM attacks using different distortion metrics are denoted by FGM-L1L_{1}, FGM-L2L_{2} and FGM-L∞L_{\infty}.

I-FGM: The iterative fast gradient method proposed in (?). The I-FGM attacks using different distortion metrics are denoted by I-FGM-L1L_{1}, I-FGM-L2L_{2} and I-FGM-L∞L_{\infty}.

Experiment Setup and Parameter Setting

Our experiment setup is based on Carlini and Wagner’s frameworkhttps://github.com/carlini/nn˙robust˙attacks. For both the EAD and C&W attacks, we use the default setting1, which implements 9 binary search steps on the regularization parameter cc (starting from 0.001) and runs I=1000I=1000 iterations for each step with the initial learning rate α0=0.01\alpha_{0}=0.01. For finding successful adversarial examples, we use the reference optimizer1 (ADAM) for the C&W attack and implement the projected FISTA (Algorithm 1) with the square-root decaying learning rate for EAD. Similar to the C&W attack, the final adversarial example of EAD is selected by the least distorted example among all the successful examples. The sensitivity analysis of the L1L_{1} parameter β\beta and the effect of the decision rule on EAD will be investigated in the forthcoming paragraph. Unless specified, we set the attack transferability parameter κ=0\kappa=0 for both attacks.

We implemented FGM and I-FGM using the CleverHans packagehttps://github.com/tensorflow/cleverhans. The best distortion parameter ϵ\epsilon is determined by a fine-grained grid search - for each image, the smallest ϵ\epsilon in the grid leading to a successful attack is reported. For I-FGM, we perform 10 FGM iterations (the default value) with ϵ\epsilon-ball clipping. The distortion parameter ϵ′\epsilon^{\prime} in each FGM iteration is set to be ϵ/10\epsilon/10, which has been shown to be an effective attack setting in (?). The range of the grid and the resolution of these two methods are specified in the supplementary material1.

The image classifiers for MNIST and CIFAR10 are trained based on the DNN models provided by Carlini and Wagner1. The image classifier for ImageNet is the Inception-v3 model (?). For MNIST and CIFAR10, 1000 correctly classified images are randomly selected from the test sets to attack an incorrect class label. For ImageNet, 100 correctly classified images and 9 incorrect classes are randomly selected to attack. All experiments are conducted on a machine with an Intel E5-2690 v3 CPU, 40 GB RAM and a single NVIDIA K80 GPU. Our EAD code is publicly available for downloadhttps://github.com/ysharma1126/EAD-Attack.

Evaluation Metrics

Following the attack evaluation criterion in (?), we report the attack success rate and distortion of the adversarial examples from each method. The attack success rate (ASR) is defined as the percentage of adversarial examples that are classified as the target class (which is different from the original class). The average L1L_{1}, L2L_{2} and L∞L_{\infty} distortion metrics of successful adversarial examples are also reported. In particular, the ASR and distortion of the following attack settings are considered:

Best case: The least difficult attack among targeted attacks to all incorrect class labels in terms of distortion.

Average case: The targeted attack to a randomly selected incorrect class label.

Worst case: The most difficult attack among targeted attacks to all incorrect class labels in terms of distortion.

Sensitivity Analysis and Decision Rule for EAD

In Table 1, during the attack optimization process the final adversarial example is selected based on the elastic-net loss of all successful adversarial examples in {x(k)}k=1I\{\mathbf{x}^{(k)}\}_{k=1}^{I}, which we call the elastic-net (EN) decision rule. Alternatively, we can select the final adversarial example with the least L1L_{1} distortion, which we call the L1L_{1} decision rule. Figure 2 compares the ASR and average-case distortion of these two decision rules with different β\beta on MNIST. Both decision rules yield 100% ASR for a wide range of β\beta values. For the same β\beta, the L1L_{1} rule gives adversarial examples with less L1L_{1} distortion than those given by the EN rule at the price of larger L2L_{2} and L∞L_{\infty} distortions. Similar trends are observed on CIFAR10 (see supplementary material1). The complete results of these two rules on MNIST and CIFAR10 are given in the supplementary material1. In the following experiments, we will report the results of EAD with these two decision rules and set β=10−3\beta=10^{-3}, since on MNIST and CIFAR10 this β\beta value significantly reduces the L1L_{1} distortion while having comparable L2L_{2} and L∞L_{\infty} distortions to the case of β=0\beta=0 (i.e., without L1L_{1} regularization).

Attack Success Rate and Distortion on MNIST, CIFAR10 and ImageNet

We compare EAD with the comparative methods in terms of attack success rate and different distortion metrics on attacking the considered DNNs trained on MNIST, CIFAR10 and ImageNet. Table 1 summarizes their average-case performance. It is observed that FGM methods fail to yield successful adversarial examples (i.e., low ASR), and the corresponding distortion metrics are significantly larger than other methods. On the other hand, the C&W attack, I-FGM and EAD all lead to 100% attack success rate. Furthermore, EAD, the C&W method, and I-FGM-L∞L_{\infty} attain the least L1L_{1}, L2L_{2}, and L∞L_{\infty} distorted adversarial examples, respectively. We note that EAD significantly outperforms the existing L1L_{1}-based method (I-FGM-L1L_{1}). Compared to I-FGM-L1L_{1}, EAD with the EN decision rule reduces the L1L_{1} distortion by roughly 47% on MNIST, 53% on CIFAR10 and 87% on ImageNet. We also observe that EAD with the L1L_{1} decision rule can further reduce the L1L_{1} distortion but at the price of noticeable increase in the L2L_{2} and L∞L_{\infty} distortion metrics.

Notably, despite having large L2L_{2} and L∞L_{\infty} distortion metrics, the adversarial examples crafted by EAD with the L1L_{1} rule can still attain 100% ASRs in all datasets, which implies the L2L_{2} and L∞L_{\infty} distortion metrics are insufficient for evaluating the robustness of neural networks. Moreover, the attack results in Table 1 suggest that EAD can yield a set of distinct adversarial examples that are fundamentally different from L2L_{2} or L∞L_{\infty} based examples. Similar to the C&W method and I-FGM, the adversarial examples from EAD are also visually indistinguishable (see supplementary material1).

Breaking Defensive Distillation

In addition to breaking undefended DNNs via adversarial examples, here we show that EAD can also break defensively distilled DNNs. Defensive distillation (?) is a standard defense technique that retrains the network with class label probabilities predicted by the original network, soft labels, and introduces the temperature parameter TT in the softmax layer to enhance its robustness to adversarial perturbations. Similar to the state-of-the-art attack (the C&W method), Figure 3 shows that EAD can attain 100% attack success rate for different values of TT on MNIST and CIFAR10. Moreover, since the C&W attack formulation is a special case of the EAD formulation in (EAD Formulation and Generalization) when β=0\beta=0, successfully breaking defensive distillation using EAD suggests new ways of crafting effective adversarial examples by varying the L1L_{1} regularization parameter β\beta. The complete attack results are given in the supplementary material1.

Improved Attack Transferability

It has been shown in (?) that the C&W attack can be made highly transferable from an undefended network to a defensively distilled network by tuning the confidence parameter κ\kappa in (4). Following (?), we adopt the same experiment setting for attack transferability on MNIST, as MNIST is the most difficult dataset to attack in terms of the average distortion per image pixel from Table 1.

Fixing κ\kappa, adversarial examples generated from the original (undefended) network are used to attack the defensively distilled network with the temperature parameter T=100T=100 (?). The attack success rate (ASR) of EAD, the C&W method and I-FGM are shown in Figure 4. When κ=0\kappa=0, all methods attain low ASR and hence do not produce transferable adversarial examples. The ASR of EAD and the C&W method improves when we set κ>0\kappa>0, whereas I-FGM’s ASR remains low (less than 2%) since the attack does not have such a parameter for transferability.

Notably, EAD can attain nearly 99% ASR when κ=50\kappa=50, whereas the top ASR of the C&W method is nearly 88% when κ=40\kappa=40. This implies improved attack transferability when using the adversarial examples crafted by EAD, which can be explained by the fact that the ISTA operation in (8) is a robust version of the C&W attack via shrinking and thresholding. We also find that setting κ\kappa too large may mitigate the ASR of transfer attacks for both EAD and the C&W method, as the optimizer may fail to find an adversarial example that minimizes the loss function ff in (4) for large κ\kappa. The complete attack transferability results are given in the supplementary material1.

Complementing Adversarial Training

To further validate the difference between L1L_{1}-based and L2L_{2}-based adversarial examples, we test their performance in adversarial training on MNIST. We randomly select 1000 images from the training set and use the C&W attack and EAD (L1L_{1} rule) to generate adversarial examples for all incorrect labels, leading to 9000 adversarial examples in total for each method. We then separately augment the original training set with these examples to retrain the network and test its robustness on the testing set, as summarized in Table 1. For adversarial training with any single method, although both attacks still attain a 100% success rate in the average case, the network is more tolerable to adversarial perturbations, as all distortion metrics increase significantly when compared to the null case. We also observe that joint adversarial training with EAD and the C&W method can further increase the L1L_{1} and L2L_{2} distortions against the C&W attack and the L2L_{2} distortion against EAD, suggesting that the L1L_{1}-based examples crafted by EAD can complement adversarial training.

Conclusion

We proposed an elastic-net regularized attack framework for crafting adversarial examples to attack deep neural networks. Experimental results on MNIST, CIFAR10 and ImageNet show that the L1L_{1}-based adversarial examples crafted by EAD can be as successful as the state-of-the-art L2L_{2} and L∞L_{\infty} attacks in breaking undefended and defensively distilled networks. Furthermore, EAD can improve attack transferability and complement adversarial training. Our results corroborate the effectiveness of EAD and shed new light on the use of L1L_{1}-based adversarial examples toward adversarial learning and security implications of deep neural networks.

Acknowledgment Cho-Jui Hsieh and Huan Zhang acknowledge the support of NSF via IIS-1719097.

References

Supplementary Material

Since the L1L_{1} penalty β∥x−x0∥1\beta\|\mathbf{x}-\mathbf{x}_{0}\|_{1} in (3) is a non-differentiable yet smooth function, we use the proximal gradient method (?) for solving the EAD formulation in (3). Define ΦZ(z)\Phi_{\mathcal{Z}}(\mathbf{z}) to be the indicator function of an interval Z\mathcal{Z} such that ΦZ(z)=0\Phi_{\mathcal{Z}}(\mathbf{z})=0 if z∈Z\mathbf{z}\in\mathcal{Z} and ΦZ(z)=∞\Phi_{\mathcal{Z}}(\mathbf{z})=\infty if z∉Z\mathbf{z}\notin\mathcal{Z}. Using ΦZ(z)\Phi_{\mathcal{Z}}(\mathbf{z}), the EAD formulation in (3) can be rewritten as

where g(x)=c⋅f(x,t)+∥x−x0∥22g(\mathbf{x})=c\cdot f(\mathbf{x},t)+\|\mathbf{x}-\mathbf{x}_{0}\|_{2}^{2}. The proximal operator Prox(x)\textnormal{Prox}(\mathbf{x}) of β∥x−x0∥1\beta\|\mathbf{x}-\mathbf{x}_{0}\|_{1} constrained to x∈p\mathbf{x}\in^{p} is

where the mapping function SβS_{\beta} is defined in (12). Consequently, using (Proof of Optimality of (8) for Solving EAD in (3)), the proximal gradient algorithm for solving (EAD Formulation and Generalization) is iterated by

Grid Search for FGM and I-FGM (Table 4)

To determine the optimal distortion parameter ϵ\epsilon for FGM and I-FGM methods, we adopt a fine grid search on ϵ\epsilon. For each image, the best parameter is the smallest ϵ\epsilon in the grid leading to a successful targeted attack. If the grid search fails to find a successful adversarial example, the attack is considered in vain. The selected range for grid search covers the reported distortion statistics of EAD and the C&W attack. The resolution of the grid search for FGM is selected such that it will generate 1000 candidates of adversarial examples during the grid search per input image. The resolution of the grid search for I-FGM is selected such that it will compute gradients for 10000 times in total (i.e., 1000 FGM operations ×\times 10 iterations) during the grid search per input image, which is more than the total number of gradients (9000) computed by EAD and the C&W attack.

Comparison of COV and EAD on CIFAR10 (Table 5)

Table 5 compares the attack performance of using EAD (Algorithm 1)) and the change-of-variable (COV) approach for solving the elastic-net formulation in (EAD Formulation and Generalization) on CIFAR10. Similar to the MNIST results in Table 1, although COV and EAD attain similar attack success rates, we find that COV is not effective in crafting L1L_{1}-based adversarial examples. Increasing β\beta leads to less L1L_{1}-distorted adversarial examples for EAD, whereas the distortion (L1L_{1}, L2L_{2} and L∞L_{\infty}) of COV is insensitive to changes in β\beta. The insensitivity of COV suggests that it is inadequate for elastic-net optimization, which can be explained by its inefficiency in subgradient-based optimization problems (?).

Figure 5 compares the average-case distortion of these two decision rules with different values of β\beta on CIFAR10. For the same β\beta, the L1L_{1} rule gives less L1L_{1} distorted adversarial examples than those given by the EN rule at the price of larger L2L_{2} and L∞L_{\infty} distortions. We also observe that the L1L_{1} distortion does not decrease monotonically with β\beta. In particular, large β\beta values (e.g., β=5⋅10−2\beta=5\cdot 10^{-2}) may lead to increased L1L_{1} distortion due to excessive shrinking and thresholding. Table 6 and Table 7 displays the complete attack results of these two decision rules on MNIST and CIFAR10, respectively.

Tables 8, 9 and 10 summarize the complete attack results of all the considered attack methods on MNIST, CIFAR10 and ImageNet, respectively. EAD, the C&W attack and I-FGM all lead to 100% attack success rate in the average case. Among the three image classification datasets, ImageNet is the easiest one to attack due to low distortion per image pixel, and MNIST is the most difficult one to attack. For the purpose of visual illustration, the adversarial examples of selected benign images from the test sets are displayed in Figures 6, 7 and 8. On CIFAR10 and ImageNet, the adversarial examples are visually indistinguishable. On MNIST, the I-FGM examples are blurrier than EAD and the C&W attack.

Complete Results on Attacking Defensive Distillation on MNIST and CIFAR10 (Tables 11 and 12)

Tables 11 and 12 display the complete attack results of EAD and the C&W method on breaking defensive distillation with different temperature parameter TT on MNIST and CIFAR10. Although defensive distillation is a standard defense technique for DNNs, EAD and the C&W attack can successfully break defensive distillation with a wide range of temperature parameters.

Complete Attack Transferability Results on MNIST (Table 13 )

Table 13 summarizes the transfer attack results from an undefended DNN to a defensively distilled DNN on MNIST using EAD, the C&W attack and I-FGM. I-FGM methods have poor performance in attack transferability. The average attack success rate (ASR) of I-FGM is below 2%. On the other hand, adjusting the transferability parameter κ\kappa in EAD and the C&W attack can significantly improve ASR. Tested on a wide range of κ\kappa values, the top average-case ASR for EAD is 98.6% using the EN rule and 98.1% using the L1L_{1} rule. The top average-case ASR for the C&W attack is 87.4%. This improvement is significantly due to the improvement in the worst case, where the top worst-case ASR for EAD is 87% using the EN rule and 85.8% using the L1L_{1} rule, while the top worst-case ASR for the C&W attack is 30.5%. The results suggest that L1L_{1}-based adversarial examples have better attack transferability.

Table 14 displays the complete results of adversarial training on MNIST using the L2L_{2}-based adversarial examples crafted by the C&W attack and the L1L_{1}-based adversarial examples crafted by EAD with the EN or the L1L_{1} decision rule. It can be observed that adversarial training with any single method can render the DNN more difficult to attack in terms of increased distortion metrics when compared with the null case. Notably, in the average case, joint adversarial training using L1L_{1} and L2L_{2} examples lead to increased L1L_{1} and L2L_{2} distortion against the C&W attack and EAD (EN), and increased L2L_{2} distortion against EAD (L1L_{1}). The results suggest that EAD can complement adversarial training toward resilient DNNs. We would like to point out that in our experiments, adversarial training maintains comparable test accuracy. All the adversarially trained DNNs in Table 14 can still attain at least 99% test accuracy on MNIST.