A Data Augmentation-based Defense Method Against Adversarial Attacks in Neural Networks

Yi Zeng, Han Qiu, Gerard Memmi, Meikang Qiu

Introduction

With the rapid development of the Deep Neural Networks (DNNs) in Computer Vision (CV), there are more and more real-world applications that rely on the DNN models to classify images or to make decisions . However, in recent years, the DNN models are well known to be vulnerable to Adversarial Examples (AE) which threats the robustness of the DNN usage . Basically, the AEs can be generated by adding carefully designed perturbations that are imperceptible to human eyes but can mislead DNN classifiers with very high accuracy .

Today, several rounds of AE attack and corresponding defense techniques have been developed as shown in . The initial research on adversarial attacks on DNN models such as Fast Gradient Sign Method (FGSM) aims at generating AEs by directly calculating the model gradients with respect to the input images. Such methods are then defeated by the defense methods based on various kinds of methods such as model distillation . Then, the improved AE attacks are proposed to combine the gradient-based approach with the optimization algorithm such as the CW aims to find the input features that made the most significant changes to the final output to mislead the DNN models. Such an optimized gradient-based approach can defeat many previous defense methods including the model distillation. Later, advanced defense methods are proposed to mitigate such attacks by obfuscating the gradients of the inference process. Specifically, some data augmentation techniques are deployed in such defense methods such as image compression, image denoising, image transformation , etc. Then, such state-of-the-art methods are then defeated by more advanced attack methods such as Backward Pass Differentiable Approximation (BPDA) that can effectively approximate the obfuscated gradients to defeat these defenses.

In this paper, we propose a novel defense method that combines several data augmentation techniques together to mitigate the adversarial attacks against on the DNN models. We propose our method, Stochastic Affine Transformation (SAT), by deploying the image translation, image rotation, and image scaling method together. Our method can be used as a preprocessing step on the input images which makes our solution agnostic on many DNN models. Firstly, our method has little influence on the DNN inference which can effectively maintain the classification accuracy of benign images. Then, intensive experimentation and comparison have been performed to show the improvement of our method compared with several previous state-of-the-art defense solutions. Moreover, our method is a lightweight preprocess-only step that can be used on resource-constrained use cases such as the Internet of Things (IoT) .

This paper includes two main contributions. (1) We design a data augmentation-based defense solution to mitigate the initial and optimized gradient-based adversarial attacks on DNN models. Our method combining several steps of data augmentation techniques can be used as a preprocessing step on input images that can effectively maintain the agnostic DNN model’s accuracy. (2) Our method can also defeat the advanced adversarial attack method such as BPDA which outperforms many previous state-of-the-art defense solutions.

This paper is organized as follows. Section 2 discusses the background information of this research including the brief definition of the adversarial examples and the previous data augmentation-based defense solutions. Section 3 presents our threat model and defense requirements. Section 4 proposes our methodology including the algorithm and the design details. Section 5 illustrates the experimentation details and evaluation results comparing with the previous state-of-the-art solutions. We then conclude in Section 6.

Research Backgrounds

In this section, we briefly introduce the background of the AEs in DNNs, the related work on AEs, and the state-of-the-art preprocessing-based defense methods based on data augmentation techniques.

AEs can be explained as imperceptible modified samples that force one or multiple DNN models outputs with wrong results. This was first highlighted by . By denoting II an input image, an adversarial example generated from it can be denoted as I~=I+δ\widetilde{I}=I+\delta, where δ\delta is the adversarial perturbation. The target model, which conducts inference for classification tasks can be denoted as ff, thus the problem of performing adversarial attacks on the target DNN model can be formulated as Eq 1.

This equation can be interpreted as an optimization task that searches for a I~\widetilde{I} based on II that can be misclassified by the target model, while keeping I~\widetilde{I} visually as similar to II as possible. The aforementioned AE generation case is untargeted which aims to mislead the DNN classifier without a pre-set wrong label. As a targeted AE generation procedure aims to attack the DNN classifier to misclassify an input II with original label ll as the pre-set wrong label l′l^{\prime}. In the concern of real-life adoption of DNN models, both cases can result in serious outcomes if the models are not protected.

Since the time this vulnerability of DNNs has been discovered, various kinds of attacks have been proposed in the past few years to help the society to better understand the nature of AEs. To sum up, past work on adversarial attacks can be classified into two main approaches including initial gradient-based and optimized gradient-based. Fast Gradient Sign Method (FGSM) is one of the most famous initial gradient-based adversarial attacks which calculates the model gradients based on the sign of the gradient of the classification loss concerning the input image. FGSM performs a one-step gradient update along the direction of the sign of gradient at each pixel under Linf⁡L_{\inf} constraints to generate AEs. Later on, variations of FGSM were introduced to better searching for the optimum AE based on a single input. Such kind of methods includes I-FGSM and MI-FGSM , aim at iteratively calculating the perturbations based on FGSM with a small step or with momentum.

Then, optimized gradient-based AE attacks are proposed to calculate the gradients based on adopting optimization algorithms to find optimal adversarial perturbations directly between the input images and output predicted labels . Such kind of attack is especially powerful in a whitebox or graybox scenario by adopting optimization algorithms to enhance the gradient calculation. Various optimized gradient-based AE attacks were proposed in recent years including Jacobian-based Saliency Map Attack (JSMA ), DeepFool , LBFGS , Carlini & Wagner (CW ), and Backward Pass Differentiable Approximation (BPDA ). We should highlight the BPDA attack here, as it invalidates dozens of existing state-of-art defense approaches in recent evaluations . The BPDA attack in a manner assumes that a defense function g(⋅)g(\cdot) maintains the property g(I)≈Ig(I)\approx I in order to preserve the functionality of the target model f(⋅)f(\cdot). Then the adversary can use g(I)g(I) on the forward pass and replace it with II on the backward pass when calculating the gradients.

2 Data Augmentation-based Preprocessing Defense Solutions

Various defensive strategies have been proposed to defeat adversarial attacks. One direction is to train a more robust model from either scratch or an existing model. Those approaches aim to rectify AEs’ malicious features by including AEs into the training set , processing all the training data , or revising the DNN topology . However, training a DNN model is very time and resource-consuming, especially for real-life cases, where models are more complicated. Besides, in real-life, DNN models are packed as closed-source applications and cannot be modified, thus those methods are not applicable. Most of all, the adversary can still adaptively generate AEs for the new models .

A more promising direction is to preprocess the input data to eliminate adversarial influence without touching the DNN model. These solutions are more suitable in the concern of real-life cases, as it is feasible, efficient, and lightweight. Thus, The preprocessing based defense is within the scope of this paper, as they do not require any laborious work with the DNN models, which made them competitive with most of the real-life defense scenarios. Below we describe some previous works and their limitations:

Feature Distillation (FD) designed a compression method based on the JPEG compression but modified the quantization step. The basic idea is to measure the importance of input features for DNNs by leveraging the statistical frequency component analysis within the DCT of JPEG. It demonstrated a huge improvement in defending adversarial attacks compared with the standard JPEG compression method .

SHIELD aims to randomize the quantization step by tuning the window size and quantization factors in the JPEG compression method. In SHIELD, the Stochastic Local Quantization (SLQ) method is used to divide an image into 8×88\times 8 blocks and applies a randomly selected JPEG compression quality (tuning quantization factors) to every block. The advantage is that the authors randomized the selective quantization steps which make the defense process different for different input images and make the adversarial attacks more difficult.

Bit-depth Reduction (BdR) performs a simple type of quantization that can remove small (adversarial) variations in pixel values from an image. In the evaluation of that work, it demonstrates a more effective result comparing to adversarial training. However, recently developed attacks are not within the scope of that work, namely the BPDA attack.

Pixel Deflection (PD) aims to add similar natural noises that are not sensitive to the DNN model. The idea of deflection is to randomly sample a pixel from an image and replace it with another randomly selected pixel from within a small square neighborhood. This could generate an artificial noise that affects little on the DNN model but can disturb the adversarial perturbations. Then, a BayesShrink denoising process is followed to recover the image content before this image is feed into the DNN model. The results of such a method are convincing since it introduces randomness into the preprocessing step and does not require any modification on the DNN model. However, the robustness of this method is significantly reduced if the attackers have knowledge of the preprocess step .

Threat Model and Defense Requirements

Untargeted attacks and targeted attacks are two major types of adversarial attacks. Untargeted attacks try to mislead the DNN models to an arbitrary label different from the correct one. On the other hand, targeted attacks only considering succeed when the DNN model predicts the input as one specific label desired by the adversary . In this paper, we only evaluate the targeted attacks. The untargeted attacks can be mitigated in the same way.

We consider a full whitebox scenario, where the adversary has full knowledge of the DNN model and the defense method, including the network architecture, exact values of parameters, hyper-parameters, and the details of the defense method. However, we assume the random numbers generated in real-time are perfect with a large entropy such that the adversary cannot obtain or guess the correct values. Such a targeted full whitebox scenario represents the strongest adversaries, as a big number of existing state-of-the-art defenses are invalidated as shown in .

As for the adversary’s capability, we assume the adversary is outside of the DNN classification system, and he is not able to compromise the inference computation or the DNN model parameters (e.g., via fault injection to cause bit-flips or backdoor attacks ). What the adversary can do is to manipulate the input data with imperceptible perturbations. In the context of computer vision tasks, he can directly modify the input image pixel values within a certain range. We use l∞l_{\infty} and l2l_{2} distortion metrics to measure the scale of added perturbations: we only allow the generated AEs to have either a maximum l∞l_{\infty} distance of 8/255 or a maximum l2l_{2} distance of 0.05 as proposed in .

2 Defense Requirements

Various life-concerned vital tasks are already implemented with DNN in real-life, e.g., video surveillance , face authentication , autonomous driving , network traffic identification , etc. Most of those cases’ inference is conducted either locally in distributed computing units or remotely with the help of cloud servers. In the case where inference procedures are conducted locally, resources are highly constrained, thus only lightweight designs of defense can protect the system without draw extra burden over those units. Moreover, for both cases, previous resolutions, e.g. adversarial training or extra modifications over the target model would be considerably costly for real-life cases (where each sample is of large size, thus any kind of retraining can be laborious). As mentioned in the previous section, most of the existing defense methods either too complicated to be compatible with real-life constraints () or not capable of effectively reduce the impact brought by adversarial attacks ().

To cope with those constraints in real-life adoption of DNNs, we believe the following properties should be taken considered when designing novel adversarial defense methods:

Accuracy-preserving: they should not affect much on the prediction accuracy of the DNN model on clean data samples that processed by those methods.

Security: they should be capable to effectively reduce effects brought by adversarial perturbations.

Lightweight: the defense method should not be too heavy to impact the devices’ or units’ performance or operations, considering the limited onboard computing capabilities and resources.

Generalization Ability: the defense method should neither require modifications over the target DNN’s structure nor require any kind of retraining.

Thus, for a better adaptation over nowadays real-life scenarios, we aim to design a preprocessing-only method to meet all those requirements. It has proved in our previous work , preprocessing-only adversarial defenses are competent enough to defense adversarial attacks, even for whitebox attacks.

Proposed Methodology

To better adapt to nowadays real-life scenarios’ adoption of DNN, we present our efficient defense method against adversarial attacks that can conduct protections on the fly, termed Stochastic Affine Transformation (SAT). Thanks to the lightweight design of SAT, this defense should be more compatible with both cloud DNN inference as well as local or edge DNN inference respecting real-life scenarios. Figure 1 illustrates an overview of adopting the SAT method in real life. The details of SAT will be present in Section 4.1. The analysis of the three hyperparameters respecting the defense efficiency is illustrated in Section 4.2.

Following the logic of adding randomness to the affine transformation without harming the classification accuracy , we propose a simple but effective way of image distortion as an adversarial example defense. The algorithm is designed based on combining several affine transformation methods. The details are illustrated in Algorithm 1.

Three basic affine transformations with randomized coefficients are bounded tother in this single procedure, namely, translation, rotation, and scaling. We add randomness to those three simple affine transformations so that the attacker cannot utilize a useful gradient to generate adversarial examples even acknowledges the details of this defense. Such is done by acquiring different coefficients that follow three uniform distribution for different samples. To be specific, there are three coefficients along with the raw image as the input of the SAT method. TT is the translation limit, RR is the rotation limit, and SS is the scaling limit. The original input will first be randomly shifted away from its original coordinates according to δx\delta_{x} and δy\delta_{y} that both follow the uniform distribution in the range (−T,T)(-T,T). Then, the data will be randomly rotated at a certain angle δr\delta_{r} that follows the uniform distribution in the range (−R,R)(-R,R). Finally, the distorted image will be acquired by scaling up or down δs\delta_{s} times, where δs\delta_{s} follows a uniform distribution in the range (1−S,1+S)(1-S,1+S).

Since only simple affine transformations and random number generator are adopted in SAT, we believe SAT is compatible with both cloud DNN inference procedures as well as localized or edge devices DNN inference procedure. This lightweight design is also hardware friendly and can conduct protections on the fly.

2 SAT Hyper-parameters

As aforementioned, there are three essential coefficients in the SAT method. In this part, we did a thorough evaluation and analysis of those three coefficients respecting the efficiency of defending adversarial attacks.

We believe a higher variance between the original data and the protected data while maintaining a high classification accuracy can help more to defend adversarial attacks, which is proved in our previous work . In this part, three different metrics are adopted to evaluate this variance, namely the l2l_{2} norm, Structural Similarity (SSIM) index, and the Peak Signal-to-Noise Ratio (PSNR). The classification accuracy (ACC) is the priority of most classification tasks thus is as well taken considerate in this part.

l2l_{2} is a widely adopted metric in deep learning domain to measure the amount of difference of two samples in the term of Euclidean distance, a higher l2l_{2} indicates a greater difference. Eq. 2 shows how l2l_{2} can be computed.

Where I′I^{\prime} and II are the two 3-channel (RGB) samples to be compared. hh and ww are the height and width of those samples respectively.

SSIM is a metric in the computer vision domain normally being adopted to measure the similarity between two images, where smaller SSIM reflects greater difference. To be specific, SSIM is based on three comparison measurements between two samples, namely luminance (ll), contrast (cc), and structure (ss). Each comparison function is elaborated in Eq 3, Eq 4, and Eq 5 respectively.

Where μ(⋅)\mu(\cdot) computes the mean of a sample, σ2(⋅)\sigma^{2}(\cdot) computes the variance of a sample. c1c1 is equal to (0.01×L)2(0.01\times L)^{2}, c2c2 is equal to (0.03×L)2(0.03\times L)^{2}. Here LL is the dynamic range of the pixel values. Finally, the SSIM can be acquired by computing the product of those three functions.

PSNR is most commonly used to measure the quality of reconstruction of lossy compression codecs, say compression, augmentation, or distortion, etc. A smaller PSNR indicates a greater difference between the two evaluating samples. PSNR can be defined via the mean squared error (MSE) between two comparing samples, which is explained in Eq 6.

Where MSE(⋅)MSE(\cdot) computes the MSE between two inputs.

We tried different values of TT and SS in the range [0.01,0.5][0.01,0.5]. Different values of RR is acquired in the range $.Thus,eachcoefficientwilltest11valuesintherespectingrange.Asfordifferentmetrics,wewillacquire1331. Thus, each coefficient will test 11 values in the respecting range. As for different metrics, we will acquire 1331(11\times 11\times 11)resultsfromdifferentcombinationsofthosedifferentvaluesofcoefficients.Figure2demonstratesthechangeofthosefourmetrics’valuewhendifferentresults from different combinations of those different values of coefficients. Figure 2 demonstrates the change of those four metrics’ value when differentT,,S,and,andR$ are adopted.

In Figure 2(a), the changes of ACC with those three coefficients varies is presented. The right side of Figure 2(a) is the color-bar that reflecting the ACC attained with respecting combinations of those three coefficients. We can learn that lower TT, SS, and RR can help the model maintain a high ACC. To ensure a high ACC, we set 95% ACC as a standard, thus those combinations with ACC below this standard would not be taken further considerations.

From Figure 2(b) we can learn that the l2l_{2} is not that sensitive with SS and RR comparing to TT in their respecting range. Combining the information provided from Figure 2(b), Figure 2(c), and Figure 2(d) we can as well acquire this similar analytical result for SSIM and PSNR. This phenomenon also reflects that the l2l_{2}, SSIM, and PSNR can all reflecting the scale of changes in a similar manner respecting affine transformations.

By overlapping Figure 2(a) with the other three figures, we can acquire a set of optimum coefficients that ensures high ACC and great variance at the same time. The set of coefficients for the following experiment is set as follows: T=0.16T=0.16, S=0.16S=0.16, and finally R=4R=4.

We compare the SAT method using those fine-tuned hyperparameters with other state-of-art adversarial defense methods, which presented in Table 1. As demonstrated, SAT can create a greater difference between raw samples and protected samples while maintaining a high ACC than other methods. We will evaluate whether this greater variance will help and how much will it help DNN models to defend adversarial attacks in the following section.

Experimentation and Evaluation

In this section, we conduct a comprehensive evaluation of the proposed technique. Various adversarial attacks are taken considered in this part: 4 kinds of standard adversarial attacks (FGSM, I-FGSM, LBFGS, and C&W) are conducted, advanced interactive gradient approximation attack, namely BPDA, is also conducted to evaluate the robustness of SAT. We compare SAT with four state-of-art defense methods publish in top-tier artificial intelligence conferences from the past two years. This section would be divided into three parts to elaborate on the settings of the experiment, efficiency over standard adversarial attacks, and the efficiency over BPDA respectively.

Tensorflow is adopted as the deep learning framework to implement the attacks and defenses. The learning rate of the C&W and BPDA attack is set to 0.1. All the experiments were conducted on a server equipped with 8 Intel I7-7700k CPUs and 4 NVIDIA GeForce GTX 1080 Ti GPU.

SAT is of general-purpose and can be applied to various models over various platforms as a preprocessing step for computer vision tasks as illustrated in Figure 1. Without the loss of generality, we choose a pre-trained Inception V3 model over the ImageNet dataset as the target model. This state-of-the-art model can reach 78.0% top-1 and 93.9% top-5 accuracy. We randomly select 100 images from the ImageNet Validation dataset for AE generation. These images can be predicted correctly by this Inception V3 model.

We consider the targeted attacks where each target label different from the correct one is randomly generated . For each different attack, we measure the classification accuracy of the generated AEs (ACC) and the attack success rate (ASR) of the targeted attack. To be noticed that the untargeted attacks are not within our scope in this work. A higher ACC or lower ASR indicates the defense is more resilient against the attacks.

For comparison, we re-implemented 4 existing solutions including FD , SHIELD , Bit-depth Reduction , and PD .

2 Evaluation on Defending Adversarial Attacks

We first evaluate the efficiency of our proposed method over standard adversarial attacks, namely FGSM, I-FGSM, C&W, and LBFGS. For FGSM and I-FGSM, AEs are generated under l∞l_{\infty} constraint of 0.03. For LBFGS and C&W, the attack process is iterated under l2l_{2} constraint and stops when all targeted AEs are found. We measure the model accuracy (ACC) and attack success rate (ASR) with the protection of SAT and other defense methods.

The results are shown in Table 2 and Table 3. For benign samples only, our proposed techniques have the smallest influence on the model accuracy comparing to past works. For defeating AEs generated by these standard attacks, the attack success rate can be kept around 0% and the model accuracy can be drastically recovered, which an efficiency against different kinds of adversarial attacks is demonstrated. To be noticed, previous work can only attain an accuracy of around 50% on samples attacked by the FGSM (ϵ=.03\epsilon=.03). SAT can recover the accuracy to 0.61%.

Comparing to the Table 1 which compares different methods’ capability of creating a variance between input and defended sample, the defense efficiency evaluated in this part shows a strong correlation with the amount of variance generated by the defense method. This has confirmed the previous conclusion in our previous work .

In a nutshell, the effectiveness of SAT against standard adversarial attack is demonstrated, as we can attain a state-of-art defense efficiency on all the evaluated attacks comparing to other methods.

3 Evaluation on Defending Advanced Adversarial Attacks

We then evaluate the effectiveness of SAT against the BPDA attack. Since the BPDA attack is an interactive attack, we record the ACC and ASR for each round for different defense methods.

The model prediction accuracy and attack success rate in each round are shown in Figure 3(a) and 3(b), respectively. We can observe that after 50 attack rounds, all other three prior solutions except FD can only keep the model accuracy lower than 5%, and attack success rates reach higher than 90%. Those defenses fail to mitigate the BPDA attack. FD can keep the attack success rate lower than 20% and the model accuracy is around 40%. This is better but still not very effective in maintaining the DNN model’s robustness.

In contrast, SAT is particularly effective against the BPDA attack. As our method can maintain an acceptable model accuracy (around 80% for 50 perturbation rounds), and restrict the attack success rate to 0 for all the record rounds. This result is as well consistent with the l2l_{2}, SSIM, and PSNR metrics compared in Table 1: the randomization effects in SAT cause greater variances between I′I^{\prime} and II, thus invalidating the BPDA attack basic assumption, which is I′≈II^{\prime}\approx I.

To sum up, the effectiveness of SAT against the BPDA attack is demonstrated, as a considerable improvement over the defense efficiency against the BPDA attack is shown comparing to previous work.

Conclusion

In this paper, we proposed a lightweight defense method that can effectively invalidate adversarial attacks, termed SAT. By adding randomness to the coefficients, we integrated three basic affine transformations into SAT. Compared with four state-of-art defense methods published in the past two years, our method clearly demonstrated a more robust and effective defense result on standard adversarial attacks. Moreover, respecting the advanced BPDA attack, SAT showed an outstanding capability of maintaining the target model’s ACC and detain the ASR to 0. This result is almost 50% better than the best result achieved by previous work against full whitebox targeted attacks.

References