Generalizable Data-free Objective for Crafting Universal Adversarial Perturbations

Konda Reddy Mopuri, Aditya Ganeshan, R. Venkatesh Babu

I Introduction

Small but structured perturbations to the input, called adversarial perturbations, are shown () to significantly affect the output of machine learning systems. Neural network based models, despite their excellent performance, are observed () to be vulnerable to adversarial attacks. Particularly, Deep Convolutional Neural Networks (CNN) based vision models () can be fooled by carefully crafted quasi-imperceptible perturbations. Multiple hypotheses attempt to explain the existence of adversarial samples, viz. linearity of the models , finite training data , etc. More importantly, the adversarial perturbations generalize across multiple models. That is, the perturbations crafted for one model fools another model even if the second model has a different architecture or is trained on a different dataset (). This property of adversarial perturbations enables potential intruders to launch attacks without the knowledge about the target model under attack: an attack typically known as black-box attack . In contrast, an attack where everything about the target model is known to the attacker is called a white-box attack. Until recently, all the existing works assumed a threat model in which the adversaries can directly feed input to the machine learning system. However, Kurakin et al. lately showed that the adversarial samples can remain misclassified even if they were constructed in physical world and observed through a sensor (e.g., camera). Given that the models are vulnerable even outside the laboratory setup , the models’ susceptibility poses serious threat to their deploy-ability in the real world (e.g., safety concerns for autonomous driving). Particularly, in case of critical applications that involve safety and security, reliable models need to be deployed to stand against the strong adversarial attacks. Thus, the effect of these structured perturbations has to be studied thoroughly in order to develop dependable machine learning systems.

Recent work by Moosavi-Dezfooli et al. presented the existence of image-agnostic perturbations, called universal adversarial perturbations (UAP) that can fool the state-of-the-art recognition models on most natural images. Their method for crafting the UAPs, based on the DeepFool attacking method, involves solving a complex optimization problem (eqn. 2) to design a perturbation. The UAP procedure utilizes a set of training images to iteratively update the universal perturbation with an objective of changing the predicted label upon addition. Similar to , Metzen et al. proposed UAP for semantic segmentation task. They extended the iterative FGSM attack by Kurakin et al. to change the label predicted at each pixel. Additionally, they craft image-agnostic perturbations to fool the system in order to predict a pre-determined target segmentation output.

However, these approaches to craft UAPs () have the following important drawbacks:

Data dependency: It is observed that the objective presented by to craft UAP requires a minimum number of training samples for it to converge and craft an image-agnostic perturbation. Moreover, the fooling performance of the resulting perturbation is proportional to the available training data (Figure 8). Similarly, the objective for semantic segmentation models (e.g., ) also requires data. Therefore, existing procedures can not craft perturbations when enough data is not provided.

Weaker black-box performance: Since information about the target models is generally not available for attackers, it is practical to study the black-box attacks. Also, black-box attacks reveal the true susceptibility of the models, while white-box attacks provide an upper bound on the achievable fooling. However, the black-box attack of UAP is significantly weaker than their white-box attack (Table VIII). Note that, in , authors have not analyzed the performance of their perturbations in the black-box attack scenario. They have assumed that the training data of the target models is known and have not considered the case in which adversary has access to only a different set of data. This amounts to performing only semi white-box attacks. Black-box attacks generally imply () that the adversary does not have access to (i) the target network architecture (including the parameters), and (ii) a large training dataset. Even in the case of semantic segmentation, since work with targeted attacks, they observed that the perturbations do not generalize to other models very well.

Task specificity: The current objectives to craft UAPs are task specific. The objectives are typically designed to suit the underlying task at hand since the concept of fooling varies across the tasks. Particularly, for regression tasks such as depth estimation and crowd counting, extending the existing approaches to craft UAPs is non-trivial.

In order to address the above shortcomings and to better analyze the stability of the models, we present a novel data-free objective to craft universal adversarial perturbations, called GD-UAP. Our objective is to craft image-agnostic perturbations that can fool the target model without any knowledge about the data distribution, such as, the number of categories, type of data (e.g., faces, objects, scenes, etc.) or the data samples themselves. Since we do not want to utilize any data samples, instead of an objective that reduces the confidence to the predicted label or flip the predicted label (as in ), we propose an objective to learn perturbations that can adulterate the features extracted by the models. Our proposed objective attempts to over-fire the neurons at multiple layers in order to deteriorate the extracted features. During the inference time, the added perturbation misfires the neuron activations in order to contaminate the representations and eventually lead to wrong prediction.

This work extends our earlier conference paper . We make the following new contributions in this paper:

We propose a novel data-free objective for crafting image-agnostic perturbations.

We demonstrate that our objective is generalizable across multiple vision tasks by extensive evaluation of the crafted perturbations across three different vision tasks covering both classification and regression.

Further, we show that apart from being data-free objective, the proposed method can exploit minimal prior information about the training data distribution of the target models in order to craft stronger perturbations.

We present comprehensive analysis of the proposed objective which includes: (a) a thorough comparison of our approach with the data-dependant counterparts, and (b) evaluation of the strength of UAPs in the presence of various defense mechanisms.

The rest of this paper is organized as follows: section II presents detailed account of related works, section III discusses the proposed data-free objective to craft image-agnostic adversarial perturbations, section IV demonstrates the effectiveness of GD-UAP to craft UAPs across various tasks, section V hosts a thorough experimental analysis of GD-UAP and finally section VI concludes the paper.

II Related Works

Szegedy et al. demonstrated that despite their superior recognition performance, neural networks are susceptible to adversarial perturbations. Subsequently, multiple other works studied this interesting and surprising property of the machine learning models. Though it is first observed with recognition models, the adversarial behaviour is noticed with models trained on other tasks such as semantic segmentation , object detection , pose estimation and deep reinforcement learning tasks . There exist multiple methods to craft these malicious perturbations for a given data sample. For recognition tasks, they range from performing simple gradient ascent on cost function to solving complex optimizations ( ). Simple and fast methods such as FGSM find the gradient of loss function to determine the adversarial perturbation. An iterative version of this attack presented in achieves better fooling via performing the gradient ascent multiple times. On the other hand, complex approaches such as and find minimal perturbation that can move the input across the learned classification boundary in order to flip the predicted label. More robust adversarial attacks have been proposed recently that transfer to real world and are invariant to general image transformations .

Moreover, it is observed that the perturbations exhibit transferability, that is, perturbations crafted for one model can fool other models with different architectures and different training sets as well (). Further, Papernot et al. introduced a practical attacking setup via model distillation to understand the black-box attack. Black-box attack assumes no information about the target model and its training data. They proposed to use a target model’s substitute to craft the perturbations.

The common underlying aspect of all these techniques is that they are intrinsically data dependent. The perturbation is crafted for a given data sample independently of others. However, recent works by Moosavi-Dezfooli et al. and Metzen et al. showed the existence of input-agnostic perturbations that can fool the models over multiple images. In , authors proposed an iterative procedure based on Deepfool attacking method to craft a universal perturbation to fool classification models. Similarly, in , authors craft universal perturbations that can affect target segmentation output. However, both these works optimize for different task specific objectives. Also, they require training data to craft the image-agnostic perturbations. Unlike the existing works, our proposed method GD-UAP presents a data-free objective that can craft perturbations without the need for any data samples. Additionally, we introduce a generic notion of fooling across multiple computer vision tasks via over-firing the neuron activations. Particularly, our objective is generalizable across various vision models in spite of differences in terms of architectures, regularizers, underlying tasks, etc.

III Proposed Approach

The pixel intensities of δ\delta are restricted by an imperceptibility constraint. Typically, it is realized as a max-norm constraint in terms of l\rotatebox90.08l_{\rotatebox{90.0}{8}} or l2l_{2} norms (e.g. ). In this paper, for all our analysis we impose l\rotatebox90.08l_{\rotatebox{90.0}{8}} norm. Thus, the aim is to find a δ\delta such that

However, the focus of the proposed work is to craft δ\delta without requiring any data samples. The data-free nature of our approach prohibits us from utilizing eqn. 2 for learning δ\delta, as we do not have access to data xx. Therefore, we instead propose to fool the CNN by contaminating the extracted representations of the input at multiple layers of the architecture. In other words, as opposed to the typical “flipping the label” objective, we attempt to “over-fire” the features extracted at multiple layers. That is, we craft a perturbation δ\delta such that it leads to additional activation firing at each layer and thereby misleading the features (filters) at the following layer. The accumulated effect of the contamination eventually leads the CNN to misclassify.

The perturbation essentially causes filters at a particular layer to spuriously fire and extract inefficacious information. Note that in the presence of data (during attack), in order to mislead the activations from retaining useful discriminative information, the perturbation (δ)(\delta) has to be highly effective. Also, the imperceptibility constraint (second part of eqn. 2) on δ\delta makes it more challenging.

Hence without utilizing any data xx, we seek an image-agnostic perturbation δ\delta that can produce maximal spurious activations at each layer of a given CNN. In order to craft such a δ\delta we start with a random perturbation and optimize for the following objective:

where li(δ)l_{i}(\delta) is the activation in the output tensor (after the non-linearity) at layer ii when δ\delta is fed to the network ff. KK is the number of layers in ff at which we maximize the activations caused by δ\delta, and ξ\xi is the max-norm limit on δ\delta.

The proposed objective computes product of activation magnitude at all the individual layers. We observed product resulting in stronger δ\delta than other forms of aggregation (e.g. sum). To avoid working with extreme values (≈0\approx 0), we apply log on the product. Note that the objective is open-ended as there is no optimum value to reach. We would ideally want δ\delta to cause as much strong disturbance at all the layers as possible, within the imperceptibility constraint. More discussion on the motivation and working of the proposed objective is presented in Section V.

III-B Implementation Details

We begin with a target network ff which is a trained CNN whose parameters are frozen and a random perturbation δ\delta. We then perform the proposed optimization to update δ\delta for causing strong activations at multiple layers in the given network. Typically, it is considered that the convolution (conv)(conv) layers learn information-extracting features which are then classifed by a series of fcfc layers. Hence, we optimize our objective only at all the convconv layers. This was empirically found to be more effective than optimizing at all layers as well. In case of advanced architectures such as GoogLeNet and ResNet , we optimize at the last layers of all the inception (or residual) blocks and the independent convconv layers. We observed that optimizing at these layers results in δ\delta with a fooling capacity similar to the one resulting from optimizing at all the intermediate layers as well (including the convconv layers within the inception/residual blocks). However, since optimizing at only the last layers of these blocks is more efficient, we perform the same.

Note that the optimization updates only the perturbation δ\delta, not the network parameters. Additionally, no image data is involved in the optimization. We update δ\delta with the gradients computed for loss in eqn. (3) iteratively till the fooling performance of the learned δ\delta gets saturated on a set of validation images. In order to validate the fooling performance of the learned δ\delta, we compose an unrelated substitute dataset (D)(D). Since our objective is not to utilize data samples from the training dataset, we randomly select 1,0001,000 images from a substitute dataset to serve as validation images. It is a reasonable assumption for an attacker to have access to 1,0001,000 unrelated images. For crafting perturbations to object recognition models trained on ILSVRC dataset , we choose random samples from Pascal VOC-2012 dataset. Similarly, for semantic segmentation models trained on Pascal VOC , we choose validation samples from ILSVRC , for depth estimation models trained on KITTI dataset we choose samples from Places-205 dataset.

III-C Exploiting additional priors

Though GD-UAP is a data-free optimization for crafting image-agnostic perturbations, it can exploit simple additional priors about the data distribution X\mathcal{X}. In this section we demonstrate how GD-UAP can utilize simple priors such as (i) mean value and dynamic range of the input, and (ii) target data samples.

Note that the proposed optimization (eqn. (3)) does not consider any information about X\mathcal{X}. We present only the norm limited δ\delta as input and maximize the resulting activations. Hence, during the optimization, input to the target CNN has a dynamic range of [−ξ, ξ][-\xi,\>\xi] (ξ=10)(\xi=10). However, during the inference time, input lies in [0, 255][0,\>255] range. Therefore, it becomes very challenging to learn perturbations that can affect the neuron activations in the presence of strong (an order higher) input signal xx. Hence, in order to make the learning easier, we may provide this useful information about the data (x∈X)(x\in\mathcal{X}), and let the optimization better explore the space of perturbations. Thus, we slightly modify our objective to craft δ\delta relative to the dynamic range of the data. We create pseudo data dd via randomly sampling from a Gaussian distribution whose mean (μ)(\mu) is equal to the mean of training data and variance (σ)(\sigma) is such that 99.9%99.9\% of the samples lie in [0, 255][0,\>255], the dynamic range of input. Thus, we solve for the following loss:

Essentially, we operate the proposed optimization in a subspace closer to the target data distribution X\mathcal{X}. In other words, dd in eqn. (4) acts as a place holder for the actual data and helps to learn perturbations which can over-fire the neuron activations in the presence of the actual data. A single Gaussian sample with twice the size of the input image is generated. Then, random crops from the Gaussian sample, augmented with simple techniques such as random cropping, blurring, and rotation are used for the optimization.

III-C2 Target data samples

Now, we modify our data-free objective to utilize samples from the target distribution X\mathcal{X} and improve the fooling ability of the crafted perturbations. Note that in the case of data availability, we can design direct objectives such as reducing confidence for the predicted label or changing the predicted label, etc. However, we investigate if our data-free objective of over-firing the activations, though is not designed to utilize data, crafts better perturbations when data is presented to the optimization. Additionally, our objective does not utilize data to manipulate the predicted confidences or labels. Rather, the optimization benefits from prior information about the data distribution such as the dynamic range, local patterns, etc., which can be provided through the actual data samples. Therefore, with minimal data samples we solve for the following optimization problem

Presenting data samples to the optimization procedure is a natural extension to presenting the dynamic range of the target data alone (section III-C1). In this case, we utilize a subset of training images on which the target CNN models are trained (similar to ).

III-D Improved Optimization

In this subsection, we present improvements to the optimization process presented in our earlier work . We observe that the proposed objective quickly accumulates δ\delta beyond the imposed max-norm constraint (ξ)(\xi). Because of the clipping performed after each iteration, the updates after δ\delta reaches the constraint are futile . To tackle this saturation, δ\delta is re-scaled to half of its dynamic range (i.e. [−5, 5][-5,\>5]). Not only does the re-scale operation allow an improved utilization of the gradients, it also retains the pattern learnt in the optimization process till that iteration.

In our previous work , the re-scale operation is done in a regular time interval of 300300 iterations. Though this re-scaling helps to learn better δ\delta, it is inefficient since it performs blind re-scaling without verifying the scope for updating δ\delta. This is specially harmful later in the learning process, when the perturbation may not be re-saturated in 300300 iterations.

Therefore, we propose an adaptive re-scaling of δ\delta based on the rate of saturation (reaching the extreme values of ±10\pm 10) in its pixel values. During the optimization, at each iteration we compute the proportion (p)(p) of the pixels in δ\delta that reached the max-norm limit ξ\xi. As the learning progresses, more number of pixels reach the max-norm limit and because of the clipping, eventually get saturated at ξ\xi. Hence, the rate of increase in pp decreases as δ\delta saturates. We compute the rate of saturation, denoted as SS, of the pixels in δ\delta after each iteration during the training. For consecutive iterations, if increase in pp is not significant (less than a pre-defined threshold θ\theta), we perform a re-scaling to half the dynamic range. We observe that this adaptive re-scaling consistently leads to better learning.

III-E Algorithmic summarization

In this subsection, for the sake of brevity we summarize the proposed approach in the form of an algorithm. Algorithm 1 presents the proposed optimization as a series of steps. Note that it is a generic form comprising of all the three variations including both data-free and with prior versions.

For ease of reference, we repeat some of the notation. FtF_{t} is the fooling rate at iteration tt, li(x)l_{i}(x) is the activation caused at layer ii of the CNN ff for an input xx, η\eta is the learning rate used for training, Δ\Delta is the gradient of the loss with respect to the input δ\delta, StS_{t} is the rate of saturation of pixels in the perturbation δ\delta at iteration tt, θ\theta is the threshold on the rate of saturation, FtF_{t} is the fooling rate, HH is the patience interval of validation for verifying the convergence of the proposed optimization.

III-F Generalized Fooling Rate (GFR)

While the notion of ‘fooling’ has been well defined for the task of image recognition, for other tasks it is unclear. Hence, in order to provide an interpretable metric to measure ‘fooling’, we introduce Generalized Fooling Rate (GFR), making it independent of the task, and dependent on the metric being used for evaluating the model’s performance.

Let MM be a metric for measuring the performance of a model for any task, where the range of MM is [0,R][0,R]. Let the metric take two inputs y^\hat{y} and yy, where y^\hat{y} is the predicted output and yy is the ground truth output, such that the performance of the model is measured as M(y^,y)M(\hat{y},y). Let yδ^\hat{y_{\delta}} be the output of the model when the input is perturbed with a perturbation δ\delta. Then, the Generalized Fooling Rate with respect to measure MM is defined as:

This definition of Generalized Fooling rate (GFR) has the following benefits:

GFR is a natural extension of ‘fooling rate’ defined for image recognition, where the fooling rate can be written as GFR(Top1)=1−Top1(yδ^,y^)GFR(Top1)=1-Top1(\hat{y_{\delta}},\hat{y}), where Top1Top1 is the Top-1 Accuracy metric.

Fooling rate should be a measure of the change in model’s output caused by the perturbation. Being independent of the ground truth yy, and dependant only on yδ^\hat{y_{\delta}} and y^\hat{y}, GFR primarily measures the change in the output. A poorly performing model which however is very robust to adversarial attacks will show very poor GFR values, highlighting its robustness.

GFR measures the performance of a perturbation in terms of the damage caused to a model with respect to a metric. This is an important feature, as tasks such as depth estimation have multiple performance measures, where some perturbation might cause harm only to some of the metrics while leaving other metrics unaffected.

For all the tasks considered in this work, we report GFR with respect to a metric as a measure of ‘fooling’.

IV GD-UAP: Effectiveness across tasks

In this section, we present the experimental evaluation to demonstrate the effectiveness of GD-UAP. We consider three different vision tasks to demonstrate the generalizability of our objective, namely, object recognition, semantic segmentation and unsupervised monocular depth estimation. Note that the set of applications include both classification and regression tasks. Also, it has both supervised and unsupervised learning setups, and various archtectures such as fully convolutional networks, and encoder-decoder networks. We explain each of the tasks separately in the following subsections.

For all the experiments, the ADAM optimization algorithm is used with the learning rate of 0.10.1. The threshold θ\theta for the rate of saturation SS is set to 10−510^{-5} and ξ\xi value of 1010 is used. Validation fooling rate FtF_{t} is measured on the substitute dataset DD after every 200200 iterations only when the threshold of the rate of saturation is crossed. If it is not crossed, FtF_{t} is measured after every 400400 iterations. Note that the algorithm specific hyper-parameters are not changed across tasks or across priors.

For all the following experiments, perturbation crafted with different priors are denoted as PNP,PRP,andPDPP_{NP},P_{RP},\text{and}P_{DP} for the No prior, Range prior, and Data prior scenario respectively. To emphasize the effectiveness of the proposed objective, we present fooling rates obtained by a random baseline perturbation. Since our learned perturbation is norm limited by ξ\xi, we sample random δ\delta from U[−ξ,ξ]\mathcal{U}[-\xi,\xi] and compute the fooling rates. In all tables in this section, the ‘Baseline’ column refers to this perturbation.

We utilized models trained on ILSVRC and Places-205205 datasets, viz. CaffeNet , VGG-F , Googlenet , VGG-1616 , VGG-1919 , ResNet-152152 . For all experiments, pretrained models are used whose weights are kept frozen throughout the optimization process.Also, in contrast to UAP , we do not use training data in the data-free scenario (sec. III-A and sec. III-C1). However, as explained earlier, we use 1,0001,000 images randomly chosen from Pascal VOC-20122012 training images as validation set (DD in Algorithm 1) for our optimization. Also, in case of exploiting additional data prior (sec. III-C2), we use limited data from the corresponding training set. For the final evaluation on ILSVRC of the crafted UAPs, 50,00050,000 images from the validation set are used. Similarly, for Places-205205 dataset, 20,50020,500 images from the validation set are used.

Table I presents the fooling rates achieved by our objective on various network architectures. Fooling rate is the percentage of test images for which our crafted perturbation δ\delta successfully changed the predicted label. Using the terminology introduced in sec. III-F, fooling rate can also be written as GFR(Top1)GFR(Top1). Higher the fooling rate, greater is the perturbation’s ability to fool and lesser is the the classifier’s robustness. Fooling rates in Table I are obtained using the mean and dynamic range prior of the training distribution (sec. III-C1). Each row in the table indicates one target model employed in the learning process and the columns indicate various models attacked using the learned perturbations. The diagonal fooling rates indicate the white-box attacking, where all the information about the model is known to the attacker. The off-diagonal rates indicate black-box attacking, where no information about the model under attack is revealed to the attacker. However, the dataset over which both the models (target CNN and the CNN under attack) are trained is same. Our perturbations cause a mean white-box fooling rate of 69.24%69.24\% and a mean black-box fooling rate of 45.13%45.13\%. Given the data-free nature of the optimization, these fooling rates are alarmingly significant. The high fooling rates achieved by the proposed approach can adversely affect the real-world deploy-ability of these models.

Figure 2 shows example image-agnostic perturbations (δ)(\delta) crafted by the proposed method. Note that the perturbations look very different for each of the target CNNs. Interestingly, the perturbations corresponding to the VGG models look similar, which might be due to their architectural similarity. Figure 3 shows sample perturbed images (x+δ)(x+\delta) for VGG-19 from ILSVRC validation set. The top row shows the clean and bottom row shows the corresponding adversarial images. Note that the adversarially perturbed images are visually indistinguishable form their corresponding clean images. All the clean images shown in the figure are correctly classified and are successfully fooled by the added perturbation. Below each image, corresponding label predicted by the model is shown. Note that the correct labels are shown in black color and the wrong ones in red.

IV-A2 Exploiting the minimal prior

In this section, we present experimental results to demonstrate how our data-free objective can exploit the additional prior information about the target data distribution as discussed in section III-C. Note that we consider two cases: (i) providing the mean and dynamic range of the data samples, denoted as range prior (sec. III-C1), and (ii) utilizing minimal data samples themselves during the optimization, denoted as data prior ( III-C2).

Table II shows the fooling rates obtained with and without utilizing the prior information. Note that all the fooling rates are computed for white-box attacking scenario. For comparison, fooling rates obtained by our previous data-free objective and a data dependent objective are also presented. Important observations to draw from the table are listed below:

Utilizing the prior information consistently improves the fooling ability of the crafted perturbations.

A simple range prior can boost the fooling rates on an average by an absolute 10%10\%, while still being data-free.

Although the proposed objective is not designed to utilize the data, feeding the data samples results in an absolute 22%22\% rise in the fooling rates. Due to this increase in performance, for all models (except ResNet-152) our method becomes comparable or even better than UAP , which is designed especially to utilizes data.

IV-B Semantic segmentation

In this subsection, we demonstrate the effectiveness of GD-UAP objective to craft universal adversarial perturbations for semantic segmentation. We consider four network architectures. The first two architectures are from FCN : FCN-Alex, based on Alexnet , and FCN-8s-VGG, based on the 16-layer VGGNet . The last two architectures are 16-layer VGGNet based DL-VGG , and DL-RN101 , which is a multi-scale architecture based on ResNet-101 .

The FCN architetures are trained on Pascal VOC-20112011 dataset , consisting 9,6109,610 training samples and the the remaining two architectures are trained on Pascal VOC-20122012 dataset , consisting 10,58210,582 training samples. However, for testing our perturbation’s performance, we only use the validation set provided by , which consist of 736736 images.

Semantic segmentation is realized as assigning a label to each of the image pixels. That is, these models are typically trained to perform pixel level classification into one of 2121 categories (including the background) using the cross-entropy loss. Performance is commonly measured in terms of mean IOU (intersection over union) computed between the predicted map and the ground truth. Extending the UAP generation framework provided in to segmentation is a non-trivial task. However, our generalizable data-free algorithm can be applied for the task of semantic segmentation without any changes.

Similar to recognition setup, we present multiple scenarios for crafting the perturbations ranging from no data to utilizing data samples from the target distribution. An interesting observation with respect to the data samples from Pascal VOC-20122012, is that, in the 10,58210,582 training samples, 65.465.4% of the pixels belong to the ‘background’ category. Due to this, when we craft perturbation using training data samples as target distribution prior, our optimization process encounters roughly 6565% pixels belonging to ‘background‘ category, and only 3535% pixels belonging to the rest 2020 categories. As a result of this data imbalance, the perturbation is not sufficiently capable to corrupt the features of pixels belonging to the categories other than ‘background’. To handle this issue, we curate a smaller set of 2,8332,833 training samples from Pascal VOC-20122012, where each sample has less than 5050% pixels belonging to ‘background’ category. We denote this as “data w/ less BG”, and only 33.533.5% of pixels in this dataset belong to the ‘background’ category. The perturbations crafted using this dataset as target distribution prior show a higher capability to corrupt features of pixels belonging to the rest 2020 categories. Since mean IOU is the average of IOU accross the 2121 categories, we further observe that perturbations crafted using “data w/ less BG” cause a substantial reduction in the mean IOU measure as well.

Table III shows the generalized fooling rates with respect to the mean IOU (GFR(mIOU)GFR(mIOU)) obtained by GD-UAP perturbations under various data priors. As explained in section III-F, the generalized fooling rate measures the change in the performance of a network with respect to a given metric, which in our case is the mean IOU. Note that, similar to the recognition case, the fooling performance monotonically increases with the addition of data priors. This observation emphasizes that the proposed objective, though being an indirect, can rightly exploit the additional prior information about the training data distribution. Also, for all the models (Other than DL-RN101), “data w/ less BG” scenario results in the best fooling rate. This can be attributed to the fact that in “data w/ less BG” scenario we reduce the data-imbalance which in turn helps to craft perturbations that fool both background and object pixels.

In Table IV we present the mean IOU metric obtained on the perturbed images learned under various scenarios along with original mean IOU obtained on clean images. It is clearly observed that the random perturbation (the baseline) is not effective in fooling the segmentation models. However, the proposed objective crafts perturbations within the same range that can significantly fool the models. We also show the mean IOU obtained by Xie et al. , an image specific adversarial perturbation crafting work. Note that since GD-UAP is an image-agnostic approach, it is unfair to expect similar performance as . Further, the mean IOU shown by for DL-VGG and DL-RN101 models (bottom 22 rows of Table IV) denote the transfer performance, i.e., black-box attacking and hence show a smaller drop of the mean IOU from that of clean images. However, they are provide as an anchor point for evaluating image-agnostic perturbations generated using GD-UAP.

Figure 4 shows sample image-agnostic adversarial perturbations learned by our objective for semantic segmentation. In Figure 4, we show the perturbations learned with “data w/ less BG” prior for all the models. Similar to the recognition case, these perturbations look different across architectures. Figures 5 shows example image and predicted segmentation outputs by FCN-Alex model for perturbations crafted with various priors. Top row shows the clean and the perturbed images. Bottom row shows the predictions for the corresponding inputs. Further, the type of prior utilized to craft the perturbation is mentioned below the predictions. Crafted perturbations are clearly successful in misleading the model to predict inaccurate segmentation maps.

Figure 6 shows the effect of perturbation on multiple networks. It shows the output maps predicted by various models for the same input perturbed with corresponding δ\delta learned with “data w/ less BG” prior. It is interesting to note from Figure 6 that for the same image, with UAPs crafted using the same prior, different networks can have very different outputs, even if their outputs for clean images are very similar.

IV-C Depth estimation

Recent works such as show an increase in use of convolutional networks for regression-based computer vision task. A natural question to ask is whether they are as susceptible to universal adversarial attacks, as CNNs used for classification. In this section, by crafting UAPs for convolutional networks performing regression, we show that they are equally susceptible to universal adversarial attacks. To the best of our knowledge, we are the first to provide an algorithm for crafting universal adversarial attacks for convolutional networks performing regression,

Many recent works like perform depth estimation using convolutional network. In , the authors introduce Monodepth, an encoder-decoder architecture which regresses the depth of given monocular input image. We craft UAP using GD-UAP algorithm for the two variants of Monodepth, Monodepth-VGG and Monodepth-ResNet50. The network is trained using KITTI dataset . In its raw form, the dataset contains 42,38242,382 rectified stereo pairs from 6161 scenes, with a typical image being 1242×3751242\times 375 pixels in size. We show results on the eigen split, introduced in , which consist of 23,48823,488 images for training and validation, and 697697 images for test. We use the same crop size as suggested by the authors of and evaluate at the input image resolution.

As in the case of object recognition, UAPs crafted by the proposed method for monodepth also show the potential to exploit priors about the data distribution. We consider three cases, (i) providing no priors (PNPP_{NP}), (ii)range prior (PRPP_{RP}), and (ii) data prior (PDPP_{DP}). For providing data priors, we randomly pick 10,00010,000 image samples from the KITTI train dataset. To attain complete independence of target data, we perform validation on a set of 1,0001,000 randomly picked images from Places-205 dataset. The optimization procedure followed is the same as in the case of the previous two task.

Table V show the performance of Monodepth-Resnet50 and Monodepth-VGGunder the presence of the various UAPs crafted by the proposed method. As can be observed from the table, the crafted UAPs have a strong impact on the performance of the network. For both the variants of monodepth, UAPs crafted with range prior, bring down the accuracy with the threshold of 1.251.25 units (δ<1.25\delta<1.25) by 25.725.7% on an average. With data priors, the crafted UAPs are able to increase the Sq Rel (an error metric) to almost 1010 times the original performance. Under the impact of the crafted UAPs, the network’s performance drops below that of the depth-baseline (Train-mean), which uses the train set mean as the prediction for all image pixels. Figure 7 shows the input-output pair for Monodepth-VGG, where the input is perturbed by the various kinds of UAPs crafted.

Table VI shows the Generalized Fooling Rates (GFR) with respect δ<1.25\delta<1.25, i.e. GFR(δ<1.25)GFR(\delta<1.25). It is observed that PRPP_{RP} has higher GFR(δ<1.25)GFR(\delta<1.25) than PDPP_{DP}. This may appear as an anomaly as PDPP_{DP}, which has access to more information, should cause stronger harm to the network than PRPP_{RP}. This is indeed reflected in terms of multiple metrics such as Abs.Rel.ErrorAbs.Rel.Error, and RMSERMSE (ref. Table V). In fact, PDPP_{DP} is able to reduce these metrics even below the values achieved by the train-set mean. This clearly shows that PDPP_{DP} is indeed stronger than PRPP_{RP} (in terms of these metrics).

However, in terms of the other metrics, such as δ<1.25\delta<1.25 and δ<1.253\delta<1.25^{3} , it is observed that PRPP_{RP} causes more harm. These metrics measure the %\% of pixels where f(x)−G(x)f(x)-G(x) (where G(x)G(x) represents the ground truth depth at xx) is lesser than pre-defined limits. In contrast, metrics such as Abs.Rel.ErrorAbs.Rel.Error, and RMSERMSE measure the overall error of the output. Hence, based on the effect of PDPP_{DP} and PRPP_{RP} on these metrics, we infer that while PDPP_{DP} shifts fewer pixels than PRPP_{RP}, it severely shifts those pixels. In fact, as shown in Figure 7, it is often noticed that PDPP_{DP} causes the network to completely miss some nearby objects, and hence predicting very high depth at such locations, whereas PRPP_{RP} causes a anomalous estimation at a higher number of pixels.

The above situation shows that the ‘fooling’ performance of a perturbation can vary based on the metric used for analysis. Further, conclusions based on a single metric may only partially reflect the truth. This motivated us to propose GFR(m)GFR(m), a metric dependant measurement of ‘fooling’, which clearly indicates the metric dependence of ‘fooling’.

V GD-UAP: Analysis and Discussion

In this section, we provide additional analysis of GD-UAP on various fronts. First, we clearly highlight the multiple advantages of GD-UAP by comparing it with other approaches. In the next subsection we provide a thorough experimental evaluation of image-agnostic perturbation in the presence of various defence mechanism. Finally, we end this section with a discussion on how GD-UAP perturbations work.

First, we compare the effectiveness of GD-UAP against the existing data-free objective proposed in . Specifically, we compare maximizing the mean versus l2l_{2} norm (energy) of the activations caused by the perturbation δ\delta (or x/d+δx/d+\delta in case of exploiting the additional priors).

Table VII shows the comparison of fooling rates obtained with both the objectives (separately) in the improved optimization setup (III-D). We have chosen 33 representative models across various generations of models (CaffeNet, VGG and ResNet) to compare the effectiveness of the proposed objective. Note that the improved objective consistently outperforms the previous one by a significant 3.18%3.18\%. Similar behaviour is observed for other vision tasks also.

V-A2 Data dependent vs. Data-free objectives

Now, we demonstrate the necessity of data dependent objective to have samples from the target distribution only. That is, methods (such as ) that craft perturbations with fooling objective (i.e. move samples across the classification boundaries) require samples from only the training data distribution during the optimization. We show that crafting with arbitrary data samples leads to significantly inferior fooling performance.

Table VIII shows the fooling rates of data dependent objective when non-target data samples are utilized in place of target samples. Experiment in which we use Places-205205 data to craft perturbations for models trained on ILSVRC is denoted as Places-205205 →\rightarrow ILSVRC and vice versa. For both the setups, a set of 10,00010,000 training images are used. Note that, the rates for the proposed method are obtained without utilizing any data (with range prior) and rates for data-free scenario can be found in Table II. Clearly the fooling rates for UAP suffer significantly, as their perturbations are strongly tied to the target data. On the other hand, for the proposed method, since it does not craft via optimizing a fooling objective, the fooling performance does not decrease. Importantly, these experiments show that the data dependent objectives are not effective when samples from the target distribution are not available. This is a major drawback as it is difficult to procure the training data in practical scenarios.

Additionally, as the data dependent objectives rely on the available training data, the ability of the crafted perturbations heavily depends on the size of the available training data. We show that the fooling performance of UAP significantly decreases as the size of the available samples decreases. Figure 8 shows the fooling rates obtained by the perturbations crafted for multiple recognition models trained on ILSVRC by UAP with varying size of samples available for optimization. We craft different UAP perturbations (using the codes provided by the authors) utilizing only 500500, 1,0001,000, 2,0002,000, 4,0004,000 and 10,00010,000 data samples and evaluate their ability to fool various models. The performance of the crafted perturbations decreases drastically (shown in different shades of blue) as the available data samples are decreased during the optimization. For comparison, fooling rates obtained by the proposed data-free objective is shown in green.

V-A3 Capacity to expoit minimal priors

An interesting question to ask is whether the data dependant approach UAP can also utilize minimal data priors. To answer this, we craft perturbations using the algorithm presented in UAP, with different data priors (including no data priors). Table IX presents the comparison of GD-UAP against UAP with different priors. The numbers are computed on the ILSVRC validation set for Googlenet. We observe that the UAP algorithm does not even converge in the absence of data, failing to create a data-free UAP. When range prior (Gaussian noise in the data range) is used to craft, the resulting perturbations demonstrate a very low fooling rate of 10.5610.56 compared to 71.4471.44 of the proposed method. This is not surprising, as a significant decrease in the performance of UAP can be observed even due to a simple mismatch of training and target data (ref Table VIII). Finally, when actual data samples are used for crafting, UAP achieves 78.578.5 success rate which is closer to 83.5483.54 of the proposed method.

From these results, we infer that since UAP is a data dependent method, it requires corresponding data samples to craft effective perturbations. When noise samples are presented for optimization, the resulting perturbations fail to generalize to the actual data samples. In contrast, GD-UAP solves for an activation objective and exploits the prior as available.

V-B Robustness of UAPs Against Defense Mechanisms

While many recent works propose novel approaches to craft UAPs, the current literature does not contain a thorough analysis of UAPs in the presence of various defense mechanisms. If simple defense techniques could render them harmless, they may not present severe threat to deployment of deep models. In this subsection, we investigate the strength of GD-UAP against various defense techniques. Particularly, we consider (1) Input Transformation Defenses, such as Gaussian blurring, Image quilting, etc., (2) Targeted Defense Against UAPs, such as Perturbation Rectifying Networks (PRN) and (3) Defense through robust architectural design such as scattering networks .

We evaluate the performance of UAPs generated from GD-UAP, as well as data-dependent approach UAP against various defense mechanisms.

For defense by input transformations, inline with and , we consider the following simple defenses: (1) 10-Crop Evaluation, (2) Gaussian Blurring, (3) Median Smoothing, (4) Bilateral Filtering (5) JPEG Compression, and (6) Bit-Depth Reduction. Further, we evaluate two sophisticated image transformations proposed in , namely, (7) TV-Minimization and (8) Image Quilting (Using code provided by the authors).

Table X presents our experimental evaluation of the defenses on the Googlenet. As we can see, while the fooling rate of all UAPs is reduced by the defenses (significantly in some cases), it is achieved at the cost of model’s accuracy on clean images. If Image Quilting, or TV-normalization is used as a defense, it is essential to train the network on quilted images, without which, a severe drop in accuracy is observed. However, in majority of the defenses, UAPs flip labels for more than 45%45\% of the images, which indicates threat to deployment. Further, as Bilateral Filtering and JPEG compression show strong defense capability at low cost to accuracy, we evaluate the performance of our GD-UAP perturbations (with range prior) on 66 classification networks in the presence of these two defenses. This is shown in Table XI. We note that when the defense mechanism significantly lowers the fooling rates, a huge price is paid in terms of % drop in Top-1 Accuracy (DAcc)(D_{Acc}), which is unacceptable. This further indicates the poor fit of input transformations as a viable defense.

V-B2 Targeted Defense Against UAPs

Now, we turn towards defenses which have been specifically engineered for UAPs. We evaluate the performance of the various UAPs, with the Perturbation Rectifying Network(PRN) as a defense. PRN is trained using multiple UAPs generated from the algorithm presented in UAP to rectify perturbed images. In Table X, we present the fooling rates obtained on Googlenet using various perturbations, using the codes provided by the authors of . While PRN is able to defend against , it shows very poor performance against our data prior perturbation. This is due to the fact that PRN is trained using UAPs generated from UAP only, indicating that PRN lacks generalizability to input-agnostic perturbations generated from other approaches. Furthermore, it is also observed that various simple input transformation defenses outperform PRN. Hence, while PRN is an important step towards defenses against UAPs, in its current form, it provides scarce security against UAPs from methods it is not trained on.

V-B3 Defense through robust architectural design

As our optimization process relies on maximizing ∣∣f(x+δ)∣∣2||f(x+\delta)||_{2} by increasing Π∣∣li(x+δ)∣∣2∀i∈{1,2,...}\Pi||l_{i}(x+\delta)||_{2}\forall i\in\{1,2,...\} (where li(.)l_{i}(.) represents the input to layer li+1l_{i+1}), one defense against our attack can be to train the network such that the change in the output to a layer lil_{i} minimally effects the output of the network f(.)f(.). One method for achieving this target can be to minimize the Lipschitz Constant KiK_{i} of each linear transformation in the network. As KiK_{i} controls the upper bound of the value of ∣∣f(x+δ)−f(x)∣∣2||f(x+\delta)-f(x)||_{2} with respect to ∣∣li(x+δ)−li(x)∣∣2||l_{i}(x+\delta)-l_{i}(x)||_{2}, minimizing Ki∀i∈{1,2,...}K_{i}\forall i\in\{1,2,...\} can lead to a stable system where minor variation in input layer do not translate to high variation in output. This would translate to low fooling rate when attacked by adversarial perturbations.

In , Bruna et al. introduce scattering network. This network consist of scattering transform, which linearize the output deformation with respect to small input deformation, ensuring that the Lipschitz constant is ≤1\leq 1. In , a hybrid approach was proposed, which uses scattering transforms in the initial layers, and learn convolutional layers (Res-blocks, in specific) on the transformed output of the scattering layers, making scattering transform based approaches feasible for Imagenet. The proposed Hybrid-network gave performance comparable to ResNet-18 and VGG-13 while containing much lesser layers.

We now evaluate the fooling rate that GD-UAP perturbations achieve in the Hybrid-Network, and compare it to fooling rate achieved on VGG-13 and ResNet-18 networks. In the Hybrid-network, ideally we would like to maximize ∣∣li(x+δ)∣∣2||l_{i}(x+\delta)||_{2} at each of the Res-Block output. However, as ∂S(x)/∂x\partial S(x)/\partial x, where S(.)S(.) represents the Scattering Transform, is non-trivial, we only perform black-box attacks on all the networks.

Table XII shows the results of the black-box attack on Hybrid Networks. Though Hybrid networks on an average decrease the fooling rate by 13%13\% when compared to other models, they still remain vulnerable.

V-C Analyzing how GD-UAP works

As demonstrated by our experimentation in section IV, it is evident that GD-UAP is able to craft highly effective perturbations for a variety of computer vision tasks. This highlights an important question about the Convolutional Neural Networks (CNNs):“How stable are the learned representations at each layer?” That is, Are the features learned by the CNNs robust to small changes in the input? As mentioned earlier ‘fooling’ refers to instability of the CNN in terms of its output, independent of the task at hand.

We depict this as a stability issue with the learned representations by CNNs. We attempt to learn the optimal perturbation in the input space that can cause maximal change in the output of the network. We achieve this via learning perturbations that can result in maximal change in the activations at all the intermediate layers of the architecture. As an example, we consider VGG-1616 CNN trained for object recognition to illustrate the working of our objective. Figure 9 shows the percentage relative change in the feature activations (∥li(x+δ)−li(x)∥2×100∥li(x)∥2)(\frac{\|l_{i}(x+\delta)-l_{i}(x)\|_{2}\times 100}{\|l_{i}(x)\|_{2}}) at various layers in the architecture. The percentage relative change in the feature activations due to the addition of the learned perturbation increases monotonically as we go deeper in the network. Because of this accumulated perturbation in the projection of the input, our learned adversaries are able to fool the CNNs independent of the task at hand. This phenomenon explains the fooling achieved by our objective. We can also observe that with the utilization of the data priors, the relative perturbation further increases which results in better fooling when the prior information is provided during the learning.

Our data-free approach, consists of increasing ∣∣li(δ)∣∣||l_{i}(\delta)|| across the layers, to increase f(δ)f(\delta). As Figure 9 shows, this crafts a perturbation, which leads to an increase in f(x+δ)f(x+\delta). One may intrepret this increase to be caused due to the locally linear nature of CNNs. i.e., f(x+δ)≈f(x)+f(δ)f(x+\delta)\approx f(x)+f(\delta). However, our experiments reveal that the feature extractor ff might not be locally linear. We observe the relation between the quantities ∣∣f(x+δ)−f(x)∣∣2||f(x+\delta)-f(x)||_{2} and ∣∣f(δ)∣∣2||f(\delta)||_{2}, where ∣∣.∣∣2||.||_{2} represents the L2L_{2}-norm, and f(.)f(.) represents the output of the last convolution layer of the network. Figure 10 presents the comparison of these two quantities for VGG-16 computed during the proposed optimization with no prior.

From the observations, we infer: (1) ∣∣f(x+δ)−f(x)∣∣2||f(x+\delta)-f(x)||_{2} is not approximately equal to ∣∣f(δ)∣∣2||f(\delta)||_{2}, and hence, ff is not observed to be locally linear, (2) However, ∣∣f(x+δ)−f(x)∣∣2||f(x+\delta)-f(x)||_{2} is strongly correlated to ∣∣f(δ)∣∣2||f(\delta)||_{2}, and our data-free optimization approach exploits this correlation between the two quantities. To summarize, our data-free optimization exploits the correlation between the quantities, ∣∣f(x+δ)−f(x)∣∣2||f(x+\delta)-f(x)||_{2} and ∣∣f(δ)∣∣2||f(\delta)||_{2}, rather than the local-linearity of feature extractor ff.

Finally, the relative change caused by our perturbations at the input to classification layer (fc8fc_{8} or softmax) can be clearly related to the fooling rates achieved for various perturbations. Table XIII shows the relative shift in the feature activations (∥li(x+δ)−li(x)∥2∥li(x)∥2)(\frac{\|l_{i}(x+\delta)-l_{i}(x)\|_{2}}{\|l_{i}(x)\|_{2}}) that are input to the classification layer and the corresponding fooling rates for various perturbations. Note that they are highly correlated, which explains why the proposed objective can fool the CNNs trained across multiple vision tasks.

VI Conclusion

In this paper, we have proposed a novel data-free objective to craft image-agnostic (universal) adversarial perturbations (UAP). More importantly, we show that the proposed objective is generalizable not only across multiple CNN architectures but also across diverse computer vision tasks. We demonstrated that our seemingly simple objective of injecting maximal “adversarial” energy into the learned representations (subject to the imperceptibility constraint) is effective to fool both the classification and regression models. Significant transfer performances achieved by our crafted perturbations can pose substantial threat to the deep learned systems in terms of black-box attacking.

Further, we show that our objective can exploit minimal priors about the target data distribution to craft stronger perturbations. For example, providing simple information such as the mean and dynamic range of the images to the proposed objective would craft significantly stronger perturbations. Though the proposed objective is data-free in nature, it can craft stronger perturbations when data is utilized.

More importantly, we introduced the idea of generalizable objectives to craft image-agnostic perturbations. It is already established that the representations learned by deep models are susceptible. On top of it, the existence of generic objectives to fool “any” learning based vision model independent of the underlying task can pose critical concerns about the model deployment. Therefore, it is an important research direction to be focused on in order to build reliable machine learning based systems.

References