Fooling Neural Network Interpretations via Adversarial Model Manipulation

Juyeon Heo, Sunghwan Joo, Taesup Moon

Introduction

As deep neural networks have made a huge impact on real-world applications with predictive tasks, much emphasis has been set upon the interpretation methods that can explain the ground of the predictions of the complex neural network models. Furthermore, accurate explanations can further improve the model by helping researchers to debug the model or revealing the existence of unintended bias or effects in the model . To that regard, research on the interpretability framework has become very active recently, for example, , to name a few. Paralleling above flourishing results, research on sanity checking and identifying the potential problems of the proposed interpretation methods has also been actively pursued recently. For example, some recent research showed that many popular interpretation methods are not stable with respect to the perturbation or the adversarial attacks on the input data.

In this paper, we also discover the instability of the neural network interpretation methods, but with a fresh perspective. Namely, we ask whether the interpretation methods are stable with respect to the adversarial model manipulation, which we define as a model fine-tuning step that aims to dramatically alter the interpretation results without significantly hurting the accuracy of the original model. In results, we show that the state-of-the-art interpretation methods are vulnerable to those manipulations. Note this notion of stability is clearly different from that considered in the above mentioned works, which deal with the stability with respect to the perturbation or attack on the input to the model. To the best of our knowledge, research on this type of stability has not been explored before. We believe that such stability would become an increasingly important criterion to check, since the incentives to fool the interpretation methods via model manipulation will only increase due to the widespread adoption of the complex neural network models.

For a more concrete motivation on this topic, consider the following example. Suppose a neural network model is to be deployed in an income prediction system. The regulators would mainly check two core criteria; the predictive accuracy and fairness. While the first can be easily verified with a holdout validation set, the second is more tricky since one needs to check whether the model contains any unfair bias, e.g., using race as an important factor for the prediction. The interpretation method would obviously become an important tool for checking this second criterion. However, suppose a lazy developer finds out that his model contains some bias, and, rather than actually fixing the model to remove the bias, he decides to manipulate the model such that the interpretation can be fooled and hide the bias, without any significant change in accuracy. (See Figure 1(a) for more details.) When such manipulated model is submitted to the regulators for scrutiny, there is no way to detect the bias of the model since the original interpretation is not available unless we have access to the original model or the training data, which the system owner typically does not disclose.

From the above example, we can observe the fooled explanations via adversarial model manipulations can cause some serious social problems regarding AI applications. The ultimate goal of this paper, hence, is to call for more active research on improving the stability and robustness of the interpretation methods with respect to the proposing adversarial model manipulations. The following summarizes the main contributions of this paper:

We first considered the notion of stability of neural network interpretation methods with respect to the proposing adversarial model manipulation.

We demonstrate that the representative saliency map based interpreters, i.e., LRP , Grad-CAM , and SimpleGradient , are vulnerable to our model manipulation, where the accuracy drops are around 2% and 1% for Top-1 and Top-5 accuracy on the ImageNet validation set, respectively. Figure 1(b) shows a concrete example of our fooling.

We show the fooled explanation generalizes to the entire validation set, indicating that the interpretations are truly fooled, not just for some specific inputs, in contrast to .

We demonstrate that the transferability exists in our fooling, e.g., if we manipulate the model to fool LRP, then the interpretations of Grad-CAM and Simple Gradient also get fooled, etc.

Related Work

Interpretation methods Various interpretability frameworks have been proposed, and they can be broadly categorized into two groups: black-box methods and gradient/saliency map based methods . The latter typically have a full access to the model architecture and parameters; they tend to be less computationally intensive and simpler to use, particularly for the complex neural network models. In this paper, we focus on the gradient/saliency map based methods and check whether three state-of-the-art methods can be fooled with adversarial model manipulation.

Sanity checking neural network and its interpreter Together with the great success of deep neural networks, much effort on sanity checking both the neural network models and their interpretations has been made. They mainly examine the stability of the model prediction or the interpretation for the prediction by either perturbing the input data or model, inspired by adversarial attacks . For example, showed that several interpretation results are significantly impacted by a simple constant shift in the input data. recently developed a more robust method, dubbed as a self-explaining neural network, by taking the stability (with respect to the input perturbation) into account during the model training procedure. has adopted the framework of adversarial attack for fooling the interpretation method with a slight input perturbation. tries to find perturbed data with similar interpretations of benign data to make it hard to be detected with interpretations. A different angle of checking the stability of the interpretation methods has been also given by , which developed simple tests for checking the stability (or variability) of the interpretation methods with respect to model parameter or training label randomization. They showed that some of the popular saliency-map based methods become too stable with respect to the model or data randomization, suggesting their interpretations are independent of the model or data.

Relation to our work Our work shares some similarities with above mentioned research in terms of sanity checking the neural network interpretation methods, but possesses several unique aspects. Firstly, unlike , which attack each given input image, we change the model parameters via fine-tuning a pre-trained model, and do not perturb the input data. Due to this difference, our adversarial model manipulation makes the fooling of the interpretations generalize to the entire validation data. Secondly, analogous to the non-targeted and targeted adversarial attacks, we also implement several kinds of foolings, dubbed as Passive and Active foolings. Distinct from , we generate not only uninformative interpretations, but also totally wrong ones that point unrelated object within the image. Thirdly, as , we also take the explanation into account for model training, but while they define a special structure of neural networks, we do usual back-propagation to update the parameters of the given pre-trained model. Finally, we note also measures the stability of interpretation methods, but, the difference is that our adversarial perturbation maintains the accuracy of the model while only focuses on the variability of the explanations. We find that an interpretation method that passed the sanity checks in , e.g., Grad-CAM, also can be fooled under our setting, which calls for more solid standard for checking the reliability of interpreters.

Adversarial Model Manipulation

We briefly review the saliency map based interpretation methods we consider. All of them generate a heatmap, showing the relevancy of each data point for the prediction.

Layer-wise Relevance Propagation (LRP) is a principled method that applies relevance propagation, which operates similarly as the back-propagation, and generates a heatmap that shows the relevance value of each pixel. The values can be both positive and negative, denoting how much a pixel is helpful or harmful for predicting the class cc. In the subsequent works, LRP-Composite , which applies the basic LRP-ϵ\epsilon for the fully-connected layer and LRP-αβ\alpha\beta for the convolutional layer, has been proposed. We applied LRP-Composite in all of our experiments.

Grad-CAM is also a generic interpretation method that combines gradient information with class activation maps to visualize the importance of each input. It is mainly used for CNN-based models for vision applications. Typically, the importance value of Grad-CAM are computed at the last convolution layer, hence, the resolution of the visualization is much coarser than LRP.

SimpleGrad (SimpleG) visualizes the gradients of prediction score with respect to the input as a heatmap. It indicates how sensitive the prediction score is with respect to the small changes of input pixel, but in , it is shown to generate noisier saliency maps than LRP.

2 Objective function and penalty terms

Our proposed adversarial model manipulation is realized by fine-tuning a pre-trained model with the objective function that combines the ordinary classification loss with a penalty term that involves the interpretation results. To that end, our overall objective function for a neural network w\bm{w} to minimize for training data \pazocalD\pazocal{D} with the interpretation method \pazocalI\pazocal{I} is defined to be

in which \pazocalLC(⋅)\pazocal{L}_{C}(\cdot) is the ordinary cross-entropy classification loss on the training data, w0\bm{w}_{0} is the parameter of the original pre-trained model, \pazocalL\pazocalF\pazocalI(⋅)\pazocal{L}_{\pazocal{F}}^{\pazocal{I}}(\cdot) is the penalty term on \pazocalDfool\pazocal{D}_{fool}, which is a potentially smaller set than \pazocalD\pazocal{D}, that is the dataset used in the penalty term, and λ\lambda is a trade-off parameter. Depending on how we define \pazocalL\pazocalFI(⋅)\pazocal{L}_{\pazocal{F}}{I}(\cdot), we categorize two types of fooling in the following subsections.

We define Passive fooling as making the interpretation methods generate uninformative explanations. Three such schemes are defined with different \pazocalL\pazocalF\pazocalI(⋅)\pazocal{L}_{\pazocal{F}}^{\pazocal{I}}(\cdot)’s: Location, Top-kk, and Center-mass foolings.

Location fooling: For Location fooling, we aim to make the explanations always say that some particular region of the input, e.g., boundary or corner of the image, is important regardless of the input. We implement this kind of fooling by defining the penalty term in (2) equals

Top-k\bm{k} fooling: In Top-kk fooling, we aim to reduce the interpretation scores of the pixels that originally had the top kk% highest values. The penalty term then becomes

in which \pazocalDfool=\pazocalD\pazocal{D}_{fool}=\pazocal{D}, and \pazocalPi,k(w0)\pazocal{P}_{i,k}(\bm{w}_{0}) is the set of pixels that had the top kk% highest heatmap values for the original model w0\bm{w}_{0}, for the ii-th data point.

Center-mass fooling: As in , the Center-mass loss aims to deviate the center of mass of the heatmap as much as possible from the original one. The center of mass of a one-dimensional heatmap can be denoted as C(\mathbf{h}_{y_{i}}^{\pazocal{I}}(\bm{w}))=\big{(}\sum_{j=1}^{d{I}}j\cdot h^{\pazocal{I}}_{y_{i},j}(\bm{w})\big{)}/\sum_{j=1}^{d{I}}h^{\pazocal{I}}_{y_{i},j}(\bm{w}), in which index jj is treated as a location vector, and it can be easily extended to higher dimensions as well. Then, with \pazocalDfool=\pazocalD\pazocal{D}_{fool}=\pazocal{D} and ∥⋅∥1\|\cdot\|_{1} being the L1L_{1} norm, the penalty term for the Center-mass fooling is defined as

2.2 Active fooling

Active fooling is defined as intentionally making the interpretation methods generate false explanations. Although the notion of false explanation could be broad, we focused on swapping the explanations between two target classes. Namely, let c1c_{1} and c2c_{2} denote the two classes of interest and define \pazocalDfool\pazocal{D}_{fool} as a dataset (possibly without target labels) that specifically contains both class objects in each image. Then, the penalty term \pazocalL\pazocalFI(\pazocalDfool;w,w0)\pazocal{L}_{\pazocal{F}}{I}(\pazocal{D}_{fool};\bm{w},\bm{w}_{0}) equals

in which the first term makes the explanation for c1c_{1} alter to that of c2c_{2}, and the second term does the opposite. A subtle point here is that unlike in Passive foolings, we use two different datasets for computing \pazocalLC(⋅)\pazocal{L}_{C}(\cdot) and \pazocalL\pazocalFI(⋅)\pazocal{L}_{\pazocal{F}}{I}(\cdot), respectively, to make a focused training on c1c_{1} and c2c_{2} for fooling. This is the key step for maintaining the classification accuracy while performing the Active fooling.

Experimental Results

For all our fooling methods, we used the ImageNet training set as our \pazocalD\pazocal{D} and took three pre-trained models, VGG19 , ResNet50 , and DenseNet121 , for carrying out the foolings. For the Active fooling, we additionally constructed \pazocalDfool\pazocal{D}_{fool} with images that contain two classes, {c1=\{c_{1}=“African Elephant”, c2=c_{2}=“Firetruck”}\}, by constructing each image by concatenating two images from each class in the 2×22\times 2 block. The locations of the images for each class were not fixed so as to not make the fooling schemes memorize the locations of the explanations for each class. An example of such images is shown in the top-left corner of Figure 3. More implementation details are in Appendix B.

2 Fooling Success Rate (FSR): A quantitative metric

In this section, we suggest a quantitative metric for each fooling method, Fooling Success Rate (FSR), which measures how much an interpretation method \pazocalI\pazocal{I} is fooled by the model manipulation. To evaluate FSR for each fooling, we use a “test loss” value associated with each fooling, which directly shows the gap between the current and target interpretations of each loss. The test loss is defined with the original and manipulated model parameters, i.e., w0\bm{w}_{0} and wfool∗\bm{w}^{*}_{\text{fool}}, respectively, and the interpreter \pazocalI\pazocal{I} on each data point in the validation set \pazocalDval\pazocal{D}_{\text{val}}; we denote the test loss for the ii-th data point (xi,yi)∈\pazocalDval(\mathbf{x}_{i},y_{i})\in\pazocal{D}_{\text{val}} as ti(wfool∗,w0,\pazocalI)t_{i}(\bm{w}^{*}_{\text{fool}},\bm{w}_{0},\pazocal{I}).

For the Location and Top-kk foolings, the ti(wfool∗,w0,\pazocalI)t_{i}(\bm{w}^{*}_{\text{fool}},\bm{w}_{0},\pazocal{I}) is computed by evaluating (3) and (4) for a single data point (xi,yi)(\mathbf{x}_{i},y_{i}) and (wfool∗,w0)(\bm{w}^{*}_{\text{fool}},\bm{w}_{0}). For Center-mass fooling, we evaluate (5), again for a single data point (xi,yi)(\mathbf{x}_{i},y_{i}) and (wfool∗,w0)(\bm{w}^{*}_{\text{fool}},\bm{w}_{0}), and normalize it with the length of diagonal of the image to define as ti(wfool∗,w0,\pazocalI)t_{i}(\bm{w}^{*}_{\text{fool}},\bm{w}_{0},\pazocal{I}). For Active fooling, we first define si(c,c′)=rs(hc\pazocalI(wfool∗),hc′\pazocalI(w0))s_{i}(c,c^{\prime})=r_{s}(\mathbf{h}^{\pazocal{I}}_{c}(\bm{w}^{*}_{\text{fool}}),\mathbf{h}^{\pazocal{I}}_{c^{\prime}}(\bm{w}_{0})) as the Spearman rank correlation between the two heatmaps for xi\mathbf{x}_{i}, generated with \pazocalI\pazocal{I}. Intuitively, it measures how close the explanation for class cc from the fooled model is from the explanation for class c′c^{\prime} from the original model. Then, we define ti(wfool∗,w0,\pazocalI)=si(c1,c2)−si(c1,c1)t_{i}(\bm{w}^{*}_{\text{fool}},\bm{w}_{0},\pazocal{I})=s_{i}(c_{1},c_{2})-s_{i}(c_{1},c_{1}) as the test loss for fooling the explanation of c1c_{1} and ti(wfool∗,w0,\pazocalI)=si(c2,c1)−si(c2,c2)t_{i}(\bm{w}^{*}_{\text{fool}},\bm{w}_{0},\pazocal{I})=s_{i}(c_{2},c_{1})-s_{i}(c_{2},c_{2}) for c2c_{2}. With above test losses, the FSR for a fooling method ff and an interpreter \pazocalI\pazocal{I} is defined as

in which 1{⋅}\bm{1}\{\cdot\} is an indicator function and RfR_{f} is a pre-defined interval for each fooling method. Namely, RfR_{f} is a threshold for determining whether the interpretations are successfully fooled or not. We empirically defined RfR_{f} as [0,0.2],[0,0.3],[0.1,1][0,0.2],[0,0.3],[0.1,1], and [0.5,2][0.5,2] for Location, Top-kk, Center-mass, and Active fooling, respectively. (More details of deciding thresholds are in Appendix C) In short, the higher the FSR metric is, the more successful ff is for the interpreter \pazocalI\pazocal{I}.

3 Passive and Active fooling results

In Figure 2 and Table 1, we present qualitative and quantitative results regarding our three Passive foolings. The followings are our observations. For the Location fooling, we clearly see that the explanations are altered to stress the uninformative frames of each image even if the object is located in the center, compare (1,5)(1,5) and (3,5)(3,5) in Figure 2 for example (a,b)(a,b) denotes the image at the aa-th row and bb-th column of the figure.. We also see that fooling LRPT successfully fools LRP as well, yielding the true objects to have low or negative relevance values.

For the Top-kk fooling, we observe the most highlighted top kk% pixels are significantly altered after the fooling, by comparing the big difference between the original explanations and those in the green colored box in Figure 2. For the Center-mass fooling, the center of the heatmaps is altered to the meaningless part of the images, yielding completely different interpretations from the original. Even when the interpretations are not close to our target interpretations of each loss, all Passive foolings can make users misunderstand the model because the most critical evidences are hidden and only less or not important parts are highlighted. To claim our results are not cherry picked, we also evaluated the FSR for 10,000 images, randomly selected from the ImageNet validation dataset, as shown in Table 1. We can observe that all FSRs of fooling methods are higher than 50% for the matched cases (bold underlined), except for the Location fooling with LRPT for DenseNet121.

Next, for the Active fooling, from the qualitative results in Figure 3 and the quantitative results in Table 2, we find that the explanations for c1c_{1} and c2c_{2} are swapped clearly in VGG19 and nearly in ResNet50, but not in DenseNet121, suggesting the relationship between the model complexity and the degree of Active fooling. When the interpretations are clearly swapped, as in (1,3)(1,3) and (2,3)(2,3) of Figure 3, the interpretations for c1c_{1} (the true class) turn out to have negative values on the correct object, while having positive values on the objects of c2c_{2}. Even when the interpretations are not completely swapped, they tend to spread out to both c1c_{1} and c2c_{2} objects, which becomes less informative; compare between the (1,8)(1,8) and (2,8)(2,8) images in Figure 3, for example. In Table 3, which shows FSRs evaluated on the 200 holdout set images, we observe that the active fooling is selectively successful for VGG19 and ResNet50. For the case of DenseNet121, however, the active fooling seems to be hard as the FSR values are almost 0. Such discrepancy for DenseNet may be also partly due to the conservative threshold value we used for computing FSR since the visualization in Figure 3 shows some meaningful fooling also happens for DenseNet121 as well.

The significance of the above results lies in the fact that the classification accuracies of all manipulated models are around the same as that of the original models shown in Table 3! For the Active fooling, in particular, we also checked that the slight decrease in Top-5 accuracy is not just concentrated on the data points for the c1c_{1} and c2c_{2} classes, but is spread out to the whole 1000 classes. Such analysis is in Appendix D. Note our model manipulation affects the entire validation set without any access to it, unlike the common adversarial attack which has access to each input data point .

Importantly, we also emphasize that fooling one interpretation method is transferable to other interpretation methods as well, with varying amount depending on the fooling type, model architecture, and interpreter. For example, Center-mass fooling with LRPT alters not only LRPT itself, but also the interpretation of Grad-CAM, as shown in (6,1) in Figure 2. The Top-kk fooling and VGG19 seem to have larger transferability than others. More discussion on the transferability is elaborated in Section 5. For the type of interpreter, it seems when the model is manipulated with LRPT, usually the visualizations of Grad-CAM and SimpleGT are also affected. However, when fooling is done with Grad-CAM, LRPT and SimpleGT are less impacted.

Discussion and Conclusion

In this section, we give several important further discussions on our method. Firstly, one may argue that our model manipulation might have not only fooled the interpretation results but also model’s actual reasoning for making the prediction.

To that regard, we employ Area Over Prediction Curve (AOPC) , a principled way of quantitatively evaluating the validity of neural network interpretations, to check whether the manipulated model also has been significantly altered by fooling the interpretation. Figure 4(a) shows the average AOPC curves on 10K validation images for the original and manipulated DenseNet121 (Top-kk fooled with Grad-CAM) models, wo\bm{w}_{o} and wfool∗\bm{w}^{*}_{\text{fool}}, with three different perturbation orders; i.e., with respect to hc\pazocalI(wo)\mathbf{h}_{c}^{\pazocal{I}}(\bm{w}_{o}) scores, hc\pazocalI(wfool∗)\mathbf{h}_{c}^{\pazocal{I}}(\bm{w}^{*}_{\text{fool}}) scores, and a random order. From the figure, we observe that wo(hc\pazocalI(wo))\bm{w}_{o}(\mathbf{h}_{c}^{\pazocal{I}}(\bm{w}_{o})) and wfool∗(hc\pazocalI(wo))\bm{w}^{*}_{\text{fool}}(\mathbf{h}_{c}^{\pazocal{I}}(\bm{w}_{o})) show almost identical AOPC curves, which suggests that wfool∗\bm{w}^{*}_{\text{fool}} has not changed much from wo\bm{w}_{o} and is making its prediction by focusing on similar parts that wo\bm{w}_{o} bases its prediction, namely, hc\pazocalI(wo)\mathbf{h}_{c}^{\pazocal{I}}(\bm{w}_{o}). In contrast, the AOPC curves of both wo(hc\pazocalI(wfool∗))\bm{w}_{o}(\mathbf{h}_{c}^{\pazocal{I}}(\bm{w}^{*}_{\text{fool}})) and wfool∗(hc\pazocalI(wfool∗))\bm{w}^{*}_{\text{fool}}(\mathbf{h}_{c}^{\pazocal{I}}(\bm{w}^{*}_{\text{fool}})) lie significantly lower, even lower than the case of random perturbation. From this result, we can deduce that hc\pazocalI(wfool∗)\mathbf{h}_{c}^{\pazocal{I}}(\bm{w}^{*}_{\text{fool}}) is highlighting parts that are less helpful than random pixels for making predictions, hence, is a “wrong” interpretation.

Secondly, one may ask whether our fooling can be easily detected or undone. Since it is known that the adversarial input example can be detected by adding small Gaussian perturbation to the input , one may also suspect that adding small Gaussian noise to the model parameters might reveal our fooling. However, Figure 4(b) shows that wo\bm{w}_{o} and wfool∗\bm{w}^{*}_{\text{fool}} (ResNet50, Location-fooled with LRPT) behave very similarly in terms of Top-1 accuracy on ImageNet validation as we increase the noise level of the Gaussian perturbation, and FSRs do not change radically, either. Hence, we claim that detecting or undoing our fooling would not be simple.

Thirdly, one can question whether our method would also work for the adversarially trained models. To that end, Figure 4(c) shows the Top-1 accuracy of ResNet50 model on \pazocalDval\pazocal{D}_{val} (i.e., ImageNet validation) and PGD(\pazocalDval\pazocal{D}_{val}) (i.e., the PGD-attacked \pazocalDval\pazocal{D}_{val}), and demonstrates that adversarially trained model can be also manipulated by our method. Namely, starting from a pre-trained wo\bm{w}_{o} (dashed red), we do the “free” adversarial training (ϵ=1.5\epsilon=1.5) to obtain wadv\bm{w}_{\text{adv}} (dashed green), then started our model manipulation with (Location fooling, Grad-CAM) while keeping the adversarial training. Note the Top-1 accuracy on \pazocalDval\pazocal{D}_{val} drops while that on PGD(\pazocalDval)(\pazocal{D}_{val}) increases during the adversarial training phase (from red to green) as expected, and they are maintained during our model manipulation phase (e.g. dashed blue). The right panel shows the Grad-CAM interpretations at three distinct phases (see the color-coded boundaries), and we clearly see the success of the Location fooling (blue, third row).

Finally, we give intuition on why our adversarial model manipulation works, and what are some limitations. Note first that all the interpretation methods we employ are related to some form of gradients; SimpleG uses the gradient of the input, Grad-CAM is a function of the gradient of the representation at a certain layer, and LRP turns out to be similar to gradient times inputs . Motivated by , Figure 5(a) illustrates the point that the same test data can be classified with almost the same accuracy but with different decision boundaries that result in radically different gradients, or interpretations. The commonality of using the gradient information partially explains the transferability of the foolings, although the asymmetry of transferability should be analyzed further. Furthermore, the level of fooling seems to have intriguing connection with the model complexity, similar to the finding in in the context of input adversarial attack. As a hint for developing more robust interpretation methods, Figure 5(b) shows our results on fooling SmoothGrad , which integrates SimpleG maps obtained from multiple Gaussian noise added inputs. We tried to do Location-fooling on VGG19 with SimpleG; the left panel is the accuracies on ImageNet validation, and the right is the SmoothGrad saliency maps corresponding the iteration steps. Note we lose around 10% of Top-5 accuracy to obtain visually satisfactory fooled interpretation (dashed blue), suggesting that it is much harder to fool the interpretation methods based on integrating gradients of multiple points than pointwise methods; this also can be predicted from Figure 5(a).

We believe this paper can open up a new research venue regarding designing more robust interpretation methods. We argue checking the robustness of interpretation methods with respect to our adversarial model manipulation should be an indispensable criterion for the interpreters in addition to the sanity checks proposed in ; note Grad-CAM passes their checks. Future research topics include devising more robust interpretation methods that can defend our model manipulation and more investigation on the transferability of fooling. Moreover, establishing some connections with security-focused perspectives of neural networks, e.g., , would be another fruitful direction to pursue.

Acknowledgements

This work is supported in part by ICT R&D Program [No. 2016-0-00563, Research on adaptive machine learning technology development for intelligent autonomous digital companion][No. 2019-0-01396, Development of framework for analyzing, detecting, mitigating of bias in AI model and training data], AI Graduate School Support Program [No.2019-0-00421], and ITRC Support Program [IITP-2019-2018-0-01798] of MSIT / IITP of the Korean government, and by the KIST Institutional Program [No. 2E29330].

References

Appendix A Back-propagation for fooling

This section describes a flow of forward and backward pass for (2). To do this, we consider a computational graph for a neural network with LL layers, as shown in the Figure 6. This graph is common for both Layer-wise Relevance Propagation (LRP) and Grad-CAM , and it can be applied to the other saliency-map based interpretation methods, such as SimpleGradient .

Appendix B Experiment details

The total number of images in \pazocalDfool\pazocal{D}_{fool} was 1,3001,300, and 1,1001,100 of them were used for the training set for our fine-tuning and the rest for validation. For measuring the classification accuracy of the models, we used the entire validation set of ImageNet, which consists of 50,00050,000 images. To measure FSR, we again used the ImageNet validation set for the Passive foolings and 200 hold-out images in \pazocalDfool\pazocal{D}_{fool} for Active fooling. We denote the validation set as \pazocalDval\pazocal{D}_{\text{val}}. The pre-trained models we used, VGG19 , ResNet50, and DenseNet121, were downloaded from torchvision, and we implemented the penalty terms given in Section 3 in Pytorch framework. All our model training and testing were done with NVIDIA GTX1080TI. The hyperparameters that used to train the models for various fooling methods and interpretations are available in Table 4.

in which mhwm_{hw} is a (h,w)(h,w) element for m\mathbf{m}. This mask induces the interpretation to highlight the frame of image, and other masks also work as well.

Appendix C Threshold determination process in FSR.

In this section, we discuss how to determine RfR_{f} to calculate FSRfI\text{FSR}_{f}{I} in (6). To decide the RfR_{f} for each fooling methods, we compared the visualization of interpretations and test loss with varying iterations, as shown in Figure 7, 8, 9, and 10. For the Location fooling in Figure 7, we can observe that the loss is gradually reduces during training process while the highlighted regions of visualization moves toward boundary. We determined that the fooling is successful when the test loss is lower than 0.2. Similarly, we decided the thresholds of other foolings (marked with orange lines in Figures 2∼\sim5) by comparing both test losses and visualizations.

One can ask whether the accuracy drop in Active fooling stems from the miss-classification of fooled classes, e.g., c1c_{1} and c2c_{2}. To refute this, we evaluated the accuracy of c1c_{1} and c2c_{2} classes with ImageNet validation dataset. Table 5 shows that the slight accuracy drop of the Actively fooled models is not caused by the fooled classes (Firetruck and African Elephant classes in our case), but by the entire classes.

Appendix E Visualizations of Passive foolings and Active fooling

Figures 11 to 17 are more qualitative results of Passive fooling and Active fooling. For Passive fooling, the content of figure is same as Figure 2. For Active fooling, we included the interpretations for c2c_{2}, which is omitted in Figure 3.