Do Input Gradients Highlight Discriminative Features?
Harshay Shah, Prateek Jain, Praneeth Netrapalli
Introduction
Interpretability methods that provide instance-specific explanations of model predictions are often used to identify biased predictions , debug trained models , and aid decision-making in high-stakes domains such as medical diagnosis . A common approach for providing instance-specific explanations is feature attribution. Feature attribution methods rank or score input coordinates, or features, in the order of their purported importance in model prediction; coordinates achieving the top-most rank or score are considered most important for prediction, whereas those with the bottom-most rank or score are considered least important.
Assumption (A): Coordinates with larger input gradient magnitude are more relevant for model prediction compared to coordinates with smaller input gradient magnitude. Sanity-checking attribution methods. Several attribution methods are based on input gradients and explicitly or implicitly assume an appropriately modified version of (A). For example, Integrated Gradients aggregate input gradients of linearly interpolated points, SmoothGrad averages input gradients of points perturbed using gaussian noise, and Guided Backprop modifies input gradients by zeroing out negative values at every layer during backpropagation. Surprisingly, unlike vanilla input gradients, popular methods that output attributions with better visual quality fail simple sanity checks that are indeed expected out of any valid attribution method . On the other hand, while vanilla input gradients pass simple sanity checks, Hooker et al. suggest that they produce estimates of feature importance that are no better than a random designation of feature importance.
Probing input gradient attributions using BlockMNIST. Our empirical findings mentioned above strongly suggest that standard models grossly violate (A). However, without knowledge of ground-truth discriminative features learned by models trained on real data, conclusively testing (A) remains elusive. In fact, this is a key shortcoming of the remove-and-retrain (ROAR) framework.
So, to further verify and better understand our empirical findings, we introduce an MNIST-based semi-real dataset, BlockMNIST, that by design encodes a priori knowledge of ground-truth discriminative features. BlockMNIST is based on the principle that for different inputs, discriminative and non-discriminative features may occur in different parts of the input. For example, in an object classification task, the object of interest can occur in different parts of the image (e.g., top-left, center, bottom-right etc.) for different images. As shown in Figure 1(a), BlockMNIST images consist of a signal block and a null block that are randomly placed at the top or bottom. The signal block contains the MNIST digit that determines the class of the image, whereas the null block, contains a square patch with two diagonals that has no information about the label. This a priori knowledge of ground-truth discriminative features in BlockMNIST data allows us to (i) validate our empirical findings vis-a-vis input gradients of standard and robust models (see fig. 1) and (ii) identify feature leakage as a reason that potentially explains why input gradients violate (A) in practice. Here, feature leakage refers to the phenomenon wherein given an instance, its input gradients highlight the location of discriminative features in the given instance as well as in other instances that are present in the dataset. For example, consider the first BlockMNIST image in fig. 1(a), in which the signal is placed in the bottom block. For this image, as shown in fig. 1(b,c), input gradients of standard models incorrectly highlight the top block because there are other instances in the BlockMNIST dataset which have signal in the top block.
Rigorously demonstrating feature leakage. In order to concretely verify as well as understand feature leakage more thoroughly, we design a simplified version of BlockMNIST that is amenable to theoretical analysis. On this dataset, we first rigorously demonstrate that input gradients of standard one-hidden-layer MLPs exhibit feature leakage in the infinite-width limit and then discuss how feature leakage results in input gradient attributions that clearly violate assumption (A).
Paper organization: Section 2 discusses related work and section 3 presents our evaluation framework, DiffROAR, to test assumption (A). Section 4 employs DiffROAR to evaluate input gradient attributions on four image classification datasets. Section 5 analyzes BlockMNIST data to differentially characterize input gradients of standard and robust models using feature leakage. Section 6 provides theoretical results on a simplified version on BlockMNIST that shed light on how feature leakage results in input gradients that violate assumption (A). Our code, along with the proposed datasets, is publicly available at https://github.com/harshays/inputgradients.
Related work
Due to space constraints, we only discuss directly related work and defer the rest to Appendix A.
Sanity checks for explanations. Several explanation methods that provide feature attributions are often primarily evaluated using inherently subjective visual assessments . Unsurprisingly, recent “sanity checks” show that sole reliance on visual assessment is misleading, as attributions can lack fidelity and inaccurately reflect model behavior. Adebayo et al. and Kindermans et al. show that unlike input gradients , other popular methods—guided backprop , gradient input , integrated gradients —output explanations which lack fidelity on image data, as they remain invariant to model and label randomization. Similarly, Yang and Kim use custom image datasets to show that several explanation methods are more likely to produce false positive explanations than vanilla input gradients. Moreover, several explanation methods based on modified backpropagation do not pass basic sanity checks . To summarize, well-known gradient-based attribution methods that seek to mitigate gradient saturation , discontinuity , and visual noise surprisingly fare worse than vanilla input gradients on multiple sanity checks.
Evaluating explanation fidelity. The black-box nature of neural networks necessitates frameworks that evaluate the fidelity or “correctness” of post-hoc explanations without knowledge of ground-truth features learned by trained models. Modification-based evaluation frameworks gauge explanation fidelity by measuring the change in model performance after masking input coordinates that a given explanation method considers most (or least) important. However, due to distribution shifts induced by input modifications, one cannot conclusively attribute changes in model performance to the fidelity of instance-specific explanations . The remove-and-retrain (ROAR) framework accounts for distribution shifts by evaluating the performance of models retrained on train data masked using post-hoc explanations. Surprisingly, contrary to findings obtained via sanity checks, experiments with the ROAR framework show that multiple attribution methods, including vanilla input gradients, are no better than model-independent random attributions that lack explanatory power . Therefore, motivated by the central role of vanilla input gradients in attribution methods, we augment the ROAR framework to understand when and why input gradients violate assumption (A).
DiffROAR evaluation framework
In this section, we introduce our evaluation framework, DiffROAR, to probe the extent to which instance-specific explanations, or feature attributions, highlight discriminative features in practice. Specifically, our framework, DiffROAR, builds upon the remove-and-retrain (ROAR) methodology to test whether feature attribution methods satisfy assumption (A) on real-world datasets.
For example, Figure 2 depicts an image and its top- unmasked variant . In this case, the attribution scheme assigns higher rank to pixels in the foreground. So, the top- unmasking operation, , highlights the monkey by retaining pixels with top- attribution ranks and zeroing out the remaining pixels that correspond to the green background.
Predictive power of unmasking schemes. The predictive power of an unmasking scheme with respect to model architecture (e.g., resnet18) can be defined as the best classification accuracy that can be attained by training a model with architecture on unmasked instances that are obtained via unmasking scheme . More formally, it can defined as follows:
Due to masking-induced distribution shifts, models with architecture that are trained using original data cannot be plugged in to estimate . The ROAR framework sidesteps this issue by retraining models on unmasked data, as similar model architectures tend to learn “similar” classifiers . Therefore, we employ the ROAR framework to estimate in two steps. First, we use unmasking scheme to obtain unmasked train and test datasets that comprise data points of the form . Then, we retrain a new model with the same architecture on unmasked train data and evaluate its accuracy on unmasked test data.
DiffROAR evaluation metric to test assumption (A). Recall that an attribution scheme maps an instance to a permutation of its coordinates that reflects the order of estimated importance in model prediction. An attribution scheme that satisfies assumption (A) must place coordinates that are more important for model prediction higher up in the the attribution order. More formally, given attribution scheme , architecture and level , we define DiffROAR as the difference between the predictive power of top- and bottom- unmasking schemes, and :
On testing assumption (A). To verify (A) for a given attribution scheme , it is necessary to evaluate whether input coordinates with higher attribution rank are more important for model prediction than coordinates with lower rank. Consequently, the ROAR-based metric in Hooker et al. , which essentially computes the top- predictive power, is not sufficient to test whether attribution methods satisfy assumption (A). Therefore, as discussed above, DiffROAR tests (A) by comparing the top- predictive power, , to the bottom- predictive power, , using multiple values of .
Testing assumption (A) on image classification benchmarks
In this section, we use DiffROAR to evaluate whether input gradient attributions of standard and adversarially robust MLPs and CNNs trained on four image classification benchmarks satisfy assumption (A). We first summarize the experiment setup and then describe key empirical findings.
Estimating DiffROAR on real data. We compute the evaluation metric, , on real datasets in four steps, as follows. First, we train a standard or robust model with architecture on the original dataset and obtain its input gradient attribution scheme . Second, as outlined in Section 3, we use attribution scheme and level (i.e., fraction of pixels to be unmasked) to extract the top- and bottom- unmasking schemes: and . Third, we apply and on the original train & test datasets to obtain top- and bottom- unmasked datasets respectively. Finally, to compute via eq. (1), we estimate top- and bottom- predictive power, and , by retraining new models with architecture on top- and bottom- unmasked datasets respectively. Also, note that we (a) average the DiffROAR metric over five runs for each model and unmasking fraction or level and (b) unmask individual image pixels without grouping them channel-wise.
Experiment setup. Now, we analyze the DiffROAR metric as a function of the unmasking fraction in order to evaluate whether input gradient attributions of models trained on four image classification benchmarks satisfy assumption (A). In particular, as shown in Figure 3, we use DiffROAR to analyze input gradients of standard and adversarially robust two-hidden-layer MLPs on SVHN & Fashion MNIST, Resnet18 on ImageNet-10, and Resnet50 on CIFAR-10. In order to calibrate our findings, we compare input gradient attributions of these models to two natural baselines: model-agnostic random attributions and input-agnostic attributions of linear models.
Input gradients of standard models. Input gradient attributions of standard MLPs trained on SVHN satisfy assumption (A), as the DiffROAR metric in Figure 3(a) is positive for all values of level < 100%. However, in Figure 3(b), the DiffROAR curves of standard MLPs trained on Fashion MNIST indicate that input gradient attributions, consistent with findings in Hooker et al. , can fare no better than model-agnostic random attributions and input-agnostic attributions of linear models vis-a-vis assumption (A). Furthermore, and rather surprisingly, the shaded area in Figure 3(c) and Figure 3(d) shows that when level < 40%, DiffROAR curves of standard Resnets trained on CIFAR-10 and Imagenet-10 are consistently negative and perform considerably worse than model-agnostic and input-agnostic baseline attributions. These results strongly suggest that on CIFAR-10 and Imagenet-10, input gradients of standard Resnets grossly violate assumption (A) and suppress discriminative features. In other words, coordinates with larger gradient magnitude have worse predictive power than coordinates with smaller gradient magnitude.
Additional results. In Appendix C, we show that our DiffROAR results are robust to choice of model architecture & SGD hyperparameters during retraining and also hold for input gradients taken with respect to cross-entropy loss. Additionally, while DiffROAR without retraining gives qualitatively similar results, they are not as consistent across architectures as with retraining, particularly for small unmasking fraction that induce non-trivial distribution shifts.
Analyzing input gradient attributions using BlockMNIST data
To verify whether input gradients satisfy assumption (A) more thoroughly, we introduce and perform experiments on BlockMNIST, an MNIST-based dataset that by design encodes a priori knowledge of ground-truth discriminative features.
BlockMNIST dataset design: The design of the BlockMNIST dataset is based on two intuitive properties of real-world object classification tasks: (i) for different images, the object of interest may appear in different parts of the image (e.g., top-left, bottom-right); (ii) the object of interest and the rest of the image often share low-level patterns such as edges that are not informative of the label on their own. We replicate these aspects in BlockMNIST instances, which are vertical concatenations of two signal and null image blocks that are randomly placed at the top or bottom with equal probability. The signal block is an MNIST image of digit or digit , corresponding to class or of the BlockMNIST image respectively. On the other hand, the null block in every BlockMNIST image, independent of its class, contains a square patch made of two horizontal, vertical, and slanted lines, as shown in Figure 1(a). It is important to note that unlike the MNIST signal block that is fully predictive of the class, the non-discriminative null block contains no information about the class. Standard as well as adversarially robust models trained on BlockMNIST data attain 99.99% test accuracy, thereby implying that model predictions are indeed based solely on the signal block for any given instance. We further verify this by noting that the predictions of trained model remain unchanged on almost every instance even when all pixels in the null block are set to zero.
Feature leakage hypothesis: Recall that the discriminative signal block in BlockMNIST images is randomly placed at the top or bottom with equal probability. Given our results in Figure 1, we hypothesize that when discriminative features vary across instances (e.g., signal block at top vs. bottom), input gradients of standard models not only highlight instance-specific features but also leak discriminative features from other instances. We term this hypothesis feature leakage.
To test our hypothesis, we leverage the modular structure in BlockMNIST to construct a slightly modified version, BlockMNIST-Top, wherein the location of the MNIST signal block is fixed at the top for all instances (see fig. 4). In this setting, in contrast to results on BlockMNIST, input gradients of standard Resnet18 and MLP models trained on BlockMNIST-Top satisfy assumption (A). Specifically, when the signal block is fixed at the top, input gradient attributions in Figure 4(b, c) clearly highlight the signal block and suppress the null block, thereby supporting our feature leakage hypothesis. Based on our BlockMNIST experiments, we believe that understanding how adversarial robustness mitigates feature leakage is an interesting direction for future work.
Additional results. In Section D.1, we (i) visualize input gradients of several BlockMNIST and BlockMNIST-Top images, (ii) introduce a quantitative proxy metric to compare feature leakage between standard and robust models, (iii) show that our findings are fairly robust to the choice and number of classes in BlockMNIST data, and (iv) evaluate feature leakage in five feature attribution methods. We also provide experiments that falsify hypotheses vis-a-vis input gradients and assumption (A) that we considered in addition to feature leakage.
Feature leakage in input gradient attributions
To understand the extent of feature leakage more thoroughly, we introduce a simplified version of the BlockMNIST dataset that is amenable to theoretical analysis. We rigorously show that input gradients of standard one-hidden-layer MLPs do not differentiate instance-specific features from other task-relevant features that are not pertinent to the given instance.
Consider distribution (6) with . There exists a max-margin classifier for in Wasserstein space (i.e., training both layers of FCN with ) given by (4), such that for all : (i) for every and (ii) for every , where denotes the block of the input gradient .
Theorem 1 guarantees the existence of a max-margin classifier such that the input gradient magnitude for any given instance is (i) a non-zero constant on each of the first task-relevant blocks, and (ii) equal to zero on the remaining noise blocks that do not contain any information about the label. However, input gradients fail at highlighting the unique instance-specific signal block over the remaining task-relevant blocks. This clearly demonstrates feature leakage, as input gradients for any given instance also highlight task-relevant features that are, in fact, not specific to the given instance. Therefore, input gradients of standard one-hidden-layer MLPs do not highlight instance-specific discriminative features and grossly violate assumption (A). In Appendix F, we present additional results that demonstrate that adversarially trained one-hidden-layer MLPs can suppress feature leakage and satisfy assumption (A).
Discussion and conclusion
In this work, we took a three-pronged approach to investigate the validity of a key assumption made in several popular post-hoc attribution methods: (A) coordinates with larger input gradient magnitude are more relevant for model prediction compared to coordinates with smaller input gradient magnitude. Through (i) evaluation on real-world data using our DiffROAR framework, (ii) empirical analysis on BlockMNIST data that encodes information of ground-truth discriminative features, and (iii) a rigorous theoretical study, we present strong evidence to suggest that standard models do not satisfy assumption (A). In contrast, adversarially robust models satisfy (A) in a consistent manner. Furthermore, our analysis in Section 5 and Section 6 indicates that feature leakage sheds light on why input gradients of standard models tend to violate (A). We provide additional discussion in Appendix B.
This work exclusively focused on “vanilla" input gradients due to their fundamental significance in feature attribution. A similarly thorough investigation that analyzes other commonly-used attribution methods is an interesting avenue for future work. Another interesting avenue for further analyses is to understand how adversarial training mitigates feature leakage in input gradient attributions.
References
Appendix A Additional related work
In this section, we briefly describe works that analyze two properties of post-hoc instance-specific explanations that are related to explanation fidelity or “correctness”. In particular, we outline recent works that study the robustness and practical utility of instance-specific explanation methods.
Robustness of explanations: Several commonly used instance-specific explanation methods lack robustness in practice. Ghorbani et al. show that instance-specific explanations and exempler-based explanations are not robust to imperceptibly small adversarial perturbations to the input. Heo et al. show that instance-specific explanations are highly vulnerable to adversarial model manipulations as well. Dombrowski et al. show that explanations lack robustness to to arbitrary manipulations and show that non-robustness stems from geometric properties of neural networks. Bansal et al. show that explanation methods are considerably sensitive to method-specific hyperparameters such as sample size, blur radius, and random seeds. Recent works promote robustness in explanations using smoothing or variants of adversarial training
Utility of explanations: A recent line of work propose evaluation frameworks to assess the practical utility of post-hoc instance-specific explanation methods via proxy downstream tasks. Chu et al. employ a randomized controlled trial to show that using explanation methods as additional information does not improve human accuracy on classification tasks. More generally, Poursabzi-Sangdeh et al. analyze the effect of model transparency (e.g., number of input features, black-box vs. white-box) on the accuracy of human decisions with respect to the task and model. Similarly, Adebayo et al. conduct a human subject study to show that subjects fail to identify defective models using attributions and instead primarily rely on model predictions. formalize the “value” of explanations as the explanation utility (i.e., as side information) in a student-teacher learning framework. In contrast to the works above, we propose an evaluation framework, DiffROAR, to evaluate the fidelity, or “correctness”, of explanations in classification tasks. In particular, using benchmark image classification tasks and synthetic data, we empirically and theoretically characterize input gradient attributions of standard as well as adversarially robust models.
Stability of explanations. Explanation stability and explanation correctness (also known as explanation fidelity) are two distinct desirable properties of explanations . That is, stability does not imply fidelity. For example, an input-agnostic constant explanation is stable but lacks fidelity. Conversely, fidelity does not imply stability—if the underlying model is itself unstable, then any correct high-fidelity explanation of that model must also be unstable. Bansal et al. and Chen et al. identify and explain why input gradients of adversarially trained models are more stable compared to those of standard models. In contrast, our work focuses on identifying and explaining why input gradients of adversarially trained models have more fidelity compared to those of standard models. Furthermore, we also take the first step towards theoretically showing that adversarial robustness can provably improve input gradient fidelity in Appendix E.
Appendix B Additional discussion
Translation invariance in BlockMNIST models. Intuitively, CNNs are translation-invariant only if the object of interest is not closer to the boundary than the receptive field of the final layer; In BlockMNIST, the digits are either close to the top boundary or the bottom boundary. Given that the receptive field of Resnets is quite large, translation invariance would not hold in this case. This is further supported by recent work , which demonstrates that “CNNs can and will exploit the absolute spatial location by learning filters that respond exclusively to particular absolute locations by exploiting image boundary effects”. We observe this phenomenon empirically in our BlockMNIST-Top experiments as well. That is, while models trained on BlockMNIST-Top data (i.e., MNIST digit in top block) attain 100% test accuracy on BlockMNIST-Top images, the accuracy of these models degrades to approximately 55% (i.e., 5% better than random chance) when evaluated on BlockMNIST-Bottom images, wherein the MNIST digit (signal) is placed in the bottom block.
Choice of removal operator in DiffROAR framework. Recall that in DiffROAR, the predictive power of a new model retrained on the unmasked dataset (i.e, data points after removal operation) is used to evaluate the fidelity of post-hoc explanation methods. Note that this approach employs retraining to account for and nullify distribution shifts induced by feature removal operators such as gaussian noise, zeros etc. Since the same removal operation is applied to unmask every image (across classes), the choice of removal operator has no effect on our DiffROAR results in Section 4. To verify this, we evaluated DiffROAR on CIFAR-10 with another removal operator in which pixels are masked/replaced by random gaussian noise (instead of zeros) and observed that the results do not change (i.e., same as in Figure 3).
Counterfactual changes vis-a-vis feature leakage. As evidenced in the BlockMNIST experiments, input gradient attributions of standard models incorporate counterfactual changes in the null block. While this phenomenon seems natural and “intuitive” in hindsight, it can be misleading in the context of feature attributions. For example, consider the typical use case for feature attributions: to highlight regions within the given instance/image that are most relevant for model prediction. Now, in the BlockMNIST setting, if input gradients leak digit-like features into the null block, then the feature attributions in the null block can be easily (mis)interpreted as the non-discriminative null patch being highly relevant for model prediction.
Comparison to results in Kim et al. . Kim et al. use the ROAR framework to conjecture that adversarial training “tilts” input gradients to better align with the data manifold. First, in contrast to Kim et al. , we thoroughly establish our DiffROAR results across datasets/architectures/hyper-parameters, revealing a significantly larger gap between the attribution quality of standard and adversarially robust models. Second, motivated by the boundary tilting hypothesis , Kim et al. use a two-dimensional synthetic dataset to empirically show that the decision boundary of robust models aligns better with the vector between the two class-conditional means. However, this empirical evidence might be misleading, as Ilyas et al. theoretically demonstrates that “this exact statement is not true beyond two dimensions” (pg. 15). Furthermore, several recent works have also provided concrete evidence to support alternative hypotheses for the existence of adversarial examples that counter the boundary tilting hypothesis that Kim et al. build upon. This discrepancy in these results motivates the need for a multipronged approach, which we adopt to empirically identify the feature leakage hypothesis using BlockMNIST and theoretically verify the hypothesis in Section 6.
Connections between adversarial robustness and data manifold: In the recent past, there have been several results showing unexpected benefits of adversarially trained models beyond adversarial robustness such as visually perceptible input gradients and feature representations that transfer better . One reason for this phenomenon widely considered in the literature is that the input data lies on a low dimensional manifold and unlike standard training, adversarial training encourages the decision boundary to lie on this manifold (i.e. alignment with data manifold). Our experiments and theoretical results on feature leakage suggest that this reasoning is indeed true for both the BlockMNIST and its simplified version presented in Section 6. Furthermore, we believe that the simplified version of BlockMNIST in eq. (6) can be used as a tool to thoroughly investigate both the benefits and potential drawbacks of adversarially trained models.
Why focus on input gradient attributions?. As discussed in Section 1, several feature attributions such as guided backprop and integrated gradients that output visually sharper saliency maps fail basic sanity checks such as model randomization and label randomization . We focus on vanilla input gradient attributions for two key reasons: (i) vanilla input gradients pass both sanity checks mentioned above and (ii) the input gradient operation is the key building block of several feature attribution methods. Our experiments and theoretical analysis are specifically designed to identify and verify feature leakage in input gradient attributions of standard models.
Comparing ROAR and DiffROAR. The following questions below illustrate key differences between ROAR and our work:
Does the framework verify assumption (A)? In Hooker et al. , the ROAR framework essentially computes the top- predictive power only, which is not sufficient to test assumption (A). In our paper, DiffROAR directly compares the top- and bottom- predictive power to test whether the given attribution method satisfies assumption (A).
Are the results in the paper conclusive? Both, ROAR and DiffROAR, make a key assumption: models retrained on unmasked datasets learn the same features as the model trained on the original dataset. Although empirically supported , this assumption makes it difficult to conclusively test assumption (A). Therefore, we empirically (Section 5) and theoretically (Section 6) verify our DiffROAR findings in settings wherein ground-truth features are known a priori.
Does the work identify why standard input gradients violate (A)? Hooker et al. do not discuss why input gradients lack explanation fidelity. In our paper, we hypothesize feature leakage as the key reason for ineffectiveness of input gradients, and validate it with empirical as well as theoretical analysis on BlockMNIST-based data
Limitations of ROAR and DiffROAR. The major limitation of ROAR and DiffROAR is the key assumption that models retrained on unmasked datasets learn the same features as the model trained on the original dataset. In the absence of ground-truth features, this assumption is empirically supported by findings that suggest that different runs of models sharing the same architecture learn similar features . Another limitation is that ROAR-based frameworks are not useful in the following setting. Consider a redundant dataset where features are either all negative (in which case label ) or all positive (in which case label ). In such cases, no feature is more or less informative than any other, so no information can be gained by ranking or removing input coordinates/features.
Appendix C Experiments on real-world datasets using DiffROAR
In this section, we first provide additional details about datasets, training, and performance of trained models vis-a-vis generalization and robustness. We also present top- and bottom- predictive power of input gradient unmasking schemes obtained via standard and robust models. Next, we show that our results on image classification benchmarks are robust to CNN architectures and SGD hyperparameters used during retraining. Then, we use DiffROAR to show that our results hold with input loss gradients, but signed input logit gradients do not satisfy assumption (A) for standard or robust models. Finally, we discuss DiffROAR results obtained without retraining and provide additional example images that are masked using input gradients of standard & robust models.
C.2 Top-k𝑘k and bottom-k𝑘k predictive power of input gradient attributions
Now, we describe the top- and bottom- predictive power curves for unmasking schemes of input gradients of standard and robust models. Recall that top- predictive power simply estimates the test accuracy of models that are retrained on datasets wherein only coordinates with top- (%) of the coordinates are unmasked in every image. The top and bottom rows in Figure 7 show how top- and bottom- predictive power of input gradient attributions of standard and robust models vary with unmasking fraction respectively. The subplots in Figure 7 show that (i) decreasing the unmasking fraction decreases top- and bottom- predictive power, and (ii) models retrained on attribution-masked datasets attain non-trivial unmasked test dataset accuracy even when a significant fraction of coordinates with the top-most and bottom-most attributions are masked.
C.3 Effect of model architecture on DiffROAR results
Recall that in Section 4, we used the DiffROAR metric to evaluate whether input gradient attributions of models trained on real-world datasets satisfy or violate assumption (A). For CNNs, we evaluated input gradient attributions of standard Resnet50 and Resnet18 models trained on CIFAR-10 and Imagenet-10 respectively. In this section, we show that our empirical findings based on these architectures extend to three other commonly-used and well-known CNN architectures: Densenet121, InceptionV3, and VGG11.
C.4 Effect of SGD Hyperparameters on DiffROAR results
C.5 Evaluating input loss gradient attributions using DiffROAR
Recall that our experiments in Section 4 evaluate whether input gradients taken w.r.t. the logit of the predicted label satisfy or violate assumption (A) on image classification benchmarks. In this section, we show that our empirical findings generalize to input loss gradients—input gradients w.r.t loss (e.g., cross-entropy)—of standard and robust models evaluated on image classification benchmarks. Specifically, we apply DiffROAR to input loss gradients of standard and robust ResNet models trained on CIFAR-10 and ImageNet-10.
Figure 10 illustrates DiffROAR curves for input loss gradient attributions on CIFAR-10 and ImageNet-10 data. In both cases, we observe that (i) input loss gradient attributions of robust models, unlike those of standard models, satisfy (A) and (ii) PGD adversarial training with larger perturbation budget increases the DiffROAR metric in a consistent manner. Recall that the magnitude in DiffROAR quantifies the extent to which the attribution order separates discriminative and task-relevant features from features that are unimportant for model prediction; see Section 3 for more information about DiffROAR.
C.6 Evaluating signed input gradient attributions using DiffROAR
In addition to input loss gradient magnitude attributions and input logit gradient magnitude attributions, our results vis-a-vis DiffROAR evaluation on image classification benchmarks extend to signed input logit gradients as well. In signed input gradient attributions, input coordinates are ranked based on where is the sign of input coordinate and is the signed input gradient value for input coordinate .
Figure 11 shows DiffROAR curves for attributions based on signed input gradients taken with respect to the logit of the predicted label. The left and right subplot evaluate DiffROAR for standard and robust (i) MLP trained on Fashion MNIST and (ii) Resnet18 models trained on CIFAR-10. Consistent with our findings in Section 4, while standard MLPs trained on Fashion MNIST fare no better than random attributions, signed input gradients of robust MLPs attain positive DiffROAR scores for all and perform considerably better than gradients of standard MLPs. Similarly, based on the DiffROAR metric, when , while signed input gradients of standard Resnet18 models perform better than absolute logit and loss gradients, signed input gradients of robust Resnet18 models continue to fare better than standard models.
C.7 The role of retraining in DiffROAR evaluation
Figure 12 shows the results on DiffROAR without retraining on the masked datasets. As we can see from the figures, the trends are not consistent across model architectures and datasets, possibly due to varying levels of distribution shift. For this reason, we employ DiffROAR with retraining as described in Section 3.
C.8 Imagenet-10 images unmasked using input gradients attributions of Resnet18 models
Recall that in Section 4, we showed that unlike input gradients of standard models, robust models consistently satisfy assumption (A). That is, input gradients of robust models highlight discriminative features, whereas input gradients of standard models tend to highlight non-discriminative features and suppress discriminative task-relevant features. In this section, we qualitatively substantiate these findings by visualizing ImageNet-10 images that are unmasked using top- and bottom- input gradient attributions of standard and robust Resnet18 models. Please note that the following visual assessments are only meant to qualitatively support findings made in Section 4 using the evaluation framework described in Section 3. As discussed in Section 3, if input gradients attain high-magnitude DiffROAR score, images unmasked using top- attributions should highlight discriminative features, whereas images unmasked using bottom- should highlight non-discriminative features.
We make two observations using Figure 13 that qualitatively support our empirical findings in Section 4. First, we observe that images unmasked using top- gradient attributions of robust models tend to highlight salient aspects of images (e.g., shape of fruit or face of monkey in Figure 13), whereas bottom- attributions often mask salient aspects of images either completely or partially. Second, images unmasked using top- and bottom- attributions using input gradients of standard models exhibit visual commonalities, supporting the fact that for standard models, DiffROAR is close to for multiple values of .
Appendix D Additional experiments on feature leakage and BlockMNIST data
In this section, we first provide additional evidence that supports the feature leakage hypothesis in the setting used in Section 5: BlockMNIST data with MNIST digits and corresponding to the signal block in class and class respectively. Then, we show that our results vis-a-vis feature leakage and BlockMNIST are robust to the choice of MNIST digits used in the signal block as well as the number of classes in the BlockMNIST classification task. Finally, we end with a brief description of experiments that we conducted in order to test another hypothesized cause to understand why input gradients of standard models tend to violate (A).
In this section, we provide (i) additional examples of BlockMNIST images and inputs gradients of standard and robust models, (ii) additional examples of BlockMNIST-Top images and input gradients, and (iii) describe a proxy metric to measure feature leakage in BlockMNIST-based data.
We further substantiate these findings using a proxy metric to quantitatively measure feature leakage in BlockMNIST-based datasets. As discussed in Section 5, in the BlockMNIST setting, we can restate assumption (A) as follows: Do input gradient attributions highlight the signal block over the null block? We measure the extent to which input gradients of a given trained model satisfies assumption (A) by evaluating the fraction of top- attributions that are placed in the null block. In Figure 16, we show that the fraction of top- attributions in the null block, when averaged over all images in the test dataset, is significantly greater for standard MLPs & CNNs than for robust MLP & CNNs. In Figure 17, we show that input gradient attributions of standard models trained on BlockMNIST-Top place significantly fewer attributions in the null block, compared to attributions of standard models trained on BlockMNIST. In both cases, the proxy metric further validates our findings vis-a-vis input gradients of standard & robust models and feature leakage.
D.2 Effect of choice and number of classes in BlockMNIST data
In this section, we show that our analysis on BlockMNIST-based datasets in Section 5 is robust to the choice and number of classes in BlockMNIST data. In particular, we reproduce our empirical findings vis-a-vis feature leakage and input gradient attributions of standard vs. robust models on three additional BlockMNIST-based tasks. In Figure 18 and Figure 19, we evaluate input gradients of standard and robust models trained on BlockMNIST and BlockMNIST-Top data, wherein the MNIST digits in class and class correspond to digits and (in the signal block) respectively. Similarly, in Figure 20 and Figure 21, we reproduce our empirical findings from Section 5 on BlockMNIST and BlockMNIST-Top data in which the MNIST digits in class and class correspond to digits and (in the signal block) respectively. In Figure 22 and Figure 23, we show that (i) input gradients of standard models violate assumption (A) due to feature leakage and (ii) adversarial training mitigates feature leakage on -class BlockMNIST and BlockMNIST-Top data, wherein each class corresponds to MNIST digit in the signal block.
D.3 Does randomness in initialization explain why input gradients violate (A)?
In this section, we investigate whether the poor quality of input gradients in standard models is due to randomness retained from the initialization. Figure 24 shows scatter plots of input gradient values over all pixels in all images before (x-axis) and after (y-axis) standard training on four image classification benchmarks. The results indicate that (i) the scale of gradients after training is at least an order of magnitude larger than those before training and (ii) the gradient values before and after training are uncorrelated. Together, these results suggest that random initialization does not have much of a role in determining the input gradients after training.
D.4 Do other feature attribution methods exhibit feature leakage?
In this section, we evaluate feature leakage in five feature attribution methods: Integrated Gradients , Layer-wise Relevance Propagation (LRP) , Guided Backprop , Smoothgrad (with standard deviation ), and Occlusion (with patch size ). First, we evaluate the aforementioned feature attribution methods on standard models trained on BlockMNIST data. As shown in Figure 25 and Figure 26, in addition to vanilla input gradients, all five feature attribution methods evaluated on standard MLPs and Resnet18 models highlight the MNIST signal block as well as the null block. Conversely, Figure 27 and Figure 28 show that when standard MLPs and Resnet18 models are trained on BlockMNIST-Top data, all feature attribution methods exclusively highlight the MNIST signal block. These results collectively indicate that similar to vanilla input gradient attributions, multiple feature attribution methods exhibit feature leakage. Furthermore, consistent with our findings on adversarial robustness vis-a-vis feature leakage, Figure 29 and Figure 30 show that feature attribution method evaluated on adversarially robust MLPs and Resnet18 model do not exhibit feature leakage on BlockMNIST data.
Appendix E Proof of Theorem 1
where we recall that is the ReLU nonlinearity.
We first prove (11). Consider the point in the support of the training distribution with . We see that:
Since . Consequently, using (9), we have that:
We can now consider to be the concatenation of for and the remaining coordinates equal to zero, which ensures that for and . We can then choose such that and . So it suffices to consider for where is the concatenation of for some for and the remaining coordinates being set to zero. Let us consider two situations separately:
Case I, : Recall from (9) the definition . First note from (9) that, and for every . If for some , then choosing with for all gives us a corresponding satisfying
So, it suffices to restrict our attention to such that for all in order to prove (10). We will now show that making all equal will further increase the value of . In order to see this, let and . Then constructing from by replacing and with ensures that while at the same time since implies . If for all , then from (9),
The maximizer of the above expression under the constraint can be seen to be when , and achieving value .
Case II, : In this case, we have from (9) that
where we used in the last step. This shows that is a max-margin classifier satisfying (4).
Gradient magnitude: For any input , we note that the input gradient is of the form for some . Consequently, the claim about the gradient magnitudes in different coordinates follows from the structure of proved above.
Appendix F Effect of adversarial training
Figure 31 empirically verifies two consequences of this conjecture. In Figure 31(a), we show that first-layer weights with large alignment with standard basis vectors also have large second-layer weights, indicating that axis-aligned first-layer weights are highly influential in the final model’s prediction. Figure 31(b) shows that the biases in first-layer ReLU units are predominantly negative.
Assuming Conjecture 1, this lemma shows that for the special case and , adversarially trained models have input gradients that reveal instance-specific features important for classification. Conjecture 1 and Lemma 1 also explain several other empirically observed properties of adversarial training such as visually perceptible input gradients and adversarial examples . In this section, we prove Lemma 1.