Do Feature Attribution Methods Correctly Attribute Features?
Yilun Zhou, Serena Booth, Marco Tulio Ribeiro, Julie Shah
Introduction
Consider the task of training a neural network to detect cancers from X-ray images, wherein the data come from two sources: a general hospital and a specialized cancer center. As can be expected, images from the cancer center contain many more cancer cases, but imagine the cancer center adds a small timestamp watermark to the top-left corner of its images. Since the timestamp is a strongly correlated with cancer presence, a model may learn to use it for prediction.
It is important to ensure the deployed model makes predictions based on genuine medical signals rather than image artifacts like watermarks. If these artifacts are known a priori, we can evaluate the model on counterfactual pairs—images with and without them—and compute prediction difference to assess their impact. However, for almost all datasets, we cannot realistically anticipate every possible artifact. As such, feature attribution methods like saliency maps (Simonyan, Vedaldi, and Zisserman 2013) are used to identify regions which are important for prediction, which humans then inspect for evidence of any artifacts. This train-and-interpret pipeline has been widely adopted in data-driven medical diagnosis (Shen et al. 2019; Mostavi et al. 2020; Si et al. 2021) and many other applications.
Crucially, this procedure assumes that the attribution methods works correctly and does not miss influential features. Is this truly the case? Direct evaluation on natural datasets is impossible as the very spurious correlations we want attribution methods to find are, by definition, unknown. Many evaluations try to sidestep this problem with proxy metrics (Samek et al. 2017; Hooker et al. 2019; Bastings, Aziz, and Titov 2019), but they are limited in various ways, notably by a lack of ground truth, as discussed in Sec. 2.2.
Instead, we propose evaluating these attribution methods on semi-natural datasets: natural datasets systematically modified to introduce ground truth information for attributions. This modification (Fig. 1) ensures that any classifier with sufficiently high performance has to rely, sometimes solely, on the manipulations. We then present desiderata, or necessary conditions, for correct attribution values; for example, features known not to affect the model’s decision should not receive attribution. The high-level idea is domain-general, and we instantiate it on image and text data to evaluate saliency maps, rationale models and attention mechanisms used to explain common deep learning architectures. We identify several failure modes of these methods, discuss potential reasons and recommend directions to fix them. Last, we advocate for testing new attribution methods against ground truth to validate their attributions before deployment.
Related Work
Feature attribution methods assign attribution scores to input features, the absolute value of which informally represents their importance to the model prediction or performance.
Saliency maps explain an image by producing of the same size, where indicates the contribution of pixel . In various works, the notion of contribution has been defined as sensitivity (Simonyan, Vedaldi, and Zisserman 2013), relevance (Bach et al. 2015), local influence (Ribeiro, Singh, and Guestrin 2016), Shapley values (Lundberg and Lee 2017), or filter activations (Selvaraju et al. 2017).
Attention mechanisms (Bahdanau, Cho, and Bengio 2015) were originally proposed to better retain sequential information. Recently they have been used as attribution values, but their their validity is under debate with different and inconsistent criteria being proposed (Jain and Wallace 2019; Wiegreffe and Pinter 2019; Pruthi et al. 2020).
Rationale models (Lei, Barzilay, and Jaakkola 2016; Bastings, Aziz, and Titov 2019; Jain et al. 2020) are inherently interpretable models for text classification with a two-stage pipeline: a selector extracts a rationale (i.e. input words), and a classifier makes a prediction based on it. The selected rationales are often regularized to be succinct and continuous.
2 Evaluation of Feature Attributions
At their core, feature attribution methods describe mathematical properties of the model’s decision function. For example, gradient describes sensitivity with respect to infinitesimal input perturbation, and SHAP describes a notion of values in a multi-player game with features as players. We associate these mathematical properties with high-level interpretations such as “feature importance”, and it is this association that requires justification.
A popular way is to assess alignment with human judgment, but models and humans can reach the same prediction while using distinct reasoning mechanisms (e.g. medical signals used by doctors and watermarks used by the model). For example, SmoothGrad (Smilkov et al. 2017) is proposed as an improvement to the original Gradient (Simonyan, Vedaldi, and Zisserman 2013) since it gives less noisy and more legible saliency maps, but it is not clear whether saliency maps should be smooth. Bastings, Aziz, and Titov (2019) evaluated their rationale model by assessing its agreement with human rationale annotation, but a model may achieve high accuracy with subtle but strongly correlated textual features such as grammatical idiosyncrasy. Covert, Lundberg, and Lee (2020) compared the feature attribution of a cancer prediction model to scientific knowledge, yet a well-performing model may rely on other signals. In general, positive results from alignment evaluation only support plausibility (Jacovi and Goldberg 2020), not faithfulness.
Another common approach successively removes features with the highest attribution values and evaluates certain metrics. One metric is prediction change (e.g. Samek et al. 2017; Arras et al. 2019; Ismail et al. 2020), but it fails to account for nonlinear interactions: for an OR function of two active inputs, the evaluation will (incorrectly) deem whichever feature removed first to be useless as its removal does not affect the prediction. Another metric is model retraining performance (Hooker et al. 2019), which may fail when different features lead to the same accuracy—as is often possible (D’Amour et al. 2020). For example, a model might achieve some accuracy by using only feature . If a retrained model using only achieves the same accuracy, the evaluation framework would (falsely) reject the ground truth attribution of due to the same re-training accuracy.
Most similar to our proposal are works that also construct semi-natural datasets with explicitly defined ground truth explanations (Yang and Kim 2019; Adebayo et al. 2020). Adebayo et al. (2020) used a perfect background correlation for a dog-vs-bird dataset, found that the model achieves high accuracy on background alone, and claimed that the correct attribution should focus solely on the background. However, we verified that a model trained on their dataset can achieve high accuracy simultaneously on foreground alone, background alone, and both combined, invalidating their ground truth claim. Similarly, Yang and Kim (2019) argue that for background classification, a label-correlated foreground should receive high attribution value, but a model could always rely solely on background with perfect label correlation. We avoid such pitfalls via label reassignment (Sec. 2), so that the model must use target features for high accuracy. Furthermore, a more subtle failure mode, in which the model can (rightfully) use the absence of information for a prediction, is avoided by our joint effective region formulation, discussed in the Remark at the end of Sec. 4.
Finally, Adebayo et al. (2018) proposed sanity checks for saliency maps by assessing their change under weight or label randomization. We establish complementary criteria for explanations by instead focusing on model-agnostic dataset-side modifications, and identify additional failure cases.
Desiderata for Attribution Values
What should the attribution values be? Although the precise values may be axiomatic, certain properties are de facto requirements if we want people to understand how a model makes a decision, verify that its reasoning process is sound, and possibly inform options for correction if it is not (c.f. the opening example in Sec. 1). For example, while LIME and SHAP define attribution differently, both would produce undeniably bad explanations if they highlight features completely ignored by the model.
Dataset Modification with Ground Truth
We now present the dataset modification procedure that lets us quantify the influence of certain features to the model. We use a running example of adding a watermark pattern to a watermark-free X-ray cancer dataset, such that the newly added watermark is guaranteed to affect the model decision.
Let and be input and output space for -class classification. Fig. 2 shows two modification steps: from an original data instance (1st column), label reassignment reduces the predictive power of existing signals (2nd column) and input manipulation introduces new predictive features (3rd and 4th columns).
Label Reassignment Our goal is to ensure that the model has to rely on certain introduced features (e.g. a watermark) to achieve a high performance. However, the model could in theory use any of the existing features (e.g. medical features) to achieve high accuracy, and thus disregard the new feature, even if it is perfectly correlated with the label. To guarantee the model’s usage of new features, we need to weaken the correlation between the original features and the labels.
We first consider label reassignment for binary classification, which is used in all experiments. During reassignment, the label is preserved with probability and flipped otherwise, so the accuracy without relying on the manipulation is at most . For the special case of , no features are informative to the label, and the performance is random in expectation. After label reassignment, a data point becomes .
Input Manipulation Next, we apply manipulations on the input according to its reassigned label . We consider a set of input manipulations, , and a manipulation function such that applies the manipulation on the input and returns the manipulated output . can include the blank manipulation that leaves the input unchanged.
Whenever a model trained on achieves expected accuracy , it is guaranteed to rely on the knowledge of manipulation, which is solely confined within the joint effective region . This gives us a straightforward, quantitative check for feature attribution methods: they should recognize the contribution inside . For our example, since only the watermark is applied to one class, corresponds to the watermarked region.
Remark It is crucial to consider the joint effective region over all manipulations for attribution values, since a model could use the absence of manipulation as a legitimate basis for decision. For example, consider an image dataset, with each image having a watermark either on the top or bottom edge correlated with the positive or negative label respectively. A model could make negative predictions based on the absence of a watermark on the top edge. In this case, the correct attribution to the top edge is within the joint ER but not within the bottom watermark ER. Current evaluations (Yang and Kim 2019; Adebayo et al. 2020) often omit this possibility by using the ER of only the manipulation applied to the target class rather than the union of all possible ERs for every class, potentially rejecting correct attributions. A more detailed explanation of how our proposed work differs from and improves upon those of Yang and Kim (2019) and Adebayo et al. (2020) is provided in App. A.
In next three sections, we experimentally compare attribution values of three types of models—saliency maps, attention mechanisms and rationale models—to those expected by the desiderata. Through the analysis, we identify their deficiencies and give recommendations for improvements.
Evaluating Image Saliency Maps
For these experiments, we simulate a common scenario where a model seemingly achieves “superhuman” performance on some hard image classification task, only for us to later find out that it exploits some image artifacts which are accidentally leaked in during the data collection process. We evaluate the extent to which several different saliency map attribution methods can identify such artifacts.
Model: We used the ResNet-34 architecture (He et al. 2016) for all experiments. The parameters are randomly initialized rather than pre-trained on ImageNet (Deng et al. 2009).
Dataset: We curate our own dataset on bird species identification. First, we train a ResNet-34 model on CUB-200-2011 (Wah et al. 2011) and identify the top four most confusing class pairs. Then, we scrape Flickr for 1,200 new images per class, center-crop all images to and mean-variance normalize using ImageNet statistics. Last, we split the 1,200 images per class into train/validation/test sets of 1000/100/100 images. Fig. 3 presents sample images, the confusion matrix for a ResNet-34 model trained on this data, and example saliency maps for a correct prediction.
Input Manipulations: We define five image manipulations which represent artifacts that could be accidentally introduced in a dataset collection process: blurring, brightness change, hue shift, pixel noise, and watermark. Fig. 4 shows the effect of for three manipulations along with the effective regions. Other manipulation types and additional details are presented in in App. B.1.
Saliency Maps: We evaluate 5 saliency map methods: Gradient (Simonyan, Vedaldi, and Zisserman 2013), SmoothGrad (Smilkov et al. 2017), GradCAM (Selvaraju et al. 2017), LIME (Ribeiro, Singh, and Guestrin 2016), and SHAP (Lundberg and Lee 2017), detailed in App. B.2.
Experiments: We set up binary classifications with pairs of easily confused species (e.g. common tern and Forester’s tern) to simulate a hard task which is made easier through the presence of artifacts. Sec. 5.4 also uses pairs of visually distinct species (e.g. common tern and fish crow).
Question: How well do saliency maps give attribution to the ground truth for (near-)perfect models?
Setup: We train 100 models, each on a random pair of similar species and a random manipulation type. We reassign labels with (i.e. totally randomly), and apply the manipulation to images of the positive post-reassignment class, leaving the negative class images unchanged.
2 Attribution vs. Test Accuracy
Setup: We use the the same setup as Sec. 5.1.
3 Attribution vs. Manipulation Visibility
Question: How well can saliency maps recognize manipulations of different visibility levels?
Setup: We conduct 100 runs, with 20 per manipulation. We further group the 20 runs into 4 groups, with 5 runs in a group using the same manipulation type and effective region but varying degrees of visibility, detailed in App. B.5. For example, the visibility for a watermark corresponds to its font size. As before, the labels are reassigned with and manipulations applied to the positive class only.
Expectation: A good saliency map should not be affected by manipulation visibility, as long as the model is objectively using it. However, different saliency maps may be better suited to detect more or less visible manipulations. For example, a less visible manipulation may be ignored by the segmentation algorithm used by LIME, while inducing sharper gradients in the decision space.
4 Attribution vs. Original Feature Correlation
Question: How does the attribution on the manipulation change if the reassigned labels are correlated with the original labels (and thus original input features) to higher or lower degrees (i.e. )?
Setup: For each manipulation, we vary the label reassignment parameter . For each , we train four models on four class pairs: two of similar species (e.g. class 4 vs. 5 in Fig. 3) and two of distinct ones (e.g. class 5 vs. 6), for a total of runs.
where , and refers to the classifier’s expected accuracy when only , only , neither, and both are available, respectively. For a classifier with accuracy , we have , , and . The formal definition of and its calculation are in App. B.6. We normalize the Shapley values to and by their sum .
For (near-)perfect classifier with , we have , and . In addition, should be close to for the distinct pair as the model can better utilize the more distinct original image features, resulting in lower attribution on manipulated features.
For watermark manipulation, SHAP shows clear decrease in attribution value as increases, while gradient also tracks the predicted range, but only for the positive class with the manipulation. This trend is not seen in other feature types, even for SHAP which approximates the Shapley values. There does not seem to be a clear difference in attribution values for similar vs. distinct species pairs either. Considering that the set of Shapley axioms is commonly accepted as reasonable, it is concerning to see that many saliency maps are inconsistent with it, and important to develop a better understanding about the underlying axiomatic assumptions (if any) made by each of them.
5 Discussion
Arguably one of the most important application of model explanation is to detect any usage of spurious correlations, but our results cast doubt on this capability from various aspects. We recommend that, before analyzing the actual model, developers should first train models that are guaranteed to use certain known features, and “dry run” the planned interpretability methods on them to make sure that these features are indeed highlighted.
Evaluating Text Attentions
It is known that certain non-semantic features can heavily influence model prediction, such as the email headers (Ribeiro, Singh, and Guestrin 2016). Plausibly, attention scores should highlight such features, and we rigorously test this with our dataset modification in this section.
Dataset: We modify the BeerAdvocate dataset (McAuley, Leskovec, and Jurafsky 2012) and further select 12,000 reviews split into train, validation, and test sets of sizes 10,000, 1,000 and 1,000 (shuffled differently for each experiment).
Question: How well can attention scores focus on highly obvious manipulations?
Setup: From our filtered dataset, we first randomly assign binary labels. For the positive reviews, we change all the article words (a / an / the) to “the”, and for the negative reviews, we change these to “a”. Thus, only these articles are correlated with the labels and constitute the effective region.
2 Misleading Non-Correlating Features
Question: When some features are known to not correlate with the label but are very similar to correlating ones, do attention scores also focus on these non-correlating ones?
Setup: Again from our filtered dataset, we apply two similar manipulations, with only one of them is correlated with the (reassigned) label. Fig. 9 details the construction of two datasets, CN and NC.
Expectation: Same as above. In particular, non-correlating articles should not be attended to.
Results: The models on both datasets achieve over 97% accuracy. Fig. 11 presents attention visualization, with more in Fig. 20 of App. C.2. The two models show very different behaviors. The CN model exclusively focuses attention on correlating articles, while the NC model behaves similarly to the previous experiment.
Observing the large variation of behaviors, we further trained the three models ten more times to see if any consistent attention pattern exists. All models achieve over 97% accuracy. Fig. 1 (left) presents the mean and standard deviation statistics for the 11 runs. The clean attention pattern by the CN model does not persist, and the model sometimes assigns higher than random weights on non-correlating articles, especially for the dataset. These results further suggests that attention weights cannot be readily and reliably interpreted as attributions without further validation.
3 Discussion
Attention is undoubtedly useful as a building block in neural networks, but their interpretation as attribution is disputed. Due to the lack of ground truth information on word-prediction correlation, past studies proposed various, and sometimes conflicting, criteria for judging the validity of attribution interpretation (Jain and Wallace 2019; Wiegreffe and Pinter 2019; Pruthi et al. 2020). However, the fundamental correctness of such proxy metrics is unclear. In our studies, we find that attentions can hardly be interpreted as attribution for model understanding and debugging purposes: for most training runs, the attention weights on correlating features at best stand out only locally, easily overwhelmed by larger global variations, setting the debate at least on the modified dataset. For natural datasets, we would unavoidably need to rely on proxy metrics, but we recommend future proposals of the metrics to be first calibrated with ground truth in a controlled setting.
Evaluating Text Rationales
Expectation: A necessary condition for a non-misleading rationale is that it should include at least one article word, regardless of selection rate. However, a desirable property of rationale is comprehensiveness (Yu et al. 2019): selecting as many article words as possible. Thus, a good rationale model should have high precision when selection rate is low and high recall when selection rate is high.
2 Misleading Non-Correlating Features
Expectation: Similar to the previous experiment, at least one correlating article word needs to be selected. However, selection of non-correlating articles is arguably more misleading than selection of other non-article words, because it suggests that these non-correlating articles also influence the prediction, even though the classifier simply ignores them.
3 Discussion
The structure of rationale models guarantees that causal relationship between the rationale features and the model prediction, but this does not necessarily imply its usefulness to model understanding. Specifically, it could highlight only barely, while including lots of non-correlating (and, in particular, misleading words such as the non-correlating articles)There are additional concerns on the unfaithfulness of rationales as Trojan explanations (Jacovi and Goldberg 2021; Zheng et al. 2021), but they were not identified in our experiments.. Indeed, our results show that rationale methods are prone to selecting misleading non-correlating features, which obfuscates the model’s reasoning process by giving more but unnecessary information to the human. The problem is more severe with RL training, possibly due to the known difficulty with REINFORCE (Williams 1992). Post-processing methods could be developed to further prune rationales to mitigate this problem.
Conclusion and Future Work
As interpretability methods, especially feature attribution ones, are increasingly deployed for quality assurance of high-stakes systems, it is crucial to ensure these methods work correctly. Current evaluations fall short—primarily due to a lack of clearly defined ground truth. Rather than evaluating explanations for models trained on natural datasets, we propose “unit tests” to assess whether feature attribution methods are able to uncover ground truth model reasoning on carefully-modified, semi-natural datasets. Surprisingly, none of our evaluated methods across vision and text domains achieve totally satisfactory performance, and we point out various future directions in Sec. 5.5, 6.3 and 7.3 to improve attribution methods.
Our dataset modification procedure closely parallels the setup for identifying and debugging model reliance on spurious correlations, which have been known to frequently affect model decisions (e.g. Ribeiro, Singh, and Guestrin 2016; Kaushik, Hovy, and Lipton 2019; Geirhos et al. 2020; Jabbour et al. 2020). Hence, the mostly negative conclusions cast doubt on this use case of interpretability methods.
An extension of the proposed evaluation procedure is to move beyond “artifact” features, which result from the manual definition of the manipulation function. Given the recent advances on generative modeling such as image inpainting (Pathak et al. 2016) and masked language prediction (Devlin et al. 2019), more realistic features could be generated, perhaps also conditioned on or guided by semantic concepts. This would make the modified dataset much more realistic looking, and thus better simulate another intended use case of interpretability: assisting scientific discovery, in which high-performing models teach humans about features of previous unknown importance.
Acknowledgement
An earlier version of this paper was presented at the 2021 NeurIPS Workshop on Explainable AI Approaches for Debugging and Diagnosis. This research is supported by the National Science Foundation (NSF) under the grant IIS-1830282. We thank the reviewers for their reviews.
References
Appendix A Relationship to Closely Related Works
In this section, we detail the similarity and difference between our proposal and two closely related ones (Yang and Kim 2019; Adebayo et al. 2020). In terms of similarity, the high-level idea is similar: for a natural dataset, we do not know individual feature contribution, but generally highly correlated and easily discriminative features should have high contribution, and we realize this notion by injecting such features directly. All three works can be seen as operationalizations of this idea.
However, this high-level idea needs two caveats, which set our paper aparts from both works. First, features that we think to be “easily discriminative” may not be be considered as such by neural nets. For example, Geirhos et al. (2018) showed that they have particular inclination toward textures. Thus, we can’t really trust them to actually pick up and use our introduced features, unless that we are assured that focusing on other features cannot achieve the performance that the model is achieving now. In this aspect, we concretely demonstrated that network trained by Adebayo et al. (2018) can achieve good performance when using the “other features” exclusively. Yang and Kim (2019) demonstrated this principle, but used out-of-distribution data, so it is not clear whether the failure is due to achieve good performance is really due to the network indeed ignoring the “other features” or due to instability of out-of-distribution extrapolation.
The second caveat is that the “other features” cannot have any information on the injected features. For example, in Fig. 4 by Adebayo et al. (2020), the saliency map that perfectly crops out the foreground dog shape could imply that the network actually uses the contour of the dog, which is inconsistent with their target conclusion that “foreground doesn’t matter in highly correlated background dataset”. In other words, background with a dog contour cropped out is very informative to the fact that the foreground is a dog. A similar cropping procedure is used by Yang and Kim (2019) as well. Another loophole is discussed in the Remark of Sec. 4, where the lack of one feature at one place could imply the presence of another feature at another place. By comparison, in our work, we define the joint effective region to contain all injected, so we are assured that the features outside of the region absolutely cannot contribute to the high performance, and then evaluate Attr% for that region.
Finally, our framing is more general. We propose our domain-agnostic framework in Sec. 3 and 4. Thus, it it straightforward to instantiate our idea to other domains such as graph, speech, or time-series data. By comparison, both works are focused on the foreground/background patch setting, and introduces necessary concepts under this context. We also proposed ways to evaluate feature-selection style attributions, more common for text models, at the end of Sec. 3.
Appendix B Additional Details and Results for Saliency Map Evaluations
We consider five image manipulation types. These manipulations are designed to simulate possible image artifacts, which an undesirable model may rely on to make decisions. Each manipulation has parameters which define the effective region and the visibility level of the manipulation effect. Some of the manipulation effects are technically stochastic, such as a watermark being placed in a random position, but the effective region captures the localized manipulation effect of all possible random instantiations. The five manipulations are described below, with examples of each manipulation and their associated effective regions shown in Fig. 14.
Peripheral blurring applies a Gaussian filter to the part of the image outside of a certain radius. It is parametrized by
the standard deviation of the Gaussian blurring filter.
Blurring could be due to either camera in motion or artistic post-processing to highlight the main subject of the image.
Central brightness shift gradually changes the brightness in the hue-saturation-brightness (HSB) space inside a certain radius, with maximal change in the center. For our experiments, the brightness change is negative, meaning that the center is dimmed. It is parametrized by
the magnitude of the brightness shift at the center.
Brightness shift could be due to times of the day, or the use of artificial light to illuminate the subject.
Striped hue shift modifies the hue (i.e. color) value of a vertical stripe in the image. From top to bottom in the stripe, the hue value is first increased and then decreased in a sinusoidal pattern. It is parametrized by
the lower position of the stripe, with the width of the stripe being (upper - lower); and
Hue shift could be due to errors in conversion of different color space encodings, which may result in color loss or distortion.
Striped noise randomly changes pixels inside a vertical stripe to a uniformly random RGB value. It is parametrized by
the lower position of the stripe, with the width of the stripe being (upper - lower); and
the probability that each pixel is replaced.
Pixel noise could be due to lossy compression or data loss during transmission.
Watermark overlays a text reading “IMGxxxx”, where “xxxx” are four random digits, to a random location inside a rectangular region. “IMG” is written in white and the digits are written in black. It is parametrized by
the upper-left coordinate of the rectangular region;
the lower-right coordinate of the rectangular region; and
Watermark is a commonly employed technique to attribute the author/organization of the image.
Normally, none of them should be expected to correlate with the label. However, especially with image scraping on the web and crowdsourced dataset construction, it is possible that some spurious correlations leak into the final dataset.
B.2 Saliency Map Methods
Gradient (Simonyan, Vedaldi, and Zisserman 2013) computes the gradient of the logit for the predicted class with respect to the input image. The three channels of gradient are summed up in absolute value to get a single channel.
SmoothGrad (Smilkov et al. 2017) averages the gradients on 50 copies of the input image , each injected with independent Gaussian noise with with and , where and are the maximal and minimal pixel values of the image.
GradCAM (Selvaraju et al. 2017) computes a saliency map from convolution filter responses. Since we use the fully convolutional ResNet-34, this method reduces to the class activation mapping (CAM) (Zhou et al. 2016).
LIME (Ribeiro, Singh, and Guestrin 2016) performs a linear regression using super-pixels of the input image. The absolute values of the coefficients are used to derive the saliency map. We use the default implementation of lime.lime_ image.LimeImageExplainer with the quickshift clustering as the super-pixel segmentation algorithm.
SHAP (Lundberg and Lee 2017) uses the idea of Shapley value (Roth 1988) for attribution. We use the GradientSHAP instantiation with the default setting of shap.GradientExplainer. We us the entire test set as the “background” data.
B.3 Attribution vs. Effective Region Size
B.4 Attribution vs. Test Accuracy
B.5 Attribution vs. Manipulation Visibility
Blurring: The visibility level is defined as the Gaussian blur standard deviation, with values of pixels, from least visible to most.
Brightness: The visibility level is defined as the magnitude of the brightness shift, with values of brightness component of the color (in the range of ), from least visible to most.
Hue: The visibility is defined as the magnitude of the hue shift, with values of hue component of the color (in the range of ), from least visible to most.
Noise: The visibility is defined as the probability that a pixel is replaced by a random value, with values of , from least visible to most.
Watermark: The visibility is defined as the font size of the watermark, with values of pixels, from least visible to most.
B.6 Attribution vs. Original Feature Correlation
As explained in Sec. 5.4, represents the expected accuracy of the model when given only feature . Note that this value should not be calculated as the model accuracy on images with every pixel but being blacked out, because such images are out of distribution where the model may exhibit unreasonable behaviors (c.f. discussion by Hooker et al. (2019)).
With balanced label distribution, means that the model has no information about the input, and thus the accuracy is 0.5. On the other hand, means that the model has full access to the input, and thus the accuracy is the normal model accuracy . In addition, we have , because the label reassignment weakens the correlation between and the label.
Appendix C Additional Results for Attention Mechanism Evaluations
Fig. 19 presents additional visualizations of the learned attention distribution of the model.
C.2 Misleading Non-Correlating Features
Fig. 20 presents additional visualizations of the learned attention distribution on the CN (left) and NC (right) datasets.
Appendix D Additional Results for Rationale Model Evaluations
Fig. 21 presents four additional reviews annotated by the “faulty” CR model showing that it consistently selects the first few words of the review.