Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement Learning

Akanksha Atrey, Kaleigh Clary, David Jensen

Introduction

Saliency map methods are a popular visualization technique that produce heatmap-like output highlighting the importance of different regions of some visual input. They are frequently used to explain how deep networks classify images in computer vision applications (Simonyan et al., 2014; Zeiler & Fergus, 2014; Springenberg et al., 2015; Ribeiro et al., 2016; Dabkowski & Gal, 2017; Fong & Vedaldi, 2017; Selvaraju et al., 2017; Shrikumar et al., 2017; Smilkov et al., 2017; Zhang et al., 2018) and to explain how agents choose actions in reinforcement learning (RL) applications (Bogdanovic et al., 2015; Wang et al., 2016; Zahavy et al., 2016; Greydanus et al., 2017; Iyer et al., 2018; Sundar, 2018; Yang et al., 2018; Annasamy & Sycara, 2019).

Saliency methods in computer vision and reinforcement learning use similar procedures to generate these maps. However, the temporal and interactive nature of RL systems presents a unique set of opportunities and challenges. Deep models in reinforcement learning select sequential actions whose effects can interact over long time periods. This contrasts strongly with visual classification tasks, in which deep models merely map from images to labels. For RL systems, saliency maps are often used to assess an agent’s internal representations and behavior over multiple frames in the environment, rather than to assess the importance of specific pixels in classifying images.

Despite their common use to explain agent behavior, it is unclear whether saliency maps provide useful explanations of the behavior of deep RL agents. Some prior work has evaluated the applicability of saliency maps for explaining the behavior of image classifiers (Samek et al., 2017; Adebayo et al., 2018; Kindermans et al., 2019), but there is not a corresponding literature evaluating the applicability of saliency maps for explaining RL agent behavior.

In this work, we develop a methodology grounded in counterfactual reasoning to empirically evaluate the explanations generated using saliency maps in deep RL. Specifically, we:

Survey the ways in which saliency maps have been used as evidence in explanations of deep RL agents.

Describe a new interventional method to evaluate the inferences made from saliency maps.

Experimentally evaluate how well the pixel-level inferences of saliency maps correspond to the semantic-level inferences of humans.

Interpreting Saliency Maps in Deep RL

Consider the saliency maps generated from a deep RL agent trained to play the Atari game Breakout. The goal of Breakout is to use the paddle to keep the ball in play so it hits bricks, eliminating them from the screen. Figure 1a shows a sample frame with its corresponding saliency. Note the high salience on the missing section of bricks (“tunnel”) in Figure 1a.

Creating a tunnel to target bricks at the top layers is one of the most high-profile examples of agent behavior being explained according to semantic, human-understandable concepts (Mnih et al., 2015). Given the intensity of saliency on the tunnel in 1a, it may seem reasonable to infer that this saliency map provides evidence that the agent has learned to aim at tunnels. If this is the case, moving the horizontal position of the tunnel should lead to similar saliency patterns on the new tunnel. However, Figures 1b and 1c show that the salience pattern is not preserved. Neither the presence of the tunnel, nor the relative positioning of the ball, paddle, and tunnel, are responsible for the intensity of the saliency observed in Figure 1a.

Examining how some of the technical details of reinforcement learning interact with saliency maps can help explain both the potential utility and the potential pitfalls of interpreting saliency maps. RL methods enable agents to learn how to act effectively within an environment by repeated interaction with that environment. Certain states in the environment give the agent positive or negative reward. The agent learns a policy, a mapping between states and actions according to these reward signals. The goal is to learn a policy that maximizes the discounted sum of rewards received while acting in the environment (Sutton & Barto, 1998). Deep reinforcement learning uses deep neural networks to represent policies. These models enable interaction with environments requiring high-dimensional state inputs (e.g., Atari games).

Consider the graphical model in Figure 2a representing the deep RL system for a vision-based game environment. Saliency maps are produced by performing some kind of intervention MM on this system and calculating the difference in logits produced by the original and modified images. The interventions used to calculate saliency for deep RL are performed at the pixel level (red node and arrow in Figure 2a). These interventions change the conditional probability distribution of “Pixels” by giving it another parent (Pearl, 2000).

Functionally, this can be accomplished through a variety of means, including changing the color of the pixel (Simonyan et al., 2014), adding a gray mask (Zeiler & Fergus, 2014), blurring a small region (Greydanus et al., 2017), or masking objects with the background color (Iyer et al., 2018). The interventions MM are used to simulate the effect of the absence of the pixel(s) on the network’s output. Note however that these interventions change the image in a way that is inconsistent with the generative process FF. They are not “naturalistic” interventions. This type of intervention produces images for which the learned network function may not be well-defined.

2 Explanations from Saliency Maps

To form explanations of agent behavior, human observers combine information from saliency maps, agent behavior, and semantic concepts. Figure 2b shows a system diagram of how these components interact. We note that semantic concepts are often identified visually from the pixel output as the game state is typically latent.

Counterfactual reasoning has been identified as a particularly effective way to present explanations of the decision boundaries of deep models (Mittelstadt et al., 2019). Humans use counterfactuals to reason about the enabling conditions of particular outcomes, as well as to identify situations where the outcome would have occurred even in the absence of some action or condition (de Graaf & Malle, 2017; Byrne, 2019). Saliency maps provide a kind of pixel-level counterfactual, but if the goal is to explain agent behavior according to semantic concepts, interventions at the pixel level seem unlikely to be sufficient.

Since many semantic concepts may map to the same set of pixels, it may be difficult to identify the functional relationship between changes in pixels and changes in network output according to semantic concepts or game state (Chalupka et al., 2015). Researchers may be interpreting differences in network outputs as evidence of differences in semantic concepts. However, changes in pixels do not guarantee changes in semantic concepts or game state.

In terms of changes to pixels, semantic concepts, and game state, we distinguish among three classes of interventions: distortion, semantics-preserving, and fat-hand (see Table 1). Semantics-preserving and fat-hand interventions are defined with respect to a specific set of semantic concepts. Fat-hand interventions change game state in such a way that the semantic concepts of interest are also altered.

The pixel-level manipulations used to produce saliency maps primarily result in distortion interventions, though some saliency methods (e.g., object-based) may conceivably produce semantics-preserving or fat-hand interventions as well. Pixel-level interventions are not guaranteed to produce changes in semantic concepts, so counterfactual evaluations that apply semantics-preserving interventions may be a more appropriate approach for precisely testing hypotheses of behavior.

Survey of Usage of Saliency Maps in Deep RL Literature

To assess how saliency maps are typically used to make inferences regarding agent behavior, we surveyed recent conference papers in deep RL. We focused our pool of papers on those that use saliency maps to generate explanations or make claims regarding agent behavior. Our search criteria consisted of examining papers that cited work that first described any of the following four types of saliency maps:

Wang et al. (2016) extend gradient-based saliency maps to deep RL by computing the Jacobian of the output logits with respect to a stack of input images.

Greydanus et al. (2017) generate saliency maps by perturbing the original input image using a Gaussian blur of the image and measure changes in policy from removing information from a region.

Iyer et al. (2018) use template matching, a common computer vision technique (Brunelli, 2009), to detect (template) objects within an input image and measure salience through changes in Q-values for masked and unmasked objects.

Most recently, attention-based saliency mapping methods have been proposed to generate interpretable saliency maps (Mott et al., 2019; Nikulin et al., 2019).

From a set of 90 papers, we found 46 claims drawn from 11 papers that cited and used saliency maps as evidence in their explanations of agent behavior. The full set of claims are given in Appendix C.

We found three categories of saliency map usage, summarized in Table 2. First, all claims interpret salient areas as a proxy for agent focus. For example, a claim about a Breakout agent notes that the network is focusing on the paddle and little else (Greydanus et al., 2017).

Second, 87% of the claims in our survey propose hypotheses about the features of the learned policy by reasoning backwards about what representation might jointly produce the observed saliency pattern and agent behavior. These types of claims either develop an a priori explanation of behavior and evaluate it using saliency, or they propose an ad hoc explanation after observing saliency to reason about how the agent is using salient areas. One a priori claim notes that the displayed score is the only differing factor between two states and evaluates that claim by noting that saliency focuses on these pixels (Zahavy et al., 2016). An ad hoc claim about a racing game notes that the agent is recognizing a time-of-day cue from the background color and acting to prepare for a new race (Yang et al., 2018).

Finally, only 7% (3 out of 46) of the claims drawn from saliency maps are accompanied by additional or more direct experimental evidence. One of these attempts to corroborate the interpreted saliency behavior by obtaining additional saliency samples from multiple runs of the game. The other two attempt to manipulate semantics in the pixel input to assess the agent’s response by, for example, adding an additional object to verify a hypothesis about memorization (Annasamy & Sycara, 2019).

2 Common Pitfalls in Current Usage

In the course of the survey, we also observed several more qualitative characteristics of how saliency maps are routinely used.

Subjectivity. Recent critiques of machine learning have already noted a worrying tendency to conflate speculation and explanation (Lipton & Steinhardt, 2018). Saliency methods are not designed to formalize an abstract human-understandable concept such as “aiming” in Breakout, and they do not provide a means to quantitatively compare semantically meaningful consequences of agent behavior. This leads to subjectivity in the conclusions drawn from saliency maps.

Unfalsiability. One hallmark of a scientific hypothesis or claim is falsifiability (Popper, 1959). If a claim is false, its falsehood should be identifiable from some conceivable experiment or observation. One of the most disconcerting practices identified in the survey is the presentation of unfalsifiable interpretations of saliency map patterns. An example: “A diver is noticed in the saliency map but misunderstood as an enemy and being shot at” (see Appendix C). It is unclear how we might falsify an abstract concept such as “misunderstanding”.

Cognitive Biases. Current theory and evidence from cognitive science implies that humans learn complex processes, such as video games, by categorizing objects into abstract classes and by inferring causal relationships among instances of those classes (Tenenbaum & Niyogi, 2003; Dubey et al., 2018). Our survey suggests that researchers infer that: (1) salient regions map to learned representations of semantic concepts (e.g., ball, paddle), and (2) the relationships among the salient regions map to high-level behaviors (e.g., tunnel-building, aiming). Researchers’ expectations impose a strong bias on both the existence and nature of these mappings.

Methodology

Our survey indicates that many researchers use saliency maps as an explanatory tool to infer the representations and processes behind an agent’s behavior. However, the extent to which such inferences are valid has not been empirically evaluated under controlled conditions.

In this section, we show how to generate falsifiable hypotheses from saliency maps and propose an intervention-based approach to verify the hypotheses generated from saliency maps. We intervene on game state to produce counterfactual semantic conditions. This provides a concrete methodology to assess the relationship between saliency and learned semantic representations.

Though saliency maps may not relate directly to semantic concepts, they may still be an effective tool for exploring hypotheses about agent behavior. As we show schematically in Figure 2b, claims or explanations informed by saliency maps have three components: semantic concepts, saliency, and behavior. Recall that our survey indicates that researchers often attempt to infer aspects of the network’s learned representations from saliency patterns. Let XX be a subset of the semantic concepts that can be inferred from the input image. Let BB represent behavior, or aggregate actions, over temporally extended sequences of frames, and let RR be a representation that is a function of some pixels that the agent learns during training.

To create scientific claims from saliency maps, we recommend using a relatively standard pattern which facilitates objectivity and falsifiability:

{concept set XX} is salient   ⟹  \implies agent has learned {representation RR} resulting in {behavior BB}.

Consider the Breakout brick reflection example presented in Section 2. The hypothesis introduced (“the agent has learned to aim at tunnels”) can be reformulated as: bricks are salient   ⟹  \implies agent has learned to identify a partially complete tunnel resulting in maneuvering the paddle to hit the ball toward that region. Stating hypotheses in this format implies falsifiable claims amenable to empirical analysis.

As indicated in Figure 2, the learned representation and pixel input share a relationship with saliency maps generated over a sequence of frames. Given that the representation learned is static, the relationship between the learned representation and saliency should be invariant under different manipulations of pixel input. We use this property to assess saliency under counterfactual conditions.

We generate counterfactual conditions by intervening on the RL environment. Prior work has focused on manipulating the pixel input. However, this does not modify the underlying latent game state. Instead, we intervene directly on game state. In the do-calculus formalism (Pearl, 2000), this shifts the intervention node in Figure 2a to game state, which leaves the generative process FF of the pixel image intact.

We employ ToyBox, a set of fully parameterized implementation of Atari games (Foley et al., 2018), to generate interventional data under counterfactual conditions. The interventions are dependent on the mapping between semantic concepts and learned representations in the hypotheses. Given a mapping between concept set XX and a learned representation RR, any intervention would require meaningfully manipulating the state in which XX resides to assess the saliency on XX under the semantic treatment applied. Saliency on x∈Xx\in X is defined as the average saliency over a bounding-boxToyBox provides transparency into the position of each object (measured at the center of sprite), which we use as the center of the bounding boxes in our experiments. around xx.

Since the learned policies should be semantically invariant under manipulations of the RL environment, by intervening on state, we can verify whether the counterfactual states produce expected patterns of saliency on the associated concept set XX. If the counterfactual saliency maps reflect similar saliency patterns, this provides stronger evidence that the observed saliency indicates the agent has learned representation R corresponding to semantic concept set XX.

Evaluation of Hypotheses on Agent Behavior

We conduct three case studies to evaluate hypotheses about the relationship between semantic concepts and semantic processes formed from saliency maps. Each case study uses observed saliency maps to identify hypotheses in the format described in Section 4. The hypotheses were generated by watching multiple episodes and noting atypical, interesting or popular behaviors from saliency maps. In each case study, we produce Jacobian, perturbation and object saliency maps from the same set of counterfactual states. We include examples of each map in Appendix A. Using ToyBox allows us to produce counterfactual states and to generate saliency maps in these altered states. The case studies are conducted on two Atari games, Breakout and Amidar.The code pertaining to the experiments can be found at https://github.com/KDL-umass/saliency_maps. The deterministic nature of both games allows some stability in the way we interpret the network’s action selection. Each map is produced from an agent trained with A2C (Mnih et al., 2016) using a CNN-based (Mnih et al., 2015) OpenAI Baselines implementation (Dhariwal et al., 2017) with default hyperparameters (see Appendix B for more details). Our choice of model is arbitrary. The emphasis of this work is on methods of explanation, not the explanations themselves.

Here we evaluate the behavior from Section 2:

Hypothesis 1: {bricks} are salient   ⟹  \implies agent has learned to {identify a partially complete tunnel} resulting in {maneuvering the paddle to hit the ball toward that region}.

To evaluate this hypothesis, we intervene on the state by translating the brick configurations horizontally. Because the semantic concepts relating to the tunnel are preserved under translation, we expect salience will be nearly invariant to the horizontal translation of the brick configuration. Figure 3a depicts saliency after intervention. Salience on the tunnel is less pronounced under left translation, and more pronounced under right translation. Since the paddle appears on the right, we additionally move the ball and paddle to the far left (Figure 3b).

Conclusion. Temporal association (e.g. formation of a tunnel followed by higher saliency) does not generally imply causal dependence. In this case, tunnel formation and salience appear to be confounded by location or, at least, the dependence of these phenomena are highly dependent on location.

Amidar is a Pac-Man-like game in which an agent attempts to completely traverse a series of passages while avoiding enemies. The yellow sprite that indicates the location of the agent is almost always salient in Amidar. Surprisingly, the displayed score is often as salient as the yellow sprite throughout the episode with varying levels of intensity. This can lead to multiple hypotheses about the agent’s learned representation: (1) the agent has learned to associate increasing score with higher reward; (2) due to the deterministic nature of Amidar, the agent has created a lookup table that associates its score and its actions. We can summarize these as follows:

Hypothesis 2: {score} is salient   ⟹  \implies agent has learned to {use score as a guide to traverse the board} resulting in {successfully following similar paths in games}.

To evaluate hypothesis 2, we designed four interventions on score:

intermittent_reset: modify the score to 0 every x∈x\in timesteps.

random_varying: modify the score to a random number between every x∈x\in timesteps.

fixed: select a score from and fix it for the whole game.

decremented: modify score to be 3000 initially and decrement score by d∈d\in at every timestep.

Figures 4a and 4b show the result of intervening on displayed score on reward and saliency intensity, measured as the average saliency over a 25x15 bounding box, respectively for the first 1000 timesteps of an episode. The mean is calculated over 50 samples. If an agent died before 1000 timesteps, the last reward was extended for the remainder of the timesteps and saliency was set to zero.

Using reward as a summary of agent behavior, different interventions on score produce different agent behavior. Total accumulated reward differs over time for all interventions, typically due to early agent death. However, salience intensity patterns of all interventions follow the original trajectory very closely. Different interventions on displayed score cause differing degrees of degraded performance (Figure 4a) despite producing similar saliency maps (Figure 4b), indicating that agent behavior is underdetermined by salience. Specifically, the salience intensity patterns are similar for the control, fixed, and decremented scores, while the non-ordered score interventions result in degraded performance. Figure 4c indicates only very weak correlations between the difference-in-reward and difference-in-saliency-under-intervention as compared to the original trajectory. Correlation coefficients range from 0.041 to 0.274, yielding insignificant p-values for all but one intervention. See full results in Appendix LABEL:app:exp2, Table LABEL:tab:quant_score.

Similar trends are noted for Jacobian and perturbation saliency methods in Appendix LABEL:app:exp2.

Conclusion. The existence of a high correlation between two processes (e.g., incrementing score and persistence of saliency) does not imply causation. Interventions can be useful in identifying the common cause leading to the high correlation.

Enemies are salient in Amidar at varying times. From visual inspection, we observe that enemies close to the player tend to have higher saliency. Accordingly, we generate the following hypothesis:

Hypothesis 3: {enemy} is salient   ⟹  \implies agent has learned to {identify enemies close to it} resulting in {successful avoidance of enemy collision}.

Without directly intervening on the game state, we can first identify whether the player-enemy distance and enemy saliency is correlated using observational data. We collect 1000 frames of an episode of Amidar and record the Manhattan distance between the midpoints of the player and enemies, represented by 7x7 bounding boxes, along with the object salience of each enemy. Figure 5a shows the distance of each enemy to the player over time with saliency intensity represented by the shaded region. Figure 5b shows the correlation between the distance to each enemy and the corresponding saliency. Correlation coefficients and significance values are reported in Table 3. It is clear that there is no correlation between saliency and distance of each enemy to the player.

Given that statistical dependence is almost always a necessary pre-condition for causation, we expect that there will not be any causal dependence. To further examine this, we intervene on enemy positions of salient enemies at each timestep by moving the enemy closer and farther away from the player. Figure 5c contains these results. Given Hypothesis 3, we would expect to see an increasing trend in saliency for enemies closer to the player. However, the size of the effect is close to 0 (see Table 3). In addition, we find no correlation in the enemy distance experiments for the Jacobian or perturbation saliency methods (included in Appendix LABEL:app:exp3).

Conclusion. Spurious correlations, or misinterpretations of existing correlation, can occur between two processes (e.g. correlation between player-enemy distance and saliency), and human observers are susceptible to identifying spurious correlations (Simon, 1954). Spurious correlations can sometimes be identified from observational analysis without requiring interventional analysis.

Discussion and Related Work

Thinking counterfactually about the explanations generated from saliency maps facilitates empirical evaluation of those explanations. The experiments above show some of the difficulties in drawing conclusions from saliency maps. These include the tendency of human observers to incorrectly infer association between observed processes, the potential for experimental evidence to contradict seemingly obvious observational conclusions, and the challenges of potential confounding in temporal processes.

One of the main conclusions from this evaluation is that saliency maps are an exploratory tool rather than an explanatory tool for evaluating agent behavior in deep RL. Saliency maps alone cannot be reliably used to infer explanations and instead require other supporting tools. This can include combining evidence from saliency maps with other explanation methods or employing a more experimental approach to evaluation of saliency maps such as the approach demonstrated in the case studies above.

The framework for generating falsifiable hypotheses suggested in Section 4 can assist with designing more specific and falsifiable explanations. The distinction between the components of an explanation, particularly the semantic concept set XX, learned representation RR and observed behavior BB, can further assist in experimental evaluation. Note that the semantic space devised by an agent might be quite different from the semantic space given by the latent factors of the environment. It is crucial to note that this mismatch is one aspect of what plays out when researchers create hypotheses about agent behavior, and the methodology we provide in this work demonstrates how to evaluate hypotheses that reflect that mismatch.

The methodology presented in this work can be easily extended to other vision-based domains in deep RL. Particularly, the framework of the graphical model introduced in Figure 2a applies to all domains where the input to the network is image data. An extended version of the model for Breakout can be found in Appendix LABEL:fig:breakout_gm.

We propose intervention-based experimentation as a primary tool to evaluate the hypotheses generated from saliency maps. Yet, alternative methods can identify a false hypothesis even earlier. For instance, evaluating statistical dependence alone can provide strong evidence against causal dependence (e.g., Case Study 3). In this work, we employ a particularly capable simulation environment (ToyBox). However, limited forms of evaluation may be possible in non-intervenable environments, though they may be more tedious to implement. For instance, each of the interventions conducted in Case Study 1 can be produced in an observation-only environment by manipulating the pixel input (Brunelli, 2009; Chalupka et al., 2015). Developing more experimental systems for evaluating explanations is an open area of research.

This work analyzes explanations generated from feed-forward deep RL agents. However, the proposed methodology is not model dependent, and aspects of the approach will carry over to recurrent deep RL agents. The proposed methodology would not work well for repeated interventions on recurrent deep RL agents due to their capacity for memorization.

Prior work has introduced alternatives to the use of saliency maps to support explanation of deep RL agents. Some of these methods also use counterfactual reasoning to develop explanations. ToyBox was developed to support experimental evaluation and behavioral tests of deep RL models (Tosch et al., 2019). Olson et al. (2019) use a generative deep learning architecture to produce counterfactual states resulting in the agent taking a different action. Others have proposed alternative methods for developing semantically meaningful interpretations of agent behavior. Juozapaitis et al. (2019) use reward decomposition to attribute policy behaviors according to semantically meaningful components of reward. Verma et al. (2018) use domain-specific languages for policy representation, allowing for human-readable policy descriptions.

Prior work in the deep network literature has evaluated and critiqued saliency maps. Kindermans et al. (2019) and Adebayo et al. (2018) demonstrate the utility of saliency maps by adding random variance in input. Seo et al. (2018) provide a theoretical justification of saliency and hypothesize that there exists a correlation between gradients-based saliency methods and model interpretation. Samek et al. (2017) and Hooker et al. (2019) present evaluations of existing saliency methods for image classification.

Conclusions

We conduct a survey of uses of saliency maps, propose a methodology to evaluate saliency maps, and examine the extent to which the agent’s learned representations can be inferred from saliency maps. We investigate how well the pixel-level inferences of saliency maps correspond to the semantic concept-level inferences of human-level interventions. Our results show saliency maps cannot be trusted to reflect causal relationships between semantic concepts and agent behavior. We recommend saliency maps to be used as an exploratory tool, not explanatory tool.

Thanks to Emma Tosch, Amanda Gentzel, Deep Chakraborty, Abhinav Bhatia, Blossom Metevier, Chris Nota, Karthikeyan Shanmugam and the anonymous ICLR reviewers for thoughtful comments and contributions. This material is based upon work supported by the United States Air Force under Contract No, FA8750-17-C-0120. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the United States Air Force.

References

Appendices

Appendix A Saliency Methods

Figure 6 shows example saliency maps of the three saliency methods evaluated in this work, namely perturbation, object and Jacobian, for Amidar.

Appendix B Model

We use the OpenAI Baselines’ implementation (Dhariwal et al., 2017) of an A2C model (Mnih et al., 2016) to train the RL agents on Breakout and Amidar. The model uses the CNN architecture proposed by Mnih et al. (2015). Each agent is trained for 40 million iterations using RMSProp with default hyperparameters (Table 4).

Appendix C Survey of Usage of Saliency Maps in Deep RL Literature

We conducted a survey of recent literature to assess how saliency maps are used to interpret agent behavior in deep RL. We began our search by focusing on work citing the following four types of saliency maps: Jacobian (Wang et al., 2016), perturbation (Greydanus et al., 2017), object (Iyer et al., 2018) and attention (Mott et al., 2019). Papers were selected if they employed saliency maps to create explanations regarding agent behavior. This resulted in selecting 46 claims from 11 papers. These 11 papers have appeared at ICML (3), NeurIPS (1), AAAI (2), ArXiv (3), OpenReview (1) and as a thesis (1). There are several model-specific saliency mapping methods that we excluded from our survey.

Following is the full set of claims. All claims are for Atari games. The Reason column represents whether an explanation for agent behavior was provided (Y/N) and the Exp column represents whether an experiment was conducted to evaluate the explanation (Y/N).

We further evaluated the effects of interventions on displayed score (Section 5) in Amidar on perturbation and Jacobian saliency maps. These results are presented in Figures LABEL:amidar_score_ex_pert and LABEL:amidar_score_ex_jac, respectively.

We also evaluated Pearson’s correlation between the differences in reward and saliency with the original trajectory for all three methods (see Table LABEL:tab:quant_score). The results support the correlation plots in Figures 4c, LABEL:amidar_score_ex_pertc and LABEL:amidar_score_ex_jacc.

E.2 Case Study 3: Amidar Enemy Distance

We further evaluated the relationship between player-enemy distance and saliency (Section 5) in Amidar on perturbation and Jacobian saliency maps. These results are presented in Figures LABEL:amidar_dist_ex_pert and LABEL:amidar_dist_ex_jac, respectively. Jacobian saliency performed the worst for the intervention-based experiment, suggesting that there is no impact of player-enemy distance on saliency.

Regression analysis between distance and perturbation and Jacobian saliency can be found in Tables LABEL:tab:quant_distance_per and LABEL:tab:quant_distance_jac respectively. The results support the lack of correlation between the observational and interventional distributions. Note, enemy 1 is more salient throughout the game compared to the other four enemies resulting in a larger interventional sample size.