Nothing Else Matters: Model-Agnostic Explanations By Identifying Prediction Invariance
Marco Tulio Ribeiro, Sameer Singh, Carlos Guestrin
Introduction
At the core of interpretable machine learning is the question of whether humans are able to make accurate predictions about a model’s behavior. Assumed in this question are three properties of the interpretable output: coverage, precision, and effort. Coverage refers to how often humans think they can predict the model’s behavior, precision to how accurate humans are in those predictions, and effort is either the up-front effort required in interpreting the model, or the effort required to make predictions about a model’s behavior.
One approach to interpretable machine learning is designing inherently interpretable models. Visualizations of these models usually have perfect coverage, but there is a trade-off between the accuracy of the model and the effort required to comprehend it - especially in complex domains like text and images, where the input space is very large, and accuracy is usually sacrificed for models that are compact enough to be comprehensible by humans. Experiments usually involve showing humans these visualizations, and measuring human precision when predicting the model’s behavior on random instances, and the time (effort) required to make those predictions .
Model-agnostic explanations avoid the need to trade off accuracy by treating the model as a black box. Explanations such as sparse linear models (henceforth called linear LIME) or gradients can still exhibit high precision and low effort (which are de-facto requirements, as there is little point in explaining a model if explanations lead to poor understanding or are too complex) even for very complex models by providing explanations that are local in their scope (i.e. not perfect coverage). However, the coverage of such explanations are not explicit, which may lead to human error. Take the example on Figure 1: we explain a prediction of a complex model, which predicts that the person described by Figure 1(a) makes less than $50K. The linear LIME explanation (Figure 1(b)) sheds some light into why, but it is not clear whether we can apply the insights from this explanation to other instances. In other words, even if the explanation is faithful locally, it is not easy to know what that local region is. Furthermore, it is not clear when the linear approximation is more or less faithful, even within the local region.
In this paper, we introduce Anchor Local Interpretable Model-Agnostic Explanations (aLIME), a system that explains individual predictions with if-then rules (similar to Lakkaraju et al. ) in a model-agnostic manner. Such rules are intuitive to humans, and usually require low effort to comprehend and apply. In particular, an aLIME explanation (or an anchor) is a rule that sufficiently “anchors” a prediction – such that changes to the rest of the instance do not matter (with high probability). For example, the anchor in Figure 1(c) states that the model will almost always predict Salary if a person is not educated beyond high school, regardless of the other features. Such explanations make their coverage very clear - they only apply when the conditions in the rule are met. We propose a method to compute such explanations that guarantees high precision with a high probability. Further, we present empirical comparison against linear LIME and qualitative evaluation on a variety of tasks (such as text/image classification and visual question answering) to demonstrate that anchors are intuitive, have high precision, and very clear coverage boundaries.
Anchors as Model-Agnostic Explanations
Let the model being explained be denoted , such that we explain individual predictions . Let an anchor be defined as a set of constraints (i.e. a rule with conjunctions), being the set of all possible constraints that are met by . For example, in Figure 1(c), c=\{\text{Education\leqHigh School}\}.
We assume we have a distribution of interest , and that we can sample from - that is, we can sample inputs where the constraints for the anchor are met. The reason to condition on is that may depend on the instance being explained (for example, see image classification in Section 4). The precision of an anchor is then defined as the expected accuracy (under ) of applying the anchor to instances that meet its constraints, formalized in Equation 1.
As argued before, high precision is a requirement of model-agnostic explanations. It is trivial to get a perfectly precise (yet useless) anchor by having the constraint set be so specific that only the example being explained meets it. In order to balance precision, coverage and effort, we optimize the objective in Equation 2, where we try to find the shortest anchor with high precision. The length of the anchor can be used as a proxy for effort, and more specific (longer) anchors will naturally have less coverage.
Algorithm: Solving Equation 2 exactly is unfeasible – precision cannot be computed exactly for arbitrary and , and finding the best has combinatorial complexity. To address the former, we approximate the precision via sampling, and solve the probably approximately correct (PAC) version of Equation 2 so that the chosen anchor will have high precision with high probability. For the latter, we employ an algorithm similar in spirit to lazy decision trees , where we construct greedily. In particular, at each step, we want to pick the constraint that dominates all other constraints in terms of precision, until the stopping criterion in Equation 2 is met. For efficiency, we want to sample as few instances as possible to make each greedy decision. We use Hoeffding bounds for the differences in precision to decide when a constraint dominates all the other constraints with high probability. This uses the same insight as Hoeffding trees , with the key difference that we can control the sampling distribution, and thus can use the bounds to sample the regions of the input space that reduce the uncertainty between the precision estimates with as few samples as possible. Due to lack of space, we omit the details of the algorithm.
Simulated Experiments
In order to evaluate the difference between linear LIME and anchor LIME (aLIME) in terms of coverage and precision, we perform simulated experiments on two UCI datasets: adult and hospital readmission. The latter is a 3-class classification problem, where the task is to predict if a patient will be readmitted to the hospital after an inpatient encounter within 30 days, after 30 days, or never.
For each dataset, we learn a gradient boosted tree classifier with trees, and generate explanations for instances in the validation dataset. We then evaluate the coverage and precision of these explanations on a separate test dataset. We use for anchor (that is, we expect precision to be close to %) unless noted otherwise, and consider that a linear LIME explanation covers every other instance within distance , where is a parameter that we vary. For aLIME, we sample from by sampling whole rows from the dataset except for the features constrained by . We evaluate explanations, chosen either at random (RP) or via Submodular Pick (SP), a procedure that picks explanations to maximize the coverage , on the validation.
We show precision-coverage plots of a single explanation () in Figures 2(a) and 3(a), where we vary for linear LIME and for aLIME. The results show that for any level of coverage, aLIME has better precision than linear LIME. Furthermore, using submodular pick greatly increases the coverage at the same precision level. Linear LIME performs particularly worse in the dataset with the highest number of dimensions (hospital readmission), where the distance degrades. We note that one of the main advantages of aLIME over linear LIME is making its coverage clear to humans - without human experiments, there is no way to know what these plots look like for linear LIME, but we can expect them to be the same for aLIME.
We vary the number of explanations the simulated user sees () in Figures 2(b), 2(c), 3(b) and 3(c). In order to keep the results comparable, we set and picked such that the average precision at for linear LIME was at least . In both datasets, aLIME is able to maintain higher precision regardless of how many explanations are shown, with coverage that dominates linear LIME. It is worth noting that both datasets are of the same data type (tabular), and are such that the behavior of the model is simple in a large part of the input space (good conditions for aLIME), as demonstrated by the top anchors that maximize coverage for each dataset in Figure 4. While these models are complex, their behavior on most of the input space (65% for adult and 81% for hospital readmission) is covered by these simple rules with high precision.
Qualitative Examples
Part-of-Speech tagging: We use a black box state-of-the-art POS tagger (http://spacy.io), and explain tag predictions for the word play in different contexts in Table 1. The anchors demonstrate that the POS picks up on the correct patterns. Furthermore, they are short and easy to understand. Anchors are particularly suited for this task, where the dimensionality is small and the behavior of good models is more easily captured by IF-THEN rules than linear models.
Image classification: We use aLIME to explain a prediction from the Inception V3 classifier on an image of a Zebra in Figure 5, where we first split the image into superpixels. The anchor in Figure 5(b) means that if we fix the non-grey superpixels, we can substitute the greyed-out superpixels by a random image, and the model will predict “zebra” around 95% of the time. To illustrate this, we display on Figure 5(c) a set of images from (i.e. where the anchor is fixed), and the model predicts “zebra”. While this choice of distribution produces images that look nothing like real images (Figure 5(c)), it makes for more robust explanations than distributions that only hide parts of the image with gray or dark patches (). This anchor demonstrates that the model picks up on a pattern that does not require a zebra to have four legs, or even a head - which is a pattern very different than the patterns humans use to detect zebras.
Visual Question Answering: Visual QA models are multi-modal, and thus can be explained in terms of the image, the question or both. Here, we find anchors on the questions, leaving the image fixed, and use a bigram language model trained on input questions as . We select two questions to explain, which are the top rows (in purple) of Figures 6(b) and 6(c). The anchors (in bold), are respectively “What” and “many”, and we show questions drawn from below the original question. The first anchor states that if “What” is in the question, the answer will be “banana” about 95% of the time, while the latter states the same about “many” and “2”, respectively – both explanations clearly indicate undesirable behavior from the model. Again, this kind of explanation is intuitive and easier to understand than a linear model, even one with high weight on the words “What” and “banana”, as one knows exactly when it applies and when it does not.
Conclusion
In this work, we argued that high precision and clear coverage bounds are very desirable properties of model-agnostic explanations. We introduced aLIME, a system is designed to produce rule-based explanations that exhibit both these properties. IF-THEN rules are intuitive and easy to understand, and identifying parts of the input that result in prediction invariance (i.e. the rest does not matter) is similar to how humans explain many of their choices. We demonstrated aLIME’s flexibility by explaining predictions from a variety of classifiers on a myriad of domains, outperforming linear explanations from LIME on simulated experiments.