"How do I fool you?": Manipulating User Trust via Misleading Black Box Explanations

Himabindu Lakkaraju, Osbert Bastani

Introduction

There has been an increasing interest in using ML models to aid decision makers in domains such as healthcare and criminal justice. In these domains, it is critical that decision makers understand and trust ML models, to ensure that they can diagnose errors and identify model biases correctly. However, ML models that achieve state-of-the-art accuracy are typically complex black boxes that are hard to understand. As a consequence, there has been a recent surge in post hoc explanation techniques for explaining black box models (?; ?; ?; ?). One of the goals of such explanations is to help domain experts detect systematic errors and biases in black box model behavior (?).

Existing techniques for explaining black boxes typically rely on optimizing fidelity—i.e., ensuring that the explanations accurately mimic the predictions of black box model (?; ?; ?). The key assumption underlying these approaches is that if an explanation has high fidelity, then biases of the black box model will be reflected in the explanation. However, it is questionable whether this assumption actually holds in practice (?). The key issue is that high fidelity only ensures high correlation between the predictions of the explanation and the predictions of the black box. There are several other challenges associated with post hoc explanations which are not captured by the fidelity metric: (i) they may fail to capture causal relationships between input features and black box predictions (?; ?), (ii) there could be multiple high-fidelity explanations for the same black box that look qualitatively different (?), and (iii) they may not be robust and can vary significantly even with small perturbations to input data (?).

These challenges increase the possibility that explanations generated using existing techniques can actually mislead the decision maker into trusting a problematic black box. However, there has been little to no prior work empirically studying if and how explanations can mislead users.

Contributions. We propose the first systematic study to explore if and how explanations of black boxes can mislead users. First, we propose a novel theoretical framework for understanding when misleading explanations can exist. We show that even if an explanation achieves perfect fidelity, it may still not reflect issues in the black box model. The key issue is that due to correlations in the features, explanations can achieve high fidelity even if they use entirely different features compared to the black box. Second, we propose a novel approach for generating potentially misleading explanations. Our approach extends the MUSE framework (?) to favor explanations that contain features that users believe are relevant and omit features that users believe are problematic. Third, we perform an extensive user study with domain experts from law and criminal justice to understand how misleading explanations impact user trust. Our results demonstrate that the misleading explanations generated using our approach can in fact increase user trust of by 9.8 times (See Figure 2). Our findings have far reaching implications both for research on ML interpretability and real-world applications of ML.

Related work. Present work on interpretable ML largely falls into three categories. First, there are approaches focused on learning predictive models that are human understandable (?; ?; ?). However, complex models such as deep neural networks and random forests typically achieve higher performance compared to interpretable models (?), so in many situations it is more desirable to use these complex models. Thus, there has been work on explaining such complex black boxes. One approach is to provide local explanations for individual predictions of the black box (?; ?; ?), which is useful when a decision maker plans to review every decision made by the black box. An alternate approach is to provide a global explanation that describes the black box as a whole, typically summarizing it using an interpretable model (?; ?), which is useful in validating the black boxes before they are deployed to automatically make decisions (i.e., without human involvement).

There has been some empirical work on studying how humans understand and trust interpretable models and explanations. For instance, Poursabzi-Sangdeh et. al. (2018) show that longer explanations are harder for humans to simulate accurately. There has also been recent work on understanding what makes explanations useful in the context of three tasks they are likely to perform given an explanation of an ML system: (i) predicting the system’s output, (ii) verifying whether the output is consistent with the explanation, and (iii) determining if and how the output would change if we change the input (?).

More closely related to our work, there has been recent work on exploring the vulnerabilities of black box explanations. For instance, there has been work demonstrating that explanations can be unstable, changing drastically even with small perturbations to inputs (?; ?). Finally, recent work has argued that black box explanations can often be misleading and can potentially lead users to trust problematic black boxes (?; ?).

In contrast, we are the first to study if and how adversarial entities could generate misleading explanations to manipulate user trust. We are also the first to explore the notion of confirmation bias in the context of black box explanations.

Problem Formulation

In this section, we introduce some notation and formalize the notions of (i) explanation of a black box model, and (ii) misleading explanation of a black box model.

Explanations. Given input data X\mathcal{X}, a set of class labels Y={1,2,⋯K}\mathcal{Y}=\{1,2,\cdots K\}, and a black box B:X→YB:\mathcal{X}\to\mathcal{Y}, our goal is to generate an explanation EE that describes the behavior of BB. Then, end users can use EE to determine whether to trust BB.

We consider an approach to explaining BB by approximating it using an interpretable model E∈EE\in\mathcal{E}. We measure the quality of this approximation using the relative error

Trustworthy black boxes & misleading explanations. We assume a workflow where the human user relies on E^\hat{E} to decide whether to trust BB. We model the human user as an oracle O:E→{0,1}\mathcal{O}:\mathcal{E}\to\{0,1\} such that

Constructing misleading explanations. Our goal is to demonstrate that misleading explanations exist. In our approach, we first devise a black box BB that we expect to be untrustworthy. This expectation is based on which features are used by the model (see Section 3). Then, we need to check if BB is actually untrustworthy (i.e., O∗(B)=0\mathcal{O}^{*}(B)=0). To do so, we choose BB to itself be an interpretable model. Then, we perform a user study where we show BB and ask if it is trustworthy, yielding O∗(B)\mathcal{O}^{*}(B). In this approach, BB is still a black box in the sense that (i) E^\hat{E} is constructed without examining the internals of BB, and (ii) users are not aware of the internals of BB when shown EE to evaluate O(E)\mathcal{O}(E).

Next, we construct an explanation EE of BB that we expect to be misleading; again, this expectation is based on which features are in the explanation (see Section 3). Then, we check if EE is indeed misleading (i.e., evaluate O(E)\mathcal{O}(E)) via a user study. Assuming we successfully constructed BB so that O∗(B)=0\mathcal{O}^{*}(B)=0, then EE is misleading if O(E)=1\mathcal{O}(E)=1. We discuss how we construct EE in Section 4 (BB is constructed similarly), and how we perform the user studies in Section 5.

Theoretical Framework

We define notions of a potentially untrustworthy black box BB and a potentially misleading explanation EE for BB. These notions are only used to guide our algorithms; once we have constructed BB and EE, we test whether BB is actually untrustworthy and EE is actually misleading via user studies. Finally, we discuss when potentially misleading explanations exist.

Quantifying user trust. We consider a simple approach to estimating whether a user trusts BB given EE. We assume their key criterion is which features are included in EE and which ones are omitted. More precisely, we assume the feature space can be decomposed into X=XD×XA×XP\mathcal{X}=\mathcal{X}_{D}\times\mathcal{X}_{A}\times\mathcal{X}_{P}, where XD\mathcal{X}_{D} corresponds to the desired features DD that the user expects to be included, XA\mathcal{X}_{A} corresponds to the ambivalent features AA for which the user is indifferent about whether they are included, and XP\mathcal{X}_{P} corresponds to the prohibited features PP that the user expects to be omitted.

Now, we say BB is potentially untrustworthy if O^∗(B)=0\hat{\mathcal{O}}^{*}(B)=0, and say EE is potentially misleading if O^(E)≠O^∗(B)\hat{\mathcal{O}}(E)\neq\hat{\mathcal{O}}^{*}(B). Figure 2 shows a potentially untrustworthy blackbox (left) and a potentially misleading explanation (right).

Existence of potentially misleading explanations. We study when potentially misleading explanations exist. First, even if an explanation has perfect fidelity, it can still be potentially misleading:

There exists a black box BB and an explanation EE of BB such that (i) EE has perfect fidelity (i.e., L(E,B)=0L(E,B)=0), and (ii) EE is potentially misleading. [See Appendix B for proof]

This result is for a specific black box and a specific explanation of that black box. Next, we study more general settings where potentially misleading explanations exist. Let E∈EE\in\mathcal{E} be the best explanation for black box BB. We focus on the case where O^∗(B)=0\hat{\mathcal{O}}^{*}(B)=0 (i.e., the black box is potentially untrustworthy), so EE is potentially misleading if O^(E)=1\hat{\mathcal{O}}(E)=1. Intuitively, potentially misleading explanations exist when the prohibited features PP can be reconstructed from the remaining ones D∪AD\cup A. In this case, a misleading explanation can internally reconstruct PP using the D∪AD\cup A. A potential concern is that even when PP can be reconstructed, it may not be possible to do so using an interpretable model. We show that an acceptable interpretable model can reconstruct PP as long as (i) an acceptable black box B+B_{+} can reconstruct PP and achieve good accuracy, and (ii) we can explain B+B_{+} using an acceptable interpretable model that achieves high fidelity. Intuitively, we expect (i) to hold when PP can be reconstructed from D∪AD\cup A, and we expect (ii) to hold since an explanation of B+B_{+} should not depend on features not in B+B_{+}.

We formalize (i) and (ii). For (i), let B+∈B+B_{+}\in\mathcal{B}_{+} be the best acceptable blackbox. The restriction error is ϵR=L(B+,B)\epsilon_{R}=L(B_{+},B). Then, (i) corresponds to ϵR≈0\epsilon_{R}\approx 0—i.e., PP can be reconstructed from D∪AD\cup A when B+B_{+} can then achieve loss similar to BB by internally reconstructing PP. For (ii), let E′∈EE^{\prime}\in\mathcal{E} be the best explanation for B+B_{+}, and let E+∈E+E_{+}\in\mathcal{E}_{+} be the best acceptable explanation of B+B_{+}. The acceptable relative error is the gap in fidelity between these two—i.e.,

Then, (ii) corresponds to ϵA≈0\epsilon_{A}\approx 0—i.e., E+E_{+} is almost as good an explanation of B+B_{+} as E′E^{\prime}. Intuitively, this assumption should hold since B+B_{+} does not use PP, so there should exist a high fidelity explanation of B+B_{+} that does not use PP.

Finally, suppose that ϵR,ϵA\epsilon_{R},\epsilon_{A} are small, and that there exists a high fidelity explanation E∈EE\in\mathcal{E} (which may not be acceptable); then, E+E_{+} is potentially misleading:

Suppose O∗(B)=0\mathcal{O}^{*}(B)=0; if L(E,B)+2ϵR+ϵA≤ϵ+L(E,B)+2\epsilon_{R}+\epsilon_{A}\leq\epsilon_{+}, then E+E_{+} is potentially misleading. [See Appendix B for proof]

Generating Misleading Explanations

Our algorithm for constructing misleading explanations of black boxes builds on the Model Understanding through Subspace Explanations (MUSE) framework (?) by incorporating additional constraints that enable us to output high fidelity explanations that include desired features and omit prohibited features.

Given a black box, MUSE produces an explanation in the form of a two-level decision set, which intuitively is a model consisting of nested if-then statements where the nesting depth is two. MUSE chooses an explanation that maximizes two objectives: (i) interpretability: easier for humans to understand, and (ii) fidelity: the explanation should mimic the behavior of the black box.

Two-level decision sets. A two-level decision set R:X→YR:\mathcal{X}\to\mathcal{Y} is a hierarchical model consisting of a set of decision sets, each of which is embedded within an outer if-then structure. The clauses within each of the two levels are unordered, so multiple rules may apply to a given example x∈Xx\in\mathcal{X}. Ties between different if-then clauses are broken according to which rules are most accurate; see (?) for details. Intuitively, the outer if-then rules can be thought of as neighborhood descriptors which correspond to different parts of the feature space, and the inner if-then rules are patterns of model behaviors within the corresponding neighborhood. Formally, a two-level decision set has form

where ci∈Yc_{i}\in\mathcal{Y} is a label, and qiq_{i} and sis_{i} are conjunctions of predicates of the form “feature∼value\text{feature}\sim\text{value}”, where ∼  ∈{=,≥,≤}\sim\;\in\{=,\geq,\leq\} is an operator; e.g., “age≥50\text{age}\geq 50” is a predicate. In particular, qiq_{i} corresponds to the neighborhood descriptor, and (si,ci)(s_{i},c_{i}) together represent the inner if-then rules with sis_{i} denoting the antecedent (i.e., the if condition) and cic_{i} denoting the consequent (i.e., the corresponding label).

Optimization problem. Below, we give an overview of the objective function of MUSE. The objective of MUSE is estimated on a given training dataset D\mathcal{D} in the context of a two-level decision set RR and a black box BB.

First, there are many measures of interpretability—e.g., explanations with fewer rules are typically easier to understand. MUSE employs seven such measures. The first four measures are the number of predicates f1(R)f_{1}(R), the feature overlap f2(R)f_{2}(R), the rule overlap f3(R)f_{3}(R), and the cover f4(R)f_{4}(R); these four measures are part of the optimization objective. The next three measures are the size g1(R)g_{1}(R), the maximum width g2(R)g_{2}(R), and the number of unique neighborhood descriptors g3(R)g_{3}(R); these three measures are included as constraints in the optimization problem. For details on the definitions of these measures, see Appendix A.1.

Second, fidelity is measured as before—e.g., the accuracy relative to BB. We use f5(R)f_{5}(R) to denote the fidelity of RR.

Finally, to construct the search space, we use frequent itemset mining (e.g., apriori (?)) to generate two sets of potential if conditions (i.e., sets of conjunctions of predicates): (i) ND\mathcal{ND} from which we can choose the neighborhood descriptors, and (ii) DL\mathcal{DL} from which we can choose the inner if-then rules. Then, the complete optimization problem is:

Optimization procedure. The optimization problem (1) is non-normal, non-negative, non-monotone, and submodular with matroid constraints (?). Exactly solving this problem is NP-Hard (?). Approximate local search provides the best known theoretical guarantees for this class of problems—i.e., (k+2+1/k+δ)−1(k+2+1/k+\delta)^{-1}, where kk is the number of constraints and δ>0\delta>0 (?).

2 Our Approach

We extend MUSE to generate potentially misleading explanations by modifying the optimization problem (1). In particular, we need to (i) ensure that none of the prohibited features PP (e.g., race) appear in the explanation (even if they are being used by the black box to make predictions), and (ii) ensure that all the desired features DD appear (even if they are not being used by the black box). Formally, let ND+⊆ND\mathcal{ND}_{+}\subseteq\mathcal{ND} denote the set of candidate if conditions for outer if clauses that do not include any prohibited attributes, and let DL+⊆DL\mathcal{DL}_{+}\subseteq\mathcal{DL} be the analog for inner if clauses. Furthermore, we also add a term to the objective that measures the number of features in XD\mathcal{X}_{D} that are part of some rule in RR:

where d∈Dd\in D is a desired feature. Maximizing this value will in turn maximize the chance that every desired attribute appears somewhere in the explanation.

Together, we use the following optimization problem to construct candidate misleading explanations:

where f6(R)=coverdesired(R)f_{6}(R)=\text{coverdesired}(R). The following theorem shows that as before, we can solve (2) with approximate local search:

(2) is non-normal, non-negative, non-monotone, and submodular, and has matroid constraints. [See Appendix B for proof]

Experimental Evaluation

Our goal is to evaluate how explanations can affect users’ trust of a black box. To this end, we first construct a black box and its explanations. Then, we perform a user study with domain experts to understand how each explanation affects user trust of the black box. All of our experiments are performed in the context of a real world application - bail decisions.

A key aspect of our approach is that the “black box” BB that we construct is itself an interpretable model. This allows us to evaluate whether BB is actually untrustworthy (i.e., O∗(B)=0\mathcal{O}^{*}(B)=0) via user studies. For the user study checking O(E)\mathcal{O}(E), we do not show users the internals of BB, so their decision of whether to trust BB is not affected by the fact that BB happens to be interpretable. Also, for an explanation EE of BB, we can check if BB is trusted given only on EE (i.e., O(E)=1\mathcal{O}(E)=1). If both of these criteria hold i.e., O∗(B)≠O(E)\mathcal{O}^{*}(B)\neq\mathcal{O}(E), then explanation EE is misleading.

Bail decisions. Our experiments focus on bail decision making, a high-stakes task. Police arrest over 10 million people each year in the U.S. (?). Soon after arrest, judges decide whether defendants should be released on bail or must wait in jail until their trial. Since cases can take several months to proceed to trial, bail decisions are consequential both for defendants as well as society. By law, a defendant should be released only if the judge believes that they will not flee or commit another crime. This decision is naturally modeled as a prediction problem.

We use a dataset on bail outcomes collected from several state courts in the U.S. between 1990-2009 (?). This dataset contains 37 features, including demographic attributes (age, gender, race), personal (e.g., married) and socio-economic information (e.g., pays rent, lives with children), current offense details (e.g., is felony), and past criminal records of about 32K defendants who were released on bail. Each defendant in the data is labeled either as risky (if he/she either fled and/or committed a new crime after being released on bail) or non-risky. The goal is to train a black box that predicts these outcomes to help judges make bail decisions. Explanations of this black box are needed to help domain experts determine whether to trust the black box.

Domain experts in user study. We carried out our study with 47 subjects. Each participant is a student enrolled in a law school at the time of our study. Each participant acknowledged having in-depth knowledge (16 participants) or at least some familiarity (31 participants) with the bail decision making process. Of the subjects, 27 self-identified as male and 20 as females; 25 are White, 15 Asian, 2 Hispanic, and 5 African American.

We split our study into two phases: (i) First, we reached out to each of the participants to determine which of the features in the bail dataset are relevant (i.e., desired) and which ones should be omitted (i.e., prohibited). We used these insights to construct our classifier and its explanations (see Section 5.1). (ii) Next, we performed the key part of our study—we reached out to all the subjects to understand how/why a particular explanation influences their trust of the black box classifier.

We discuss how we construct our black box (designed to be untrustworthy) and its explanations (some of which are designed to be misleading). We surveyed the domain experts to identify desired and prohibited features, and then used this information to construct our classifier and explanations. We generate an untrustworthy black box BB by explicitly including prohibited features and omitting desired features, and generate misleading explanations for BB by explicitly including desired features and/or omitting prohibited features.

Identifying prohibited and desired features. We surveyed all our 47 subjects to identify prohibited and desired features. Each participant is shown all 37 features in the bail dataset, and is asked to indicate which ones are relevant and which ones should be omitted when predicting if a defendant is risky and should not be released on bail. Figure 2 shows the 5 features (xx-axis) ranked as the most prohibited (left) and the most desired (right) ones by the participants. It also shows how many participants voted for each feature (yy-axis). Race and gender stand out unanimously as the top prohibited features; prior jail incarcerations (PJI) and prior failure to appear (PFTA) If a defendant has failed to appear in the past, that means they failed to show up for court dates and is deemed a flight risk. are the top desired features. In both cases, the first two features received significantly more votes compared to all the other features, so we use race and gender as prohibited features, and use PJI and PFTA as desired features in all subsequent experiments.

Black box and explanations. We use the identified prohibited and desired features to construct our black box and its explanations. At a high level, our approach is to construct a black box that is designed to be untrustworthy to the domain experts should they be familiar with its inner workings, and construct high-fidelity explanations of this black box designed to mislead them into trusting the black box.

To this end, we randomly shuffle the bail dataset and split it into train (70%), test (25%), and validation (5%) sets. We employ our framework with different parameter settings to construct both the black box and its explanations. We leverage the validation set and a coordinate descent style tuning procedure similar to that of MUSE to set the hyperparameters λ1,λ2,...,λ6\lambda_{1},\lambda_{2},...,\lambda_{6} (?).

We first construct a black box BB that uses race and gender (prohibited) and does not use PJI and PFTA (desired); thus, BB is most likely untrustworthy to the domain experts should they examine its internal workings. We use our framework to build BB; while designed to construct explanations, it can be applied to build an interpretable classifier by replacing the black box labels B(x)B(x) (for each x∈Xx\in\mathcal{X}) with the corresponding ground truth label yy. We use desired features D={PJI,PFTA}D=\{\text{PJI},\text{PFTA}\} and prohibited features P={race,gender}P=\{\text{race},\text{gender}\}. The resulting black box BB, shown in Figure 2 (left), is an interpretable two-level decision set; its accuracy on the held-out test set is 83.28%.

We then use our framework to construct three different high-fidelity explanations E1,E2,E3E_{1},E_{2},E_{3} of BB, as follows: (i) E1E_{1} does not use either prohibited features or desired features (i.e., we use P={race,gender,PJI,PFTA}P=\{\text{race},\text{gender},\text{PJI},\text{PFTA}\} and D=∅D=\varnothing), (ii) E2E_{2} uses both prohibited and desired features (i.e., we use P=∅P=\varnothing and D={race,gender,PJI,PFTA}D=\{\text{race},\text{gender},\text{PJI},\text{PFTA}\}, and (iii) E3E_{3} uses desired features but not prohibited features (i.e., we use P={race,gender}P=\{\text{race},\text{gender}\} and D={PJI,PFTA}D=\{\text{PJI},\text{PFTA}\}. We show E3E_{3} in Figure 2 (right);

A potential concern is that our goal is to study how qualititative aspects of each explanation (e.g., which features appear) affects whether a user trusts BB; however, the fidelity of an explanation can also affect user trust. Thus, it is important to control for fidelity beforehand. To this end, we estimate the fidelity of each explanation on the held-out test set; the fidelities for E1,E2,E3E_{1},E_{2},E_{3} are 97.3%, 98.9%, and 98.2% respectively. These values are all very similar; thus, differences in whether the user trusts or mistrusts BB must be due to the structure of the explanations rather than their fidelities.

2 Human Evaluation of Trust in Black Box

Next, we performed a user study with the domain experts to understand how our different explanations E1,E2,E3E_{1},E_{2},E_{3} affect user trust of the same black box model BB.

User study design. We designed an online user study in which 41 of the 47 domain experts that we recruited participated. Remaining 6 participants were used to explore how interactive explanations can affect user trust. Each participant was randomly chosen to be shown either the black box BB (with fidelity 100%) or one of the explanations E1,E2,E3E_{1},E_{2},E_{3} (with their corresponding fidelities). Including the black box BB is critical since it allows us to estimate the baseline trust O∗(B)\mathcal{O}^{*}(B)—i.e., whether users trust BB if they understand its internals. Each participant was instructed beforehand that the explanations they see are only correlational, not causal. Participants were allowed to take as much time as they wanted to complete the study.

Each participant was asked (i) to answer the following yes/no question: “Below is an explanation generated by state-of-the-art ML for a particular black box designed to assist judges in bail decisions. Based on this explanation, would you trust the underlying model enough to deploy it?”, and (ii) a follow-up descriptive question to explain why they decided to trust or mistrust the black box.

Results and discussion. Figure 3 shows the results of our user study. Each of the bars corresponds to either the black box or one of the explanations (xx-axis). We show the corresponding user trust, measured as the fraction of participants who responded that they trust the underlying black box—i.e., answered yes to the question above (yy-axis).

As can be seen, only 9.1% of the participants who saw the actual black box trusted it (blue), establishing our baseline that the black box is not trustworthy. Next, we discuss users who only saw one of the explanations of the black box. First, only 10% of the participants who saw E2E_{2} (brown), which includes race and gender as well as PJI and PFTA, trusted the underlying black box. On the other hand, 70% and 88% of participants who saw E1E_{1} (yellow) and E3E_{3} (purple), respectively, trusted the underlying black box. The prohibited features race and gender do not appear in E1E_{1} or E3E_{3}; in addition, E3E_{3} includes the desired features PJI and PFTA.

These results show that E1E_{1} and E3E_{3} are misleading users—i.e., they lead the user to trust a black box, while users find the actual black box untrustworthy. Since BB and E2E_{2} both include race and gender, participants are unwilling to trust the black box in these two cases. On the other hand, race and gender do not appear in E1E_{1} and E3E_{3}, and in these cases users are very likely to trust the underlying black box. These results are in spite of the clear warning we show to participants saying that the explanations shown are not causal. Furthermore, participants who see E3E_{3} appear to trust the underlying black box more frequently than those who see E1E_{1}, most likely since the desired attributes PJI and PFTA are used by E3E_{3}.

Finally, we analyzed the reasons participants gave for their responses. They are consistent with our findings—i.e., user trust appears to primarily be driven by whether the race and gender features appear in the explanation shown.

Discussion & Conclusions

We carried out the first systematic study of if and how explanations of black boxes can mislead users and affect user trust, including a novel theoretical framework for understanding when misleading explanations can exist, a novel approach for generating explanations that are likely to be misleading, and an extensive user study with domain experts from law and criminal justice to understand how misleading explanations impact user trust. We find that user trust can be manipulated by high-fidelity, misleading explanations. These misleading explanations exist since prohibited features (e.g., race or gender) can be reconstructed based on correlated features (e.g., zip code). Thus, adversarial actors can fool end users into trusting an untrustworthy black box—e.g., one that employs prohibited attributes to make decisions. We consider two ways to address this challenge.

First, recent research (?) has advocated for thinking about explanations as an interactive dialogue where end users can query or explore different explanations (called perspectives) of the black box. In fact, MUSE is designed for interactivity—e.g., a judge can ask MUSE “How does the black box make predictions for defendants of different races and/or genders?”, and it would return an explanation that only uses race and/or gender on outer if-then clauses. We performed another user study with 6 domain experts from our participant pool to study their trust in the underlying black box BB when they could explore various explanations of BB using MUSE, and found that only 16.7% of the participants (1 out of 6) trusted BB. This value is much closer to the baseline trust (9.1%).

Second, there has been recent work on capturing causal relationships between input features and black box predictions (?; ?). Explanations relying on correlations not only may be misleading (?), but have also been shown to lack robustness (?), and causal explanations may address these issues.

References

Appendix A Algorithm

We describe how each of our interpretability objectives are measured given a two-level decision set RR (with MM rules), a black box BB, and a training set D={x1,...,xN}⊆X\mathcal{D}=\{x_{1},...,x_{N}\}\subseteq\mathcal{X}. These measures are summarized in Table 2.

The first four measures are structural properties of RR. First, we want to minimize the size of RR, which is the number of triples (q,s,c)(q,s,c) in RR. Second, we want to minimize the maximum width of RR, which is the maximum over (q,s,c)(q,s,c) in RR of the quantities width(s)\text{width}(s) and width(q)\text{width}(q), where the width of an if-then rule is the number of predicates that occur in its condition. Third, we want to minimize the total number of predicates in RR, which is the sum over (q,s,c)(q,s,c) in RR of width(s)+width(q)\text{width}(s)+\text{width}(q). Fourth, we want to minimize the number of decision sets in RR, which is the number of unique neighborhood descriptors qq in R\mathcal{R}.

The next measures are are semantic properties of RR. Intuitively, this objective captures the idea that outer if-then clauses (i.e., neighborhood descriptors) and inner if-then rules have different semantic meanings. To make the distinction more clear, the overlap between the features that appear in outer and inner if-then rules should be minimized. In particular, for each pair (q,s)(q,s) of outer and inner if-then rules, we sum up the number of features that occur in both qq and ss; we want to minimize this quantity.

The final measure captures a property specific to decision sets. For decision sets, multiple rules may apply for a given example x∈Xx\in\mathcal{X}; A rule applies to xx if xx satisfies its condition. i.e., rules can be ambiguous. To maximize interpretability, for most examples xx, the rules should be unambiguous—i.e., only one rule should apply for a given xx. First, we want to minimize the rule overlap, which is the number of extra rules that apply. Second, we want to maximize cover, which counts the number of instances in the dataset that satisfy some rule in R\mathcal{R}.

where Wmax{W}_{\text{max}} is the maximum width of any rule in either candidate sets. To ensure that the objective is non-negative, we have subtracted each measure from its upper bound.

Appendix B Proofs of Theorems

where p0=N(0,1)p_{0}=\mathcal{N}(0,1). In other words, x1x_{1} is a standard Gaussian random variable, x1x_{1} and x2x_{2} are perfectly correlated, and the outcome is 1 if x2≥0x_{2}\geq 0 and 0 otherwise. Next, consider a black box

i.e., BB achieves zero loss. Since BB uses the prohibited feature x2x_{2}, it is probably untrustworthy—i.e., O^∗(B)=0\hat{\mathcal{O}}^{*}(B)=0. Similarly, consider an explanation

Since this explanation uses the desired feature and not the prohibited feature, it is acceptable; thus, it is probably misleading—i.e., O^(E)≠O^∗(B)\hat{\mathcal{O}}(E)\neq\hat{\mathcal{O}}^{*}(B). Finally, note that

Thus, EE achieves perfect fidelity, as claimed. ∎

Proof of Theorem 3.2. First, we have the following decomposition of the relative error: for any F,F′,F′′:X→YF,F^{\prime},F^{\prime\prime}:\mathcal{X}\to\mathcal{Y},

This result follows since for any y,y′,y′′∈Yy,y^{\prime},y^{\prime\prime}\in\mathcal{Y},

where the first line follows since by definition, E′E^{\prime} maximizes error relative to B+B_{+} over E∈EE\in\mathcal{E}, and the second line follows by the definition of ϵA\epsilon_{A}. Now, again by our decomposition of relative error, we have

where the last line follows since relative error is symmetric. Putting these three inequalities together, we have

where the second line follows by our assumption in the theorem statement. Since E+∈E+E_{+}\in\mathcal{E}_{+}, by definition of O^\hat{\mathcal{O}}, we have O^(E+)=1\hat{\mathcal{O}}(E_{+})=1, as claimed. ∎.

Proof of Theorem 4.1. If at least one term in a linear combination is non-normal (resp., non-monotone), then the entire linear combination is non-normal (resp., non-monotone). Given that the objective in (1) is already non-normal (resp., non-monotone), then it follows that the objective in (2) it is likewise non-normal (resp., non-monotone). In particular, coverdesired computes how many of the desired features DD appear in RR. By definition, this value cannot be negative. Since the objective in (1) is non-negative and coverdesired(R)\text{coverdesired}(R) is non-negative, so the objective in (2) is also non-negative. The non-monotone property follows similarly. Next, we did not add any new constraints to (2), and the constraints in (1) are known to follow a matroid structure. Thus, (2) also has matroid constraints.

Finally, note that coverdesired denotes the number of desired features that appear in RR. This function clearly has diminishing returns—i.e., more desired attributes will be covered when we add a new rule to a smaller set of rules compared to a larger set. Therefore, this function is submodular. Since the objective in (1) is submodular and coverdesired is submodular, it follows that the objective in (2) is also submodular since a linear combination of submodular functions is submodular. ∎