Local Rule-Based Explanations of Black Box Decision Systems

Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Dino Pedreschi, Franco Turini, Fosca Giannotti

Introduction

Popular magazines and newspapers are full of commentaries about algorithms taking critical decisions that heavily impact on our life and society, from granting a loan to finding a job or driving our car. The worry is not only due to the increasing automation of decision making, but mostly to the fact that the algorithms are opaque and their logic unexplained. The main cause for this lack of transparency is that often the algorithm itself has not been directly coded by a human but it has been generated from data through machine learning. Machine learning allows building predictive models which map user features into a class (outcome or decision), obtained by generalizing from a training set of examples. This learning process is made possible by the digital records of past decisions and classification outcomes, typically provided by human experts and decision makers. The process of inferring a classification model from examples cannot be controlled step by step because the size of training data and the complexity of the learned model are too big for humans. This is how we got trapped in a paradoxical situation in which, on one side, the legislator defines new regulations requiring that automated decisions should be explained to affected peopleWe refer here to the so-called ”right to explanation” established in the European General Data Protection Regulation (GDPR), entering into force in May 2018. while, on the other side, even more sophisticated and obscure algorithms for decision making are generated (wachter2017right, ; goodman2016eu, ).

The lack of transparency in algorithms generated through machine learning grants to them the power to perpetuate or reinforce forms of injustice by learning bad habits from the data. In fact, if the training data contains a number of biased decision records, or misleading classification examples due to data collection mistakes or artifacts, it is likely that the resulting algorithm inherits the biases and recommends discriminatory or simply wrong decisionswww.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing (Barocas2016, ; Berk2017, ). The inability of obtaining an explanation for what one considers a biased decision is a profound drawback of learning from big data, limiting social acceptance and trust on its adoption in many sensitive contexts. Starting from (pedreshi2008discrimination, ) a rich literature has been flourishing on discrimination discovery and avoidance. Some of the ideas developed in that context can be reinterpreted for addressing the more general problem of explaining the logic driving a decision taken by an obscure algorithm, which is precisely the problem tackled in this paper.

In particular, in this paper we address the problem of explaining the decision outcome taken by an obscure algorithm by providing “meaningful explanations of the logic involved” when automated decision making takes place, as prescribed by the GDPR. The decision system can be obscure because based on a deep learning approach, or because of inaccessibility of the source code, or other reasons. We perform our research under some specific assumptions. First, we assume that an explanation is interesting for a user if it clarifies why a specific decision pertaining that user has been made, i.e., we aim for local explanations, not general, global, descriptions of how the overall system works (guidotti2018survey, ). Second, we assume that the vehicle for offering explanations should be as close as possible to the language of reasoning, that is logic. Thus, we are also assuming that the user can understand elementary logic rules. Finally, we assume that the black box decision system can be queried as many times as necessary, to probe its decision behavior to the scope of reconstructing its logic; this is certainly the case in a legal argumentation in court, or in an industrial setting where a company wants to stress-test a machine learning component of a manufactured product, to minimize risk of failures and consequent industrial liability. On the other hand, we make no assumptions on the specific algorithms used in the obscure classifier: we aim at an agnostic explanation method, one that works analyzing the input-output behavior of the black box system, disregarding its internals.

We propose a solution to the black box outcome explanation problem suitable for relational, tabular data, called LORE (for LOcal Rule-based Explanations). Given a black box binary predictor bb and a specific instance xx labeled with outcome yy by bb, we build a simple, interpretable predictor by first generating a balanced set of neighbor instances of the given instance xx through an ad-hoc genetic algorithm, and then extracting from such a set a decision tree classifier. A local explanation is then extracted from the obtained decision tree. The local explanation is a pair composed by (i) a logic rule, corresponding to the path in the tree that explains why xx has been labeled as yy by bb, and (ii) a set of counterfactual rules, explaining which conditions should be changed by xx so to invert the class yy assigned by bb. For example, from the compas dataset (Barocas2016, ; Berk2017, ) we may have the following explanation: the rule {age≤39,race=African−American,recidivist=True}→High Risk\{\mathit{age}{\leq}39,\mathit{race}{=}\mathit{African{-}American},\mathit{recidivist}{=}\mathit{True}\}{\rightarrow}\mathit{High\ Risk} and the counterfactuals {age>40},{race=Native−American}\{\mathit{age}{>}40\},\{\mathit{race}{=}\mathit{Native{-}American}\}.

The intuition behind our method, common to other approaches, such as LIME (ribeiro2016should, ), and Anchor (ribeiro2018anchors, ) is that the decision boundary for the black box can be arbitrarily complex over the whole data space, but in the neighborhood of a data point there is a high chance that the decision boundary is clear and simple, hence amenable to be captured by an interpretable model. The novelty of our method lies in (i) a focused procedure, based on genetic algorithm, to explore the decision boundary in the neighborhood of the data point, which produces a high-quality training data to learn the local decision tree, and (ii) a high expressiveness of the proposed local explanations, which surpasses state-of-the-art methods providing not only succinct evidence why a data point has been assigned a specific class, but also counterfactuals suggesting what should be different in the vicinity of the data point to reverse the predicted outcome. We propose extensive experiments to assess both quantitatively and qualitatively the accuracy of our explanation method.

In the rest of this paper, after describing the state of the art in the field of explanation of black box decision models (Section 2), we offer a formalization of the problem by defining the notions of black box outcome explanation, explanation through interpretable models, and local explanation (Section 3). We then define our method LORE in Section 4. Section 5 is devoted to the experiments, the set up of which requires the definition of appropriate validation measures. We critically compare local versus global explanations, rule-based versus linear explanations, different types of rule-based explanations with respect to the state of the art, and discuss the advantages of genetic algorithms for neighborhood generation. Conclusions and future research directions are discussed in Section 6.

Related Work

Recently, the research of methods for explaining black box decision systems has caught much attention (guidotti2018survey, ). A large number of papers propose approaches for understanding the global logic of the black box by providing an interpretable classifier able to mimic the obscure decision system. Generally, these methods are designed for explaining specific black box models, i.e., they are not black box agnostic. Decision trees have been adopted to globally explain neural networks (craven1996extracting, ; krishnan1999extracting, ) and tree ensembles (hara2016making, ; tan2016tree, ). Classification rules have been widely used to explain neural networks (johansson2004accuracy, ; augasta2012reverse, ; andrews1995survey, ) but also to understand the global behavior of SVMs (fung2005rule, ; nunez2002rule, ). Only few methods for global explanation are agnostic with respect to the black box (lou2012intelligible, ; henelius2014peek, ). In the cases in which the training set is available, classification rules are also widely used to avoid building black boxes by directly designing a transparent classifier (guidotti2018survey, ) which is locally or globally interpretable on its own (wang2015falling, ; lakkaraju2016interpretable, ; malioutov2017learning, ).

Other approaches, more related to the one we propose, address the problem of explaining the local behavior of a black box (guidotti2018survey, ). In other words, they provide an explanation for the decision assigned to a specific instance. In this context there are two kinds of approaches: the model-dependent approaches and the agnostic ones. In the first category most of the papers aim at explaining neural networks and base their explanation on saliency masks, i.e., a subset of the instance that explains what is mainly responsible for the prediction (xu2015show, ; zhou2016learning, ). Examples of salient mask are parts of an image, or words or sentences in a text. On the other hand, agnostic approaches provide explanations for any type of black box. In (ribeiro2016should, ) the authors present LIME, which starts from instances randomly generated in the neighborhood of the instance to be explained. The method infers from them linear models as comprehensible local predictors. The importance of a feature in the linear model represents the explanation finally provided to the user. As a limitation of the approach, a random generation of the neighborhood does not take into account density of black box outcomes in the neighborhood instances. Hence, the linear classifiers inferred from them may not correctly characterize outcome values as a function of the predictive features. We will instead use a genetic algorithm that exploits the black box for instance generation.

Extensions of LIME using decision rules (called anchors) and program expression trees are presented in (ribeiro2018anchors, ) and (singh2016programs, ) respectively. (ribeiro2018anchors, ) uses a bandit algorithm that randomly constructs the anchors with the highest coverage and respecting a precision threshold. (singh2016programs, ) adopts a simulated annealing approach that randomly grows, shrinks, or replaces nodes in an expression tree. The neighborhood generation process adopted is the same as in LIME. Another crucial weak point of those approaches, is the need for user-specified parameters for desired explanations: the number of features (ribeiro2016should, ), the level of precision, the maximum expression tree depth (ribeiro2018anchors, ). Our approach is instead parameter-free.

Concerning the counterfactual part of our notion of explanation, (wachter2017counterfactual, ) computes a counterfactual for an instance xx by solving an optimization problem over the space of instances. The solution is an instance x′x^{\prime} close to xx but with different outcome assigned by the black boxIf instead of a black box, we are given a machine learning model, this problem is known as the inverse classification problem (DBLP:journals/jcst/AggarwalCH10, ).. Our approach provides a more abstract notion of counterfactuals, consisting of logic rules rather than flips of feature values. Thus, the user is given not only a specific example of how to obtain actionable recourse (e.g., how to improve application for getting a benefit), but also an abstract characterization of its neighboorhood instances with reversed black box outcome.

To the best of our knowledge, in the literature there is no work proposing a black box agnostic method for local decision explanation based on both decision and counterfactual rules.

Problem and Explanations

Let us start recalling basic notation on classification of tabular data. Afterwards, we define the black box outcome explanation problem, and the notion of explanation for which we propose a solution.

A predictor or classifier, is a function b:X(m)→Yb:\mathcal{X}^{(m)}\rightarrow\mathcal{Y} which maps data instances (tuples) xx from a feature space X(m)\mathcal{X}^{(m)} with mm input features to a decision yy in a target space Y\mathcal{Y}. We write b(x)=yb(x)=y to denote the decision yy predicted by bb, and b(X)=Yb(X)=Y as a shorthand for {b(x) ∣ x∈X}=Y\{b(x)\ |\ x\in X\}=Y. We restrict here to binary decisions. An instance xx consists of a set of mm attribute-value pairs (ai,vi)(a_{i},v_{i}), where aia_{i} is a feature (or attribute) and viv_{i} is a value from the domain of aia_{i}. The domain of a feature can be continuous or categorical. A predictor can be a machine learning model, a domain-expert rule-based system, or any combination of algorithmic and human knowledge processing. We assume that a predictor is available as a software function that can be queried at will. In the following, we denote by bb a black box predictor, whose internals are either unknown to the observer or they are known but uninterpretable by humans. Examples include neural networks, SVMs, ensemble classifiers, or a composition of data mining, legacy software, and hard-coded expert systems. Instead, we denote with cc an interpretable predictor, whose internal processing yielding a decision c(x)=yc(x)=y can be given a symbolic interpretation understandable by a human. Examples of such predictors include rule-based classifiers, decision trees, decision sets, and rational functions.

Given a black box predictor bb and an instance xx, the black box outcome explanation problem consists in providing an explanation ee for the decision b(x)=yb(x)=y. We approach the problem by learning an interpretable predictor cc that reproduces and accurately mimes the local behavior of the black box. An explanation of the decision is then derived from cc. By local, we mean focusing on the behavior of the black box in the neighborhood of the specific instance xx, without aiming at providing a single description of the logic of the black box for all possible instances. The neighborhood of xx is not given, but rather it has to be generated as part of the explanation process. However, we assume that some knowledge is available about the characteristics of the feature space X(m)\mathcal{X}^{(m)}, in particular the ranges of admissible values for the domains of features and, possibly, the (empirical) distribution of features. Nothing is instead assumed about the process of constructing the black box bb. Let us formalize the problem, and the approach based on interpretable models.

Let bb be a black box, and xx an instance whose decision b(x)b(x) has to be explained. The black box outcome explanation problem consists in finding an explanation e∈Ee\in E belonging to a human-interpretable domain EE.

Let c=ζ(b,x)c=\zeta(b,x) be an interpretable predictor derived from the black box bb and the instance xx using some process ζ(⋅,⋅)\zeta(\cdot,\cdot). An explanation e∈Ee\in E is obtained through cc, if e=ε(c,x)e=\varepsilon(c,x) for some explanation logic ε(⋅,⋅)\varepsilon(\cdot,\cdot) which reasons over cc and xx.

One point is still missing: which is a comprehensible domain EE of explanations? We will define an explanation ee as a pair of objects:

The first component r=p→yr=p\rightarrow y is a decision rule describing the reason for the decision value y=c(x)y=c(x) The second component Φ\Phi is a set of counterfactual rules, namely the minimal number of changes in the feature values of xx that would reverse the decision of the predictor. Let us consider as an example the following explanation for a loan request for user x={(age=22),(job=none),(amount=10k),(car=no)x=\{(\mathit{age}{=}22),(\mathit{job}=\mathit{none}),(\mathit{amount}{=}\mathit{10k}),(\mathit{car}{=}\mathit{no}):

Here, the decision deny\mathit{deny} is due to the age lower than 25, the absence of job and an amount greater than 5k (see component rr). For changing the decision instead it is required either an age higher than 25 and a smaller amount, or owning a clerk job and a car (see component Φ\Phi). Details are provided in the rest of the section.

In a decision rule (simply, a rule) rr of the form p→yp\rightarrow y, the decision yy is the consequence of the rule, while the premise pp is a boolean condition on feature values. We assume that pp is the conjunction of split conditions sc\mathit{sc} of the form a∈[v1,v2]a\in[v_{1},v_{2}], where aa is a feature and v1,v2v_{1},v_{2} are values in the domain of aa extended with Using ±∞\pm\infty we can model with a single notation typical univariate split conditions, such as equality (a=va=v as a∈[v,v]a\in[v,v]), upper bounds (a≤va\leq v as a∈[−∞,v]a\in[-\infty,v]), strict lower bounds (a>va>v as a∈[v+ϵ,∞]a\in[v+\epsilon,\infty] for a sufficiently small ϵ\epsilon). However, since our method is parametric to a decision tree induction algorithm, split conditions can also be multivariate, e..g, a≤b+va\leq b+v for a,ba,b features (as in oblique decision trees). ±∞\pm\infty. An instance xx satisfies rr, or rr covers xx, if the boolean condition pp evaluates to true for xx, i.e., if sc(x)\mathit{sc}(x) is true for every sc∈p\mathit{sc}\in p. For example, the rule {age≤25,job=none}→deny\{\mathit{age}{\leq}25,\mathit{job}{=}\mathit{none}\}{\rightarrow}\mathit{deny} is satisfied by x0={(age=22),(job=none)}x_{0}=\{(\mathit{age}{=}22),(\mathit{job}{=}\mathit{none})\} and not satisfied by x1={(age=22),(job=clerk)}x_{1}=\{(\mathit{age}{=}22),(\mathit{job}{=}\mathit{clerk})\}. We say that rr is consistent with cc, if c(x)=yc(x)=y for every instance xx that satisfies rr. Consistency means that the rule specifies some conditions for which the predictor makes a specific decision. When the instance xx for which we have to explain the decision satisfies pp, the rule p→yp\rightarrow y represents a motivation for taking a decision value, i.e., pp locally explains why bb returned yy.

Consider now a set δ\delta of split conditions. We denote the update of pp by δ\delta as p[δ]=δ∪{(a∈[v1,v2])∈p ∣ ∄[w1,w2],(a∈[w1,w2])∈δ}p[\delta]=\delta\cup\{(a\in[v_{1},v_{2}])\in p\ |\ \not\exists[w_{1},w_{2}],(a\in[w_{1},w_{2}])\in\delta\}. Intuitively, p[δ]p[\delta] is the logical condition pp with ranges for attributes overwritten as stated in δ\delta, e.g. {age≤25,job=none}[age>25]\{\mathit{age}{\leq}25,\mathit{job}{=}\mathit{none}\}[\mathit{age}{>}25] is {age>25,job=none}\{\mathit{age}>25,\mathit{job}{=}\mathit{none}\}. A counterfactual rule for pp is a rule of the form p[δ]→y^p[\delta]\rightarrow\hat{y}, for y^≠y\hat{y}\neq y. We call δ\delta a counterfactual. Consistency is meaningful also for counterfactual rules, meaning that the rule is an instance of the decision logic of cc. A counterfactual δ\delta describes what features to change and how to change them to get an outcome different from yy. Since cc predicts either yy or y^\hat{y}, if such changes are applied to the given instance xx, the predictor cc will return a different decision. Continuing the example before, changing the age feature of x0x_{0} to any value greater than 2525 will change the predicted outcome of cc from deny\mathit{deny} to grant\mathit{grant}. An expected property of a consistent counterfactual rule p[δ]→y^p[\delta]\rightarrow\hat{y} is that it should be minimal w.r.t. xx. Minimality is measuredSuch a measure can be extended to exploit additional knowledge on the feature domains in order not to generate invalid or unrealistic rules. E.g., the split condition age≤25\mathit{age}\leq 25 appears closer than age>30\mathit{age}>30 for an instance with age=26\mathit{age}=26. However, it is not actionable: an individual cannot lower her age, or change her race or gender. w.r.t. the number of split conditions in p[δ]p[\delta] not satisfied by xx. Formally, we define nf(p[δ],x)=∣{sc∈p[δ] ∣ ¬sc(x)}∣\mathit{nf}(p[\delta],x)=|\{\mathit{sc}\in p[\delta]\ |\ \neg\mathit{sc}(x)\}| (where nf(⋅,⋅)\mathit{nf}(\cdot,\cdot) stands for the number of falsified split conditions), and, when clear from the context, we simply write nf\mathit{nf}. For example, {age<25,job=clerk}→grant\{\mathit{age}<25,\mathit{job}{=}\mathit{clerk}\}\rightarrow\mathit{grant} is a counterfactual with two conditions falsified by x0x_{0}. It is not minimal if the counterfactual {age>25,job=none}→y^\{\mathit{age}{>}25,\mathit{job}{=}\mathit{none}\}\rightarrow\hat{y}, with only one falsified condition, is consistent for cc. In summary, a counterfactual rule p[δ]→y^p[\delta]\rightarrow\hat{y} is a (minimal) motivation for reversing a decision value.

We are now in the position to formally introduce the notion of explanation that we are able to provide.

Let xx be an instance, and c(x)=yc(x)=y be the decision of cc. A local explanation e=⟨r,Φ⟩e=\langle r,\Phi\rangle is a pair of: a decision rule r=(p→y)r=(p\rightarrow y) consistent with cc and satisfied by xx; and, a set Φ={p[δ1]→y^,…,p[δv]→y^}\Phi=\{p[\delta_{1}]\rightarrow\hat{y},\dots,p[\delta_{v}]\rightarrow\hat{y}\} of counterfactual rules for pp consistent with cc.

This definition completes the elements of the black box outcome explanation problem. A solution to the problem will then consists of: (i) computing an interpretable predictor cc for a given black box bb and an instance xx, i.e., defining the function ζ(⋅,⋅)\zeta(\cdot,\cdot) according to Definition 3.2; (ii) deriving a local explanation from cc and xx, i.e., defining the explanation logic ε(⋅,⋅)\varepsilon(\cdot,\cdot) according to Definition 3.2.

Proposed Method

We propose LORE (LOcal Rule-based Explanations, Algorithm 1) as a solution to the black box outcome explanation problem. An interpretable predictor cc is built for a given black box bb and instance xx by first generating a set of NN neighbor instances of xx through a genetic algorithm, and then extracting from such a set a decision tree cc. A local explanation, consisting of a single rule rr and a set of counterfactual rules Φ\Phi, is then derived from the structure of cc.

The goal of this phase is to identify a set of instances ZZ, with feature characteristics close to the ones of xx, that is able to reproduce the local decision behavior of the black box bb. Since the objective is to learn a predictor, the neighborhood should be flexible enough to include instances with both decision values, namely Z=Z=∪Z≠Z=Z_{=}\cup Z_{\neq} where instances z∈Z=z\in Z_{=} are such that b(z)=b(x)b(z)=b(x), and instances z∈Z≠z\in Z_{\neq} are such that b(z)≠b(x)b(z)\neq b(x). In Algorithm 1, we extract balanced subsets Z=Z_{=} and Z≠Z_{\neq} (lines 2–3), and then put Z=Z=∪Z≠Z=Z_{=}\cup Z_{\neq} (line 4). This task differs from approaches to instance selection (DBLP:journals/air/Olvera-LopezCTK10, ), based on genetic algorithms (DBLP:journals/kbs/TsaiEC13, ) (also specialized for decision trees (geneticselection2009, )), in that their objective is to select a subset of instances from an available training set. In our case, instead we cannot assume that the training set used to train bb is available, or not even that bb is a supervised machine learning predictor for which a training set exists. Our task is instead similar to instance generation in the field of active learning (DBLP:journals/kais/FuZL13, ), also including evolutionary approaches (DBLP:journals/ijamc-igi/DerracGH10, ). We adopt an approach based on a genetic algorithm which generates z∈Z=∪Z≠z\in Z_{=}\cup Z_{\neq} by maximizing the following fitness functions:

where d:X(m)→d:\mathcal{X}^{(m)}\rightarrow is a distance function, Itrue=1I_{\mathit{true}}=1, and Ifalse=0I_{\mathit{false}}=0. The first fitness function looks for instances zz similar to xx (term 1−d(x,z)1-d(x,z)), but not equal to xx (term Ix=zI_{x=z}) for which the black box bb produces the same outcome as xx (term Ib(x)=b(z)I_{b(x)=b(z)}). The second one leads to the generation of instances zz similar to xx, but not equal to it, for which bb returns a different decision. Intuitively, for an instance z0z_{0} such that b(x)≠b(z0)b(x)\neq b(z_{0}) and x≠z0x\neq z_{0}, it turns out fitness=x(z0)<1\mathit{fitness}^{x}_{=}(z_{0})<1. For any instance z0z_{0} such that b(x)=b(z0)b(x)=b(z_{0}), instead, we have fitness=x(z0)≥1\mathit{fitness}^{x}_{=}(z_{0})\geq 1. Finally, fitness=x(x)=1\mathit{fitness}^{x}_{=}(x)=1. Thus, maximization of fitness=x(⋅)\mathit{fitness}^{x}_{=}(\cdot) occurs necessarily for instances different from xx and whose prediction is equal to b(x)b(x).

Like neural networks, genetic algorithms (holland1992adaptation, ) are based on the biological metaphor of evolution. They have three distinct aspects. (i) The potential solutions of the problem are encoded into representations that support the variation and selection operations. In our case these representations, generally called chromosomes, correspond to instances in the feature space Xm.\mathcal{X}^{m}. (ii) A fitness function evaluates which chromosomes are the “best life forms”, that is, most appropriate for the result. These are then favored in survival and reproduction, thus shaping the next generation according to the fitness function. In our case, these instances correspond to those similar to xx, according to distance d(⋅,⋅)d(\cdot,\cdot), and with the same/different outcome returned by the black box bb depending on the fitness function fitness=x\mathit{fitness}^{x}_{=} or fitness≠x\mathit{fitness}^{x}_{\neq}. (iii) Mating (called crossover) and mutation produce a new generation of chromosomes by recombining features of their parents. The final generation of chromosomes, according to a stopping criterion, is the one that best fit the solution.

Algorithm 2 generates the neighborhoods Z=Z_{=} and Z≠Z_{\neq} of xx by instantiating the evolutionary approach of (back2000evolutionary, ). Using the terminology of (DBLP:journals/ijamc-igi/DerracGH10, ), it is an instance of generational genetic algorithms for evolutionary prototype generation. However, prototypes are a condensed subset of a training set that enable optimization in predictor learning. We aim instead at generating new instances that separate well the decision boundary of the black box bb. The usage of classifiers within fitness functions of genetic algorithms can be found in (eshelman1991chc, ; baluja1994population, ; wu2006optimal, ; cano2005stratification, ). However, the classifier they use is always the one for which the population must be selected or generated for and not another one (the black box) like in our case. Algorithm 2 first initializes the population P0P_{0} with NN copies of the instance xx to explain. Then it enters the evolution loop that begins by selection of the Pi+1P_{i+1} population having the highest fitness score. After that, the crossover operator is applied on a proportion of Pi+1P_{i+1} according to the pc\mathit{pc} probability, the resulting and the untouched individuals are placed in Pi+1′P^{\prime}_{i+1}. We use a two-point crossover which selects two parents and two crossover features at random, and then swap the crossover feature values of the parents (see Figure 2). Thereafter, a proportion of Pi+1′P^{\prime}_{i+1}, determined by pm\mathit{pm}, is mutated and placed in Pi+1′′P^{\prime\prime}_{i+1}. The unmutated individuals are also added to Pi+1′′P^{\prime\prime}_{i+1}. Mutation consists of replacing features values at random according to the empirical distributionIn experiments, we derive such a distribution from the test set of instances to explain. of a feature (see Figure 2). Individuals in Pi+1′′P^{\prime\prime}_{i+1} are evaluated according to the fitness function, and the evolution loop continues until GG generations are completed. Finally, the best individuals according to the fitness function are returned. Algorithm 2 is run twice, once using the fitness function fitness=x\mathit{fitness}^{x}_{=} to derive neighborhood instances Z=Z_{=} with the same decision as xx, and once using fitness≠x\mathit{fitness}^{x}_{\neq} to derive neighborhood instances Z≠Z_{\neq} with different decision as xx. Finally, we set Z=Z=∪Z≠Z=Z_{=}\cup Z_{\neq}.

Figure 3 shows an example of neighborhood generation for a black box consisting of a random forest model and a bi-dimensional feature space. The figure contrasts uniform random generation around a specific instance xx (starred) to our genetic approach. The latter yields a neighborhood that is denser in the boundary region of the predictor. The density of generated instances will be a key factor in extracting local (interpretable) predictors.

A key element in the definition of the fitness functions is the distance d(x,z)d(x,z). We account for the presence of mixed types of features by a weighted sum of simple matching coefficient for categorical features, and of the normalized Euclidean distancereference.wolfram.com/language/ref/NormalizedSquaredEuclideanDistance.html for continuous features. Formally, assuming hh categorical features and m−hm-h continuous ones, we use:

Our approach is parametric to d(⋅,⋅)d(\cdot,\cdot), and it can readily be applied to improved heterogeneous distance functions (DBLP:journals/prl/McCaneA08, ). Empirical results with different distance functions are reported in Section 5.3.

2. Local Rule-Based Classifier and Explanation Extraction

Given the neighborhood ZZ of xx, the second step is to build an interpretable predictor cc trained on the instances z∈Zz\in Z labeled with the black box decision b(z)b(z). Such a predictor is intended to mimic the behavior of bb locally in the ZZ neighborhood. Also, cc must be interpretable, so that an explanation for xx (decision rule and counterfactuals) can be extracted from it. The LORE approach considers decision tree classifiers due to the following reasons: (i) decision rules can naturally be derived from a root-leaf path in a decision tree; and, (ii) counterfactuals can be extracted by symbolic reasoning over a decision tree. For a decision tree cc, we derive an explanation e=⟨r,Φ⟩e=\langle r,\Phi\rangle as follows. The decision rule r=(p→y)r=(p\rightarrow y) is formed by including in pp the split conditions on the path from the root to the leaf node that is satisfied by the instance xx, and setting y=c(x)y=c(x). By construction, the rule rr is consistent with cc and satisfied by xx. Consider now the counterfactual rules in Φ\Phi. Algorithm 3 looks for all paths in the decision tree cc leading to a decision y^≠y\hat{y}\neq y. Fix one of such paths, and let qq be the conjunction of split conditions in it. Again by construction, q→y^q\rightarrow\hat{y} is a counterfactual rule consistent with cc. Notice that the counterfactual δ\delta for which q=p[δ]q=p[\delta] has not to be explicitly computedHowever, it can be done as follows. Consider the path from the leaf of pp to the leaf of qq. When moving from a child to a father node, we retract the split condition. E.g., a≤v2a\leq v_{2} is retracted from {a∈[v1,v2]}\{a\in[v_{1},v_{2}]\} by adding a∈[v1,+∞]a\in[v_{1},+\infty] to δ\delta. When moving from a father node to a child, we add the split condition to δ\delta. – this is a benefit of using decision trees. Among all such qq’s, only the ones with minimum number of split conditions sc\mathit{sc} not satisfied by xx (line 4 of Algorithm 3) are kept in Φ\Phi. As an example, consider the decision tree in Figure 4, and the instance x={(age,22),(job,clerk),(income,800)}x=\{(\mathit{age},22),(\mathit{job},\mathit{clerk}),(\mathit{income},800)\} for which the decision deny\mathit{deny} (e.g., of a loan) has to be explained. The path followed by xx is the leftmost one in the tree. The decision rule extracted from the path is {age≤25,job=clerk,income≤900}→deny\{\mathit{age}\leq 25,\mathit{job}{=}\mathit{clerk},\mathit{income}\leq 900\}\rightarrow\mathit{deny}. There are four paths leading to the opposite decision: q1q_{1} == {age≤25,job=clerk,income>900}\{\mathit{age}\leq 25,\mathit{job}{=}\mathit{clerk},\mathit{income}>900\}, q2q_{2} == {17<age≤25,job=other}\{17<\mathit{age}\leq 25,\mathit{job}=\mathit{other}\}, q3q_{3} == {age>25,income≤1500,job=other}\{\mathit{age}>25,\mathit{income}\leq 1500,\mathit{job}=\mathit{other}\}, and q4q_{4} == {age>25,income>1500}\{\mathit{age}>25,\mathit{income}>1500\}. It turns out: nf(q1,x)=1\mathit{nf}(q_{1},x)=1, nf(q2,x)=1\mathit{nf}(q_{2},x)=1, nf(q3,x)=2\mathit{nf}(q_{3},x)=2, and nf(q4,x)=2\mathit{nf}(q_{4},x)=2. Thus, Φ={q1→deny,q2→deny}\Phi=\{q_{1}{\rightarrow}\mathit{deny},q_{2}{\rightarrow}\mathit{deny}\}.

As a further output, LORE computes a counterfactual instance starting from a counterfactual rule q→y^q\rightarrow\hat{y} and xx. Among all possible instances that satisfy qq, we choose the one that minimally changes attributes from xx according to qq. This is done by looking at the split conditions falsified by xx: {sc∈q ∣ ¬sc(x)}\{\mathit{sc}\in q\ |\ \neg\mathit{sc}(x)\}, and modifying the lower/upper bound in sc\mathit{sc} that is closer to the value in xx. As an example, for the above path q1q_{1}, the counterfactual instance of xx is {(age,22),(job,clerk),(income,900+ϵ)}\{(\mathit{age},22),(\mathit{job},\mathit{clerk}),(\mathit{income},900+\epsilon)\}, and for q2q_{2} is {(age,22),(job,other),(income,800)}\{(\mathit{age},22),(\mathit{job},\mathit{other}),(\mathit{income},800)\}.

Experiments

LORE has been developed in PythonThe source code and datasets will be available at anonimized url. The experiments were performed on Ubuntu 16.04.1 LTS 64 bit, 32 GB RAM, 3.30GHz Intel Core i7, using, for the genetic neighborhood generation, the deap library (DEAP_JMLR2012, ), and for decision tree induction (the interpretable predictor), the yadt system (ruggieri2004yadt, ), which is a C4.5 implementation with multi-way splits of categorical attributes. After presenting the experimental setup, we report next: (i) some analyses on the effect of the genetic algorithm parameters for the neighborhood generation; (ii) evidence that the local genetic neighborhood is more effective than a global approach; (iii) a qualitative and quantitative comparison with naïve baselines and state of the art competitors.

We ran experiments on three real-world tabular datasets: adult, compas and germanhttps://archive.ics.uci.edu/ml/datasets/adult, https://github.com/propublica/compas-analysis, https://archive.ics.uci.edu/ml/datasets/statlog+(german+credit+data). In each of them, an instance represents attributes of an individual person. All datasets includes both categorical and continuous features.

The adult dataset from UCI Machine Learning Repository, includes 48,84248,842 instances with demographic information like age, workclass, marital-status, race, capital-loss, capital-gain etc. The income divides the population into two classes “¡=50K” and “¿50K”.

The compas dataset from ProPublica contains the features used by the COMPAS algorithm for scoring defendants and their risk (Low, Medium and High), for over 10,00010,000 individuals. We considered two classes “Low-Medium” and “High” risk, and we use the following features: age, sex, race, priors_count, days_b_scree ning_arrest, length_of_stay, c_charge_degree, is_recid.

In the german dataset from UCI Machine Learning Repository each person of the 1,0001,000 entries is classified as a “good” or “bad” creditor according to attributes like age, sex, checking_account, credit_amount, duration, purpose, etc.

We experimented the following predictors as black boxes: support vector machines with RBF kernel (SVM), random forests with 100 trees (RF), and multi-layer neural networks with ‘lbfgs’ solver (NN). Implementations of the predictors are from the scikit-learn library. Unless differently stated, default parameters were used for both the black boxes and the libraries of LORE. Missing values were replaced by the mean for continuous features and by the mode for categorical ones.

Each dataset was randomly split into train (80% instances), and test (20% instances). The former is used to train black box predictors. The latter, denoted by XX, is the set of instances for which the black box decision have to be explained. In the following, for some fixed set of instances, we denote by YY the set of decisions provided by the interpretable predictor cc and by Y^\hat{Y} the decisions provided by the black box bb on the same set.

We consider the following properties in evaluating the mimic performances of the decision tree cc inferred by LORE and of the explanations returned by it against the black box classifier bb:

fidelity(Y,Y^)∈\mathit{\textbf{{fidelity}}}(Y,\hat{Y})\in. It compares the predictions of cc and of the black box bb on the instances ZZ used to train cc (doshi2017towards, ). It answer the question: how good is cc at mimicking bb?

l-fidelity(Y,Y^)∈\mathit{\textbf{{l-fidelity}}}(Y,\hat{Y})\in. It compares the predictions of cc and bb on the instances ZxZ_{x} covered by the decision rule in a local (hence “l-”) explanation for xx. It answers the question: how good is the decision rule at mimicking bb?

cl-fidelity(Y,Y^)∈\mathit{\textbf{{cl-fidelity}}}(Y,\hat{Y})\in. It compares the predictions of cc and bb on the instances ZxZ_{x} covered by the counterfactual rules in a local explanation for xx.

hit(y,y^)∈{0,1}\mathit{\textbf{{hit}}}(y,\hat{y})\in\{0,1\}. It compares the predictions of cc and bb on the instance xx under analysis. It returns 11 if y=c(x)y=c(x) is equal to y^=b(x)\hat{y}=b(x), and otherwise.

c-hit(y,y^)∈{0,1}\mathit{\textbf{{c-hit}}}(y,\hat{y})\in\{0,1\}. It compares the predictions of cc and bb on a counterfactual instance of xx built from counterfactual rules in a local explanation of xx.

We measure the first three of them by the f1-measure (tan2005introduction, ). Aggregated values of f1 and hit/c-hit are reported by averaging them over the the set of test instances x∈Xx\in X.

2. Analysis of Neighborhood Generation

We analyze here the impact of the number of generations GG and size of neighborhood NN on the performances of instance generation and on the size complexity of the LORE output. We report only results for german dataset, since we get similar results for the other ones. The other parameters of Algorithm 2 (probabilities of crossover pc\mathit{pc} and mutation pm\mathit{pm}) are set with the default values of 0.50.5 and 0.20.2 respectively (back2000evolutionary, ). Figure 5 shows in the top plots the value of fitness functions and measures of sizes of local classifier cc (decision tree depth), of decision rule (size of the antecedent pp), and of counterfactual rules (number nf\mathit{nf} of falsified split conditions). The bottom plots show fidelity (f1-measure) and hit (rate) as well as running times of neighborhood generation. Fixed N=1000N=1000, after 1010 generations, the fitness function converges around the optimal value (top left), fidelity is almost maximized (bottom left), and also the measures of sizes (top right) become stable and small. We then set G=10G\text{=}10 in all other experiments. Figure 5-(bottom right) shows instead that the size NN of the neighborhood instances to be generated is relevant for the hit\mathit{hit} rate but not for fidelity\mathit{fidelity}. By taking into account also the running time (right side scale of the bottom plots), a good trade-off is obtained by setting N=1000N\text{=}1000.

3. Comparing Distance Functions

A key element of the neighborhood generation is the distance function used by the genetic algorithm. A legitimate question is whether the results of the approach are affected by the choice of the distance function adopted (see Section 4.1). For instance, (wachter2017counterfactual, ) presents considerable differences in their output of counterfactual instance varying the choice of the distance in their stochastic optimization approach. Table 1 reports basic measures contrasting the normalized Euclidean distance adopted by LORE with cosine and min-max distance on german dataset. The table does not highlight any considerable difference. This can be justified by the fact that, following instance generation, there are phases, such as decision tree building, that abstract instances to patterns, resulting in resilience against variability due to the distance function adopted.

4. Validation of Local Explanations

We now compare our local approach with a global approach, and discuss alternative neighborhood instance generation methods.

Extracting a predictor from the neighborhood of an instance is a winning strategy, if contrasted to an approach that builds a single predictor from all instances in the test set, i.e., Z=XZ=X. In particular, this means that the interpretable predictor will be the same for all instances in the test set. Let compare our approach with such a global approach. Table 2 reports the mean and standard deviation values of hit\mathit{hit}, fidelity\mathit{fidelity}, fairness\mathit{fairness} and tree depth for each dataset aggregating over the results of the various black boxes. While for hit\mathit{hit} both LORE and global obtain similar high performances, for the other scores LORE considerably overtakes global. In particular, the size and depth of the decision tree of the global approach may lead to explanations (decision rules and counterfactuals) more complex to understand than those returned by the proposed local approach LORE.

Comparing Neighborhood Generations

After concluding that “local is better than global”, we now show that our genetic programming approach improves over the following baselines in the generation of neighborhoods:

crn returns as ZZ the k=100k=100 instances from XX (the test set) that are closest to xx;

rnd augment the output of crn with additional randomly generated instances so that a stratified ZZ is obtained;

ris starting from the output of rnd performs the instance selection procedurehttp://contrib.scikit-learn.org/imbalanced-learn CNN (DBLP:journals/air/Olvera-LopezCTK10, );

ros starting from the output of rnd performs a random oversampling to balance the decision outcomes in ZZ.

Table 3 reports the aggregated evaluation measures over the various black boxes and datasets. LORE overtook the performance of all the other neighbors generators. Intuitively, this means that LORE’s genertic programming approach contributes more than the other methods in capturing/explaining the behavior of the black box, both for direct and counterfactual decisions. Such a conclusion is reinforced by Figure 6, which shows the box plots of the distributions of fidelity\mathit{fidelity}, l-fidelity\mathit{l{\text{-}}fidelity} and cl-fidelity\mathit{cl\text{-}fidelity}, and some summary data on the size of decision trees (∣c∣|c|), of decision rule premises (∣p∣|p|), and of the number of falsified split conditions in counterfactual rules (nf\mathit{nf}). LORE has the highest mean and median f1-measures (high mimic of the black box), the smallest interquartile ranges (low variability of results), and the lowest complexity sizes. Only for cl-fidelity\mathit{cl\text{-}fidelity} LORE has the largest variability, but a median value that is higher than the 90th90^{th} percentile of the competitors.

5. Comparison with the State-of-Art

In this section we compare our approach with the state of the art.

We present first a quantitative and qualitative comparison with the linear explanations of LIMEhttps://github.com/marcotcr/lime (ribeiro2016should, ). A first crucial difference is that in LIME, the number of features composing an explanation is an input parameter that must be specified by the user. LORE, instead, automatically provides the user with an explanation including only the features useful to justify the black box decision. This is a clear improvement over LIME. In experiments, unless otherwise stated, we vary the number of features of LIME explanations from two to ten and we consider the performance with the highest score.

Table 4 reports the mean and standard deviation of hit\mathit{hit} for each black box predictor and dataset. Moreover, Figure 7 details the box plots of fidelity\mathit{fidelity} (top) and l-fidelity\mathit{l{\text{-}}fidelity} (bottom). Results show that LORE definitely outperforms LIME under various viewpoints. Regarding the hit\mathit{hit} score, even when LORE is worse than LIME, it has a score close to 11. For RF black box, instead, LIME performs considerably worse than LORE. The box plots show that, in addition, LORE has better (local) fidelity scores and is more robust than LIME, which, on the contrary, exhibits very high variability in the neighborhood of the instance to explain (i.e., for l-fidelity\mathit{l{\text{-}}fidelity}). This can be tracked back to the genetic instance generation of LORE. Figure 8 reports a multidimensional scaling of the neighborhood of a sample instance xx generated by the two approaches. LORE computes a dense and compact neighborhood. The instances generated by LIME, instead, can be very distant from each other and always with a low density around xx.

We claim that the explanations provided by LORE are more abstract and comprehensible than the ones of LIME. Consider the example in Figure 9. The top part reports a LORE local explanation for an instance xx from the german dataset. The central part is a LIME explanation. Weights are associated to the categorical values in the instance xx to explain, and to continuous upper/lower bounds where the bounding values are taken from xx. Each weight tells the user how much the decision would have changed for different (resp., smaller/greater) values of a specific categorical (resp., continuous) feature. In the example, the weight 0.110.11 has the following meaning (ribeiro2016should, ): if the duration in months had been higher than the value it is for xx, the prediction would have been, on average, 0.11 less “0” (or 0.11 more “1”). A not very easy logic to follow when compared to a single decision rule which characterize the contextual conditions for the decision of the black box. Another major advantage of our notion of explanation consists of the set of counterfactual rules. LIME provides a rough indication of where to look for a different decision: different categorical values or lower/higher continuous values of some feature. LORE’s counterfactual rules provide high-level and minimal-change contexts for reversing the outcome prediction of the black box.

5.2. Rules vs Anchors for Explanations

A recent extension of LIME is the Anchorhttps://github.com/marcotcr/anchor approach (ribeiro2018anchors, ). It provides explanations in the form of decision rules, called anchors. Rules are computed by incrementally adding equality conditions in the premise, while an estimate of the rule precision is above a minimum threshold (set to 95%). Such an estimation relies on neighborhood generation through pure-exploration multi-armed bandit.

On a qualitative level of comparison, the Anchor approach requires the apriori discretization of continuous features, while the decision rule of LORE benefits of the capabilities of decision tree to split continuous features. Contrast, for instance, the example rules in Figure 9. Moreover, the approach of Anchor does not clearly extend to compute counterfactuals.

Let us compare now the two approaches on a quantitative level. Figure 10 reports the average precision of decision rules, where the precision of a rule is the fraction of instances in the neighborhood set that are correctly classified by the rule. Although LORE does not require to set the level of precision as parameter, the rule precision is on average high and very similar to that one obtained by Anchor, which is by construction at least 95%. This can be attributed to the performances of the decision tree induction algorithm, and of the instance generation procedure which produces balanced neighborhoods Z−Z_{-} and Z+Z_{+}. Figure 10 also shows the average coverage of decision rules, where the coverage of a rule is the fraction of instances to explain covered by the rule. As reported in (ribeiro2018anchors, ), large values of coverage are preferable, since this means that the set of decision rules produced over the instances to explain can be condensed/restricted to a subset of it. LORE shows a consistently better coverage than Anchor. Finally, we compare the stability of the two approaches with respect to randomness introduced in the neighboorhood generation. We measure stability using the Jaccard coefficient of feature sets used in the 10 decision rules computed for a same instances in 10 runs of the system. Table 5 reports mean and standard deviation of the Jaccard coefficient. LORE has a better stability than Anchor for all datasets and black boxes.

Conclusion

We have proposed a local black box agnostic explanation approach based on logic rules. LORE builds an interpretable predictor for a given black box and instance to be explained. The local interpretable predictor, a decision tree, is trained on a dense set of artificial instances similar to the one to explain generated by a genetic algorithm. The decision tree enables the extraction of a local explanation, consisting of a single rule for the decision and a set of counterfactual rules for the reversed decision. An ample experimental evaluation of the proposed approach has demonstrated the effectiveness of the genetic neighborhood procedure that leads LORE to outperform the proposals in the state of the art. A number of extensions and additional experiments can be mentioned as future work. First, LORE now works tabular data. An interesting future research direction is to make the method suitable for image and text data, for example by applying a pre-processing step for extracting semantic tags/concepts that may be mapped to a tabular format. Second, another study might be focused on the possibility to derive a global description of the black box bottom-up by composing the local explanations and minimizing the size (complexity) of the global description. Third, research lab experiments would be useful for evaluating the human comprehensibility of the provided explanations. Finally, LORE explanations can be used for identifying data and/or algorithmic biases. After the local explanations are retrieved, it would be interesting to develop an approach for deriving an unbiased dataset for safely training the obscure classifier, or to prevent the black box from introducing an algorithmic bias.

References