Explanation by Progressive Exaggeration

Sumedha Singla, Brian Pollack, Junxiang Chen, Kayhan Batmanghelich

Introduction

With the explosive adoption of deep learning for real-world applications, explanation and model interpretability have received substantial attention from the research community (Kim, 2015; Doshi-Velez & Kim, 2017; Molnar, 2019; Guidotti et al., 2019). Explaining an outcome of a model in high stake applications, such as medical diagnosis from radiology images, is of paramount importance to detect hidden biases in data (Cramer et al., 2018), evaluate the fairness of the model (Doshi-Velez & Kim, 2017), and build trust in the system (Glass et al., 2008). For example, consider evaluating a computer-aided diagnosis of Alzheimer’s disease from medical images. The physician should be able to assess whether or not the model pays attention to age-related or disease-related variations in an image in order to trust the system. Given a query, our model provides an explanation that gradually exaggerates the semantic effect of one class, which is equivalent to traversing the decision boundary from side to another.

Although not always clear, there are subtle differences between interpretability and explanation (Turner, 2016). While the former mainly focuses on building or approximating models that are locally or globally interpretable (Ribeiro et al., 2016), the latter aims at explaining a predictor a-posteriori. The explanation approach does not compromise the prediction performance. However, a rigorous definition for what is a good explanation is elusive. Some researchers focused on providing feature importance (e.g., in the form of a heatmap (Selvaraju et al., 2017)) that influence the outcome of the predictor. In some applications (e.g., diagnosis with medical images) the causal changes are spread out across a large number of features (i.e., large portions of the image are impacted by a disease). Therefore, a heatmap may not be informative or useful, as almost all image features are highlighted. Furthermore, those methods do not explain why a predictor returns an outcome. Others have introduced local occlusion or perturbations to the input (Zhou et al., 2014; Fong & Vedaldi, 2017) by assessing which manipulations have the largest impact on the predictors. There is also recent interest in generating counterfactual inputs that would change the black box classification decision with respect to the query inputs (Goyal et al., 2019; Liu et al., 2019). Local perturbations of a query are not guaranteed to generate realistic or plausible inputs, which diminishes the usefulness of the explanation, especially for end users (e.g., physicians). We argue that the explanation should depend not only on the predictor function but also on the data. Therefore, it is reasonable to train a model that learns from data as well as the black-box classifier (e.g., (Chang et al., 2019; Dabkowski & Gal, 2017; Fong & Vedaldi, 2017)).

Our proposed method falls into the local explanation paradigm. Our approach is model agnostic and only requires access to the predictor values and its gradient with respect to the input. Given a query input to a black-box, we aim at explaining the outcome by providing plausible and progressive variations to the query that can result in a change to the output. The plausibility property ensures that perturbation is natural-looking. A user can employ our method as a “tuning knob” to progressively transform inputs, traverse the decision boundary from one side to the other, and gain understanding about how the predictor makes a decision. We introduce three principles for an explanation function that can be used beyond our application of interest. We evaluate our method on a set of benchmarks as well as real medical imaging data. Our experiments show that the counterfactually generated samples are realistic-looking and in the real medical application, satisfy the external evaluation. We also show that the method can be used to detect bias in training of the predictor.

Method

We view the (visual) explanation of the black-box as a generative process that produces an input for the black-box that slightly perturbs current prediction (f(x)+δf({\mathbf{x}})+\delta) while remaining plausible and realistic. By repeating this process towards each end of the binary classification spectrum, we can traverse the prediction space from one end to the other and exaggerate the underlying effect. We conceptualize the traversal from one side of the decision boundary to the other as walking across a data manifold, Mx{\mathcal{M}}_{x}. We assume the walk has a fixed step size and each step of the walk makes δ\delta change to the posterior probability of the the classifier, ff. Since the output of ff is bounded between $,wecantakeat−most, we can take at-most\lfloor\frac{1}{\delta}\rfloorsteps.Eachpositive(negative)stepincreases(decreases)theposteriorprobabilityofthepreviousstep.Weassumethatthereisalow−dimensionalembeddingspace(steps. Each positive (negative) step increases (decreases) the posterior probability of the previous step. We assume that there is a low-dimensional embedding space (\mathcal{M}_{z})thatencodesthewalk.Anencoder,) that encodes the walk. An encoder,E:{\mathcal{M}}_{x}\rightarrow{\mathcal{M}}_{z},mapsaninput,, maps an input,{\mathbf{x}},fromthedatamanifold,, from the data manifold,\mathcal{M}_{x},totheembeddingspace.Agenerator,, to the embedding space. A generator,G:{\mathcal{M}}_{z}\rightarrow{\mathcal{M}}_{y}$, takes both the embedding coordinate and the number of steps and maps it back to the data manifold (see Figure1).

Data Consistency: perturbed samples generated by If{\mathcal{I}}_{f} should lie on the data manifold, Mx{\mathcal{M}}_{x}, to be consistent with real data. In other words, the generated samples should look realistic when compared to other samples.

Compatibility with ff: changing the second argument in If(x,⋅){\mathcal{I}}_{f}({\mathbf{x}},\cdot) should produce the desired outcome from classifier ff, i.e., f(If(x,δ))≈f(x)+δf({\mathcal{I}}_{f}({\mathbf{x}},\delta))\approx f({\mathbf{x}})+\delta.

Self Consistency: Applying reverse perturbation should bring x{\mathbf{x}} back to its original form i.e., If(If(x,δ),−δ)=x{\mathcal{I}}_{f}({\mathcal{I}}_{f}({\mathbf{x}},\delta),-\delta)={\mathbf{x}}. Also, applying setting δ\delta to zero should return the query, i.e., If(x,0)=x{\mathcal{I}}_{f}({\mathbf{x}},0)={\mathbf{x}}.

Each criterion is enforced via a loss function which are discussed in the following sections.

We adopt the Generative Adversarial Networks (GANs) framework for our model (Goodfellow et al., 2014). The GANs implicitly model the underlying data distribution by setting up a min-max game between generative (G)G) and discriminative (DD) networks:

where z{\mathbf{z}} and PzP_{\mathbf{z}} are the noise distribution and the corresponding canonical distribution. There has been significant progress toward improving GANs stability as well as sample quality (Brock et al., 2019; Karras et al., 2019). The advantage of GANs is that they produce realistic-looking samples without an explicit likelihood assumption about the underlying probability distribution. This property is appealing for our application.

Furthermore, we need to provide the desired amount of perturbation to the black-box, ff. Hence, we use a Conditional GAN (cGAN) that allows the incorporation of a context as a condition to the GAN (Mirza & Osindero, 2014; Miyato & Koyama, 2018). To define the condition, we fix the step size, δ\delta, and descritize the walk which effectively cuts the posterior probability range of the predictor (i.e., $)into) into\lfloor\frac{1}{\delta}\rfloorequally−sizedbins.Hence,onecanviewtheperturbationfromequally-sized bins. Hence, one can view the perturbation fromf({\mathbf{x}})totof({\mathbf{x}})+\deltaaschangingthebinindexfromthecurrentvalueas changing the bin index from the current valuec_{f}({\mathbf{x}},0)totoc_{f}({\mathbf{x}},\delta)wherewherec_{f}({\mathbf{x}},\delta)returnsthebinindexofreturns the bin index off({\mathbf{x}})+\delta.Weuse. We usec_{f}({\mathbf{x}},\delta)$ as a condition to the cGAN.

The cGAN optimizes the following loss function:

where cc denotes a condition. Instead of generating random samples from PzP_{\mathbf{z}}, we use the output of an encoder, E(x)E({\mathbf{x}}), as input to the generator. Finally, the explainer function is defined as:

Our architecture is based on Projection GAN (Miyato & Koyama, 2018), a modification of cGAN. An advantage of the Projection GAN is that it scales well with the number of classes allowing δ→0\delta\rightarrow 0. The Projection GAN imposes the following structure on the discriminator loss function:

where LcGAN(D,G^){\mathcal{L}}_{\text{cGAN}}(D,\hat{G}) indicates the loss function in Eq. 1 when G^\hat{G} is fixed, ϕ(⋅)\bm{\phi}(\cdot) and ψ(⋅)\psi(\cdot) are networks producing vector (feature) and scalar outputs respectively. The r(c∣x)r(c|{\mathbf{x}}) is a conditional ratio function which will be discussed in Section 2.2.

2 Compatibility with the black box

In our model, the condition cc is an ordered variable i.e., cf(x,δ1)<cf(x,δ2)c_{f}({\mathbf{x}},\delta_{1})<c_{f}({\mathbf{x}},\delta_{2}) when δ1<δ2\delta_{1}<\delta_{2}. Therefore, we adapt the first term in Eq. 3 to account for ordinal multi-class regression by transforming cf(x,δ)c_{f}({\mathbf{x}},\delta) into ⌊1δ⌋−1\lfloor\frac{1}{\delta}\rfloor-1 binary classification terms (Frank & Hall, 2001):

where ϕ(⋅)\bm{\phi}(\cdot) is the feature network in Eq. 3 and vi{\mathbf{v}}_{i}’s are parameters. We also need to ensure that plugging xδ{\mathbf{x}}_{\delta} into f(⋅)f(\cdot) yields f(x)+δf({\mathbf{x}})+\delta (i.e., compatible with ff). This condition is enforce by a Kullback–Leibler (KL) divergence loss term. Adding the KL loss and the conditional ratio function we arrive at the following loss:

While the first term is a function of both GG and DD, the second term influences only the generator GG.

3 Self Consistency

We use a reconstruction loss term to enforce encoder-decoder consistency and satisfy the identity constraint of x=If(x,0){\mathbf{x}}={\mathcal{I}}_{f}({\mathbf{x}},0),

We also require that the perturbation is reversible (i.e., If(If(x,δ),−δ)=x{\mathcal{I}}_{f}({\mathcal{I}}_{f}({\mathbf{x}},\delta),-\delta)={\mathbf{x}}). We use a cycle-consistency (Zhu et al., 2017) loss to reconstruct the input from its corresponding perturbed image,

Note that the conditions for the generators in Eq. 5 and 6 are the same. However, in the former, we are reconstructing the input x{\mathbf{x}} from its latent space, but in the latter, we perturb xδ{\mathbf{x}}_{\delta} from the bin index cf(x,δ)c_{f}({\mathbf{x}},\delta) back to original bin index cf(x,0)c_{f}({\mathbf{x}},0).

4 Objective Functions

We adapted the hinge version of the adversarial loss for LcGAN(G,D)\mathcal{L}_{\text{cGAN}}(G,D).

where λcGAN,λf,λrec\lambda_{\text{cGAN}},\lambda_{f},\lambda_{\text{rec}} are the hyper-parameters that balance the importance of the loss terms.

Related Work

Our work broadly relates to literature in interpretation methods that are designed to provide a visual explanation of the decisions made by a black-box function ff, for a given query sample x{\mathbf{x}}.

Perturbation-based methods: These methods provide interpretation by showing what minimal changes are required in x{\mathbf{x}} to induce a desirable output of ff. Some methods employed image manipulation via the removal of image patches (Zhou et al., 2014) or the occlusion of image regions (Zhou et al., 2014) to change the classification score. Recently, the use of influence function, as proposed by (Koh & Liang, 2017) are applied as a form of data perturbation to modify a classifier’s response. The authors in (Fong & Vedaldi, 2017) proposed the use of optimal perturbation, defined as removing the smallest possible image region in x{\mathbf{x}} that results in the maximum drop in classification score. In another approach, (Chang et al., 2019) proposed a generative process to find and fill the image regions that correspond to the largest change in the decision output of a classifier. To switch the decision of a classifier, (Goyal et al., 2019) suggested generating counterfactuals by replacing the regions of x{\mathbf{x}} with patches from images with a different class label. All of the aforementioned works perform pixel- or patch-level manipulation to x{\mathbf{x}}, which may not result in natural-looking images. In contrast, our model enforces that the perturbed data be consistent with the unperturbed data to ensure that the perturbation is plausible. Furthermore, our method can be applied to general data and is not restricted to the imaging domain.

Saliency map-based methods: Saliency maps explains the decision of ff on x{\mathbf{x}} by highlighting the relevant regions of xx. Some earlier work in this direction(Simonyan et al., 2013; Springenberg et al., 2015; Bach et al., 2015) focuses on computing the gradient of the target class with respect to x{\mathbf{x}} and considers the image regions with large gradients as most informative. Building on this work, the class activation map (CAM) (Zhou et al., 2016) and its generalized version Grad-CAM Selvaraju et al. (2017) and other variants such as LPR (Bach et al., 2015) use a linear or non-linear combination of the activation layers to derive relevance score for every pixel in an image. These gradient-based methods are not model-agnostic and require access to intermediate layers. Recently, Adebayo et al. (2018) have shown that some saliency methods are independent both of the model and of the data generating process. We used their propose evaluation to validate our interpretation model. The saliency maps are also prone to adversarial attacks as shown by Ghorbani et al. (2019) and Kindermans et al. (2017). Furthermore, if the causal effect of a class is distributed across an image, which is the case in radiology images, the saliency approaches highlight large sections of the image, which greatly reduce the usefulness of the interpretation.

Generative explanation-based methods: These are interpretation models that uses a generative process to produce visual explanations. The contrastive explanations method (CEM) (Dhurandhar et al., 2018) generates explanations that show minimum regions in x{\mathbf{x}} which must be present/absent for a particular classification decision. In another work, (Liu et al., 2019; Joshi et al., 2019; Samangouei et al., 2018) generates explanations that highlight what features should be changed in x{\mathbf{x}} so that the classifier confidence in the prediction is strengthen (prototype) or weakened (counterfactual). Our approach is aligned with these latter lines of work, although our method and model architecture is different. Our method allows for the gradual change of the class effect, and our consistency criteria result in high-quality feasible perturbation in x{\mathbf{x}}. We rigorously evaluate our method on real medical imaging applications, in addition to the curated computer vision datasets.

Experiments

We set up six experiments to evaluate our method. First, we assess if our method satisfies the three criteria of the explainer function introduced in Section 2. We report both qualitative and quantitative results. Second, we apply our method on a medical image diagnosis task. We use external domain knowledge about the disease to perform a quantitative evaluation of the explanation. Third, we train two classifiers on biased and unbiased data and examine the performance of our method in identifying the bias. While our method does not produce a saliency map, in our forth experiment, we use the two counterfactual samples on the boundary $$ to generate a saliency map and compare it with the other methods. In fifth experiment, we demonstrate the scenario in which the classifier under inspection is multi-label. And the last experiment evaluate our model in human experiments. In Appendix A, we show further experiments on more target classes and an ablation study, to show the relative importance of each of the three criteria of the explainer function.

Our experiments are conducted on the CelebA (Liu et al., 2015) and CheXpert (Irvin et al., 2019) datasets. CelebA contains 200K celebrity face images, each with forty attribute labels. We considered binary classifier trained on the “smiling” and “young” attributes. CheXpert is a medical dataset containing 224K chest x-ray images from 65K patients and has labels for fourteen radio-graphic observations. We considered Cardiomegaly as the target class for generating explanations. All images are re-sized to 128×128128\times 128 before processing.

Figure 2 reports the qualitative results on three datasets. Given a query image x{\mathbf{x}} at inference time, our model generates a series of images xδ{\mathbf{x}}_{\delta} as visual explanations, which gradually increase the posterior probability f(xδ)f({\mathbf{x}}_{\delta}) (top label). We show results for three prediction tasks: smiling or not-smiling, young or old, and Cardiomegaly or healthy. The values on the top of each figure report the f(xδ)f({\mathbf{x}}_{\delta})’s. For Cardiomegaly, we show the outlines of the heart as well as its normalized size (values inside the parenthesis), which is indicative of the disease.

Data Consistency: The generated explanations are synthesized variations of the query image. To quantitatively compare their visual quality, we consider Fréchet Inception Distance (FID) (Heusel et al., 2017). We compared our results against the counterfactual explanations produced by xGEM (Joshi et al., 2018). The details of the xGEM model are given in appendix A.2. We divided the real and fake (i.e., generated explanations) images into two groups (on either boundary of f(x)∈f({\mathbf{x}})\in) and reported the FID for each group and the overall score. Our method significantly outperforms xGEM, producing crisper and more realistic-looking images. xGEM is based on variational autoencoder (VAE) which are known to produce blurry images (see Figure 9).

Compatibility with the black-box ff: To quantify whether the generation process is aligned with the desire perturbation δ\delta, we plotted the expected outcome f(x)+δf({\mathbf{x}})+\delta against the actual response of the classifier for the generated explanations, f(xδ)f({\mathbf{x}}_{\delta}). Figure 3 shows how our model performs when generating a series of explanations starting from a wide range of initial query images. The performance is almost perfect for Young/Old, but less so for more challenging classification problems such as Smiling or Cardiomegaly. The plot also validates that we are producing perturb images covering the entire classification range, $$. Appendix A.3 shows additional result from CelebA dataset.

Identity preservation: The generated explanations should differ only in semantic features associated with the target class, while retaining the identity of the query image. We extracted the latent embedding for real images (E(x)E({\mathbf{x}})) and their corresponding explanations (E(xδ)E({\mathbf{x}}_{\delta})), for different values of δ\delta. We calculated latent space closeness as the percentage of the times, xδ{\mathbf{x}}_{\delta} is closest to the query image x{\mathbf{x}} as compared to other generated explanations

where, m∈If(X−{x},δ)){\mathbf{m}}\in{\mathcal{I}}_{f}({\mathcal{X}}-\{{\mathbf{x}}\},\delta)) is the set of explanations generated for all the real images excluding the query image x{\mathbf{x}}. Another, popular approach to quantify identity of two face images, is to perform face verification. We used state-of-the-art face recognition model trained on VGGFace2 dataset (Cao et al., 2018) as feature extractor for both real images and their corresponding fake explanations. For face verification, we calculated the closeness between real and fake image as cosine distance between their feature vectors. The faces were considered as verified i.e., fake explanation have same identity as real image, if the distance is below 0.5. Table 2 summarizes the results.

Our method achieved high performance on localized attribute “smiling”, which alters a relatively small region of the face image as compare to attribute “age” which affects the entire face. Medical images like chest x-ray have very fine grain details which are difficult to preserve in the generative process of GAN. Our explainer function preserves the high level features like shape and size of the lung, but it struggles to retain the low level features like anatomy of the breast and shape of the collar bones. Also, it should be noted that both the datasets have multiple images for same person, but we ignore this information in our analysis and treat each image as a different identity. We compared our performance against xGEM (Joshi et al., 2018). VAE explicitly minimizes for latent space closeness. The generated explanation by xGEM were blurry version of the query image. Hence, although they were close to query image in latent space, but they didn’t preserve the identity of the individual as shown in face verification task and is evident in Figure 9 in appendix A.2. In comparison, our model achieved good performance on both the tasks.

2 Counterfactual evaluation on medical data

Cardiomegaly refers to an abnormal enlargement of the heart (Brakohiapa et al., 2017). To understand the explanations derived for Cardiomegaly target class, we overlaid the heart segmentation over the x-ray image and visualize the gradual change in heart size. The heart segmentation is shown as outlines in Figure 2, with their corresponding heart size (top values in parentheses). The heart segmentation is derive by training a UNet (Ronneberger et al., 2015) model on the segmentation in chest radiograph (SCR) dataset (van Ginneken et al., 2006). We registered x{\mathbf{x}} with its associated xδ{\mathbf{x}}_{\delta} and applied the resulting transformation to the heart masks of x{\mathbf{x}} to derive the heart masks for xδ{\mathbf{x}}_{\delta}.

For population-level analysis, we plotted the average heart size of xδ{\mathbf{x}}_{\delta} vs the condition used for generation (f(x)+δf({\mathbf{x}})+\delta) in Figure 4 (a). The plot shows a positive correlation between the heart size and the response of the classifier f(x)f({\mathbf{x}}), which agrees with the definition of Cardiomegaly. To better understand the results, we divided the population into two groups, the first group (xh{\mathbf{x}}^{h}; f(xh)<0.1f({\mathbf{x}}^{h})<0.1) consists of real images of healthy x-rays, and the second group (xc{\mathbf{x}}^{c}; f(xc)>0.9f({\mathbf{x}}^{c})>0.9) contains real images of abnormal x-rays positive for Cardiomegaly. For xh{\mathbf{x}}^{h} we generated counterfactual as xδc{\mathbf{x}}_{\delta}^{c} such that f(xδc)>0.9f({\mathbf{x}}_{\delta}^{c})>0.9. Similarly, counterfactuals for xc{\mathbf{x}}^{c} are derived as xδh{\mathbf{x}}_{\delta}^{h} such that f(xδh)<0.1f({\mathbf{x}}_{\delta}^{h})<0.1. In Figure 4 (b), we show the distribution of heart size in the four groups. We reported the dependent t-test statistics for paired samples (xh and xδc)\left({\mathbf{x}}^{h}\text{ and }{\mathbf{x}}_{\delta}^{c}\right), (xc and xδh)\left({\mathbf{x}}^{c}\text{ and }{\mathbf{x}}_{\delta}^{h}\right). A significant p-value ≪0.001\ll 0.001 rejected the null hypothesis (i.e., that the two groups have similar distributions). We also reported the independent two-sample t-test statistics for healthy (xh and xδh, p-value >0.01)\left({\mathbf{x}}^{h}\text{ and }{\mathbf{x}}_{\delta}^{h}\text{, p-value }>0.01\right) and abnormal (xc and xδc, p-value <0.01)\left({\mathbf{x}}^{c}\text{ and }{\mathbf{x}}_{\delta}^{c}\text{, p-value }<0.01\right) populations. Given higher p-values, we cannot reject the null hypothesis of identical average distributions with high confidence. Our model derived explanations successfully captured the change in heart size while generating counterfactual explanations.

3 Saliency Map

Saliency maps show the importance of each pixel of an image in the context of classification. Our method is not designed to produce saliency maps as a continuous score for every feature of the input. We extract an approximate saliency map by quantifying the regions that changed the most when comparing explanations at the opposing ends of the classification spectrum. For each query image, we generated two visual explanations corresponding to the two extremes of the decision boundary \big{(}f({\mathbf{x}}_{\delta})=0\text{ and }f({\mathbf{x}}_{\delta})=1\big{)}. The absolute difference between these explanations is our saliency map. Figure 5 shows the saliency map obtain from our method and its comparison with popular gradient based methods. We restricted the saliency maps obtained from different methods to have positive values and normalize them to range . Subjective, the saliency maps produced by our method are very localized and are comparable to the other methods.

We adapted the metric introduced in (Samek et al., 2016) to compare the different saliency maps. In an iterative procedure, we progressively replace a percentage of the most relevant pixels in an image (as given by the saliency map) with random values sampled from a uniform distribution. We observe the corresponding change in the classification performance as shown in Figure 4 (c). All the methods experienced a drop in the accuracy of the classifier with increase in the fraction of perturb pixels. The saliency maps produced by our model is significantly better than random maps and are comparable to the other saliency map methods. It should be noted that, there are many ways to quantify important regions in a image, using the series of explanations generated by our method. We didn’t optimize to find the best saliency map and showed results for one such method.

4 Bias Detection

5 Evaluating Class Discrimination

In multi-label settings, multiple labels can be true for a given image. In this test, we evaluated the sensitivity of our generated explanations to the class being explained. We consider a classifier trained to identify multiple attributes: young, smiling, black-hair, no-beard and bangs in face images from CelebA dataset. We used our model to generate explanations while considering one of the attributes as the target. Ideally, an explanation model trained to explain a target attribute should produce explanations consistent with the query image on all the attributes beside the target. Figure 7 plots the fraction of the generated explanations, that have flipped in source attribute as compared to the query image. Each column represents one source attribute. Each row is one run of our method to explain a given target attribute.

6 Human Evaluation

We used Amazon Mechanical Turk (AMT) to conduct human experiments to demonstrate that the progressive exaggeration produced by our model is visually perceivable to humans. We presented AMT workers with three tasks. In the first task, we evaluated if humans can detect the relative order between two explanations produced for a given image. We ask the AMT workers, “Given two images of the same person, in which image is the person younger (or smiling more)?” (see Figure 8). We experimented with 200 query images and generated two pairs of explanations for each query image (i.e., 400 hits). The first pair (easy) imposed the two images are samples from opposite ends of the explanation spectrum (counterfactuals), while the second pair (hard) makes no such assumption.

In the second task, we evaluated if humans can identify the target class for which our model has provided the explanations. We ask the AMT workers, “What is changing in the images? (age, smile, hair-style or beard)”. We experimented with 100 query images from each of the four attributes (i.e., 400 hits). In the third task, we demonstrate that our model can help the user to identify problems like possible bias in the black-box training. Here, we used the same setting as in the second task but also showed explanations generated for a biased classifier. We ask the AMT workers, “What is changing in the images? (smile or smile and gender)” (see Figure 8). We generated explanations for 200 query images each, from a biased-classifier (fBiasedf_{\text{Biased}}) explainer from Section 4.4 and an unbiased classifier (fNo-biasedf_{\text{No-biased}}) explainer (i.e., 400 hits). In all the three tasks, we collected eight votes for each task, evaluated against the ground truth, and used the majority vote for calculating accuracy.

We summarize our results in Table 4. In the first task, the annotators achieved high accuracy for the easy pair when there was a significant difference among the two explanation images, as compared to the hard pair when the two explanations can have very subtle differences. Overall, the annotators were successful in identifying the relative order between the two explanation images.

In the second task, the annotators were generally successful in correctly identifying the target class. The target class “bangs” proved to be the most difficult to identify, which was expected. The generated images for “bangs” were qualitatively, the most subtle. For the third task, the correct answer was always the target class i.e., “smile”. In the case of biased classifier explainer, the annotators selected “Smile and Gender” 12.5% of the times. The gradual progression made by the explainer for a biased classifier was very subtle and was changing large regions of the face as compared to the unbiased explainer. The difference is much more visible when we compare the explanation generated for the same query image for a biased and no-biased classifier, as in Figure 6. But in a realistic scenario, the no-biased classifier would not be available to compare against. Nevertheless, the annotators detected bias at roughly the same level of accuracy as our classifier (Table 3). Future work could improve upon bias detection.

Conclusion

In this paper, we proposed a novel interpretation method that explains the decision of a black-box classifier by producing natural-looking, gradual perturbations of the query image, resulting in an equivalent change in the output of the classifier. We evaluated our model on two very different datasets, including a medical imaging dataset. Our model produces high-quality explanations while preserving the identity of the query image. Our analysis shows that our explanations are consistent with the definition of the target disease without explicitly using that information. Our method can also be used to generate a saliency map in a model agnostic setting. In addition to the interpretability advantages, our proposed method can also identify plausible confounding biases in a classifier.

References

Appendix A Appendix

The architecture for the generator and discriminator is adapted from Miyato & Koyama (2018). The image encoding learned by encoder E(x)E({\mathbf{x}}) is fed into the generator. The condition cf(x,δ)c_{f}({\mathbf{x}},\delta) is passed to each resnet block in the generator, using conditional batch normalization. The generator has five resnet blocks, where each block consists of BN-ReLU-Conv3-BN-ReLU-Conv3. BN is batch normalization, ReLU is activation function, and Conv3 is the convolution filter. The encoder function uses the same structure but downsamples the image. The discriminator function has five resnet blocks, each of which has the form ReLU-Conv3-ReLU-Conv3.

A.2 xGEM implementation

We refers to Joshi et al. (2019) for the implementation of xGEM. First, VAE is trained to generate face images. The VAE used is available at:https://github.com/LynnHo/VAE-Tensorflow. All settings and architectures were set to default values. The original code generates an image of dimension 64x64. We extended the given network to produce an image with dimensions 128x128. The pre-trained VAE is then extended to incorporate the cross-entropy loss for flipping the label of the query image. The model evaluates the cross-entropy loss by passing the generated image through the classifier. Figure 9 shows the qualitative difference between the explanations generated by our proposed method and xGEM.

A.3 Extended results for evaluating the criteria of the explainer

Here, we provide results for four more prediction tasks on celebA dataset: no-beard or beard, heavy makeup or light makeup, black hair or not back hair, and bangs or no-bangs. Figure 10 shows the qualitative results, an extended version of results in Figure 2. We evaluated the results from these prediction tasks for compatibility with black-box ff (see Figure 11), data consistency and self consistency (see Table 5).

A.4 Ablation Study

Our proposed model has three types of loss functions: adversarial loss from cGAN, KL loss, and reconstruction loss. The three losses enforce the three properties of our proposed explainer function: data consistency, compatibility with ff, and self-consistency, respectively. In the ablation study, we quantify the importance of each of these components by training different models, which differ in one hyper-parameter while rest are equivalent (λcGAN=1\lambda_{\text{cGAN}}=1, λf=1\lambda_{f}=1 and λrec=100\lambda_{\text{rec}}=100). For data consistency, we evaluate Fréchet Inception Distance (FID). FID score measures the visual quality of the generated explanations by comparing them with the real images. We show results for two groups. In the first group, we consider real and fake images where the classifier has high confidence in presence of the target label i.e., f(xδ),f(x)∈[0.9,1.0]f({\mathbf{x}}_{\delta}),f({\mathbf{x}})\in[0.9,1.0]. In second group, the target label is absent i.e., f(xδ),f(x)∈[0.0,0.1)f({\mathbf{x}}_{\delta}),f({\mathbf{x}})\in[0.0,0.1). We also report an overall score by considering all the real and generated explanations together. For compatability with ff we plotted the desired output of the classifier i.e., f(x)+δf({\mathbf{x}})+\delta against the actual output of the classifier f(xδ)f({\mathbf{x}}_{\delta}) for the generated explanations. For self consistency, we calculated the Latent Space Closeness (LSC) measure and Face verification accuracy (FVA). LSC quantifies the fraction of the population in which the generated explanation is nearest to the query image than any other generated explanation in embedding space. FVA measures the percentage of the instances in which the query image and generated explanation have the same face identity as per the model trained on VGGFace2. For the ablation study, we consider the prediction task of young vs old on the CelebA dataset. Figure 12 shows the results for compatibility with ff. Table 6 summarizes the results for data consistency and self-consistency.