Fairwashing Explanations with Off-Manifold Detergent
Christopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller, Pan Kessel
Introduction
Explanation methodsSee (Samek et al., 2019) and references therein for a detailed overview. are increasingly adopted by machine learning practitioners and incorporated into standard deep learning libraries (Kokhlikyan et al., 2019; Alber et al., 2019; Ancona et al., 2018). The interest in explainability is partly driven by the hope that explanations can act as proof for a sensible, fair, and trustworthy decision-making process(Aïvodji et al., 2019; Lapuschkin et al., 2019). As an example, a bank could provide explanations for its rejection of a loan application. By doing so, the bank can demonstrate that the decision was not based on illegal or ethically questionable features. It can furthermore provide feedback to the customer. In some situations, an explanation of an algorithmic decision may even be required by law.
However, this hope is based on the assumption that explanations faithfully reflect the underlying mechanisms of the algorithmic decision. In this work, we demonstrate unequivocally that this assumption should not be made carelessly because explanations can be easily manipulated.
Briefly put, the manipulability of explanations arises from the fact that the data manifold is typically low-dimensional compared to its high-dimensional embedding space. The training process only determines the classifier in directions along the manifold. However, many explanation methods are mainly sensitive to directions orthogonal to the data manifold. Since these directions are undetermined by training, they can be changed at will.
This theoretical insight allows us to propose a modification to explanation methods which make them significantly more robust with respect to such manipulations. Namely, the explanation is projected along tangential directions of the data manifold. We show, both theoretically and experimentally, that these tangent-space-projected (tsp) explanations are indeed significantly more robust. We thereby establish a novel and exciting connection between the fields of explainability and manifold learning.
In summary, our main contributions are as follows:
Using differential geometry, we establish theoretically that popular explanation methods can be easily manipulated.
We validate our theoretical predictions in detailed experiments for various explanation methods, classifier architectures, and datasets, as well as for different tasks.
We propose a modification to existing explanation methods which make them more robust with respect to these manipulations.
In doing so, we relate explainability to manifold learning.
This work was crucially inspired by (Heo et al., 2019). In this reference, adversarial model manipulation for explanations is proposed. Specifically, the authors empirically show that one can train models such that they have structurally different explanations while suffering only a very mild drop in classification accuracy compared to their unmanipulated counterparts. For example, the adversarial model manipulation can change the positions of the most relevant pixels in each image or increase the overall sum of relevances in a certain subregion of the images. Contrary to their work, we analyze this problem theoretically. Our analysis leads us to demonstrate a stronger form of manipulability. Namely, the model can be manipulated such that it structurally reproduces arbitrary target explanations while keeping all class probabilities the same for all data points. Our theoretical insights not only illuminate the underlying reasons for the manipulability but also allow us to develop modifications of existing explanation methods which make them more robust. Another approach (Kindermans et al., 2019) adds a constant shift to the input image, which is then eliminated by changing the bias of the first layer. For some methods, this leads to a change in the explanation map. Contrary to our approach, this requires a shift in the data. In (Adebayo et al., 2018), explanation maps are changed by randomization of (some of) the network weights. This is different to our method as it dramatically changes the output of the network and is proposed as a consistency check of explanations. In (Dombrowski et al., 2019) and (Ghorbani et al., 2019), it is shown that explanations can be manipulated by an infinitesimal change in input while the output of the network is approximately unchanged. Contrary to this approach, we manipulate the model and keep the input unchanged.
2 Explanation Methods
We note that, by convention, explanation maps are usually calculated with respect to the classifier before applying the final softmax non-linearity (Kokhlikyan et al., 2019; Alber et al., 2019; Ancona et al., 2018). Throughout the paper, we will therefore denote this function as .
We use the following explanation methods:
Gradient: The map is used and quantifies how infinitesimal perturbations in each pixel change the prediction (Simonyan et al., 2014; Baehrens et al., 2010).
x Grad: This method uses the map (Shrikumar et al., 2017). For linear models, the exact contribution of each pixel to the prediction is obtained.
Integrated Gradients: This method defines
where is a suitable baseline. We refer to the original reference (Sundararajan et al., 2017) for more details.
Layer-wise Relevance Propagation (LRP): This method (Bach et al., 2015; Montavon et al., 2017) propagates relevance backwards through the network. In our experiments, we use the following setup: for the output layer, relevance is given by
which is then propagated backwards through all layers but the first using the -rule
where denotes the positive weights of the -th layer, is the activation vector of the -th layer, and is a small constant ensuring numerical stability. For the first layer, we use the -rule to account for the bounded input domain
where and are the lower and upper bounds of the input domain respectively. For theoretical analysis, we consider the -rule in all layers for simplicity. This rule is obtained by substituting in (1). We refer to the resulting method as -LRP.
This choice of methods is necessarily not exhaustive. However, it covers two classes of attribution methods, i.e. propagation and gradient-based explanations. Furthermore, the chosen methods are widely used in practice (Kokhlikyan et al., 2019; Alber et al., 2019; Ancona et al., 2018).
Manipulation of Explanations
In this section, we will theoretically deduce that explanation methods can be arbitrarily manipulated by adversarially training a model.
In the following, we will briefly summarize the basic tools of differential geometry before applying them in the context of explainability in the next section. For additional technical details, we refer to Appendix A.1.
A -dimensional submanifold is a subset of which is itself a -dimensional manifold. is called the embedding manifold of . A properly embedded submanifold is a submanifold embedded in which is also closed as a set.
With these definitions, we can now state a crucial theorem for our theoretical analysis. In Appendix A.1, we show that:
where denotes the restriction of on the submanifold . Furthermore, the derivative of the extension is given by
Technical details not withstanding, this theorem states that a function defined on a submanifold can be extended to the entire embedding manifold . The extension’s derivatives orthogonal to the submanifold can be freely chosen.
This theorem is a generalization of the well-known submanifold extension lemma (see, for example, Lemma 5.34 in (Lee, 2012)) in that it not only shows that an extension exists but also that one has control over the gradient of the extension . While we could not find such a statement in the literature, we suspect that it is entirely obvious to differential geometers but typically not needed for their purposes.
2 Explanation Manipulation: Theory
We stress that this assumption is also known as the manifold conjecture and is expected to hold across a wide range of machine learning tasks. We refer to (Goodfellow et al., 2016) for a detailed discussion.
Under this assumption, the following theorem can be derived for the Gradient, , and -LRP methods (only the proof for the Gradient method is given; see Appendix 2 for other methods):
In particular, both classifiers have the same train, validation, and test loss.
where denotes the mean-squared error and .
Proof: By Theorem 1, we can find a function which agrees with on the data manifold but has the derivative
for all . By definition, this is its gradient explanation .
As explained in Appendix A.2.1, we can assume without loss of generality that for . We can furthermore rescale the target map such that for . This rescaling is merely conventional as it does not change the relative importance of any input component with respect to the others. It then follows that
Intuition: Somewhat roughly, this theorem can be understood as follows: two models, which behave identically on the data, need to only agree on the low-dimensional submanifold . The gradients ”orthogonal” to the submanifold are completely undetermined by this requirement. By the manifold assumption, there are however much more ”orthogonal” than ”parallel” directions and therefore the explanation is largely controlled by these. We can use this fact to closely reproduce an arbitrary target while keeping the function’s values on the data unchanged.
We stress however that there are a number of non-trivial differential geometric arguments needed in order to make these statements rigorous and quantitative. For example, it is entirely non-trivial that an extension to the embedding manifold exists for arbitrary choice of target explanation. This is shown by Theorem 1 whose proof is based on a differential geometric technique called partition of the unity subordinate to an open cover. See Appendix A.1 for details.
3 Explanation Manipulation: Methods
We can now define a modified classifier by
and therefore have the same train, validation, and test error. However, the gradient explanations are now given by
Since the can be chosen freely, we can modify the explanations arbitrarily in directions orthogonal to the data submanifold (parameterized by the normal vectors ). Similar statements can be shown for other explanation methods and we refer to the Appendix A.3 for more details.
As we will discuss in Section 2.4, one can use these tricks even for data which does not (initially) lie on a hyperplane.
4 Explanation Manipulation: Practice
In this section, we will demonstrate manipulation of explanations experimentally. We will first discuss applying logistic regression to credit assessment and then proceed to the case of deep neural networks in the context of image classification. The code for all our experiments is publicly available at https://github.com/fairwashing/fairwashing.
In the following, we will suppose that a bank uses a logistic regression algorithm to classify whether a prospective client should receive a loan or not. The classification uses the features where
and is the income of the applicant. Normalization is chosen such that the features are of the same order of magnitude. Details can be found in the Appendix B.
We then define a logistic regression classifier by choosing the weights , i.e. female applicants are severely discriminated against. The discriminating nature of the algorithm may be detected by inspecting, for example, the gradient explanation maps .
However, the bank can easily ”fairwash” the explanations, i.e. hide the fact that the classifier is sexist. This can be done by adding new features which are linearly dependent on the previously used features. As a simple example, one could add the applicant’s paid taxes as a feature. By definition, it holds that
where we assume that there is a fixed tax rate of on all income. The features used by the classifier are now . By (13), all data samples obey
This example is merely an (oversimplified) illustration of a general concept: for each additional feature which linearly depends on the previously used features, a condition of the form (14) for some normal vector is obtained. We can then construct a classifier with arbitrary explanation along each of these normal vectors.
We will now experimentally demonstrate the practical applicability of our methods in the context of image classification with deep neural networks.
Datasets: We consider the MNIST, FashionMNIST, and CIFAR10 datasets. We use the standard training and test sets for our analysis. The data is normalized such that it has mean zero and standard deviation one. We sum the explanations over the absolute values of its channels to get the relevance per pixel. The resulting relevances are then normalized to have a sum of one.
Figure 2 illustrates this for examples from the FashionMNIST and CIFAR10 test sets. We stress that we use a single model for Gradient, xGrad, and Integrated Gradient methods which demonstrates that the manipulation generalizes over all considered gradient-based methods.
Robust Explanations
Having demonstrated both theoretically and experimentally that explanations are highly vulnerable to model manipulation, we will now use our theoretical insights to propose explanation methods which are significantly more robust under such manipulations.
In this section, we will define a robuster gradient explanation method. Appendix C discusses analogous definitions for other methods.
As explained in Section 2.1, we can decompose the tangent space of the embedding manifold as follows . Let be the projection on the first summand of this decomposition. We stress that the form of the projector depends on the point but we do not make this explicit in order to simplify notation. We can then define:
The tangent-space-projected (tsp) explanation field is a vector field on the data manifold . It associates to each , the tangent-space-projected (tsp) explanation given by
Intuitively, the tsp-explanation is the explanation of the model projected on the ”tangential directions” of the data manifold.
On the other hand, the components tangential to the manifold agree
It can therefore be expected that tsp-explanations are significantly more robust compared to their unprojected counterparts .
For other explanation methods, the corresponding tsp-explanations may be obtained using a slightly modified projector . We refer to Appendix C for more details.
2 TSP Explanations: Methods
Flat Submanifolds and Logistic Regression: Recall from Section 2.3 that for a logistic regression model with gradient explanation , we can define a manipulated model
We discuss the case of other explanation methods in the Appendix C.1.
General Case: In many practical applications, we do not know the explicit form of the projection matrix . In these situations, we propose to construct by one of the following two methods:
Autoencoder method: the hyperplane method requires that the data manifold is sufficiently densely sampled, i.e. the nearest neighbors are small deformations of the data point itself. In order to estimate tangent space for datasets without this property, we use techniques from the well-established field of manifold learning. Following (Shao et al., 2018), we train an autoencoder on the dataset and then perform an SVD decomposition of the Jacobian of decoder ,
The underlying motivation for this procedure is reviewed in Appendix C.2.
After one of these methods is used to estimate the projector for a given , the corresponding tsp-explanation can be easily computed by .
3 TSP Explanations: Practice
In this section, we will apply tsp-explanations to the examples of Section 2.4 and show that they are significantly more robust under model manipulations.
From the arguments of the previous section, it follows that the explanations of the manipulated and original model agree. We indeed confirm this experimentally, see Figure 4. We refer to the Appendix B for more details.
For MNIST and FashionMNIST, we use the hyperplane method to estimate the tangent space. For CIFAR10, we find that the manifold is not densely sampled enough and we therefore use the autoencoder method. This is computationally expensive and takes about 48h using four Tesla P100 GPUs. We refer to Appendix D for more details.
Figure 5 shows the tsp-explanations for the examples of Figure 2. The explanation maps of the original and manipulated model show a high degree of visual similarity. This suggests the manipulation occurred mainly in directions orthogonal to the data manifold (as the tsp-explanations are obtained from the original explanations by projecting out the corresponding components). This is also confirmed quantitatively, see Appendix D. Furthermore, tsp-explanations tend to be considerably less noisy than their unprojected counterparts (see Figure 5 vs 2). This is expected from our theoretical analysis: consider gradient explanations for concreteness. Their components orthogonal to the data manifold are undetermined by training and are therefore essentially chosen at random. This fitting noise is projected out in the tsp-explanation which results in a less noisy explanation.
We refer to Appendix D for more detailed discussion.
Conclusion
A central message of this work is that widely-used explanation methods should not be used as proof for a fair and sensible algorithmic decision-making process. This is because they can be easily manipulated as we have demonstrated both theoretically and experimentally. We propose modifications to existing explanation methods which make them more robust with respect to such manipulations. This is achieved by projecting explanations on the tangent space of the data manifold. This is exciting because it connects explainability to the field of manifold learning. For applying these methods, it is however necessary to estimate the tangent space of the data manifold. For high-dimensional datasets, such as ImageNet, this is an expensive and challenging task. Future work will try to overcome this hurdle. Another promising direction for further research is to apply the methods developed in this work to other application domains such as natural language processing.
Acknowledgements
We thank the reviewers for their valuable feedback. P.K. is greatly indebted to his mother-in-law as she took care of his sick son and wife during the final week before submission. We acknowledge Shinichi Nakajima for stimulating discussion. K-R.M. was supported in part by the German Ministry for Education and Research (BMBF) under Grants 01IS14013A-E, 01GQ1115, 01GQ0850, 01IS18025A and 01IS18037A. This work is also supported by the Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (No. 2017-0-001779), as well as by the Research Training Group ”Differential Equation- and Data-driven Models in Life Sciences and Fluid Dynamics (DAEDALUS)” (GRK 2433) and Grant Math+, EXC 2046/1, Project ID 390685689 both funded by the German Research Foundation (DFG).
References
Appendix A Proofs
We first recall a few basic definitions and theorems necessary for the proof of Theorem 1. Our presentation will be necessarily brief as it can hardly replace a course on differential geometry. However we closely follow (Lee, 2012) to which we refer for a more detailed and complete presentation.
An embedded submanifold of is a subset that is itself a manifold (with respect to the subspace topology) endowed with a smooth structure with respect to which the inclusion map is a smooth embedding. If is closed as a set, the submanifold is called properly embedded.
We say that a submanifold satisfies the local k-slice condition if each point is contained in the domain of a chart for which is a single k-slice in .
An embedded -dimensional submanifold satisfies the local -slice condition.
We refer to Theorem 5.8 of (Lee, 2012) for a proof.
Let M be a smooth manifold and an embedded submanifold. A vector field along assigns to each a vector .
For each , we can decompose the tangent space , where is the orthogonal complement of .
A standard tool for extending functions from a local coordinate patch to the entire manifold is given by the following definition:
and :
:
is locally finite, i.e. , such that for only finitely many values of .
It can be shown that for any open cover of a manifold , a partition of the unity subordinate to this cover exists. We refer to Theorem 2.23 of (Lee, 2012) for a proof.
Our main theorem is a generalization of the well-known submanifold extension lemma (see, for example, Lemma 5.34 in (Lee, 2012)). While we could not find such a generalization in the literature, we suspect that it is entirely obvious to differential geometers but typically not needed for their purposes. We now state this main theorem before giving a proof:
Proof: Since is embedded, there exists a slice chart for each . We extend in by the smooth map
By the definition of a slice chart, for . Therefore, it follows that
Let be a partition of unity subordinate to the open cover .We note that is open since is closed. We define
For , it holds that and thus because . Since the collection of supports of the is locally finite, is smooth.
The gradient of at can be straightforwardly calculated. For , one obtains
We note that sum and differentiation commute due to the local finiteness of the partition . Using , it follows that . We thus have derived that
The first term vanishes due to . For the last term, we use that for it holds that . As a result, we derive that
A.2 Theorem 2
As noted in the main text, a global rescaling of the explanation maps is merely conventional. A natural convention is to bound the explanations such that for all . For the gradient map, this can be ensure by defining where (since by assumption ). In particular, all target explanation maps are then chosen to obey this bound. For convenience, we can absorb rescaling in the classifier by redefining . As a result, we always choose the convention that without loss of generality.
More generally, let denote any bounded explanation method
We note that all considered explanation maps obey
From this, it follows that any bounded explanation method can be assumed to be bounded by because this can be ensured by an irrelevant rescaling. We again adopt the convention in which this rescaling factor is absorbed in .
A.2.2 Proofs for other Explanations
In this appendix, we will proof Theorem 2 for and -LRP.
The explanation of is given by . The mean-squared error between target and model explanation is then given by
Using the fact that we can assume without loss of generalityWe note that the necessary rescaling of is not in conflict with the shift to ensure because the latter condition is scale-invariant. and that we can rescale arbitrarily, it then follows
We assume that the network uses relu non-linearities. In fact, LRP can be shown to be theoretically well-motivated under this assumption by using Deep Taylor Decomposition (Montavon et al., 2017).
It can be shown that -LRP can be mathematically reformulated as
and on affine linear functions as the standard gradient . We refer to the Appendix A of (Ancona et al., 2018) for a proof. By our assumption, all non-linearities are relu and therefore obey
where is the Heaviside step function. This coincides with normal gradient operator . This observation was, to the best of our knowledge, first made in (Ancona et al., 2018). Therefore, the proof for applies verbatim for this method as well.
A.3 Flat Manifolds and other Explanation Methods
It was shown in the main text that one can always construct a model
which agrees with for all datapoints but has gradient explanation map
By choosing appropriately, we can always set components of corresponding to orthogonal directions of the data to an arbitrary , i.e.
where we have normalized such that it has unit norm. For , we can similarly choose
As already discussed in Appendix A.2, valid explanations map have to be zero in components for which the corresponding input component are vanishing. As a result, one only needs to set to a non-vanishing value if . Thus, the expression above is well-defined for all valid explanation maps. The corresponding statement for -LRP method can be proven completely analogously.
We also note that -LRP and IntGrad coincide with the xGrad method for logistic regression. For the latter, one has to choose a vanishing baseline point . The generalization to non-vanishing baselines is however straightforward by substituting .
Appendix B Credit Risk using other Explanation Methods
The bars in Figures 6 and 8 show the average explanation map with error bars as standard deviations. We only show explanation maps for positive classification results (examples where credit was given). All explanation maps are normalized to have .
Appendix C TSP-Explanations
For the method, we let the projection operator act only on the gradient factor of the explanation map, i.e.
This is equivalent to redefining the projection matrix to
and applying this redefined projection operator on the unprojected map , i.e.
Analogously, we define for the IntGrad method
where projects on the tangent space of the point at which the corresponding gradient is calculated. In practice however, we cannot guarantee that all the corresponding points lie on the data manifold . We therefore propose to use the projection operator for the data point instead. We find empirically that this leads to robuster explanations. This definition can again be reformulated in terms of a redefinition of the projection operator in complete analogy to the case of .
For the LRP method, we propose to use the generalized projection matrix (26) since -LRP is equivalent to for relu activations (see Appendix A.2) but we also find empirically that the standard projection matrix on the data manifold leads to more robust explanations.
The corresponding statement for -LRP can be proven analogously. The same is true for IntGrad if one assumes that all intermediate point as well as the baseline point are on the data manifold.
C.2 Autoencoder Method
In the following, we will first show how the proposed procedure for estimating tangent space arises from certain asymptotic limit of autoencoders.
An asymptotically-trained autoencoder with encoder and has zero reconstruction error, i.e.
where is a continuous probability density describing the data. Furthermore, the decoder maps on the data manifold , i.e.
The latter condition arises from the fact that we want the decoder to generate data samples from latent representations. We note there is good theoretical and experimental evidence that these conditions hold asymptotically for (at least some of the) popular autoencoder architectures, in particular Variational Autoencoders \citeappkingma2014auto.
For a continuous data distribution , it holds that
i.e. every datapoint is perfectly reconstructed.
This theorem then immediately implies that:
The decoder of an asymptotically-trained autoencoder is surjective on the data manifold .
Proof: Assume the contrary, then there exists a such that : . But by the previous theorem, it has to hold that obeys since the autoencoder has vanishing reconstruction error.
We stress however that we do not have a rigorous proof for this outside of the asymptotic regime discussed above. We furthermore want to remark that our thinking was heavily inspired by the discussion in (Shao et al., 2018) which uses very similar techniques. Last but not least, there are a number of alternative approaches in the literature to estimate tangent space. Notable examples include Contractive Autoencoders \citeapprifai2011manifold and semi-supervised GANs \citeappkumar2017improved. It would be interesting to compare these approaches to the one taken in this paper but we leave this to future work.
Appendix D Details on Experiments
For FashionMNIST and MNIST, we used a convolutional network with two groups of convolution with 20 and 50 filters of size respectively, relu activation and max-pooling over , followed by a dense layer with outputs, a relu activation, and finally another dense layer with outputs down to the number of classes (). We used VGG16 (Simonyan & Zisserman, 2015) for experiments on CIFAR10.
All images were normalized to mean and standard deviation within the training set over all pixels. For CIFAR10 training, we padded all images with 4 pixels of each side in every dimension, and then randomly cropped back to the original size of .
The original models for FashionMNIST and MNIST were trained from scratch using standard SGD with a learning rate of and a momentum of . The original VGG-16 model for CIFAR10 was trained also trained using standard SGD, but with a learning rate of , momentum of and weight decay of .
All manipulated models on all datasets were trained using Adam \citeappkingma2015adam by fine-tuning the original model with a fixed learning rate of until convergence. We set the weighting factor of the loss function (11) to . We use the same hyperparameters for manipulating tsp-explanations to ensure fair comparison. To ensure our results do not depend on a specific weighting factor , we demonstrate the same experiment shown in Figure 3 with in Figure 10.
The target explanation map used in our experiments is shown in Figure 11.
The accuracies, MSE and KL-divergence of the original and adversarially trained models are documented in Tables 1, 2 and 3 respectively.
In the following, we briefly summarize the procedure used to estimate tangent space for the various datasets.
CIFAR10: We use the autoencoder method described in the main text. This is because the manifold is not densly sampled enough for the hyperplane method, see Figure 13. We normalize the data as described above and split it by class. A separate autoencoder is trained for each class for three epochs using the Adam optimizer with a learning rate of . We use a same VQ-VAE architecture as in this examplehttps://github.com/deepmind/sonnet/blob/master/sonnet/examples/vqvae_example.ipynb. After training, the Jacobian is calculated by backpropagation for each data sample . We note that this could be sped up by forward-mode differentiation. We then perform an SVD-decomposition of the result and tune the number of singular components ensuring good reconstruction.
D.1 FashionMNIST
D.1.2 Additional Distance Metrics for Quantitative Comparison
D.2 MNIST
D.2.2 Quantitative Comparison
D.3 CIFAR10
D.3.2 Quantitative Comparison
Appendix E Pixel-flipping
We compare the original explanations with the respective TSP-explanations using pixel-flipping \citeappevaluating. This metric measures how fast the network confidence declines when removing features with highest relevance. The pixels are inpainted using the telea-method \citeapptelea to alleviate uncontrolled behaviour of the classifier off the manifold. Our result clearly show that tsp-methods perform well on this metric.