Towards Best Practice in Explaining Neural Network Decisions with LRP
Maximilian Kohlbrenner, Alexander Bauer, Shinichi Nakajima, Alexander Binder, Wojciech Samek, Sebastian Lapuschkin
I Introduction
In recent years, deep neural networks (DNN) have become the state of the art method in many different fields, but are mainly applied as black-box predictors. Since understanding the decisions of artificial intelligence systems is crucial in numerous scenarios and partially demanded by lawe.g. via the “right to explanation” proclaimed in the General Data Protection Regulation of the European Union , neural network interpretability has been established as an important and active research area. Consequently, many approaches to explaining neural network decisions have been proposed in recent years, e.g. . The Layer-wise Relevance Propagation (LRP) framework has proven successful at providing a meaningful intuition and measurable quantities describing a network’s feature processing and decision making . LRP attributes relevance scores to the model inputs or intermediate neurons by decomposing a model output of interest. The method follows the principles of relevance conservation and proportional decomposition. Therefore, attributions computed with LRP maintain a strong connection to the predictor output. While early applications of LRP administer a single decomposition rule uniformly to all layers of a model , more recent work describes a trend towards assigning specific decomposition rules purposedly to layers wrt. function and position within the network . This trend has tacitly emerged and formulates a best practice for applying LRP. Under qualitative evaluation, the attribution maps resulting from this current approach seem to be more robust against the well-known effects of shattered gradients and demonstrate an increased discriminativity between different target classes compared to the uniform application of a single rule.
However, recent literature applying LRP-rules in a layer-dependent manner do not justify the beneficial effects of this novel variant quantitatively, but only based on human observation. In this paper, we design and conduct a series of experiments in order to verify whether a layer-specific application of different decomposition rules actually constitutes an improvement above earlier descriptions and applications of LRP . That is, we measure and compare capabilities of various methods from explainable AI — with a focus on earlier and more recent approaches to LRP — to precisely localize the ground-truth objects in images via attribution of relevance scores. Our experiments are conducted on popular computer vision data sets with ground truth object localizations, the ImageNet and PascalVOC datasets, using different neural network models.
II Feedforward Neural Networks and LRP
Feedforward neural networks constitute a popular architecture type, ranging from simple multi-layer perceptrons and shallower convolutional architectures such as the LeNet-5 to deeper and more complex Inception and VGG-like architectures . These types of neural network commonly use ReLU non-linearities and first pass information through a stack of convolution and pooling layers, followed by several fully connected layers. The good performance of feedforward architectures in numerous problem domains, and the availability as pre-trained models makes them a valuable standard architecture in neural network design.
Consequently, feedforward networks have been subject to investigations in countless contributions towards neural network interpretability, including applications of LRP , which finds its mathematical foundation in Deep Taylor Decomposition (DTD) .
The most basic attribution rule of LRP (to which we will refer to as LRPz) is defined as
and performs a proportional decomposition of a given upper layer relevance value at some layer () and neuron to obtain lower layer relevance scores for neurons at layer (), wrt. to the localized preactivations and their respective aggregations at the layer output. Here, the localized preactivations describe quantities propagated through the model during prediction time, e.g. and within a neural network layer with learned weight parameters . Note that Eq. (1) is conservative between layers and in general maintains an equality at any layer () of the model.
Further purposed LRP-rules beyond Eq. (1) are introduced in , which can be understood as advancements thereof:
So does the LRPε decomposition rule add a signed and small constant to the denominator in order to prevent divisions by zero and to diminish the effect of recessive (e.g. weak and noisy) mappings to the relevance decomposition.
The LRPαβ-rule performs and then merges separate decompositions for the activatory () and inhibitory () parts of the forward pass
Here, the non-negative parameter permits a weighting of relevance distribution towards activations and inhibitions. The parameter is given implicitly s.t. in order to uphold conservativity of relevance between layers. The commonly used parameter can be derived from DTD and has been rediscovered in ExcitationBackprop .
Later work introduces LRP♭read: “flat”, as in the musical ., a decomposition rule which spreads the relevance of a neuron uniformly across all its inputs. This rule assumes and in Eq. (1) only for backpropagating given relevance scores to lower layers , and has seen application in the input layer(s) of neural networks. The LRP♭-rule provides invariance to the decomposition process wrt. to translations in the input domain and effectively propagates relevance scores of higher layer neurons — encoding “explanations” of more abstract concepts — towards the input via the neurons’ receptive fields, without further transformation. Note that the LRP♭ decomposition rule is thus unsuitable for decomposing fully connected layers.
Earlier applications of LRP (e.g. ) did use one single decomposition rule uniformly over the whole network, which often resulted in suboptimal “explanations” of model behavior . So are network-wide applications of LRPz (in the following denoted as LRPz, in order to distinguish this specific configuration of LRP from the rule LRPz) and network-wide applications of LRPε (denoted as LRPε) respectively identical and highly similar to GradientInput (GI) in ReLU-activated DNNs . LRPz and LRPε demonstrate — albeit working well for shallower convolutional models such as the LeNet-5 or simpler fully-connected networks — the effect of gradient shattering as overly complex attributions for deeper models (cf. Fig. 1(a)). A Network-wide application of LRPαβ (denoted as LRPαβ) demonstrates robustness against gradient shattering and produces visually pleasing attribution maps, however is lacking in class- or object discriminativity . By separately considering activatory and inhibitory mappings during the decomposition process, LRPαβ tends to attribute relevance to similar sets of input features activating sequences of neurons throughout the network, regardless of the output class chosen for relevance decomposition (cf. Fig. 1(b)). Further, LRPαβ introduces the constraint of strictly positive layer activations , which is in general not guaranteed, especially at the (logit) output of a model. A dissatisfaction of this constraint may result in a sign inversion of all backpropagated relevance scores.
II-B A Current Best Practice for LRP
A recent trend among XAI researchers and practitioners employing LRP is the use of a composite strategy of rule applications for decomposing the prediction of a neural network . That is, different parts of the DNN are decomposed using purposed rules, which in combination are robust against gradient shattering while sustaining object discriminativity. Common among these works is the utilization of LRPε with (or just LRPz) to decompose fully connected layers close to the model output, followed by an application of LRPαβ to the underlying convolutional layers (usually with ). Here, the separate decomposition of the positive and negative forward mappings complements the localized feature activation of convolutional filters activated by, and feeding into ReLUs. A final decomposition step within the convolution layers near the input uses the LRP♭-rule. Most commonly this rule (or alternatively the DTD-rule defined in context of Deep Taylor Decompositon ) is applied to the input layer only. In summary, we here describe this pattern of rule application as LRPCMP (for CoMPosite). Fig. 1 provides a qualitative overview of the effect of LRPCMP in contrast to other parameterizations and methods, which we will further discuss in Sec. IV-B. Note that the option to apply the LRP♭ decomposition to the first layers near the input (instead of only the first one layer) provides control over the local and semantic scale of the computed attributions (see Fig. 1(e)-(g)). Previous works profit from this option for comparing DNNs of varying depth, and differently configured convolutional stacks , or by increasing readability of attributions maps aligned to the requirements of human investigators .
III Metric and Assumptions
The declared purpose of LRP is to precisely and quantitatively inform about the (image or intermediate) features which contribute towards or against the decision of a model wrt. to a specific predictor output . While the recent LRPCMP exhibits improved properties above previous variants of LRP by eyeballing, an objective verification requires quantification. The visual object detection setting, as it is described by the Pascal VOC (PVOC) or ImageNet datasets — both of which include object bounding box annotations — delivers an optimal experimental setting for this purpose.
An assumed ideal model would, in such a setting, exhibit true object understanding by only predicting based on the object itself. A good and representative attribution method should therefore reflect that object understanding of the model closely i.e. by marking (parts of) the shown object as relevant and disregarding visual features not representing the object itself. Similar to , we therefore rely on a measure based on localization of attribution scores. In the following, we will evaluate LRPCMP against other methods and variants of LRP on ImageNet using a pre-trained VGG-16 network, and on PVOC 2007 using a pre-trained (on PVOC 2012) CaffeNet model . Both models perform well on their respective task and have been obtained from https://modelzoo.co/ .
III-B Verifying Object-centricity During Prediction
In practice, both datasets can not be assumed to be free from contextual biases (cf. ), and in both settings models are trained to categorize images rather than localize objects. Still, we (necessarily) assume that the models we use dominantly base their decision on the target object, as opposed to the image context.
We verify our hypothesis in Fig. 2, by showing for both models and datasets the reaction of the corresponding predictor to the occlusion of the object area vs. the occlusion of the image background. That is, for each image of the respective dataset, we leverage the available bounding box annotations and compute partially occluded versions where either the object area or class-specific image background (i.e. the non-object area) are replaced with mean color values per corresponding pixel and dataset. We then measure the for the ground truth label(s) of based on the network’s logit outputs, and plot this value as a function of relative bounding box size. Fig. 2 shows the average values and standard deviation for per bounding box size (discretized into 100 uniform bins) when replacing either the object (area within the bounding box) or the context (rest of the image).
Occluding the object area consistently leads to a sharper decrease in the output for the specific class. The trend is especially evident for smaller objects. This supports our claim that the networks base their decision mainly on the object itself.
III-C Attribution Localization as a Quantitative Measure
This gives us a performance criterion for attribution methods in object detection and classification. In order to track the fraction of the total amount of relevance that is attributed to the object, we use the inside-total relevance ratio without, and a weighted variant within consideration of the object size:
Correctly locating small objects is more difficult than locating image-sized objects. Since the ratio is always greater than or equal to 1 and increases for smaller objects, puts additional emphasis on measuring the outcome for small bounding box sizes. In both cases, higher values indicate larger fractions of relevance attributed to the object area (and not background), and therefore are the desirable outcome.
IV Experiments and Results
We perform our experiments on both the ImageNet and the PVOC 2007 datasets, since both collections provide large numbers of ground truth object bounding boxes.
For PVOC, we compute attribution maps for all samples (approx. ) from PVOC 2007, using a model which has been pre-trained on the multi label setting of PVOC 2012 . The respective model performs with a mean AP of on PVOC 2007. Since PVOC describes a multi label setting, multiple classes can be present in the same image. We therefore evaluate and once for each unique existing pair of class sample , yielding approximately measurements. Images with a higher number of (smaller) bounding boxes thus effectively have a stronger impact on the results than images with larger (and fewer), image-filling objects, while at the same time describing a more difficult setting. Many of the objects shown in PVOC images are not centered. In order to use all available object information in our evaluation, we rescale the input images to the network’s desired input shape, avoiding the (partial) cropping of objects.
On ImageNet (2012 version), bounding box information does only exist for the validation samples (displaying one class per image) and can be downloaded from the official websitehttp://www.image-net.org/challenges/LSVRC/2012. We evaluate a pre-trained VGG-16 model from the keras model zoo, obtained via the iNNvestigate toolbox. The model performs with a top-5 accuracy on the ImageNet test set. For all images the shortest side is rescaled to fit the model input and the longest side is center-cropped to obtain a quadratic input shape. Bounding box information is adjusted correspondingly.
For computing attribution maps, we make use of existing XAI software packages, depending on the models’ formats. That is, for the VGG-16 model we use the Keras and Tensorflow based iNNvestigate toolbox. For the PVOC data and the CaffeNet architecture, we compute attributions using the Caffe based LRP Toolbox .
Both XAI packages support the same functionality regarding LRP, yet differ in the provided selection of other attribution methods. Our study, however, shall be focussed on the beneficial or detrimental effects between the variants of LRP used in literature.
We complement the results with Guided Backprop (GB) and for ImageNet with Pattern Attribution (PA) only available in iNNvestigate. On both datasets, we evaluate attributions for the ground truth class labels, independent of the network prediction.
IV-B Qualitative Observations
Fig. 1 exemplarily shows attribution maps computed with different methods based on the VGG-16 model, for two object classes present in the ImageNet labels and the input image; “Bernese Mountain Dog” and “Tiger Cat”. Attributions in Figs. 1(a)-(d) result from uniform rule application to the whole network. Next to applications of LRPz and LRPαβ with , this includes Guided Backprop and Pattern Attribution . Neither of these maps demonstrate class-discriminativeness and prominently attribute scores to the same areas, regardless of the target class chosen for attribution. LRPz additionally shows the effects of gradient shattering in a highly complex attribution structure due to its equivalence to GI. Such attributions would be difficult to use and juxtapose in further algorithmic or manual analyses of model behavior.
To the right, attribution maps in Figs. 1(e)-(g) correspond to variants of LRPCMP, which apply different decomposition rules depending on layer type and position. In Fig. 1(e), the LRP♭-rule is not applied at all, while in Fig. 1(f) it is used for the first three convolutional layers, and the whole convolutional stack — including pooling layers — in Fig. 1(g). Both heatmaps in Fig. 1(e) and Fig. 1(f) use . Here altogether, the visualized attribution maps correspond more to an “intuitive expectation” of how relevance should be attributed compared to Figs. 1(a)-(d), assuming a model predicts based on object understanding. Figs. 1(e)-(g) demonstrate the change in scale and semantic, from attributions to local features to a very coarse localization map, with changing placements of the LRP♭-rule. Further, it becomes clear that with an application of the LRPαβ-rule in upper layers, object localization is lost (see Fig. 1(b) vs. Fig. 1(g)), while an application in lower layers avoids issues related to gradient shattering, as shown in Figs. 1(e)-(f) compared to Fig. 1(a).
Note that the special case shown in Fig. 1(g) is highly similar to an application of the Class Activation Mapping (CAM) method in the fully connected part of the model, however replaces the upsampling over the model’s convolutional stack of the CAM approach with the LRP♭ decomposition based approach of the LRP framework, and is thus naturally capable of distributing negative relevance scores.
Note that the VGG-16 network used here never has been trained in a multi-label setting. Despite only receiving one object category per input sample, it has learned to distinguish between different object types shown in the same image, e.g. that a dog is not a cat. This in turn reflects well in the attribution maps computed after the LRPCMP pattern.
Further examples akin to Fig. 1 are given in the Appendix.
IV-C Quantitative Results
Figs. 3(a) and (b) show the average in-total ratio as a function of bounding box size, discretized over 100 equally spaced intervals, for PVOC 2007 and ImageNet. Averages for and over the whole (and partial) datasets can be found in Tab. I. Large values indicate more precise attribution to the relevant object.
The inside-total relevance ratio highly depends on the size of the bounding box. In addition to the average and as an aggregate over all classes and images, we also report and , the average values over all objects whose bounding box does not span more than and times the area of the whole image respectively. The assumed Baseline is the uniform attribution of relevance over the whole image, which is outperformed by all methods.
LRPz performs noticeably worse on ImageNet than on PVOC, which we trace back to the significant difference in model depth (13 vs 21 layers) affecting gradient shattering. We omit LRPε in Tab. I due to the identity in results to LRPz. LRPαβ has the tendency to attribute to all shown objects (via generally neuron-activating features) and suffers from the multiple object classes per image in PVOC, where ImageNet shows only one class. Also, the similarity of attributions between PA and LRPαβ with observed in Fig. 1 seem consistent on ImageNet and result in close measurements in Tab. I.
Tab. I demonstrates that LRPCMP clearly outperforms other methods consistently on large datasets. That is, the increased precision in attribution to relevant objects is especially evident in the presence of smaller bounding boxes in . This can also be seen in and in Tab. I and the left parts of Figs. 3(a) and (b), where a majority of the image shows contextual information or other classes. Once bounding boxes become (significantly) larger and cover over of the image, all methods converge towards perfect performance, as expected. In both settings, LRPCMP:α2+♭ yields the best results, while overall the composite strategy is more effectful than a fine tuning of decomposition rule parameters.
IV-D Conclusion
In this study, we discuss a recent development in the application of Layer-wise Relevance Propagation. We summarize this emerging strategy of a composite application of multiple purposed decomposition rules as LRPCMP and juxtapose its effects to previous approaches to LRP and other methods, which uniformly apply a single decomposition rule to all layers of the model. For the first time, our results show that LRPCMP does not only yield measurably more representative attribution maps, but also provides a solution against gradient shattering affecting previous approaches, and improves properties related to object localization and class discrimination via attribution. Moreover, LRPCMP is able to precisely attribute negative relevance scores to class-contradicting features while requiring only one modified backward pass though the model, using established tools from the LRP framework. The discussed beneficial effects are demonstrated qualitatively and verified quantitatively at hand of two large and widely used computer vision datasets.
References
Appendix
In Fig. 4 we provide further illustrative examples similar to Fig. 1, using different input images and object classes.