Concept Activation Regions: A Generalized Framework For Concept-Based Explanations
Jonathan Crabbé, Mihaela van der Schaar
Introduction
Deep learning models are both useful and challenging. Their utility is reflected in their increasing contributions to sophisticated tasks such as natural language processing , computer vision and scientific discovery . Their challenging nature can be attributed to their inherent complexity. State of the art deep models typically contain millions to billions parameters and, hence, appear as black-boxes to human users. The opacity of black-box models make it difficult to: anticipate how models will perform at deployment ; reliably distil knowledge from the models and earn the trust of stakeholders in high-stakes domains . With the aim of increasing the transparency of black-box models, the field of explainable AI (XAI) developed . We can broadly divide XAI methods in 2 categories: 1 Methods that restrict the model’s architecture to enable explanations. Examples include attention models that motivate their predictions by highlighting features they pay attention to and prototype-based models that motivate their predictions by highlighting relevant examples from their training set . 2 Post-hoc methods that can be used in a plug-in fashion to provide explanation for a pre-trained model. Examples include feature importance methods (also known as feature attribution or saliency methods) that highlight features the model is sensitive to ; example importance methods that identify influential training examples and hybrid methods combining the two previous approaches . In this work, we focus on a different type of explanation methods known as concept-based explanations. Let us now summarize the relevant literature to contextualize our own contribution.
Related work. Concept-based explanations were first formalized with the concept activation vector (CAV) formalism . Given a concept specified by the user (e.g. stripes in an image), linear classifiers are used as probes to assess whether a deep model’s representation space separates examples where a concept is present (concept positive examples) from examples where a concept is absent (concept negative examples). A CAV is then extracted from the linear classifier’s weights. With this CAV, it is possible to provide post-hoc explanations such as the sensitivity of a model’s prediction to the presence/absence of a concept. It goes without saying that several concepts are needed to explain the prediction of a deep model. To formalize this idea, existing works have used concept basis decomposition and sufficient statistics . Although most applications of CAV involve image data, we note that the formalism has been successfully applied to time-series data . While concepts are typically specified by the user, early works in computer vision have been undertaken to discover concepts in the form of meaningful image segmentations . The main criticism against the CAV formalism is that it requires concept positive examples to be linearly separable from concept negative examples . This is because linear separability of concept sets is a restrictive criterion that the model is not explicitly trained to fulfil. To address this issue, works like concept whitening transformations and concept bottleneck models propose to do away with the post-hoc nature of concept-based explanations. These methods introduce new neural network architectures that permit to train the models with concept labels. We stress that this requires the set of concepts to be specified before training the model. This assumes that we know what concepts are relevant for a model to solve a downstream task a-priori. This is not the case whenever we train a model to solve a task for which little or no knowledge is available. In this setup, it seems more appropriate to train a model for the task first and, then, perform a post-hoc analysis of the model to determine the concepts that were relevant in providing a solution. Furthermore, recent concerns have emerged regarding the reliability of interpretations provided by these altered model architectures .
Contributions. In this work, our purpose is to retain the flexible post-hoc nature of concept-based explanations without assuming that the concept sets are linearly separable. To that aim, we introduce concept activation regions (CARs), an extension of the CAV framework illustrated in Figure 1. 1 Generalized formalism. As a substitute to linear separability, we propose in Section 2.1 to adapt the smoothness assumption from semi-supervised learning . Intuitively, this more general assumption only requires positive and negative examples to be scattered across distinct clusters in the model’s representation space. In practice, this generalization is implemented by substituting CAV’s linear classifiers by kernel-based support vector classifiers. We demonstrate that choosing radial kernels leads to CAR classifiers that are invariant under isometries of the latent space. This permits to assign identical explanations to latent spaces characterized by the same geometry. Moreover, we show in Section 3.1.1 that our CAR classifiers yield a substantially more accurate description of how concepts are distributed in deep model’s representation spaces. 2 Better global explanations. With the nonlinear decision boundaries of our support vector classifiers, there is no obvious way to adapt the notion of concept activation vectors. Since CAV’s concept importance (TCAV score) is computed with these vectors, we need an alternative approach. In Section 2.2, we propose to define concept importance by building on our smoothness assumption. Concretely, a concept is important for a given example if the model’s representation for this example lies in a cluster of concept positive representations. With this characterization, we define TCAR scores that are the CAR equivalent of TCAV scores. In Section 3.1.2, we demonstrate that TCAR scores lead to global explanations that are more consistent with concept annotations provided by humans. 3 Concept-based feature importance. In Section 2.3, we argue that our CAR formalism permits to assign concept-specific feature importance scores for each example fed to the neural network. We verify empirically in Section 3.1.3 that those feature importance scores reflect meaningful concept associations. Finally, we illustrate in Section 3.2 how these contributions permit to establish that deep models implicitly discover known scientific concepts.
Concept Activation Regions (CARs)
A better assumption. Although the representations from Figure 2 do not linearly separate the concept sets, we note that concept positive and negative examples are scattered across distinct clusters. In this way, a model that produces these representations appears to make a difference between presence and absence of the concept. From this angle, we could consider that the concept is well encoded in the representation space geometry. To formalize this more general notion of concept set separability, we adapt the smoothness assumption originally formulated in semi-supervised learning .
A concept is encoded in the latent space if is smooth with respect to the concept. This means that we can separate into a concept activation region (CAR) where the concept is mostly present (i.e. ) and a region where the concept is mostly absent (i.e. ). If two points in a high-density region of the latent space are close to each other, then we should have or .
We can easily see that linear separability is trivially included in this assumption: it corresponds to the case where the CAR and are separated by a hyperplane. Conversely, our assumption does not require the CAR and to be separated by a hyperplane for a concept to be relevant, as illustrated in Figure 3.
There are two crucial components in the previous assumption that need to be detailed: what do we mean by density and how do we extract a CAR from .
2 Detecting Concepts
The concept density for an example can similarly be defined as .
This density function is not necessarily positive. Indeed, we have assigned a positive contribution for examples from and a negative one for those of . The idea is that whenever the density of is higher around . Conversely, whenever the density of is higher around . Finally, if is isolated from positive and negative examples or if the density of balances the density of around .
Global explanations. Our CAR formalism permits to extend those local (i.e. sample-wise) considerations globally. Let us start by defining the equivalent of TCAV scores described in Section 2.1. This score aims at understanding how the model relates classes with concepts. We define the TCAR score as the fraction of examples that have class and whose representation lies in the CAR : . Note that , where corresponds to no overlap and to a full overlap. In Appendix B, we extend this approach to measure the overlap between two concepts.
Latent space isometries invariance. In many applications such as clustering or data visualization , the only relevant geometrical information of the representation space is the distance between every pair of points . Informally, we say that two representation spaces are isometric if they assign the same distance to each pair of points. In the aforementioned applications, two isometric representation spaces are therefore indistinguishable from one another. Since concept-based explanations similarly describe the representation space geometry, one might require similar invariance to hold in this context. In Appendix D, we show that our CAR formalism provides such guarantee if is a radial kernel . To the best of our knowledge, this type of analysis has not been performed in the context of concept-based explanation methods. We believe that future works in this domain would greatly benefit from this type of insight.
3 Concepts and Features
Experiments
The code to reproduce all the experiments from this section is available at https://github.com/JonathanCrabbe/CARs and https://github.com/vanderschaarlab/CARs.
Our purpose is to empirically validate the formalism described in the previous section. We have several independent components to evaluate: 1 the concept classifier used to detect the CARs , 2 the global explanations induced by the TCAR values and 3 the feature importance scores induced by the concept densities .
Datasets. We perform our experiments on 3 datasets. 1 The MNIST dataset consists of grayscale images, each representing a digit. We train a convolutional neural network (CNN) with 2 layers to identify the digit of each image. 2 The MIT-BIH Electrocardiogram (ECG) dataset consists of univariate time series with time steps, each representing a heartbeat cycle. We train a CNN with 3 layers to determine whether each heartbeat is normal or abnormal. 3 The Caltech-UCSD Birds-200 (CUB) dataset consists of coloured images of various sizes, each representing a bird from one of the species present in the dataset. We fine-tune an Inceptionv3 neural network to identify the species each bird belongs to among the possible choices.
Concepts. For each dataset, we study the models through the lens of several well-defined concepts that are provided by human annotations. 1 For MNIST: we use concepts that correspond to simple geometrical attributes of the images: Loop (positive images include a loop), Vertical/Horizontal Line (positive images include a vertical/horizontal line) and Curvature (positive images contain segments that are not straight lines). 2 For ECG: we use concepts defined by cardiologists to better characterize abnormal heartbeats : Premature Ventricular, Supraventricular, Fusion Beats and Unknown. The exact definition for each of these concepts is beyond the scope of this paper. We simply note that annotations for those concepts are available in the ECG dataset. 3 For CUB: we use concepts that correspond to visual attributes of the birds (e.g. their size, the colour of their wings, etc.). We use the same procedure as to extract those concepts from the CUB dataset. We stress that, in each case, we selected concepts that can be unambiguously associated to the classes that are predicted by the models. Hence, it is reasonable to expect those concepts to be salient for the models. In each case, we sample the positive and negative sets and from the model’s training sets. For more details on the concepts and the models, please refer to Appendix E.
Methodology. The purpose of this experiment is to assess if the concept regions identified by our CAR classifier generalize well to unseen examples. Each of the models described above are endowed with several representation spaces (one per hidden layer). For several of those latent spaces, we fit our CAR classifier (SVC with radial basis function kernel) to discriminate the concept sets for each concept . These two sets have a size and are sampled from the model’s training set. The classifier is then evaluated by computing its accuracy on a holdout balanced concept set of size 100 sampled from the model’s testing set. For MNIST and ECG, we repeat this experiment 10 times for each concept and let the sets vary on each run. For comparison, we perform the same experiment with a linear CAV classifier as a benchmark. We report the overall (all the concepts together) accuracy in Figure 4.
Analysis. The CAR classifier substantially outperforms the CAV classifier. Note that this advantage is even more striking in the representation spaces associated to the deeper (last) DNN layers. This can be better understood through the lens of Cover’s theorem : the linear separation underlying CAV is usually easier to achieve in higher dimensional spaces, which corresponds to the shallower (first) DNN layers in this case. When the dimension of the latent space becomes comparable to the size of the concept sets, linear separation often fails to maintain a high accuracy. By contrast, the SVC underlying CARs manage to maintain high accuracy through the more flexible notion of concept smoothness defined in Assumption 2.1. This suggests that concepts can be well encoded in the geometry of the latent space even when accurate linear separability is not possible. We also note that the accuracy CAR classifiers seems to increase with the representation’s depth. This is consistent with the behaviour of class probes .
Statistical significance. The statistical significance of the concept classifiers is evaluated with the permutation test from . All of them are statistically significant with p-value except for some concepts classifiers that are fitted with the layers Mixed5d and Mixed6e of the CUB Inceptionv3 model. We note that those classifiers do not generalize well in Figure 4(c). This suggests that deeper networks are required to identify more challenging concepts correctly.
Take-away 1: CAR classifiers better capture how concepts are spread across representation spaces.
1.2 Consistency of global explanations
Analysis. The TCAR scores better correlate with the true presence of concepts. This difference can be understood by looking at the examples from Figure 5. We note that TCAV scores tend to predict nonexistent associations (e.g. yellow wings for American crows) and miss existing associations (e.g. curvature for digit 2). For ECG, we note that TCAR does capture the fact that some concepts (like fusion beats) are less represented within the class, while TCAV does not. In all of these cases, TCAV explanations might give the impression that the model did not learn meaningful class-concept associations. The TCAR analysis leads to the opposite conclusion. Since we have established that TCAR is built upon more accurate concept classifiers, it seems that models indeed learn concepts as intended, in spite of what TCAV explanations suggest.
Take-away 2: TCAR scores more faithfully reflect the true association between classes and concepts.
1.3 Coherency of concept-based feature importance
Analysis. First, we observe that CAR-based feature importance correlates weakly with vanilla feature importance ( for all datasets). This confirms that desideratum 1 is fulfilled. Then, we note that most of the CAR-based feature importance scores are decorrelated or weakly correlated with each other. The counterexamples that we observe indeed correspond to concepts that can be identified with similar features. A first example is the positive correlation between the loop and the curvature concepts for MNIST (), both concepts are generally associated to curved symbols. Another example is the negative correlation between striped and solid back patterns for CUB birds (), those concepts are mutually exclusive and identified by inspecting the bird’s back. Those examples support that CAR-based feature importance fulfils desideratum 2.
Take-away 3: CAR-based feature importance is concept-specific and captures concept associations.
2 Use Case: Machine Learning Model Rediscovering Known Medical Concepts
We will now describe a use case of the CAR formalism introduced in this paper. We stress that the literature already contains numerous use cases of CAV concept-based explanations, especially in the medical setting . Since CARs generalize CAVs, it goes without saying that they apply to these use cases. Rather than repeating existing usage of concept-based explanations, we discuss an alternative use case motivated by recent trends in machine learning. With the successes of deep models in various scientific domains , we witness an increasing overlap between scientific discovery and machine learning. While this new trend opens up fascinating opportunities, it comes with a set of new challenges. The evaluation of machine learning models is arguably one of the most important of these challenges. In a scientific context, the canonical machine learning approach to validate models (out-of-sample generalization) is likely to be insufficient. Beyond generalization on unseen data, the scientific validity of a model requires consistency with established scientific knowledge . We propose to illustrate how our CARs can be used in this context.
Dataset. We use the data collected with the Surveillance, Epidemiology, and End Results (SEER) Program. The dataset contains a US population-based cohort of 171,942 men diagnosed with non-metastatic prostate cancer between Jan 1, 2000, and Dec 31, 2016. Each patient is described by age, the results of a prostate-specific antigen blood test (PSA), the clinical stage of its tumour and two Gleason scores (primary and secondary). Each patient also has a label that indicates if they died because of their prostate cancer. We train a multilayer perceptron (MLP) to predict the patient’s mortality on of the data and test on the remaining . For a more detailed description of the data and the model, please refer to Appendix F.
Concepts. Doctors use an established grading system to predict how likely the cancer is to spread . This system assigns to each patient a grade between 1 and 5. The probability that the cancer spreads increases with this grade. It can be computed from the two Gleason scores (more details in Appendix F). We can consider each of these 5 grades as a concept. We note that this grade is not explicitly part of the input features of the MLP that we trained. Our purpose is to assess if the MLP implicitly discovered those grades in order to predict the patient’s mortality. If this happens to be the case, this would demonstrate that the MLP is in-line with existing medical knowledge.
Analysis. Let us summarize the findings for each of the above points. 1 All of the CAR classifiers generalize well on the test set (their accuracy ranges from to ). This strongly suggests that the MLP implicitly separates patients with different grades in representation space. 2 Figure 7(a) demonstrates that the model associates higher grades with higher mortality. This is in line with the clinical interpretation of the grades. 3 Figure 7(b) suggests that the Gleason scores constitute the most important features overall (highest quantiles) for the model to discriminate between grades. This is consistent with the fact that grades are computed with the Gleason scores. We conclude that the MLP implicitly identifies the patient’s grades with high accuracy and in a way that is consistent with the medical literature.
Take-away 4: The CAR formalism can reliably support scientific evaluation of a model.
Conclusion
We introduced Concept Activation Regions, a new framework to relax the linear separability assumption underlying the Concept Activation Vector formalism. We showed that our framework guarantees crucial properties, such as invariance with respect to latent symmetries. Through extensive validation on several datasets, we verified that 1 Concept Activation Regions better capture the distribution of concepts across the model’s representation space, 2 The resulting global explanations are more consistent with human annotations and 3 Concept Activation Regions permit to define concept-specific feature importance that is consistent with human intuition. Finally, through a use case involving prostate cancer data, we show the neural network can implicitly rediscover known scientific concepts, such as the prostate cancer grading system.
Many important points that were not covered in the main paper can be found in the appendices. In Appendix A, we discuss how to tune the various hyperparameters of the CAR classifiers. In Appendix B, we explain how to generalize concept activation vectors to nonlinear decision boundaries. In Appendix G, we show that the CAR explanations are robust to adversarial perturbations and background shifts. In Appendix H, we demonstrate that CAR explanations can be used to relate abstract concepts discovered by self-explaining neural networks with human concepts. Finally, Appendix I illustrates how CAR explanations allow us to probe language models.
We believe that our extended concept explainability framework opens up many interesting avenues for future work. A first one would be to probe state of the art neural networks with an approach similar to Section 3. In particular, it would be interesting to analyse if improving model performance is associated with a better encoding of human concepts. A second one would be to analyze how concept discovery can benefit from our generalized notion of concept activation. A more fine-grained characterization of a model’s latent space is likely to improve the surfaced concept. A third one, as suggested by Section 3.2, would be to use concept-based explanations to make a better scientific assessment of neural networks. Indeed, consistency with well-established knowledge is crucial for a scientific model to be accepted.
Acknowledgments and Disclosure of Funding
The authors are grateful to Fergus Imrie, Yangming Li and the 3 anonymous NeurIPS reviewers for their useful comments on an earlier version of the manuscript. Jonathan Crabbé is funded by Aviva and Mihaela van der Schaar by the Office of Naval Research (ONR), NSF 172251.
References
Checklist
Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes] All the claims from the abstract and Section 1 are verified empirically in Section 3. The theoretical claims are demonstrated in Appendices C and D.
Did you describe the limitations of your work? [Yes] The assumption for our concept-based explanations to be valid are clearly stated in Assumption 2.1. Our method only applies to neural networks, which is stated in Section 2.
Did you discuss any potential negative societal impacts of your work? [Yes] Our work improves the transparency of deep neural networks. This has many beneficial societal impacts, as described in Section 1. We do not see any potential negative societal impact to this approach. We have carefully reviewed the points from the ethical guideline and none of the mentioned negative impact seems to apply to our work.
Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes] We have read the ethics review guidelines and confirm that our paper conforms to them.
If you are including theoretical results…
Did you state the full set of assumptions of all theoretical results? [Yes] The assumptions in Propositions C.1 and D.1 are clearly stated in Appendices C and D respectively.
Did you include complete proofs of all theoretical results? [Yes] The proofs for Propositions C.1 and D.1 are given in Appendices C and D respectively.
Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes] The full code is available at https://github.com/JonathanCrabbe/CARs and https://github.com/vanderschaarlab/CARs. The implementation of our method closely follows Algorithms 1, 2 and 4 in the appendices. All the details to reproduce the experimental results are given in Appendices E and F.
Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes] All the training details are specified in Appendices E and F.
Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [Yes] Figure 4 measures the accuracy of our method over several runs. All the runs are aggregated together in the form of a box-plot.
Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes] Our computing resources are described in Appendices E and F.
If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…
If your work uses existing assets, did you cite the creators? [Yes] We have included a citation for the MNIST dataset , for the ECG dataset , for the CUB dataset , for the SEER dataset and for the InceptionV3 model .
Did you mention the license of the assets? [Yes] The licenses are mentioned in Appendices E and F.
Did you include any new assets either in the supplemental material or as a URL? [N/A] We have used only existing assets.
Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [N/A] All the datasets that we use are publicly available.
Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [N/A] The medical datasets we use are public and have been de-identified.
If you used crowdsourcing or conducted research with human subjects…
Did you include the full text of instructions given to participants and screenshots, if applicable? [N/A]
Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A]
Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [N/A]
Appendix A CAR Classifiers
In this appendix, we provide some details about our CAR classifiers.
Implementation. To determine the concept activation regions, we fit a SVC for each concept . This process is described in Algorithm 1.
Choosing the kernel. Algorithm 1 requires the user to specify a kernel function . To select this kernel, a good approach is to train several SVCs with different kernels and see how each SVC generalizes on a validation set. We illustrate this process with MNIST in Figure 8. We note that the Gaussian RBF kernel outperforms the other kernels on the validation set, although the Matern kernel achieves perfect accuracy on the training set. This highlights the importance of evaluating a CAR classifier on a held-out dataset to make sure that the related CAR offers a good description of how concepts are distributed in the latent space . In our experiments, we found that Gaussian RBF kernels are often the most interesting option.
Concept set size. Algorithm 1 requires the user to specify concepts sets and . What size should we choose for these concept sets? It seems logical that larger concepts sets are more likely to yield more accurate CARs. To study this experimentally, we propose to fit several concept classifiers for MNIST by varying the size of their training concept sets and . We report the results in Figure 9. As we can see, the curves flatten above . Increasing the concept sets size beyond this point does not improve the accuracy of the resulting CAR classifier. We recommend to acquire concept examples until the performance of the CAR classifier stabilizes. In our experiments, we found that examples is often sufficient to obtain accurate CAR classifiers.
Tuning hyperparameters. In the case where the user desires a CAR classifier that generalizes as well as possible, tuning these hyperparameters might be useful. We propose to tune the kernel type, kernel width and error penalty of our CAR classifiers for each concept by using Bayesian optimization and a validation concept set:
Repeat 3-5 for a predetermined number of trials.
We applied this process to the CAR accuracy experiment (same setup as in Section 3.1.1 of the main paper) to tune the CAR classifiers for the CUB concepts. Interestingly, we noticed no improvement with respect to the CAR classifiers reported in the main paper: tuned and standard CAR classifier have an average accuracy of for the penultimate Inception layer. This suggests that the accuracy of CAR classifiers is not heavily dependant on hyperparameters in this case.
Appendix B TCAR Global Explanations
In this appendix, we provide some details about the TCAR scores.
Implementation. When the CAR classifiers are available, they permit to compute TCAR scores through Algorithm 2.
TCAR between concepts. Up until now, we have discussed TCAR scores that indicate how models relate classes to concepts. It is possible to define a similar score to estimate how models relate two concepts with each other. Given a set of examples, we define the TCAR score associated to the concepts as the ratio
Again, corresponds to no overlap and describes a perfect overlap. We note that the concept-concept TCAR score can be interpreted as a Jaccard index between the sets and . This score is symmetric with respect to the concepts: . The computation of this score is done as in Algorithm 3.
TCAR between MNIST concepts. As an illustration, we compute concept-concept TCAR scores for the MNIST concepts and report the results in Figure 10. We see that the model relates concepts that tend to appear together (e.g. curvature and loop). Hence, concept-concept TCAR scores can serve as a proxy for the concept semantics encoded in a model’s representation space. We note the similarity with the correlation between concept-based feature importance illustrated in Figure 6. The main difference is that concept-concept TCAR scores do not explicitly refer to input features.
In this way, all the interpretation provided by the CAV formalism are also available in the CAR formalism.
Appendix C CAR Feature Importance
In this appendix, we provide some details about our concept-based feature importance.
With this property, we can interpret features with as those that tend to increase the concept density. This means that those features are important for the feature extractor to map the example in a region of the representation space where the concept is present. Hence, those are features that are important to identify a given concept . Conversely, features with tend to decrease the concept density and therefore brings the example in a region of the representation space where the concept is absent. These features can therefore be interpreted as important to reject the presence of a given concept .
Input baseline choice. The choice of baseline input has a notable effect on feature importance methods . What constitutes a good input baseline is problem dependant. Intuitively, should correspond to an input where no information is present . Let us now explain how this information removal is achieved with the datasets that we use in our experiments. 1 For MNIST, we chose a black image as an input baseline. This is because MNIST images have a black background an the information comes from white pixels that represent the digits. 2 For ECG, we chose a constant time series as an input baseline. This is because ECG time series are normalized ( for all ) and the pulse information comes from time steps where the time series is non-vanishing . 3 For CUB, choosing a baseline is more complicated. The reason for this is that the background colour changes from one image to the other and rarely corresponds to a single colour. To address this, we proceed as in the literature and select a baseline that corresponds to a blurred version of the image we wish to explain: , where is a Gaussian filter of width and denotes the convolution operation. We note that this baseline depends on which input we want to explain. In our implementation, we use to have images that are significantly blurred. 4 For SEER, we chose a constant vector as an input baseline. This is because all continuous features are standardized and all categorical features are one-hot encoded.
Appendix D CAR Latent Isometry Invariance
In this appendix, we prove that CAR explanations are invariant under isometries of the latent space when built with a radial kernel. Let us first rigorously define the notion of isometry between two vector spaces.
Let and be two normed vector spaces. An isometry from to is a map such that for all :
We say that the two spaces and are isometric if there exists a bijective isometry from to .
An explanation method is invariant to latent space isometries if applying a bijective isometry to the model’s latent space does not affect the explanations produced by the method. To make this more formal, we write the model as , where is the inverse of the bijective isometry . In this setup, we could produce explanations with CARs by making the following replacements for the feature extractor and the label map: and . The explanations are defined to be invariant to latent space isometries if they are unaffected by this replacement. It is legitimate to expect this since the previous replacement leads to the same model and substitutes the latent space by an isometric latent space .
Let us now discuss the isometry invariance of CAR explanations. First, we recall that CARs are defined through a kernel . Not all kernels lead to isometry invariant CAR explanations. We will show that it holds for a family of kernels known as radial kernels.
For each concept , we note that CAR explanations exclusively rely on the concept density from Definition 2.1 and the associated support vector classifier . Hence, it is sufficient to show that these two functions are invariant under isometry.
Concept Density. We start with the concept density . We note that the kernel function can be applied to vectors from as for all . Applying CAR in the latent space isometric to corresponds to using this kernel to compute an alternative density . Let us fix an example . Under the isometry, the concept density for this example transforms as . Let us show that this is in fact an invariance:
We deduce that the concept density is invariant under isometry: .
SVC. The proof is more involved for the SVC . We assume that the reader is familiar with the standard theory of SVC. If this is not the case, please refer e.g. to Chapter 7 of . For the sake of notation, we will abbreviate and by and respectively. Similarly, we abbreviate and by and respectively. The SVC concept classifier can be written as
where and are the indices of the support vectors. Under isometry, the SVC classification for a latent vector transforms as with and
We will now show that so that the SVC is invariant under isometries. We proceed in 3 steps.
1 We show that for any and :
By injecting this in (5), we are able to make the following replacements: and for all .
2 We show that and for all . To that aim, we note that and maximize the objective
Since this objective is identical to the one from the original SVC and the constraints are unaffected by the isometry, we deduce that the solution to this convex optimization problem is identical to the solution of (D). Hence, we have that and for all . Again, we can make these replacements in (5).
3 We show that . First, we note that the support vector indices are invariant under isometry: and similarly . Hence we have
By making the replacements from points 1, 2 and 3 in (5), we deduce that . ∎
Appendix E Empirical Evaluation
This appendix provides useful details to reproduce the empirical evaluation from Section 3.
Computing Resources. All the empirical evaluations were run on a single machine equipped with a 18-Core Intel Core i9-10980XE CPU and a NVIDIA RTX A4000 GPU. The machine runs on Python 3.9 and Pytorch 1.10.2 .
Dataset licenses. The MNIST dataset dataset is made available under the terms of the Creative Commons Attribution-Share Alike 3.0 License. The ECG dataset dataset is made available under the terms of the Open Data Commons Attribution License v1.0. The CUB dataset is made available for non-commercial research and educational purposes.
Models. The detailed architecture of the models are provided in Tables 2, 3 and 4. The InceptionV3 architecture is the same as in the literature . We use its official Pytorch implementation.
Data Split. All the datasets are naturally split in training and testing data. In the ECG dataset, the different types of abnormal heartbeats are imbalanced (e.g. the fusion beats constitute only of the training set). Hence, we create a synthetic training set with balanced concepts using SMOTE . As in , the CUB dataset is augmented by using random crops and random horizontal flips.
Model Fitting. In fitting each model, we use the test set as a validation set since our purpose is not to obtain the models with the best generalization but simply models that perform well on a set of examples we wish to explain (here the examples from the test set). All the models are trained to minimize the cross-entropy between their prediction and the true labels. The hyperparameters are as follows. 1 For MNIST we use a Adam optimizer with batches of 120 examples, a learning rate of , a weight decay of for epochs with patience . 2 For ECG we use a Adam optimizer with batches of 300 examples, a learning rate of , a weight decay of for epochs with patience . 3 For CUB, we use a stochastic gradient descent optimizer with batches of 64 examples, a learning rate of , a weight decay of for epochs with patience .
Concepts. The concept mapping between MNIST classes and concepts is provided in Table 5. For the ECG and the CUB datasets, the presence/absence of a concept for each example is readily available in the dataset.
Concept classifiers. All the concept classifiers are implemented with scikit-learn . For CAR classifiers, we fit a SVC with Gaussian RBF kernel and default hyperparameters from scikit-learn. For CAV classifiers, we fit a linear classifier with a stochastic gradient descent optimizer with learning rate and a tolerance of for epochs and the remaining default hyperparameters from scikit-learn.
Statistical significance. The statistical significance test from Section 3.1.1 is performed with the scikit-learn implementation of the permutation test. For MNIST and ECG, we consider 100 permutations per concept. For CUB, this test is more expensive since the latent spaces are high-dimensional. We consider only 25 permutations per concept in that case.
Concept-based feature importance. In the experiment from Section 2.3, we use Captum’s implementation of Integrated Gradients with default parameters. In the case of CUB, storing the feature importance scores for each concept and for the whole test set requires a prohibitive amount of memory. To avoid this problem, we select concepts and subsample 50 positive and 50 negative examples per concept from the test set. This corresponds to a set of examples. We compute the feature importance for these examples only.
Alternative architecture. We extended the analysis of Section 3.1.1 to a ResNet-50 architecture. We fine-tune the ResNet model on the CUB dataset and reproduced the experiment from Section 3.1.1 with this new architecture. In particular, we fit a CAR and a CAV classifier on the penultimate layer of the ResNet. We then measure the accuracy averaged over the CUB concepts. This results in accuracy for CAR classifiers and accuracy for CAV classifiers. We deduce that CAR classifiers are highly accurate to identify concepts in the penultimate ResNet layer. As in the main paper, we observe that CAR classifiers outperform CAV classifiers, although the gap is smaller than for the Inception-V3 neural network. We deduce that our CAR formalism extends beyond the architectures explored in the paper and we hope that CAR will become widely used to interpret any more architectures.
Appendix F Use Case
This appendix provides useful details to reproduce the use case from Section 3.2.
Computing Resources. The use case was run on a single machine equipped with a 18-Core Intel Core i9-10980XE CPU and a NVIDIA RTX A4000 GPU. The machine runs on Python 3.9 and Pytorch 1.10.2 .
Dataset license. The SEER dataset is made available under the terms of the SEER Research Data Use Agreement.
Model. The detailed architecture of the model is provided in Table 6.
Data split. We randomly split the whole SEER dataset into a training set ( of the data) and a test set (the remaining ). Since patients with a death outcome are in minority (less than ), we oversample them to obtain a balanced training set.
Model fitting. In fitting the model, we use the test set as a validation set since our purpose is not to obtain the model with the best generalization but simply a model that performs well on a set of examples we wish to explain (here the examples from the test set). The model is trained to minimize the cross-entropy between its prediction and the true labels. We use a Adam optimizer with batches of 500 examples, a learning rate of , a weight decay of for epochs with patience .
Concepts. The concepts correspond to prostate cancer grades. Those grades can be computed from the Gleason score as follows :
It goes without saying that the model is trained with the Gleason scores only, not with the grades.
Concept classifiers. We fit a SVC with linear kernel and default hyperparameters from scikit-learn.
Concept-based feature importance. We use Captum’s implementation of Integrated Gradients with default parameters.
Appendix G Explanation Robustness
We observe that the TCAR scores keep a high correlation with the true proportion of examples that exhibit the concept even when all the test examples are adversarially perturbed. We conclude that TCAR explanations are robust with respect to adversarial perturbations in this setting.
Appendix H Using CAR to Understand Unsupervised Concepts
Our CAR formalism adapts to a wide variety of neural network architectures. In this appendix, we use CAR to analyze the concepts discovered by a self explaining neural network (SENN) trained on the MNIST dataset. As in , we use a SENN of the form
Where and are respectively the activation and the relevance of the synthetic concept discovered by the SENN model. We follow the same training process as . This yields a set of concepts explaining the predictions made by the SENN .
When this correlation increases, the concepts and tend to be relevant together more often. We report the correlation between each pair in Table 8.
SENN Concept 2 correlates well with the Vertical Line Concept.
SENN Concept 3 correlates well with the Horizontal Line Concept
SENN Concept 5 correlates well with the Loop Concept.
SENN Concepts 1 and 4 are not well covered by our concepts.
The above analysis shows the potential of our CAR explanations to better understand the abstract concepts discovered by SENN models. We believe that the community would greatly benefit from the ability to perform similar analyses for other interpretable architectures, such as disentangled VAEs.
Appendix I Using CAR with NLP
CAR is a general framework and can be used in a wide variety of domains that involve neural networks. In the main paper, we show that CAR provides explanations for various modalities:
We assess the generalization performance of the CAR classifier on a holdout concept set made of concept positive and negative sentences (60 sentences in total). The CAR classifier has an accuracy of on this holdout dataset. This suggests that the concept is smoothly encoded in the model’s representation space, which is consistent with the importance of positive adjectives to identify positive reviews. We deduce that our CAR formalism can be used in a NLP setting. We believe that using CARs to analyze large-scale language model would be an interesting study that we leave for future work.
Appendix J Increasing Explainability at Training Time
Improving neural networks explainability at training time constitutes a very interesting area of research but is beyond the scope of our paper. That said, we believe that our paper indeed contains insights that might be the seed of future developments in neural network training. As an illustration, we consider an important insight from our paper: the fact that the accuracy of concept classifiers seems to increase with the depth of the layer for which we fit a classifier. In the main paper, this is mainly reflected in Figure 4. This observation has a crucial consequence: it is not possible to reliably characterize the shallow layers in terms of the concepts we use.
In order to improve the explainability of those shallow layers, one could leverage the recent developments in contrastive learning. The purpose of this approach would be to separate the concept set and in the representation space corresponding to a shallow layer of the neural network. A practical way to implement this would be to follow . Assume that we want to separate concept positives and negatives in the representation space induced by the shallow feature extractor . As , one can use a projection head and enforce the separation of the concept sets through the contrastive loss
Appendix K Further Examples
In this appendix, we provide several examples to illustrate the experiments from Section 3.
Concept classifiers. The accuracy of the concept classifiers for all the MNIST and ECG concepts are given in Figures 11 and 12. Each box-plot is built with 10 random seeds where the concept sets and are allowed to vary. We observe that the CAR classifiers are more accurate for each of the observed concepts
Global explanations. The global concept explanations for all the MNIST concepts and some of the CUB concepts are given in Figures 13 and 14. As the examples from Section 3.1.2, we see that TCAR explanations are more consistent with the human concept annotations.
Saliency maps. Examples of concept-based an vanilla saliency maps for MNIST, ECG and CUB examples are given in Figures 15 , 16 , 17 , 18 and 19. As explained in the main paper, we observe that concept-based saliency maps are indeed distinct form vanilla saliency maps. Furthermore, saliency maps for different concepts are not interchangeable. In the CUB case, we note that concept are not always identified with the minimal amount of features (e.g. some of the saliency maps from Figure 19 highlight pixels that do not always belong to the bird’s breast). This surprising observation seems to occur even for concept-bottleneck models that are explicitly trained to recognize concepts . We believe that this could be improved by training concept classifiers with images that include a segmentation highlighting the concept of interest (e.g. the bird’s breast). We leave this idea for future works.