Concept Activation Regions: A Generalized Framework For Concept-Based Explanations

Jonathan Crabbé, Mihaela van der Schaar

Introduction

Deep learning models are both useful and challenging. Their utility is reflected in their increasing contributions to sophisticated tasks such as natural language processing , computer vision and scientific discovery . Their challenging nature can be attributed to their inherent complexity. State of the art deep models typically contain millions to billions parameters and, hence, appear as black-boxes to human users. The opacity of black-box models make it difficult to: anticipate how models will perform at deployment ; reliably distil knowledge from the models and earn the trust of stakeholders in high-stakes domains . With the aim of increasing the transparency of black-box models, the field of explainable AI (XAI) developed . We can broadly divide XAI methods in 2 categories: 1 Methods that restrict the model’s architecture to enable explanations. Examples include attention models that motivate their predictions by highlighting features they pay attention to and prototype-based models that motivate their predictions by highlighting relevant examples from their training set . 2 Post-hoc methods that can be used in a plug-in fashion to provide explanation for a pre-trained model. Examples include feature importance methods (also known as feature attribution or saliency methods) that highlight features the model is sensitive to ; example importance methods that identify influential training examples and hybrid methods combining the two previous approaches . In this work, we focus on a different type of explanation methods known as concept-based explanations. Let us now summarize the relevant literature to contextualize our own contribution.

Related work. Concept-based explanations were first formalized with the concept activation vector (CAV) formalism . Given a concept specified by the user (e.g. stripes in an image), linear classifiers are used as probes to assess whether a deep model’s representation space separates examples where a concept is present (concept positive examples) from examples where a concept is absent (concept negative examples). A CAV is then extracted from the linear classifier’s weights. With this CAV, it is possible to provide post-hoc explanations such as the sensitivity of a model’s prediction to the presence/absence of a concept. It goes without saying that several concepts are needed to explain the prediction of a deep model. To formalize this idea, existing works have used concept basis decomposition and sufficient statistics . Although most applications of CAV involve image data, we note that the formalism has been successfully applied to time-series data . While concepts are typically specified by the user, early works in computer vision have been undertaken to discover concepts in the form of meaningful image segmentations . The main criticism against the CAV formalism is that it requires concept positive examples to be linearly separable from concept negative examples . This is because linear separability of concept sets is a restrictive criterion that the model is not explicitly trained to fulfil. To address this issue, works like concept whitening transformations and concept bottleneck models propose to do away with the post-hoc nature of concept-based explanations. These methods introduce new neural network architectures that permit to train the models with concept labels. We stress that this requires the set of concepts to be specified before training the model. This assumes that we know what concepts are relevant for a model to solve a downstream task a-priori. This is not the case whenever we train a model to solve a task for which little or no knowledge is available. In this setup, it seems more appropriate to train a model for the task first and, then, perform a post-hoc analysis of the model to determine the concepts that were relevant in providing a solution. Furthermore, recent concerns have emerged regarding the reliability of interpretations provided by these altered model architectures .

Contributions. In this work, our purpose is to retain the flexible post-hoc nature of concept-based explanations without assuming that the concept sets are linearly separable. To that aim, we introduce concept activation regions (CARs), an extension of the CAV framework illustrated in Figure 1. 1 Generalized formalism. As a substitute to linear separability, we propose in Section 2.1 to adapt the smoothness assumption from semi-supervised learning . Intuitively, this more general assumption only requires positive and negative examples to be scattered across distinct clusters in the model’s representation space. In practice, this generalization is implemented by substituting CAV’s linear classifiers by kernel-based support vector classifiers. We demonstrate that choosing radial kernels leads to CAR classifiers that are invariant under isometries of the latent space. This permits to assign identical explanations to latent spaces characterized by the same geometry. Moreover, we show in Section 3.1.1 that our CAR classifiers yield a substantially more accurate description of how concepts are distributed in deep model’s representation spaces. 2 Better global explanations. With the nonlinear decision boundaries of our support vector classifiers, there is no obvious way to adapt the notion of concept activation vectors. Since CAV’s concept importance (TCAV score) is computed with these vectors, we need an alternative approach. In Section 2.2, we propose to define concept importance by building on our smoothness assumption. Concretely, a concept is important for a given example if the model’s representation for this example lies in a cluster of concept positive representations. With this characterization, we define TCAR scores that are the CAR equivalent of TCAV scores. In Section 3.1.2, we demonstrate that TCAR scores lead to global explanations that are more consistent with concept annotations provided by humans. 3 Concept-based feature importance. In Section 2.3, we argue that our CAR formalism permits to assign concept-specific feature importance scores for each example fed to the neural network. We verify empirically in Section 3.1.3 that those feature importance scores reflect meaningful concept associations. Finally, we illustrate in Section 3.2 how these contributions permit to establish that deep models implicitly discover known scientific concepts.

Concept Activation Regions (CARs)

A better assumption. Although the representations from Figure 2 do not linearly separate the concept sets, we note that concept positive and negative examples are scattered across distinct clusters. In this way, a model that produces these representations appears to make a difference between presence and absence of the concept. From this angle, we could consider that the concept is well encoded in the representation space geometry. To formalize this more general notion of concept set separability, we adapt the smoothness assumption originally formulated in semi-supervised learning .

A concept c∈[C]c\in[C] is encoded in the latent space H\mathcal{H} if H\mathcal{H} is smooth with respect to the concept. This means that we can separate H=Hc⨆H¬c\mathcal{H}=\mathcal{H}^{c}\bigsqcup\mathcal{H}^{\neg c} into a concept activation region (CAR) Hc\mathcal{H}^{c} where the concept cc is mostly present (i.e. ∣g(Pc)⋂Hc∣≫∣g(Nc)⋂Hc∣\left|\boldsymbol{g}(\mathcal{P}^{c})\bigcap\mathcal{H}^{c}\right|\gg\left|\boldsymbol{g}(\mathcal{N}^{c})\bigcap\mathcal{H}^{c}\right|) and a region H¬c\mathcal{H}^{\neg c} where the concept cc is mostly absent (i.e. ∣g(Nc)⋂H¬c∣≫∣g(Pc)⋂H¬c∣\left|\boldsymbol{g}(\mathcal{N}^{c})\bigcap\mathcal{H}^{\neg c}\right|\gg\left|\boldsymbol{g}(\mathcal{P}^{c})\bigcap\mathcal{H}^{\neg c}\right|). If two points h1,h2∈H\boldsymbol{h}_{1},\boldsymbol{h}_{2}\in\mathcal{H} in a high-density region of the latent space are close to each other, then we should have h1,h2∈Hc\boldsymbol{h}_{1},\boldsymbol{h}_{2}\in\mathcal{H}^{c} or h1,h2∈H¬c\boldsymbol{h}_{1},\boldsymbol{h}_{2}\in\mathcal{H}^{\neg c}.

We can easily see that linear separability is trivially included in this assumption: it corresponds to the case where the CAR Hc\mathcal{H}^{c} and H¬c\mathcal{H}^{\neg c} are separated by a hyperplane. Conversely, our assumption does not require the CAR Hc\mathcal{H}^{c} and H¬c\mathcal{H}^{\neg c} to be separated by a hyperplane for a concept to be relevant, as illustrated in Figure 3.

There are two crucial components in the previous assumption that need to be detailed: what do we mean by density and how do we extract a CAR Hc\mathcal{H}^{c} from H\mathcal{H}.

2 Detecting Concepts

The concept density for an example x∈X\boldsymbol{x}\in\mathcal{X} can similarly be defined as ρc[g(x)]\rho^{c}[\boldsymbol{g}(\boldsymbol{x})].

This density function is not necessarily positive. Indeed, we have assigned a positive contribution for examples from Pc\mathcal{P}^{c} and a negative one for those of Nc\mathcal{N}^{c}. The idea is that ρc(h)>0\rho^{c}(\boldsymbol{h})>0 whenever the density of g(Pc)\boldsymbol{g}(\mathcal{P}^{c}) is higher around h\boldsymbol{h}. Conversely, ρc(h)<0\rho^{c}(\boldsymbol{h})<0 whenever the density of g(Nc)\boldsymbol{g}(\mathcal{N}^{c}) is higher around h\boldsymbol{h}. Finally, ρc(h)≈0\rho^{c}(\boldsymbol{h})\approx 0 if h\boldsymbol{h} is isolated from positive and negative examples or if the density of g(Pc)\boldsymbol{g}(\mathcal{P}^{c}) balances the density of g(Nc)\boldsymbol{g}(\mathcal{N}^{c}) around h\boldsymbol{h}.

Global explanations. Our CAR formalism permits to extend those local (i.e. sample-wise) considerations globally. Let us start by defining the equivalent of TCAV scores described in Section 2.1. This score aims at understanding how the model relates classes with concepts. We define the TCAR score as the fraction of examples that have class kk and whose representation lies in the CAR Hc\mathcal{H}^{c}: TCARkc=\nicefrac∣g(Dk)⋂Hc∣∣Dk∣\textrm{TCAR}^{c}_{k}=\nicefrac{{\left|\boldsymbol{g}(\mathcal{D}_{k})\bigcap\mathcal{H}^{c}\right|}}{{\left|\mathcal{D}_{k}\right|}}. Note that TCARkc∈\textrm{TCAR}^{c}_{k}\in, where corresponds to no overlap and 11 to a full overlap. In Appendix B, we extend this approach to measure the overlap between two concepts.

Latent space isometries invariance. In many applications such as clustering or data visualization , the only relevant geometrical information of the representation space H\mathcal{H} is the distance ∥h1−h2∥H\left\|\boldsymbol{h}_{1}-\boldsymbol{h}_{2}\right\|_{\mathcal{H}} between every pair of points (h1,h2)∈H2(\boldsymbol{h}_{1},\boldsymbol{h}_{2})\in\mathcal{H}^{2}. Informally, we say that two representation spaces are isometric if they assign the same distance to each pair of points. In the aforementioned applications, two isometric representation spaces are therefore indistinguishable from one another. Since concept-based explanations similarly describe the representation space geometry, one might require similar invariance to hold in this context. In Appendix D, we show that our CAR formalism provides such guarantee if κ\kappa is a radial kernel . To the best of our knowledge, this type of analysis has not been performed in the context of concept-based explanation methods. We believe that future works in this domain would greatly benefit from this type of insight.

3 Concepts and Features

Experiments

The code to reproduce all the experiments from this section is available at https://github.com/JonathanCrabbe/CARs and https://github.com/vanderschaarlab/CARs.

Our purpose is to empirically validate the formalism described in the previous section. We have several independent components to evaluate: 1 the concept classifier used to detect the CARs Hc\mathcal{H}^{c}, 2 the global explanations induced by the TCAR values and 3 the feature importance scores induced by the concept densities ρc\rho^{c}.

Datasets. We perform our experiments on 3 datasets. 1 The MNIST dataset consists of 28×2828\times 28 grayscale images, each representing a digit. We train a convolutional neural network (CNN) with 2 layers to identify the digit of each image. 2 The MIT-BIH Electrocardiogram (ECG) dataset consists of univariate time series with 187187 time steps, each representing a heartbeat cycle. We train a CNN with 3 layers to determine whether each heartbeat is normal or abnormal. 3 The Caltech-UCSD Birds-200 (CUB) dataset consists of coloured images of various sizes, each representing a bird from one of the 200200 species present in the dataset. We fine-tune an Inceptionv3 neural network to identify the species each bird belongs to among the 200200 possible choices.

Concepts. For each dataset, we study the models through the lens of several well-defined concepts that are provided by human annotations. 1 For MNIST: we use C=4C=4 concepts that correspond to simple geometrical attributes of the images: Loop (positive images include a loop), Vertical/Horizontal Line (positive images include a vertical/horizontal line) and Curvature (positive images contain segments that are not straight lines). 2 For ECG: we use C=4C=4 concepts defined by cardiologists to better characterize abnormal heartbeats : Premature Ventricular, Supraventricular, Fusion Beats and Unknown. The exact definition for each of these concepts is beyond the scope of this paper. We simply note that annotations for those concepts are available in the ECG dataset. 3 For CUB: we use C=112C=112 concepts that correspond to visual attributes of the birds (e.g. their size, the colour of their wings, etc.). We use the same procedure as to extract those concepts from the CUB dataset. We stress that, in each case, we selected concepts that can be unambiguously associated to the classes that are predicted by the models. Hence, it is reasonable to expect those concepts to be salient for the models. In each case, we sample the positive and negative sets Pc\mathcal{P}^{c} and Nc\mathcal{N}^{c} from the model’s training sets. For more details on the concepts and the models, please refer to Appendix E.

Methodology. The purpose of this experiment is to assess if the concept regions Hc\mathcal{H}^{c} identified by our CAR classifier generalize well to unseen examples. Each of the models described above are endowed with several representation spaces (one per hidden layer). For several of those latent spaces, we fit our CAR classifier (SVC with radial basis function kernel) to discriminate the concept sets Pc,Nc\mathcal{P}^{c},\mathcal{N}^{c} for each concept c∈[C]c\in[C]. These two sets have a size Nc=200N^{c}=200 and are sampled from the model’s training set. The classifier is then evaluated by computing its accuracy on a holdout balanced concept set Tc\mathcal{T}^{c} of size 100 sampled from the model’s testing set. For MNIST and ECG, we repeat this experiment 10 times for each concept and let the sets Pc,Nc,Tc\mathcal{P}^{c},\mathcal{N}^{c},\mathcal{T}^{c} vary on each run. For comparison, we perform the same experiment with a linear CAV classifier as a benchmark. We report the overall (all the concepts together) accuracy in Figure 4.

Analysis. The CAR classifier substantially outperforms the CAV classifier. Note that this advantage is even more striking in the representation spaces associated to the deeper (last) DNN layers. This can be better understood through the lens of Cover’s theorem : the linear separation underlying CAV is usually easier to achieve in higher dimensional spaces, which corresponds to the shallower (first) DNN layers in this case. When the dimension dHd_{H} of the latent space becomes comparable to the size NcN^{c} of the concept sets, linear separation often fails to maintain a high accuracy. By contrast, the SVC underlying CARs manage to maintain high accuracy through the more flexible notion of concept smoothness defined in Assumption 2.1. This suggests that concepts can be well encoded in the geometry of the latent space even when accurate linear separability is not possible. We also note that the accuracy CAR classifiers seems to increase with the representation’s depth. This is consistent with the behaviour of class probes .

Statistical significance. The statistical significance of the concept classifiers is evaluated with the permutation test from . All of them are statistically significant with p-value <.05<.05 except for some concepts classifiers that are fitted with the layers Mixed5d and Mixed6e of the CUB Inceptionv3 model. We note that those classifiers do not generalize well in Figure 4(c). This suggests that deeper networks are required to identify more challenging concepts correctly.

Take-away 1: CAR classifiers better capture how concepts are spread across representation spaces.

1.2 Consistency of global explanations

Analysis. The TCAR scores better correlate with the true presence of concepts. This difference can be understood by looking at the examples from Figure 5. We note that TCAV scores tend to predict nonexistent associations (e.g. yellow wings for American crows) and miss existing associations (e.g. curvature for digit 2). For ECG, we note that TCAR does capture the fact that some concepts (like fusion beats) are less represented within the class, while TCAV does not. In all of these cases, TCAV explanations might give the impression that the model did not learn meaningful class-concept associations. The TCAR analysis leads to the opposite conclusion. Since we have established that TCAR is built upon more accurate concept classifiers, it seems that models indeed learn concepts as intended, in spite of what TCAV explanations suggest.

Take-away 2: TCAR scores more faithfully reflect the true association between classes and concepts.

1.3 Coherency of concept-based feature importance

Analysis. First, we observe that CAR-based feature importance correlates weakly with vanilla feature importance (∣r∣<.25\left|r\right|<.25 for all datasets). This confirms that desideratum 1 is fulfilled. Then, we note that most of the CAR-based feature importance scores are decorrelated or weakly correlated with each other. The counterexamples that we observe indeed correspond to concepts that can be identified with similar features. A first example is the positive correlation between the loop and the curvature concepts for MNIST (r=.71r=.71), both concepts are generally associated to curved symbols. Another example is the negative correlation between striped and solid back patterns for CUB birds (r=−.86r=-.86), those concepts are mutually exclusive and identified by inspecting the bird’s back. Those examples support that CAR-based feature importance fulfils desideratum 2.

Take-away 3: CAR-based feature importance is concept-specific and captures concept associations.

2 Use Case: Machine Learning Model Rediscovering Known Medical Concepts

We will now describe a use case of the CAR formalism introduced in this paper. We stress that the literature already contains numerous use cases of CAV concept-based explanations, especially in the medical setting . Since CARs generalize CAVs, it goes without saying that they apply to these use cases. Rather than repeating existing usage of concept-based explanations, we discuss an alternative use case motivated by recent trends in machine learning. With the successes of deep models in various scientific domains , we witness an increasing overlap between scientific discovery and machine learning. While this new trend opens up fascinating opportunities, it comes with a set of new challenges. The evaluation of machine learning models is arguably one of the most important of these challenges. In a scientific context, the canonical machine learning approach to validate models (out-of-sample generalization) is likely to be insufficient. Beyond generalization on unseen data, the scientific validity of a model requires consistency with established scientific knowledge . We propose to illustrate how our CARs can be used in this context.

Dataset. We use the data collected with the Surveillance, Epidemiology, and End Results (SEER) Program. The dataset contains a US population-based cohort of 171,942 men diagnosed with non-metastatic prostate cancer between Jan 1, 2000, and Dec 31, 2016. Each patient is described by age, the results of a prostate-specific antigen blood test (PSA), the clinical stage of its tumour and two Gleason scores (primary and secondary). Each patient also has a label that indicates if they died because of their prostate cancer. We train a multilayer perceptron (MLP) to predict the patient’s mortality on 90%90\% of the data and test on the remaining 10%10\%. For a more detailed description of the data and the model, please refer to Appendix F.

Concepts. Doctors use an established grading system to predict how likely the cancer is to spread . This system assigns to each patient a grade between 1 and 5. The probability that the cancer spreads increases with this grade. It can be computed from the two Gleason scores (more details in Appendix F). We can consider each of these 5 grades as a concept. We note that this grade is not explicitly part of the input features of the MLP that we trained. Our purpose is to assess if the MLP implicitly discovered those grades in order to predict the patient’s mortality. If this happens to be the case, this would demonstrate that the MLP is in-line with existing medical knowledge.

Analysis. Let us summarize the findings for each of the above points. 1 All of the CAR classifiers generalize well on the test set (their accuracy ranges from 90%90\% to 100%100\%). This strongly suggests that the MLP implicitly separates patients with different grades in representation space. 2 Figure 7(a) demonstrates that the model associates higher grades with higher mortality. This is in line with the clinical interpretation of the grades. 3 Figure 7(b) suggests that the Gleason scores constitute the most important features overall (highest quantiles) for the model to discriminate between grades. This is consistent with the fact that grades are computed with the Gleason scores. We conclude that the MLP implicitly identifies the patient’s grades with high accuracy and in a way that is consistent with the medical literature.

Take-away 4: The CAR formalism can reliably support scientific evaluation of a model.

Conclusion

We introduced Concept Activation Regions, a new framework to relax the linear separability assumption underlying the Concept Activation Vector formalism. We showed that our framework guarantees crucial properties, such as invariance with respect to latent symmetries. Through extensive validation on several datasets, we verified that 1 Concept Activation Regions better capture the distribution of concepts across the model’s representation space, 2 The resulting global explanations are more consistent with human annotations and 3 Concept Activation Regions permit to define concept-specific feature importance that is consistent with human intuition. Finally, through a use case involving prostate cancer data, we show the neural network can implicitly rediscover known scientific concepts, such as the prostate cancer grading system.

Many important points that were not covered in the main paper can be found in the appendices. In Appendix A, we discuss how to tune the various hyperparameters of the CAR classifiers. In Appendix B, we explain how to generalize concept activation vectors to nonlinear decision boundaries. In Appendix G, we show that the CAR explanations are robust to adversarial perturbations and background shifts. In Appendix H, we demonstrate that CAR explanations can be used to relate abstract concepts discovered by self-explaining neural networks with human concepts. Finally, Appendix I illustrates how CAR explanations allow us to probe language models.

We believe that our extended concept explainability framework opens up many interesting avenues for future work. A first one would be to probe state of the art neural networks with an approach similar to Section 3. In particular, it would be interesting to analyse if improving model performance is associated with a better encoding of human concepts. A second one would be to analyze how concept discovery can benefit from our generalized notion of concept activation. A more fine-grained characterization of a model’s latent space is likely to improve the surfaced concept. A third one, as suggested by Section 3.2, would be to use concept-based explanations to make a better scientific assessment of neural networks. Indeed, consistency with well-established knowledge is crucial for a scientific model to be accepted.

Acknowledgments and Disclosure of Funding

The authors are grateful to Fergus Imrie, Yangming Li and the 3 anonymous NeurIPS reviewers for their useful comments on an earlier version of the manuscript. Jonathan Crabbé is funded by Aviva and Mihaela van der Schaar by the Office of Naval Research (ONR), NSF 172251.

References

Checklist

Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes] All the claims from the abstract and Section 1 are verified empirically in Section 3. The theoretical claims are demonstrated in Appendices C and D.

Did you describe the limitations of your work? [Yes] The assumption for our concept-based explanations to be valid are clearly stated in Assumption 2.1. Our method only applies to neural networks, which is stated in Section 2.

Did you discuss any potential negative societal impacts of your work? [Yes] Our work improves the transparency of deep neural networks. This has many beneficial societal impacts, as described in Section 1. We do not see any potential negative societal impact to this approach. We have carefully reviewed the points from the ethical guideline and none of the mentioned negative impact seems to apply to our work.

Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes] We have read the ethics review guidelines and confirm that our paper conforms to them.

If you are including theoretical results…

Did you state the full set of assumptions of all theoretical results? [Yes] The assumptions in Propositions C.1 and D.1 are clearly stated in Appendices C and D respectively.

Did you include complete proofs of all theoretical results? [Yes] The proofs for Propositions C.1 and D.1 are given in Appendices C and D respectively.

Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes] The full code is available at https://github.com/JonathanCrabbe/CARs and https://github.com/vanderschaarlab/CARs. The implementation of our method closely follows Algorithms 1, 2 and 4 in the appendices. All the details to reproduce the experimental results are given in Appendices E and F.

Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes] All the training details are specified in Appendices E and F.

Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [Yes] Figure 4 measures the accuracy of our method over several runs. All the runs are aggregated together in the form of a box-plot.

Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes] Our computing resources are described in Appendices E and F.

If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…

If your work uses existing assets, did you cite the creators? [Yes] We have included a citation for the MNIST dataset , for the ECG dataset , for the CUB dataset , for the SEER dataset and for the InceptionV3 model .

Did you mention the license of the assets? [Yes] The licenses are mentioned in Appendices E and F.

Did you include any new assets either in the supplemental material or as a URL? [N/A] We have used only existing assets.

Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [N/A] All the datasets that we use are publicly available.

Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [N/A] The medical datasets we use are public and have been de-identified.

If you used crowdsourcing or conducted research with human subjects…

Did you include the full text of instructions given to participants and screenshots, if applicable? [N/A]

Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A]

Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [N/A]

Appendix A CAR Classifiers

In this appendix, we provide some details about our CAR classifiers.

Implementation. To determine the concept activation regions, we fit a SVC sκcs^{c}_{\kappa} for each concept c∈[C]c\in[C]. This process is described in Algorithm 1.

Choosing the kernel. Algorithm 1 requires the user to specify a kernel function κ\kappa. To select this kernel, a good approach is to train several SVCs sκcs^{c}_{\kappa} with different kernels κ\kappa and see how each SVC generalizes on a validation set. We illustrate this process with MNIST in Figure 8. We note that the Gaussian RBF kernel outperforms the other kernels on the validation set, although the Matern kernel achieves perfect accuracy on the training set. This highlights the importance of evaluating a CAR classifier on a held-out dataset to make sure that the related CAR offers a good description of how concepts are distributed in the latent space H\mathcal{H}. In our experiments, we found that Gaussian RBF kernels are often the most interesting option.

Concept set size. Algorithm 1 requires the user to specify concepts sets Pc\mathcal{P}^{c} and Nc\mathcal{N}^{c}. What size Nc=∣Pc∣=∣Nc∣N^{c}=\left|\mathcal{P}^{c}\right|=\left|\mathcal{N}^{c}\right| should we choose for these concept sets? It seems logical that larger concepts sets are more likely to yield more accurate CARs. To study this experimentally, we propose to fit several concept classifiers for MNIST by varying the size NcN^{c} of their training concept sets Pc\mathcal{P}^{c} and Nc\mathcal{N}^{c}. We report the results in Figure 9. As we can see, the curves flatten above Nc>200N^{c}>200. Increasing the concept sets size beyond this point does not improve the accuracy of the resulting CAR classifier. We recommend to acquire concept examples until the performance of the CAR classifier stabilizes. In our experiments, we found that Nc=200N^{c}=200 examples is often sufficient to obtain accurate CAR classifiers.

Tuning hyperparameters. In the case where the user desires a CAR classifier that generalizes as well as possible, tuning these hyperparameters might be useful. We propose to tune the kernel type, kernel width and error penalty of our CAR classifiers sκcs^{c}_{\kappa} for each concept c∈[C]c\in[C] by using Bayesian optimization and a validation concept set:

Repeat 3-5 for a predetermined number of trials.

We applied this process to the CAR accuracy experiment (same setup as in Section 3.1.1 of the main paper) to tune the CAR classifiers for the CUB concepts. Interestingly, we noticed no improvement with respect to the CAR classifiers reported in the main paper: tuned and standard CAR classifier have an average accuracy of (93±.2)%(93\pm.2)\% for the penultimate Inception layer. This suggests that the accuracy of CAR classifiers is not heavily dependant on hyperparameters in this case.

Appendix B TCAR Global Explanations

In this appendix, we provide some details about the TCAR scores.

Implementation. When the CAR classifiers are available, they permit to compute TCAR scores through Algorithm 2.

TCAR between concepts. Up until now, we have discussed TCAR scores that indicate how models relate classes to concepts. It is possible to define a similar score to estimate how models relate two concepts with each other. Given a set D⊂X\mathcal{D}\subset\mathcal{X} of examples, we define the TCAR score associated to the concepts c1,c2∈[C]c_{1},c_{2}\in[C] as the ratio

Again, TCAR=0\textrm{TCAR}=0 corresponds to no overlap and TCAR=1\textrm{TCAR}=1 describes a perfect overlap. We note that the concept-concept TCAR score can be interpreted as a Jaccard index between the sets g(D)⋂Hc1\boldsymbol{g}(\mathcal{D})\bigcap\mathcal{H}^{c_{1}} and g(D)⋂Hc2\boldsymbol{g}(\mathcal{D})\bigcap\mathcal{H}^{c_{2}}. This score is symmetric with respect to the concepts: TCARc1,c2=TCARc2,c1\textrm{TCAR}^{c_{1},c_{2}}=\textrm{TCAR}^{c_{2},c_{1}}. The computation of this score is done as in Algorithm 3.

TCAR between MNIST concepts. As an illustration, we compute concept-concept TCAR scores for the MNIST concepts and report the results in Figure 10. We see that the model relates concepts that tend to appear together (e.g. curvature and loop). Hence, concept-concept TCAR scores can serve as a proxy for the concept semantics encoded in a model’s representation space. We note the similarity with the correlation between concept-based feature importance illustrated in Figure 6. The main difference is that concept-concept TCAR scores do not explicitly refer to input features.

In this way, all the interpretation provided by the CAV formalism are also available in the CAR formalism.

Appendix C CAR Feature Importance

In this appendix, we provide some details about our concept-based feature importance.

With this property, we can interpret features i∈[dX]i\in[d_{X}] with ai(ρc∘g,x)>0a_{i}(\rho^{c}\circ\boldsymbol{g},\boldsymbol{x})>0 as those that tend to increase the concept density. This means that those features are important for the feature extractor g\boldsymbol{g} to map the example in a region of the representation space H\mathcal{H} where the concept is present. Hence, those are features that are important to identify a given concept c∈[C]c\in[C]. Conversely, features i∈[dX]i\in[d_{X}] with ai(ρc∘g,x)<0a_{i}(\rho^{c}\circ\boldsymbol{g},\boldsymbol{x})<0 tend to decrease the concept density and therefore brings the example x\boldsymbol{x} in a region of the representation space H\mathcal{H} where the concept is absent. These features can therefore be interpreted as important to reject the presence of a given concept c∈[C]c\in[C].

Input baseline choice. The choice of baseline input xˉ\bar{\boldsymbol{x}} has a notable effect on feature importance methods . What constitutes a good input baseline is problem dependant. Intuitively, xˉ\bar{\boldsymbol{x}} should correspond to an input x∈X\boldsymbol{x}\in\mathcal{X} where no information is present . Let us now explain how this information removal is achieved with the datasets that we use in our experiments. 1 For MNIST, we chose a black image xˉ=0\bar{\boldsymbol{x}}=\boldsymbol{0} as an input baseline. This is because MNIST images have a black background an the information comes from white pixels that represent the digits. 2 For ECG, we chose a constant xˉ=0\bar{\boldsymbol{x}}=\boldsymbol{0} time series as an input baseline. This is because ECG time series are normalized (xt∈x_{t}\in for all t∈[dX]t\in[d_{X}]) and the pulse information comes from time steps where the time series is non-vanishing xt≠0x_{t}\neq 0. 3 For CUB, choosing a baseline is more complicated. The reason for this is that the background colour changes from one image to the other and rarely corresponds to a single colour. To address this, we proceed as in the literature and select a baseline xˉ\bar{\boldsymbol{x}} that corresponds to a blurred version of the image we wish to explain: xˉ(x)=Gσ⊗x\bar{\boldsymbol{x}}(\boldsymbol{x})=\boldsymbol{G}_{\sigma}\otimes\boldsymbol{x}, where Gσ\boldsymbol{G}_{\sigma} is a Gaussian filter of width σ\sigma and ⊗\otimes denotes the convolution operation. We note that this baseline depends on which input x∈X\boldsymbol{x}\in\mathcal{X} we want to explain. In our implementation, we use σ=50\sigma=50 to have images that are significantly blurred. 4 For SEER, we chose a constant vector xˉ=0\bar{\boldsymbol{x}}=\boldsymbol{0} as an input baseline. This is because all continuous features are standardized and all categorical features are one-hot encoded.

Appendix D CAR Latent Isometry Invariance

In this appendix, we prove that CAR explanations are invariant under isometries of the latent space when built with a radial kernel. Let us first rigorously define the notion of isometry between two vector spaces.

Let (H,∥⋅∥H)(\mathcal{H},\left\|\cdot\right\|_{\mathcal{H}}) and (H′,∥⋅∥H′)(\mathcal{H}^{\prime},\left\|\cdot\right\|_{\mathcal{H}^{\prime}}) be two normed vector spaces. An isometry from H\mathcal{H} to H′\mathcal{H}^{\prime} is a map τ:H→H′\boldsymbol{\tau}:\mathcal{H}\rightarrow\mathcal{H}^{\prime} such that for all h1,h2∈H\boldsymbol{h}_{1},\boldsymbol{h}_{2}\in\mathcal{H}:

We say that the two spaces (H,∥⋅∥H)(\mathcal{H},\left\|\cdot\right\|_{\mathcal{H}}) and (H′,∥⋅∥H′)(\mathcal{H}^{\prime},\left\|\cdot\right\|_{\mathcal{H}^{\prime}}) are isometric if there exists a bijective isometry τ\boldsymbol{\tau} from H\mathcal{H} to H′\mathcal{H}^{\prime}.

An explanation method is invariant to latent space isometries if applying a bijective isometry τ\boldsymbol{\tau} to the model’s latent space H\mathcal{H} does not affect the explanations produced by the method. To make this more formal, we write the model as f=l∘g=l∘τ−1∘τ∘g\boldsymbol{f}=\boldsymbol{l}\circ\boldsymbol{g}=\boldsymbol{l}\circ\boldsymbol{\tau}^{-1}\circ\boldsymbol{\tau}\circ\boldsymbol{g}, where τ−1\boldsymbol{\tau}^{-1} is the inverse of the bijective isometry τ\boldsymbol{\tau}. In this setup, we could produce explanations with CARs by making the following replacements for the feature extractor and the label map: g↦ττ∘g\boldsymbol{g}\overset{\boldsymbol{\tau}}{\mapsto}\boldsymbol{\tau}\circ\boldsymbol{g} and l↦τl∘τ−1\boldsymbol{l}\overset{\boldsymbol{\tau}}{\mapsto}\boldsymbol{l}\circ\boldsymbol{\tau}^{-1}. The explanations are defined to be invariant to latent space isometries if they are unaffected by this replacement. It is legitimate to expect this since the previous replacement leads to the same model f↦τf\boldsymbol{f}\overset{\boldsymbol{\tau}}{\mapsto}\boldsymbol{f} and substitutes the latent space by an isometric latent space H↦τH′=τ(H)\mathcal{H}\overset{\boldsymbol{\tau}}{\mapsto}\mathcal{H}^{\prime}=\boldsymbol{\tau}(\mathcal{H}).

Let us now discuss the isometry invariance of CAR explanations. First, we recall that CARs are defined through a kernel κ\kappa. Not all kernels lead to isometry invariant CAR explanations. We will show that it holds for a family of kernels known as radial kernels.

For each concept c∈[C]c\in[C], we note that CAR explanations exclusively rely on the concept density ρc\rho^{c} from Definition 2.1 and the associated support vector classifier sκcs_{\kappa}^{c}. Hence, it is sufficient to show that these two functions are invariant under isometry.

Concept Density. We start with the concept density ρc\rho^{c}. We note that the kernel function κ\kappa can be applied to vectors from H′\mathcal{H}^{\prime} as κ′(h1′,h2′)≡χ(∥h1′−h2′∥H′)\kappa^{\prime}(\boldsymbol{h}^{\prime}_{1},\boldsymbol{h}^{\prime}_{2})\equiv\chi(\left\|\boldsymbol{h}^{\prime}_{1}-\boldsymbol{h}^{\prime}_{2}\right\|_{\mathcal{H}^{\prime}}) for all (h1′,h2′)∈H′(\boldsymbol{h}^{\prime}_{1},\boldsymbol{h}^{\prime}_{2})\in\mathcal{H}^{\prime}. Applying CAR in the latent space H′\mathcal{H}^{\prime} isometric to H\mathcal{H} corresponds to using this kernel to compute an alternative density ρ′c\rho^{\prime c}. Let us fix an example x∈X\boldsymbol{x}\in\mathcal{X}. Under the isometry, the concept density for this example transforms as ρc∘g(x)↦τρ′c∘τ∘g(x)\rho^{c}\circ\boldsymbol{g}(\boldsymbol{x})\overset{\boldsymbol{\tau}}{\mapsto}\rho^{\prime c}\circ\boldsymbol{\tau}\circ\boldsymbol{g}(\boldsymbol{x}). Let us show that this is in fact an invariance:

We deduce that the concept density is invariant under isometry: ρc∘g(x)↦τρc∘g(x)\rho^{c}\circ\boldsymbol{g}(\boldsymbol{x})\overset{\boldsymbol{\tau}}{\mapsto}\rho^{c}\circ\boldsymbol{g}(\boldsymbol{x}).

SVC. The proof is more involved for the SVC sκcs^{c}_{\kappa}. We assume that the reader is familiar with the standard theory of SVC. If this is not the case, please refer e.g. to Chapter 7 of . For the sake of notation, we will abbreviate g(xc,n)\boldsymbol{g}(\boldsymbol{x}^{c,n}) and g(x¬c,n)\boldsymbol{g}(\boldsymbol{x}^{\neg c,n}) by hc,n\boldsymbol{h}^{c,n} and h¬c,n\boldsymbol{h}^{\neg c,n} respectively. Similarly, we abbreviate τ∘g(xc,n)\boldsymbol{\tau}\circ\boldsymbol{g}(\boldsymbol{x}^{c,n}) and τ∘g(x¬c,n)\boldsymbol{\tau}\circ\boldsymbol{g}(\boldsymbol{x}^{\neg c,n}) by h′c,n\boldsymbol{h}^{\prime c,n} and h′¬c,n\boldsymbol{h}^{\prime\neg c,n} respectively. The SVC concept classifier can be written as

where Sc={n∈[Nc]∣αc,n≠0}\mathcal{S}^{c}=\{n\in[N^{c}]\mid\alpha^{c,n}\neq 0\} and S¬c={n∈[Nc]∣α¬c,n≠0}\mathcal{S}^{\neg c}=\{n\in[N^{c}]\mid\alpha^{\neg c,n}\neq 0\} are the indices of the support vectors. Under isometry, the SVC classification for a latent vector h∈H\boldsymbol{h}\in\mathcal{H} transforms as sκc(h)↦τsκ′′c(h′)s^{c}_{\kappa}(\boldsymbol{h})\overset{\boldsymbol{\tau}}{\mapsto}s^{\prime c}_{\kappa^{\prime}}(\boldsymbol{h}^{\prime}) with h′=τ(h)\boldsymbol{h}^{\prime}=\boldsymbol{\tau}(\boldsymbol{h}) and

We will now show that sκc(h)=sκ′′c(h′)s^{c}_{\kappa}(\boldsymbol{h})=s^{\prime c}_{\kappa^{\prime}}(\boldsymbol{h}^{\prime}) so that the SVC is invariant under isometries. We proceed in 3 steps.

1 We show that κ′[h1′,h2′]=κ[h1,h2]\kappa^{\prime}\left[\boldsymbol{h}^{\prime}_{1},\boldsymbol{h}^{\prime}_{2}\right]=\kappa\left[\boldsymbol{h}_{1},\boldsymbol{h}_{2}\right] for any (h1,h2)∈H2(\boldsymbol{h}_{1},\boldsymbol{h}_{2})\in\mathcal{H}^{2} and h1′=τ(h1),h2′=τ(h2)\boldsymbol{h}^{\prime}_{1}=\boldsymbol{\tau}(\boldsymbol{h}_{1}),\boldsymbol{h}^{\prime}_{2}=\boldsymbol{\tau}(\boldsymbol{h}_{2}):

By injecting this in (5), we are able to make the following replacements: κ′[h′,h′c,n]=κ[h,hc,n]\kappa^{\prime}\left[\boldsymbol{h}^{\prime},\boldsymbol{h}^{\prime c,n}\right]=\kappa\left[\boldsymbol{h},\boldsymbol{h}^{c,n}\right] and κ′[h′,h′¬c,n]=κ[h,h¬c,n]\kappa^{\prime}\left[\boldsymbol{h}^{\prime},\boldsymbol{h}^{\prime\neg c,n}\right]=\kappa\left[\boldsymbol{h},\boldsymbol{h}^{\neg c,n}\right] for all n∈[Nc]n\in[N^{c}].

2 We show that α′c,n=αc,n\alpha^{\prime c,n}=\alpha^{c,n} and α′¬c,n=α¬c,n\alpha^{\prime\neg c,n}=\alpha^{\neg c,n} for all n∈[Nc]n\in[N^{c}]. To that aim, we note that α′c\boldsymbol{\alpha}^{\prime c} and α′¬c\boldsymbol{\alpha}^{\prime\neg c} maximize the objective

Since this objective is identical to the one from the original SVC and the constraints are unaffected by the isometry, we deduce that the solution to this convex optimization problem is identical to the solution of (D). Hence, we have that α′c,n=αc,n\alpha^{\prime c,n}=\alpha^{c,n} and α′¬c,n=α¬c,n\alpha^{\prime\neg c,n}=\alpha^{\neg c,n} for all n∈[Nc]n\in[N^{c}]. Again, we can make these replacements in (5).

3 We show that β′=β\beta^{\prime}=\beta. First, we note that the support vector indices are invariant under isometry: S′c={n∈[Nc]∣α′c,n≠0}={n∈[Nc]∣αc,n≠0}=Sc\mathcal{S}^{\prime c}=\{n\in[N^{c}]\mid\alpha^{\prime c,n}\neq 0\}=\{n\in[N^{c}]\mid\alpha^{c,n}\neq 0\}=\mathcal{S}^{c} and similarly S′¬c=S¬c\mathcal{S}^{\prime\neg c}=\mathcal{S}^{\neg c}. Hence we have

By making the replacements from points 1, 2 and 3 in (5), we deduce that sκc(h)=sκ′′c(h′)s^{c}_{\kappa}(\boldsymbol{h})=s^{\prime c}_{\kappa^{\prime}}(\boldsymbol{h}^{\prime}). ∎

Appendix E Empirical Evaluation

This appendix provides useful details to reproduce the empirical evaluation from Section 3.

Computing Resources. All the empirical evaluations were run on a single machine equipped with a 18-Core Intel Core i9-10980XE CPU and a NVIDIA RTX A4000 GPU. The machine runs on Python 3.9 and Pytorch 1.10.2 .

Dataset licenses. The MNIST dataset dataset is made available under the terms of the Creative Commons Attribution-Share Alike 3.0 License. The ECG dataset dataset is made available under the terms of the Open Data Commons Attribution License v1.0. The CUB dataset is made available for non-commercial research and educational purposes.

Models. The detailed architecture of the models are provided in Tables 2, 3 and 4. The InceptionV3 architecture is the same as in the literature . We use its official Pytorch implementation.

Data Split. All the datasets are naturally split in training and testing data. In the ECG dataset, the different types of abnormal heartbeats are imbalanced (e.g. the fusion beats constitute only 0.7%0.7\% of the training set). Hence, we create a synthetic training set with balanced concepts using SMOTE . As in , the CUB dataset is augmented by using random crops and random horizontal flips.

Model Fitting. In fitting each model, we use the test set as a validation set since our purpose is not to obtain the models with the best generalization but simply models that perform well on a set of examples we wish to explain (here the examples from the test set). All the models are trained to minimize the cross-entropy between their prediction and the true labels. The hyperparameters are as follows. 1 For MNIST we use a Adam optimizer with batches of 120 examples, a learning rate of 10−310^{-3}, a weight decay of 10−510^{-5} for 5050 epochs with patience 1010. 2 For ECG we use a Adam optimizer with batches of 300 examples, a learning rate of 10−310^{-3}, a weight decay of 10−510^{-5} for 5050 epochs with patience 1010. 3 For CUB, we use a stochastic gradient descent optimizer with batches of 64 examples, a learning rate of 10−310^{-3}, a weight decay of 4⋅10−54\cdot 10^{-5} for 1,0001,000 epochs with patience 5050.

Concepts. The concept mapping between MNIST classes and concepts is provided in Table 5. For the ECG and the CUB datasets, the presence/absence of a concept for each example is readily available in the dataset.

Concept classifiers. All the concept classifiers are implemented with scikit-learn . For CAR classifiers, we fit a SVC with Gaussian RBF kernel and default hyperparameters from scikit-learn. For CAV classifiers, we fit a linear classifier with a stochastic gradient descent optimizer with learning rate 10−210^{-2} and a tolerance of 10−310^{-3} for 1,0001,000 epochs and the remaining default hyperparameters from scikit-learn.

Statistical significance. The statistical significance test from Section 3.1.1 is performed with the scikit-learn implementation of the permutation test. For MNIST and ECG, we consider 100 permutations per concept. For CUB, this test is more expensive since the latent spaces are high-dimensional. We consider only 25 permutations per concept in that case.

Concept-based feature importance. In the experiment from Section 2.3, we use Captum’s implementation of Integrated Gradients with default parameters. In the case of CUB, storing the feature importance scores for each concept and for the whole test set requires a prohibitive amount of memory. To avoid this problem, we select C=6C=6 concepts and subsample 50 positive and 50 negative examples per concept from the test set. This corresponds to a set of 600600 examples. We compute the feature importance for these examples only.

Alternative architecture. We extended the analysis of Section 3.1.1 to a ResNet-50 architecture. We fine-tune the ResNet model on the CUB dataset and reproduced the experiment from Section 3.1.1 with this new architecture. In particular, we fit a CAR and a CAV classifier on the penultimate layer of the ResNet. We then measure the accuracy averaged over the CUB concepts. This results in (89±1)%(89\pm 1)\% accuracy for CAR classifiers and (87±1)%(87\pm 1)\% accuracy for CAV classifiers. We deduce that CAR classifiers are highly accurate to identify concepts in the penultimate ResNet layer. As in the main paper, we observe that CAR classifiers outperform CAV classifiers, although the gap is smaller than for the Inception-V3 neural network. We deduce that our CAR formalism extends beyond the architectures explored in the paper and we hope that CAR will become widely used to interpret any more architectures.

Appendix F Use Case

This appendix provides useful details to reproduce the use case from Section 3.2.

Computing Resources. The use case was run on a single machine equipped with a 18-Core Intel Core i9-10980XE CPU and a NVIDIA RTX A4000 GPU. The machine runs on Python 3.9 and Pytorch 1.10.2 .

Dataset license. The SEER dataset is made available under the terms of the SEER Research Data Use Agreement.

Model. The detailed architecture of the model is provided in Table 6.

Data split. We randomly split the whole SEER dataset into a training set (90%90\% of the data) and a test set (the remaining 10%10\%). Since patients with a death outcome are in minority (less than 3%3\%), we oversample them to obtain a balanced training set.

Model fitting. In fitting the model, we use the test set as a validation set since our purpose is not to obtain the model with the best generalization but simply a model that performs well on a set of examples we wish to explain (here the examples from the test set). The model is trained to minimize the cross-entropy between its prediction and the true labels. We use a Adam optimizer with batches of 500 examples, a learning rate of 10−310^{-3}, a weight decay of 10−510^{-5} for 500500 epochs with patience 5050.

Concepts. The concepts correspond to prostate cancer grades. Those grades can be computed from the Gleason score as follows :

It goes without saying that the model is trained with the Gleason scores only, not with the grades.

Concept classifiers. We fit a SVC with linear kernel and default hyperparameters from scikit-learn.

Concept-based feature importance. We use Captum’s implementation of Integrated Gradients with default parameters.

Appendix G Explanation Robustness

We observe that the TCAR scores keep a high correlation with the true proportion of examples that exhibit the concept even when all the test examples are adversarially perturbed. We conclude that TCAR explanations are robust with respect to adversarial perturbations in this setting.

Appendix H Using CAR to Understand Unsupervised Concepts

Our CAR formalism adapts to a wide variety of neural network architectures. In this appendix, we use CAR to analyze the concepts discovered by a self explaining neural network (SENN) trained on the MNIST dataset. As in , we use a SENN of the form

Where gs(x)g_{s}(\boldsymbol{x}) and θs(x)\theta_{s}(\boldsymbol{x}) are respectively the activation and the relevance of the synthetic concept s∈[S]s\in[S] discovered by the SENN model. We follow the same training process as . This yields a set of S=5S=5 concepts explaining the predictions made by the SENN f:X→Yf:\mathcal{X}\rightarrow\mathcal{Y}.

When this correlation increases, the concepts ss and cc tend to be relevant together more often. We report the correlation between each pair (s,c)(s,c) in Table 8.

SENN Concept 2 correlates well with the Vertical Line Concept.

SENN Concept 3 correlates well with the Horizontal Line Concept

SENN Concept 5 correlates well with the Loop Concept.

SENN Concepts 1 and 4 are not well covered by our concepts.

The above analysis shows the potential of our CAR explanations to better understand the abstract concepts discovered by SENN models. We believe that the community would greatly benefit from the ability to perform similar analyses for other interpretable architectures, such as disentangled VAEs.

Appendix I Using CAR with NLP

CAR is a general framework and can be used in a wide variety of domains that involve neural networks. In the main paper, we show that CAR provides explanations for various modalities:

We assess the generalization performance of the CAR classifier on a holdout concept set made of Nc=30N^{c}=30 concept positive and negative sentences (60 sentences in total). The CAR classifier has an accuracy of 87%87\% on this holdout dataset. This suggests that the concept cc is smoothly encoded in the model’s representation space, which is consistent with the importance of positive adjectives to identify positive reviews. We deduce that our CAR formalism can be used in a NLP setting. We believe that using CARs to analyze large-scale language model would be an interesting study that we leave for future work.

Appendix J Increasing Explainability at Training Time

Improving neural networks explainability at training time constitutes a very interesting area of research but is beyond the scope of our paper. That said, we believe that our paper indeed contains insights that might be the seed of future developments in neural network training. As an illustration, we consider an important insight from our paper: the fact that the accuracy of concept classifiers seems to increase with the depth of the layer for which we fit a classifier. In the main paper, this is mainly reflected in Figure 4. This observation has a crucial consequence: it is not possible to reliably characterize the shallow layers in terms of the concepts we use.

In order to improve the explainability of those shallow layers, one could leverage the recent developments in contrastive learning. The purpose of this approach would be to separate the concept set Pc\mathcal{P}^{c} and Nc\mathcal{N}^{c} in the representation space H\mathcal{H} corresponding to a shallow layer of the neural network. A practical way to implement this would be to follow . Assume that we want to separate concept positives and negatives in the representation space H\mathcal{H} induced by the shallow feature extractor g:X→H\boldsymbol{g}:\mathcal{X}\rightarrow\mathcal{H}. As , one can use a projection head p:H→Z\textbf{p}:\mathcal{H}\rightarrow\mathcal{Z} and enforce the separation of the concept sets through the contrastive loss

Appendix K Further Examples

In this appendix, we provide several examples to illustrate the experiments from Section 3.

Concept classifiers. The accuracy of the concept classifiers for all the MNIST and ECG concepts are given in Figures 11 and 12. Each box-plot is built with 10 random seeds where the concept sets Pc\mathcal{P}^{c} and Nc\mathcal{N}^{c} are allowed to vary. We observe that the CAR classifiers are more accurate for each of the observed concepts

Global explanations. The global concept explanations for all the MNIST concepts and some of the CUB concepts are given in Figures 13 and 14. As the examples from Section 3.1.2, we see that TCAR explanations are more consistent with the human concept annotations.

Saliency maps. Examples of concept-based an vanilla saliency maps for MNIST, ECG and CUB examples are given in Figures 15 , 16 , 17 , 18 and 19. As explained in the main paper, we observe that concept-based saliency maps are indeed distinct form vanilla saliency maps. Furthermore, saliency maps for different concepts are not interchangeable. In the CUB case, we note that concept are not always identified with the minimal amount of features (e.g. some of the saliency maps from Figure 19 highlight pixels that do not always belong to the bird’s breast). This surprising observation seems to occur even for concept-bottleneck models that are explicitly trained to recognize concepts . We believe that this could be improved by training concept classifiers with images that include a segmentation highlighting the concept of interest (e.g. the bird’s breast). We leave this idea for future works.