CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model Behavior

Eldar David Abraham, Karel D'Oosterlinck, Amir Feder, Yair Ori Gat, Atticus Geiger, Christopher Potts, Roi Reichart, Zhengxuan Wu

Introduction

Explaining model behavior has emerged as a central goal within ML. In NLP, models have grown in size and complexity, and while they have become increasingly successful, they have also become more opaque [Lipton, 2018, Pearl, 2019], raising concerns about trust [Guidotti et al., 2018, Jacovi and Goldberg, 2020], safety [Amodei et al., 2016, Otte, 2013], and fairness [Goodman and Flaxman, 2017, Hardt et al., 2016]. These concerns will persist if these models remain “black-boxes”.

Seeking to open the black-box, researchers have developed methods that try to explain model behavior [Bastings et al., 2021, Feder et al., 2021b, Gehrmann et al., 2020, Lundberg and Lee, 2017, Ribeiro et al., 2016]. However, there is no consensus about how to evaluate such methods to allow robust comparisons. This is not surprising, since such evaluations require very rich empirical data. Intuitively, we would like to (1) intervene on model inputs, to modify specific concepts without changing other correlated information, (2) observe the effects this has on model predictions, and, finally, (3) assess explanation methods for their ability to accurately predict these effects.

The absence of interventional data, or even an agreed-upon non-interventional benchmark, has created an environment in which explanation methods are often evaluated individually, and without comparison to alternatives. Attempts have been made to conduct comparative evaluations [Feder et al., 2021b, Goyal et al., 2020, Pruthi et al., 2022], but only with synthetic, simplified datasets. Furthermore, these attempts do not define a unified evaluation approach, nor do they seek to contribute benchmark datasets that support such evaluations.

In this paper, we seek to overcome this obstacle by introducing CEBaB (Causal Estimation-Based Benchmark). Table 1 summarizes the structure of CEBaB with a toy example: beginning with a review text from the OpenTable website, we crowdsourced edits of the original text that are designed to meet a specific goal, such as changing the food rating in the original text to negative or unknown. All of the resulting edits were validated by five crowdworkers and each full text was evaluated by five crowdworkers for its overall sentiment. CEBaB is grounded in 2,299 original reviews, which were expanded via this editing procedure to a total of 15,089 texts, targeting four different aspect-level concepts (food, service, ambiance, noise) with three potential labels (positive, negative, and unknown, i.e., not expressed in the review), and each full text was labeled on a five-star scale.

We focus on using CEBaB to compare concept-based explanation methods. This allows us to go beyond the effect of individual tokens to study how more abstract concepts (in our case, aspect-level sentiment) contribute to model predictions (about the overall sentiment of the text). Our proposed metrics center around assessing concept-based explanation methods for their ability to accurately estimate causal concept effects [Goyal et al., 2020], allowing us to isolate the effect of individual concepts.

More specifically, we use CEBaB to measure the causal effects of particular variables in a causal graph, and we cast each explanation method as a causal estimator of these measurements. For example, suppose our causal graph of the data says that all four of our aspect-level categories will affect a reviewer’s overall rating. To estimate the effect of positive food quality on the predicted overall rating from a classifier, we need to compare examples with high food quality to those with low quality, holding all other aspects constant. Such pairs of examples are normally not observed, but this is precisely what CEBaB provides. With CEBaB, we can directly compare the actual change in model predictions with the change that a concept-based explanation method predicts.

In our experiments, we evaluate five leading concept-based explanation methods: CONEXP [Goyal et al., 2020], TCAV [Kim et al., 2018], ConceptSHAP [Yeh et al., 2020], INLP [Ravfogel et al., 2020], CausaLM [Feder et al., 2021b], and S-Learner [Künzel et al., 2019]. These methods make a wide range of different assumptions about how much access we have to the model’s internal structure, and they also diverge in the degree to which they account for the causal nature of the concept effect estimation problem. Remarkably, CEBaB reveals that most methods cannot beat a simple baseline. Indeed, this negative result emphasizes the value in our primary contribution of providing the data and metrics that enables a direct comparison of explanation methods.

Previous Work

Benchmark datasets have propelled ML forward by creating shared metrics that predictive models can be evaluated on [Hu et al., 2020, Kiela et al., 2021, Wang et al., 2018, 2019]. Unfortunately, benchmarks that are suitable for assessing the quality of model explanations are still uncommon [Feder et al., 2021a, Hooker et al., 2019]. Previous work on comparing explanation methods has generally only correlated the performance of a given explainability method with others, without ground-truth comparisons [DeYoung et al., 2020, Hase and Bansal, 2020, Hooker et al., 2019, Samek et al., 2021].

Other works that do compare to some ground-truth either employ a non-causal evaluation scheme [Kim et al., 2018], use causal evaluation metrics which do not capture performance on individual examples [Tenney et al., 2020], evaluate on synthetic counterfactuals and rule-based augmentations [Feder et al., 2021b, Tenney et al., 2020], or are tailored for a specific explanation method and hard to generalize [Yeh et al., 2020]. To the best of our knowledge, CEBaB is the first large-scale naturalistic causal benchmark with interventional data for NLP.

Explanation Methods and Causality

Probing is a relatively new technique for understanding what model internal representations encode. In probing, a small supervised [Conneau et al., 2018, Tenney et al., 2019] or unsupervised [Clark et al., 2019, Manning et al., 2020, Saphra and Lopez, 2019] model is used to estimate whether specific concepts are encoded at specific places in a network. While probes have helped illuminate what models (especially pretrained ones) have learned from data, Geiger et al. show with simple analytic examples that probes cannot reliably provide causal explanations for model behavior.

Feature importance methods can also be seen as explanation methods [Molnar, 2020]. Many methods in this space are restricted to input features, but gradient-based methods can often quantify the relative importance of hidden states as well [Binder et al., 2016, Shrikumar et al., 2017, Springenberg et al., 2014, Zeiler and Fergus, 2014]. The Integrated Gradients method of Sundararajan et al. has a natural causal interpretation stemming from its exploration of baseline (counterfactual) inputs [Geiger et al., 2021]. However, even where these methods can focus on internal states, it remains difficult to connect their analyses with real-world concepts that do not reduce to simple properties of inputs.

Intervention-based methods involve modifying inputs or internal representations and studying the effects that this has on model behavior [Lundberg and Lee, 2017, Ribeiro et al., 2016]. Recent methods perturb input or hidden representations to create counterfactual states that can then be used to estimate causal effects [Elazar et al., 2021, Finlayson et al., 2021, Soulos et al., 2020, Vig et al., 2020, Geiger et al., 2021]. However, these methods are prone to generating implausible inputs or network states unless the interventions are carefully controlled [Geiger et al., 2020].

Generating counterfactual texts automatically remains challenging and is still a work-in-progress [Calderon et al., 2022]. To overcome this problem, another class of approaches proposes to manipulate the representation of the text with respect to some concept, rather than the text itself [Elazar et al., 2021, Feder et al., 2021b, Ravfogel et al., 2020]. These methods fall into the category of concept-based explanations and we discuss two of them extensively in §3.

Estimating Concept Effects with CEBaB

We now define the core metrics that we use to evaluate different explanation methods. Figure 1 provides a high-level view of the causal process we are envisioning. The process begins with an exogenous variable UU representing a state of the world. For CEBaB, we can imagine that the value of UU is a state of affairs uu of a person evaluating a restaurant in a particular way. uu contributes to a review variable XX, with the value xx of XX mediated by uu and by mediating concepts C1,…CkC_{1},\ldots C_{k}, which correspond to the four aspect-level categories in CEBaB (food, service, ambiance, and noise), each of which can have values c∈{positive,negative,unknown}c\in\{\text{positive},\text{negative},\text{unknown}\}. The review xx is processed by a model that outputs a vector of scores over classes (sentiment labels in CEBaB).

Our central goal is to use CEBaB to evaluate explanation methods themselves. CEBaB supports many approaches to such evaluation. In this paper, we adopt an approach based on individual-level rather than average effects. This makes very rich use of the counterfactual text and associated labels provided by CEBaB. The starting point for this metric is the Individual Causal Concept Effect:

For a neural network N\mathcal{N} and feature function ϕ\phi, the individual causal concept effect of changing the value of concept CC from cc to c′c^{\prime} for state of affairs uu in an underlying data generation process G\mathcal{G} is

ICaCE is a theoretical quantity. In practice, we use the Empirical Individual Causal Concept Effect.

For a neural network N\mathcal{N} and feature function ϕ\phi, the empirical individual causal concept effect of changing the value of concept CC from cc to c′c^{\prime} for state of affairs uu is

where (xuC=c,xuC=c′)(x^{C=c}_{u},x^{C=c^{\prime}}_{u}) is a tuple of inputs originating from uu with the concept CC set to the values cc and c′c^{\prime}, respectively.

The ICaCE^Nϕ\widehat{\text{ICaCE}}_{\mathcal{N}_{\phi}} for a pair of examples (xuC=c,xuC=c′)(x^{C=c}_{u},x^{C=c^{\prime}}_{u}) is simply the difference between the output score vectors for the two cases. With CEBaB, we can easily calculate these values because we have clusters of examples that are tied to the same reviewing situation uu and express different concept values.

For assessing an explanation method E\mathcal{E}, we compare ICaCE values with those returned by E\mathcal{E}. Our core metric is the ICaCE-Error:

For a neural network N\mathcal{N}, feature function ϕ\phi and distance metric Dist\mathsf{Dist}, the ICaCE-Error of an explanation method E\mathcal{E} for changing the value of concept CC from cc to c′c^{\prime} is:

We present results for three choices of Dist\mathsf{Dist} which vary in their ability to model the direction and magnitude of effects. These choices give subtly different but largely converging results, as detailed in Section 6 and reported more fully in Appendix D.

Aggregating Individual Causal Concept Effect

It is often useful to also have a direct estimate of a model’s ability to capture concept-level causal effects. For this, we employ an aggregating version of ICaCE^\widehat{\text{ICaCE}}, the Empirical Causal Concept Effect:

For a neural network N\mathcal{N} and feature function ϕ\phi, the empirical causal concept effect of changing the value of concept CC from cc to c′c^{\prime} in dataset D\mathcal{D} is

This is an empirical estimator of the Causal Concept Effect (CaCE) of Goyal et al. . It estimates, in general, how the classifier predictions change for a given concept and intervention direction.

Estimating Real-World Causal Effect of Aspect Sentiment on Overall Sentiment

We can also estimate ground truth causal effects in CEBaB by simply using its labels directly. There are again a variety of ways that this could be done. We opt for the one that makes the richest use of the structures afforded by CEBaB. For perspicuity, in parallel to the neural network-based ICaCE^\widehat{\text{ICaCE}} (Definition 2), we define the Empirical Individual Treatment Effect for our dataset:

The empirical individual treatment effect of changing the value of concept CC from cc to c′c^{\prime} in CEBaB is

where ff is a simple look-up procedure that retrieves the overall sentiment labels for CEBaB examples.

We aggregate over these values by taking their average, in parallel to what we do for network predictions (Definition 4). This yields the Empirical Average Treatment Effect (ATE^\widehat{\text{ATE}}) for CEBaB.

Alternative Metrics

In Appendix A in our supplementary materials, we consider alternative formulations of the core metrics with causal concept effects and absolute causal concept effects, relating them to the different questions they engage with. We opt for the individual causal concept effect in our central metric (Definition 3), taking the central question to be what caused an ML model to produce an output for an actual input created from a real-world process.

Evaluated Explanation Methods

We compare several model explanation methods that share three main characteristics. First, they are all suitable for NLP models and have been used in the literature for generating model explanations in the form of estimated effects on model predictions. Second, they all provide concept-level explanations, for a pre-defined list of human-interpretable concepts (e.g., how sensitive a restaurant review rating classifier is to language related to food quality). This approach is also forward-looking, allowing more researchers to construct new hypotheses (i.e., concepts we have not collected labels for) and estimate their effect on the predictor. Third, all of the tested methods are model-agnostic, meaning that they separate the explanation from the model. At the same time, these methods differ in five important ways, as summarized Table 2.

We now turn to reviewing the explanation methods that we later compare on CEBaB (§6). In our mathematical formulas, we employ a unified notation for all methods, to make the definitions more accessible and easier to integrate into our experimental set-up. Assume we have a classifier N\mathcal{N} (which outputs a probability vector) and feature function ϕ\phi, and we want to compute the effect on Nϕ(xuC=c)\mathcal{N}_{\phi}(x^{C=c}_{u}) of changing the value of concept CC from cc to c′c^{\prime} using an unseen test set (D,Y)(\mathcal{D},Y).

The gold labels of CEBaB are the difference between the logits for some original review xuC=c{x}_{u}^{C=c} and ground-truth counterfactual xuC=c′{x}_{u}^{C=c^{\prime}}. As a baseline, we sample an original review xu′C=c′{x}_{u^{\prime}}^{C=c^{\prime}} with the same aspect-labels as the xuC=c′{x}_{u}^{C=c^{\prime}} and use it as an approximate counterfactual:

We do this sampling using predicted aspect labels from the aspect-level sentiment analysis models described in Appendix C.

Conditional Expectation (CONEXP)

Goyal et al. propose a baseline where the effect of a concept CC is the average difference in predictions on examples with different values of CC.

where DC=c\mathcal{D}^{C=c} and DC=c′\mathcal{D}^{C=c^{\prime}} are subsets of D\mathcal{D} where CC takes values cc and c′c^{\prime}, respectively. To predict an effect, this method only relies on CC, cc, and c′c^{\prime}, resulting in an estimate that does not depend on the specific input text itself.

Conditional Expectation Learner (S-Learner)

We adapt S-Learner, a popular method for estimating the Conditional Average Treatment Effect (CATE) [Künzel et al., 2019]. To estimate causal concept effects, our S-Learner trains a logistic regression model E\mathcal{E} to predict N(ϕ(x))\mathcal{N}(\phi(x)) using the values of all the labeled concepts of example xx, denoted by x′x^{\prime}.This training approach, where an explainer model is fit to predict the output of the original model, shares the intuition of LIME, the widely used explanation method Ribeiro et al. , but for concept-level effects. Then, during inference, we compute an individual effect for example pair (xuC=c,xuC=c′)({x}_{u}^{C=c},{x}_{u}^{C=c^{\prime}}) by comparing the output of the model Ex\mathcal{E}_{x} on this pair:

At inference time, S-Learner assumes access to all aspect-level labels x′x^{\prime}, which might not always be available. To alleviate this issue, we instead predict the aspect-level labels x′x^{\prime} from the original text xx using models described in Appendix C.

TCAV

Kim et al. use Concept Activation Vectors (CAVs), which are semantically meaningful directions in the embedding space of ϕ\phi. Our adapted version of Testing with CAVs (TCAV) outputs a vector measuring the sensitivity of each output class kk to changes towards the direction of a concept vCv_{C} at the point of the embedded input. It is computed as:

where KK is the number of classes and vCv_{C} is a linear separator learned to separate concept CC in the embedding space of ϕ\phi.

ConceptSHAP

Yeh et al. propose this expansion to SHAP [Lundberg and Lee, 2017], to generate concept-based explanation based on Shapley values [Shapley, 1953]. Given a complete (i.e., such that the accuracy it achieves on a test set is higher than some threshold β\beta) set of mm concepts {C1,…,Cm}\{C_{1},\ldots,C_{m}\}, ConceptSHAP calculates the contribution of each concept to the final prediction. Our adapted version outputs a vector for each C∈{C1,…,Cm}C\in\{C_{1},\ldots,C_{m}\} and xx. We justify this modification and provide implementation details in Appendix H.

CausaLM

Feder et al. [2021b] estimate the causal effect of a binary concept CC on the model’s predictions by adding auxiliary adversarial tasks to the language representation model in order to learn a counterfactual representation ϕCCF(x)\phi^{\text{CF}}_{C}(x), while keeping essential information about potential confounders (control concepts). Their method outputs the text representation-based individual treatment effect (TReITE), which is computed as:

where ϕCCF\phi^{\text{CF}}_{C} denotes the learned counterfactual representation, where the information about concept CC is not present, and N′\mathcal{N}^{\prime} is a classifier trained on this counterfactual representation. A key feature of CausaLM is its ability to control for confounding concepts (if modeled).As in Feder et al. [2021b], we control for the most correlated potential confounder. An inherent drawback of this technique is that it can only estimate interventions well for c′=Unknownc^{\prime}=\text{Unknown}, since the counterfactual representation is only trained to remove a concept CC.

Iterative Nullspace Projection (INLP)

Ravfogel et al. remove a concept from a representation vector by repeatedly training linear classifiers that aim to predict that attribute from the representations and projecting the learned representations on their null-space. Similar to CausaLM, INLP also estimates the TReATE (Equation 10) and can only estimate interventions for c′=Unknownc^{\prime}=\text{Unknown}.

The CEBaB Dataset

Table 1 provides an intuitive overview of the structure of CEBaB. In the editing phase of dataset creation, crowdworkers modified an existing OpenTable review in an effort to achieve a specific aspect-level goal while holding all other properties of the original text constant. Our aspect-level categories are food, ambiance, service, and noise. In the validation phrase, crowdworkers labeled each example relative to each aspect as ‘Positive’, ‘Negative’, or ‘Can’t tell’ (Unknown). Having five labels per example allows us to infer a majority label or reason in terms of the full label distributions. In the rating phase, each full text was labeled using a common five-star scale, again by five crowdworkers.

We began with 2,299 original reviews from OpenTable (related to 1,084 restaurants) and expanded them, via the above editing procedure, into a total of 15,089 texts. The distribution of normalized edit distances has peaks around 0.28 and 0.77, showing that workers made non-trivial changes to the originals, and even often had to make substantial changes to achieve the editing goal. (See Appendix B for the full distribution.)

Table 3 summarizes the resulting label distributions, where an example has label yy if at least 3 of the 5 labelers chose yy, otherwise it is in the ‘no majority’ category. 99% of aspect-level edits have a majority label that corresponds to the editing goal, and 88% of the texts have a review-level majority label on the five-star scale. Overall, these percentages show that workers were extremely successful in achieving their editing goals and that edits have systematic effects on overall sentiment.

The central goal of CEBaB is to create edit pairs: pairs of examples that come from the same original text and differ only in their labels for a particular aspect. For example, in Table 1, the first two ‘food edit’ cases form an edit pair, since they come from the same original text and differ only in their food label. Original texts can also contribute to edit pairs; the original text in Table 1 forms an edit pair with each of the texts it is related to by edits. Table 3(c) summarizes the distribution of edit pairs, and Table 3(d) reports the ground-truth ATE^\widehat{\text{ATE}} values (§3).

We release the dataset with fixed train/dev/test splits. In creating these splits, we enforce two high-level constraints. The first is our ‘grouped’ requirement: for each original review tt, all texts that are related to tt via editing occur in the same split as tt. This ensures that models are not evaluated on examples that are related by editing to those they have seen in training. Second, if any text tt in a group received a ‘no majority’ label, then the entire group containing tt is put in the train set. This ensures that there is no ambiguity about how to evaluate models on dev and test examples.

Once these high-level conditions were imposed, the examples were sampled randomly to create the splits. This allows that individual workers can contribute edited texts across splits. This minor compromise was necessary to ensure that we could have large dev and test splits. Appendix C in our supplementary materials shows that worker identity has negligible predictive power.

There are two versions of the train set: inclusive and exclusive. The inclusive train set contains all original and edited non-dev/test texts (11,728 texts). The exclusive version samples exactly one train text from each set of texts that are related by editing (1,755 examples). The rationale is that models trained with an original review as well as its edited counterparts may explicitly learn causal effects trivially by aggregating learning signals across inputs. Our exclusive train split prevents this, which helps facilitate fair comparisons between explanation methods and better resembles a real-world setting.

Our dataset is released publicly in JSON format and is available in the Hugging Face datasets library. It includes restaurant metadata, full rating distributions, and anonymized worker ids. Appendix B in our supplementary materials provides additional details on the dataset construction, including the prompts used by the crowdworkers, the number of workers per task, worker compensation, and a sample of examples with ratings to help convey the nature of workers’ edits and the overall quality of the resulting texts and labels. In addition, Appendix C reports on a wide range of classifier experiments at the aspect-level and text-level that show that models perform well on CEBaB classification tasks, which bolsters the claim that CEBaB is a reliable tool for assessing explanation methods.

Experiments and Results

For each experiment, we fine-tune a pretrained language model to predict the overall sentiment of all restaurant reviews from our exclusive OpenTable train set. Since the goal of our work is not to achieve state-of-the-art performance, but rather to compare explanation methods and demonstrate the usage of CEBaB, we test the ability of methods to explain commonly used models, trained with standard experimental configurations.

In the main text, we report results for bert-base-uncased fine-tuned as a five-way classifier. Appendix D includes results for GPT-2, RoBERTa, and an LSTM, fine-tuned on binary, 3-way and 5-way versions of the sentiment task. All results, including the ground-truth effect that depends on the specific instance of a model, are averaged across 5 seeds.

To evaluate the intrinsic capacity of a model to capture causal effects, we report the CaCE^\widehat{\text{CaCE}} values, as in Definition 4. The results for bert-base-uncased are given in Table 4. They are intuitive and well-aligned with the ATE^\widehat{\text{ATE}} estimates in Table 3(d), indicating that the model has captured the real-world effects.

Our primary assessment of the evaluation methods is given in Figure 2, again focusing on a five-way bert-base-uncased model as representative of our results. We provide values based on cosine, L2, and normdiff as the value of Dist in Definition 3. The cosine-distance metric measures if the estimated and observed effect have the same direction but does not take the magnitudes of the effects into account. The L2-distance measures the Euclidian norm of the difference of the observed and estimated effect. Both the direction and magnitude of the effects influence this metric. To only compare the magnitudes, we use the normdiff-distance, which computes the absolute difference between the Euclidean norms of the observed and estimated effects, thus completely ignoring the directions of both effects.

Remarkably, our approximate counterfactual baseline proves to be the best method at capturing both the direction and magnitude of the effects. The fact that a simple baseline method beats almost all other methods indicates that we need better explanation methods if we are going to capture even relatively simple causal effects like those given by CEBaB.

Recall from Table 2 that the compared methods require different levels of access to concept labels at inference time. Approximate counterfactuals and S-Learner have access to both the direction of the intervention and the predicted test-time aspect labels, enabling them to outperform CONEXP, which has access to only the direction of the intervention, and TCAV, ConceptSHAP, and CausaLM, which have access to neither the intervention direction nor test-time aspect labels.

The INLP method ties with the best method for the cosine metric, despite having access to neither intervention directions nor test-time aspect labels. Perhaps this method could be extended to make use of this additional information and decisively improve upon our approximate counterfactual baseline.

While CausaLM and INLP both estimate the effect of removing a concept from an input, INLP uses linear probes to guide interventions on the original model, while CausaLM trains an entirely new model with an auxiliary adversarial objective. The direct use of the original model is something INLP shares with the approximate counterfactual baseline; it seems that a tight connection to the original model may underlie success on CEBaB.

Conclusion

Our main contributions in this paper are twofold. First, we introduced CEBaB, the first benchmark dataset to support comparing different explanation methods against a single ground-truth with human-created counterfactual texts and multiply-validated concept labels for aspect-level and overall sentiment. Using this resource, one can isolate the true causal concept effect of aspect-level sentiment on any trained overall sentiment classifier. CEBaB provides a level playing field on which we can compare a variety of explanation methods that differ in their assumptions about their access to the model, their computational demands, their access to ground-truth concept labels at inference time, and their overall conception of the explanation problem. Furthermore, the evaluated methods make absolutely no use of CEBaB’s counterfactual train set. In turn, we hope that CEBaB will facilitate the development of explanation methods that can take advantage of the very rich counterfactual structure CEBaB provides across all its splits.

Second, we have provided an in-depth experimental analysis of how well multiple model explanation methods are able to capture the true concept effect. A naive baseline that approximates counterfactuals through sampling achieves the best performance, with INLP and S-Learner being the only other methods that achieves state-of-the art on any metric. While CEBaB is only grounded in one task, sentiment analysis alone is enough to produce starkly negative results that should serve as a call to action for NLP researchers aiming to explain their models.

Acknowledgments and Disclosure of Funding

This research is supported in part by a grant from Meta AI. Karel D’Oosterlinck was supported through a doctoral fellowship from the Special Research Fund (BOF) of Ghent University. We thank our crowdworkers for their invaluable contributions to CEBaB.

References

Supplementary Materials

Appendix A Causal Concept Effects and Metrics for Explanation Methods

Data do not materialize out of thin air. Rather, data are generated from real-world processes with complex causal structures we do not observe directly. Causal inference is the task of estimating theoretical causal effect quantities.

When estimating causal effects, researchers commonly measure the average treatment effect, which is the difference in mean outcomes between the treatment and control groups [Rubin, 1974]. Formally, we define the average treatment effect of binary treatment TT on an outcome YY under a data generation process G\mathcal{G} that represents the unknown details of the real-world.

The ATE is a theoretical quantity we cannot compute in practice, since we do not have access to G\mathcal{G} nor can we observe both interventions for the same subject.

However, we are concerned with estimating the causal effect of variables representing non-binary concepts in real-world systems, on data in an appropriate format for processing by a modern AI model that predicts vector encoding probability distributions over outputs.

Let N\mathcal{N} be a neural network outputting a probability vector, where its kk-th entry represents the probability to predict the kk-th class, and let ϕ\phi be a feature representation (e.g., BERT embedding). In the context of model explanations, we will define the tools needed to answer three questions:

Given a real-world circumstance uu that led to input data xuC=cx_{u}^{C=c}, what is expected effect of a concept CC changing from value cc to value c′c^{\prime} on the model output of Nϕ\mathcal{N}_{\phi} provided input data xuC=cx_{u}^{C=c}?

What is the expected effect of a concept CC changing from value cc to value c′c^{\prime} on the output of the model Nϕ\mathcal{N}_{\phi} provided input data XX across real-world circumstances UU?

What is the magnitude of the expected effect of a changing the concept CC on the output of the model Nϕ\mathcal{N}_{\phi} provided input data XX across real-world settings UU?

For example, in the context of CEBaB, we might ask

Given a real-world dining experience uu with good food quality (Cfood=+C_{\text{food}}=+) that led to a restaurant review xuCfood=+x^{C_{\text{food}}=+}_{u}, what is the effect of changing the food quality CfoodC_{\text{food}} from Cfood=+C_{\text{food}}=+ to Cfood=−C_{\text{food}}=- on the output of an overall-sentiment text classifier Nϕ\mathcal{N}_{\phi} provided a review of the dining experience?

What is the expected effect of changing the food quality CfoodC_{\text{food}} from positive ++ to negative −- on the output of the model Nϕ\mathcal{N}_{\phi} across real-world dining experiences that lead to restaurant reviews?

What is the magnitude of the expected effect of a changing food quality CfoodC_{\text{food}} on the output of the model Nϕ\mathcal{N}_{\phi} across real-world dining experiences that lead to restaurant reviews?

Each of the above questions requires the estimation of a different theoretical quantity. In respect to the order of the questions, these quantities are the individual causal concept effect, the causal concept effect, and the absolute causal concept effect.

We believe the most practical question in explainable AI is: why does this model have this output behavior for an actual input. For this reason, our focus in the main text is individual causal concept effects. We define our central metric that captures the performance of an explainer on CEBaB as the average error on individual causal effect predictions (Definition 3).

We do not evaluate the ability of explainers to evaluate the causal concept effect or the absolute causal concept effect.

For an exogenous setting uu that led to concept CC taking on value cc and the creation of input data xuC=cx_{u}^{C=c}, the individual causal concept effect of a concept CC changing from value cc to c′c^{\prime} in a data generation process G\mathcal{G} on a neural network N\mathcal{N} with feature representation ϕ\phi is

The causal concept effect is the effect in general, meaning there is no input data generated from a fixed exogenous real-world setting:

The absolute causal concept effect estimate of the magnitude of the effect a concept has on a classifier output, regardless the concept values. We aggregate over all possible intervention values in the following way

where C{C} is the set of all possible values for concept in addition to denoting the concept itself.We take the absolute value since CaCENϕ(G,C,c,c′)=−CaCENϕ(G,C,c′,c)\text{CaCE}_{\mathcal{N}_{\phi}}(\mathcal{G},C,c,c^{\prime})=-\text{CaCE}_{\mathcal{N}_{\phi}}(\mathcal{G},C,c^{\prime},c), and these cancel each other in the summation.

A.2 Empirical Estimates

Similar to the ATE, causal concept effects are theoretical quantities we can only estimate in reality. To perform such estimates, we need a dataset consisting of pairs (xuc,xuc′)∈D(x^{c}_{u},x^{c^{\prime}}_{u})\in\mathcal{D} that are drawn from a data generation process G\mathcal{G}. A major contribution of this work is crowdsourcing such a dataset, CEBaB. These pairs allow us to compute empirical estimations of (individual) causal concept effects.

For an exogenous setting uu, the empirical individual causal concept effect of a concept CC changed from value cc to c′c^{\prime}, for D\mathcal{D} sampled from G\mathcal{G}, on a neural network N\mathcal{N} trained on a feature representation ϕ\phi is

Given a full dataset D\mathcal{D} of such pairs, we can estimate the causal concept effect

And also the absolute causal concept effect

Notice that the only difference between causal concept effects (Definition 7) and empirical causal concept effects (Definition 8) is that we change the expectation taken over G\mathcal{G} to be the average over a dataset D∼G\mathcal{D}\sim\mathcal{G}.

A.3 Explainer Errors

Given a dataset D\mathcal{D} and an explainer ENϕ(xuc,c′)\mathcal{E}_{\mathcal{N}_{\phi}}(x^{c}_{u},c^{\prime}) that predicts individual causal concept effects ICACENϕ(xuc,c′)\textit{ICACE}_{\mathcal{N}_{\phi}}(x^{c}_{u},c^{\prime}), we define metrics capturing the ability of E\mathcal{E} to estimate causal effects by simple computing the averaged distance between our explainer and the empirical causal effect

The average distance between the explainer and the empirical individual causal concept effects.

The distance between the average of explainer outputs and the empirical causal concept effect

The distance between the average magnitude of explainer outputs and the empirical absolute causal effect

where ∥⋅∥\|\cdot\| is some distance metric and DC\mathcal{D}_{C} is the subset of data where CC is the concept changed and DCc→c′\mathcal{D}_{C}^{c\to c^{\prime}} is the subset of data where CC is the concept changed from value cc to value c′c^{\prime}.

In the main text, we use the ICaCE-Error as our primary evaluation metric.

Appendix B CEBaB

Our supplementary materials contain a full Datasheet for CEBaB as a separate markdown document.

Table 5 gives an overview of the metadata associated with the original review texts in CEBaB.

B.2 Crowdworkers

A total of 254 workers participated in our experiments. All of them come from a pool of workers whom we prequalified to participate in our tasks based on the work they did for us on previous crowdsourcing projects. Thus, we expected that they would do high quality work, and they more than lived up to our expectations, as indicated by the high degree of success they achieved when editing and the high degree of consensus they reached about how to label examples.

There are a total of 642 instances of 15,0006 for which, despite our best efforts, a worker validated an example that they themselves created during the editing phase. Removing the contributions of these workers affects the majority in only 24 cases, with no clear pattern to the changes, so we kept all the validation labels in order to ensure that every example has give responses.

B.3 Editing Phase

A total of 183 workers participated in this phase. Workers were paid US$0.25 per example. Figure 3 shows the annotation interface that workers used when changing the target aspect’s sentiment to either ‘Positive’ or ‘Negative’, and Figure 4 shows the interface where the task was to hide the target aspect’s sentiment.

Figure 5 summarizes the distribution of edit distances between original and edited texts. These distances are calculated at the character-level and normalized by the length of the original or review, whichever is longer.

B.4 Validation Phase

A total of 174 workers participated in this phase. Workers were paid US$0.35 per batch of 10 examples. Figure 6 shows the annotation interface that workers used.

B.5 Review-level Rating Phase

A total of 155 workers participated in this phase. Workers were paid US$0.35 per batch of 10 examples. Figure 7 shows the annotation interface that workers used.

B.6 Randomly Selected Examples

Table 6 provides a random sample of edit pairs from CEBaB’s dev set.

B.7 Five-way Empirical ATE for CEBaB

Table 7 provides the binary ATE^\widehat{\text{ATE}} values for CEBaB. These can be compared with the corresponding five-way values in Table 3(d) in the main text.

B.8 Edit variability

In the editing phase we ask human annotators to produce edits of an original review with regard to some concept. This is inherently a noisy process, which may impact the quality of our final benchmark. The CEBaB dataset features a modest set of paired edits (176 pairs in total). Each of these pairs contains two edits, starting from the same original sentence and edit goal, which results in two different edited sentences. Like all sentences in CEBaB, these edits were labeled for their review score by human annotators.

Figure 8(a) shows the distribution of the difference in final review majorities produces by these paired edits. Most paired edits differ at most by one star in their final majority rating, indicating that in general there is some noise associated with the editing procedure, but this does not have a major impact on the final review score. Figure 8(b) shows the same distribution when we consider the average review score an edit received, as opposed to the majority score. If we consider these average scores, most of the paired edits differ only slightly in their resulting review score.

Figures 9a-c shows the distribution of this pairwise review score in more detail. In an idealized setting without variability, the distribution would be centered around the diagonal of the heatmap. When going from 5-way classification to ternary and binary classification, the variability introduced by the edits becomes less relevant with regard to the final review majority label.

Appendix C CEBaB Modeling Experiments

This section reports on standard classifier-based experiments with CEBaB, aimed at providing a sense for the dataset when it is used as a standard supervised sentiment dataset. We report experiments on the aspect-level and review-level ratings. In addition, we present evidence that author identity does not have predictive value.

We rely on the Hugging Face transformers library.https://github.com/huggingface/transformers [Wolf et al., 2019] We train our models with 4 Nvidia 2080 Ti RTX 11GB GPUs on a single node machine. We use a maximum sequence length of 128 with a fix batch size of 32 with a initial learning rate of 2e−52e^{-5}. We run each experiment 5 times with distinct random seeds. We train our models with a minimum epoch number of 5 with our largest training set. We linearly scale our training epoch number by the size of the training set. We skip hyperparameter tuning for optimized task performance as our goal for this paper is to evaluate explanation methods. We release all of our models on Huggingface Dataset Hub.

C.2 Models

We include 4 different types of models, including BERT (bert-base-uncased) [Devlin et al., 2019], RoBERTa (roberta-base) [Liu et al., 2019], GPT-2 (gpt2) [Radford et al., 2019], as well as LSTM with dot-attention [Luong et al., 2015]. Our LSTM model uses bert-base-uncased tokenizer for simplicity. We initialize the embeddings of tokens for our LSTM using fastText [Joulin et al., 2016]. We reconfigure the classification head all other models the same classification head as in RoBERTa as a non-linear multilayer perceptron (MLP).We implemented T5 (t5-base; [Raffel et al., 2019]) as a text-to-text model with the goal of treating predicted tokens as class labels. However, this raised unanticipated implementation questions concerning how to post-process multi-token class labels (e.g., “very positive”) for use in our explainer methods. As a result, we have elected to leave the T5 results out of the current draft, but we intend to include them in the next version once they have been more thoroughly vetted.

C.3 Multi-class Sentiment Analysis Benchmark

We report model performance results under 3 training conditions: Binary Classification, where we label reviews with 1 star and 2 star ratings as negative, reviews with 4 star and 5 star as positive, and 3-star reviews are dropped; Ternary Classification, where we add another neutral class for reviews with 3 star ratings; and 5-way Classification, where each star rating by itself is considered as a class. We leave out reviews in the train set in the ‘no majority’ category. (Dev and Test do not contain any such examples.) Table 8 shows the performance results for our models under different conditions. Our results suggest that RoBERTa has the edge over others across all evaluated tasks.

C.4 Aspect-based Sentiment Analysis Benchmark

Our dataset can be naturally used as an aspect-based sentiment analysis (ABSA) benchmark. For each sentence, it may contain up to 4 aspects with respect to the reviewing restaurant. As ABSA benchmarks are usually small and sparse with missing labels, our dataset provides validated aspect-based labels, and is one of the largest human validated ABSA benchmark.

To evaluate model performance, we adapt standard finetuning approach for ABSA benchmarks as proposed by Sun et al. . Instead of single sentence classification, we add another auxiliary sentence representing the aspect. For instance, to predict the label for the ‘food’ aspect for “the food here is good but not the service”, we append a single aspect token with a separator, and construct our input sentence as “the food here is good but not the service [SEP] food”. Table 8 shows the performance results for our models under different conditions.

C.5 Author Identity Prediction

One potential artifact of our benchmark is edited sentence may expose author identity, which may result in artifact in interpreting model performance. To quantify this potential artifact, we train models to predict author identities based on the sentences. We create author identity prediction dataset by aggregating our dataset by anonymized worker ids. We then split the dataset into train/dev with a 4-to-1 ratio. For model training, we finetune RoBERTa for 5 epochs with a batch size of 32, a learning rate of 2e−52e^{-5}, and a maximum sequence length of 128. Note that we only consider top-k annotators ranked by their contributions (i.e., number of examples in our dataset). Table 9 shows the performance results of our finetuned models with a random classifier. Our results suggest that potential artifacts may exist but only for a limited extend.

Appendix D Additional Results

In this section, we report additional results for bert-base-uncased, roberta-base, gpt-2, and an LSTM, fine-tuned on binary, ternary and 5-way versions of the sentiment task. These models are described in Appendix C. Table 10 summarizes all the results.

We refer to the results section in the main text for an explanation of the different metrics considered. Which metric is best depends on the final use-case and whether it is more important to estimate the direction or the magnitude of the effect.

Figure 10 shows the results for the ICaCE-Error with the cosine distance metric. The explanation methods that take the direction of the intervention into account (Approx, CONEXP, S-Learner) are the clear winners across all different models considered. S-Learner marginally wins across the most settings, but the conceptually simple Approx baseline is a close second. The strong performance of this simple baseline across the board suggests that most methods perform subpar, and that there is potential value in developing better concept-based model explanation methods.

Both TCAV and ConceptSHAP struggle to achieve better-than-random performance across all settings. Further analysis is needed to exactly understand why these methods are struggling.

Some additional trends emerge that require more analysis to fully understand. For example, Approx generally increases in performance when evaluated on more fine-grained classification settings, while CONEXP is typically worse here.

ICaCE-normdiff

Figure 11 shows the results for the ICaCE-Error with the normdiff distance metric. In general, it is more difficult for explanation methods to estimate the magnitude of the intervention effect when the task increases in complexity. For a given explanation method and model, best results are often achieved for the binary classification problem.

The conceptually simple Approx baseline wins across the board. S-Learner is only able to match its performance a few times. While previous results already showed that most of the methods fall behind the Approx baseline, the results are particularly striking for this metric.

While S-learner and CONEXP were somewhat comparable on the cosine metric, their differences become clear on the normdiff metric: S-Learner is better at estimating the magnitude of the intervention.

An interesting trend can be observed for TCAV, which has good performance on the binary task but becomes worse than random when evaluated on the ternary and 5-way settings. ConceptSHAP is the only method that consistently breaks the upward trend when going from ternary to the 5-way setting. More analysis is needed to understand both these phenomena.

ICaCE-L2

Figure 12 shows the results for the ICaCE-Error with the L2 distance metric. Because this metric takes both the scale and direction of the effect into account, it is slightly harder to interpret. In general, the performance drops when evaluated on more fine-grained classification settings.

Again, the Approx baseline is a strong contestant, but on this metric the results are more varied. S-Learner is consistently the best at producing the closest explanation in Euclidian distance to the real effect for the 5-way setting.

Appendix E CausaLM

The CausaLM algorithm was originally designed to estimate the average treatment effect of a high-level concept on pre-trained language models. Its output estimator is the textual representation averaged treatment effect (TReATE), which is computed as:

where ϕCCF\phi^{\text{CF}}_{C} denotes the learned counterfactual representation that information about concept CC is not present, N′\mathcal{N}^{\prime} is a classifier trained on this counterfactual representation, and D\mathcal{D} is a dataset.

However, for comparison on the CEBaB data, we require the estimation of individual causal concept effects (ICaCE). To allow a fair comparison, we swap the TReATE output estimator with TReITE (Equation 10). The only difference between these estimators is that in TReITE we remove the average across D\mathcal{D}, and output the estimated effect of individual examples.

E.2 Implementation details

For all counterfactual models, we optimize using the Adam optimizer with lr=2e-5, epochs=3, batch_size=48, and the relative weight of the adversarial task, λ\lambda, is set to 0.10.1.

For both the factual models and fine-tuning phase, we optimize using the Adam optimizer with lr=1e-3, epochs=50, and batch_size=256. The differences in hyperparameter values is due to the different architectures we employ; for the counterfactual models we train the entire language model (ϕ\phi), and for the factual models and the fine-tuning phase we freeze the embedding weights (ϕ\phi) and train only the classification head (N\mathcal{N}).

All CausaLM models were trained using 2 Nvidia GTX 1080 Ti 12GB GPUs.

Appendix F INLP

The INLP algorithm was originally designed to debias word embeddings by iteratively projecting them onto the null-space of some protected attribute (concept). However, INLP may serve as an estimation method similar to CausaLM, with the two following crucial differences. First, its lack of ability to control for potential confounders. Second, it operates on the representation rather than on the actual model weights. Since CausaLM and INLP share common characteristics, their output estimators are computed in the same way. See §E for extended details.

F.2 Implementation details

In order to guard for a \sayprotected attribute (concept), INLP determines whether this concept is present in an embedding or not by learning a linear separator in the embedding space. Following the practice suggested in the original paper, we choose our linear separator to be an SVM learned using SGD with α=0.01\alpha=0.01, ε=0.001\varepsilon=0.001, and max_iter=1000. Logistic regression showed similar behavior. We project the representation to the null-space with respect to the concept 1010 times. In fact, and similarly to the original paper, we converge to random accuracy of predicting the concept from the counterfactual representation after 4-5 iterations.

For all concepts, the classification head on top of the language model that trained to predict the overall sentiment labels trains for 5 epochs using the Adam optimizer with lr=2e-5.

Appendix G TCAV

The Testing with Concept Activation Vectors (TCAV) explanation method was originally designed to count the percentage of test inputs from dataset D\mathcal{D} that are positively influenced by some high-level concept. It outputs a count over the number of examples that are change towards the direction of concept CC, and computed as:

where kk is some class index and vCv_{C} is a linear direction in the activation space, given by the coefficients of a linear separator trained to distinguish between examples that include or exclude the concept CC.

While TCAV’s output is a count over examples, we use the raw sensitivity (directional derivative). This approach is supported by the authors of the original paper: \sayone could also use a different metric that considers the magnitude of the conceptual sensitivities Kim et al. . Also, since TCAV operates on the gradients of a model’s logits but the ICaCEs are the difference of two probability vectors, we normalize its outputs by taking Tanh.

G.2 Implementation details

To learn the Concept Activation Vector (CAV, i.e., a linear direction in the activation space of ϕ\phi), we train a linear separator to distinguish between examples that include the concept (labeled positive or negative) and examples that do not include it (labeled unknown). When learning CAVs, we drop all CEBaB train examples that are not labeled for aspect (concept) or do not have a majority with respect to the aspect.

Identically to the original paper, our CAV linear separator is an SVM learned using SGD with α=0.01\alpha=0.01, ε=0.001\varepsilon=0.001 and max_iter =1000=1000.

Appendix H ConceptSHAP

The original ConceptSHAP algorithm takes a complete set of concepts C∈{C1,...,Cm}C\in\{C_{1},...,C_{m}\} (such that its completeness score in Equation 25 is higher than some threshold) and outputs the relative contribution to the test accuracy of each CiC_{i}. It outputs an estimator given by the following formula

where η\eta is a scoring function operating on sets of concepts that output accuracy ratios.

Similarly to the other methods, if η\eta outputs accuracy ratios, then the output of ConceptSHAP is not a suitable estimator for ICaCE. Our straightforward adaptation for ConceptSHAP is to make η\eta output class probabilities for classes instead of accuracy ratios.

Our adapted version outputs a vector for each C∈{C1,…,Cm}C\in\{C_{1},\ldots,C_{m}\} and xx according to the following equation:

Yeh et al. calculate concept directions vCjv_{C_{j}} automatically by learning a neural network classifier. To allow for a fair comparison between ConceptSHAP and the other evaluated methods, we use the concept activation vectors vC1,…,vCmv_{C_{1}},\ldots,v_{C_{m}} as the input concepts (similarly to those used in Kim et al. ).

In addition, in the original paper the authors learn the concepts vCv_{C} automatically, by using a carefully constructed loss function. To allow a fair comparison, we learn the concept vector by exploiting our labeled aspects (concepts), in a way similar to TCAV. See Section G.2 for more details.

H.2 Completeness Scores of Treatment Concepts

Given a feature representation ϕ\phi and a classification head N\mathcal{N}, the completeness score is defined by:

For all models, the completeness we get for the set of concepts S={ambiance,food,service,noise}S=\{\text{ambiance},\text{food},\text{service},\text{noise}\} is larger than 0.90.9.

H.3 Hyperparameters

The hyperparameters for CAV are identical to those of TCAV (Section G.2). To calculate η\eta and the completeness score, we follow the original paper and set gg to be a two-layer perceptron with 500 hidden units, learned using Adam optimizer for 50 epochs, employing lr=1e-2 and batch_size=128.