Counterfactual Interventions Reveal the Causal Effect of Relative Clause Representations on Agreement Prediction
Shauli Ravfogel, Grusha Prasad, Tal Linzen, Yoav Goldberg
Introduction
The success of neural language models, both in NLP tasks and as cognitive models, has fueled targeted evaluation of these models’ word prediction accuracy on a range of syntactically complex constructions Linzen et al. (2016); Gauthier et al. (2020); Warstadt et al. (2020); Mueller et al. (2020); Marvin and Linzen (2018). What are the internal representations that support such sophisticated syntactic behavior? In this paper, we tackle this question using an intervention-based approach Woodward (2005). Our method, AlterRep, is designed to study whether a model uses a particular linguistic feature in a manner which is consistent with the grammar of the language. The method involves two steps: first, it generates counterfactualWe use the word counterfactual as it is used when referring to counterfactual examples Verma et al. (2020): an altered version of an element that is similar to the original element in all aspects except one. contextual word representations by altering the neural network’s representation of the linguistic feature under consideration; and second, it characterizes the change in the model’s word prediction behaviour that results from replacing the original word representations with their counterfactual variants. If the resulting change in word prediction aligns with predictions from linguistic theory, we can infer that the model uses the feature under consideration in a manner consistent with the grammar of the language.
We demonstrate the utility of AlterRep using relative clauses (RCs). According to the grammar of English, to correctly determine whether the masked verb in 1 should be singular or plural, a model must recognize that the masked verb is outside the RC the officers love, and should therefore agree with the subject of the main clause (the skater, which is singular), rather than with the subject of the RC (the officers, which is plural).
. The skater the officers love [MASK] happy.
To investigate whether a neural model uses RC boundary representations as predicted by the grammar of English, we use AlterRep to generate two counterfactual representations of the masked verb: one which encodes (incorrectly) the verb is inside the RC, and another which encodes (correctly) that the verb is outside the RC. Crucially, the difference between the counterfactual and original representations is minimal: the aspects of the representation which do not encode information about RC boundaries remain unchanged. Therefore, if the model uses RC boundary information as dictated by the grammar of English—and if our method successfully identifies the way in which RC boundary information is represented by the model—we expect the incorrect counterfactual to cause the masked verb to incorrectly agree with the noun inside the RC, and the correct counterfactual to cause agreement with the noun outside the RC, correctly.
We report experiments applying this logic to BERT variants of different sizes Devlin et al. (2019); Turc et al. (2019). We found that while all layers of the BERT variants encoded information about RC boundaries, only the information encoded in the middle layers was used in a manner consistent with the grammar of English. This contrast highlights the pitfalls of drawing behavioral conclusions from probing results alone, and motivates causal approaches such as AlterRep.
For BERT-base, we also found that counterfactual representations learned solely from one type of RC influenced the model’s predictions in sentences containing other RC types, suggesting that this model encodes information about RC boundaries in an abstract manner that generalizes across different RC types. Going beyond our case study of RC representations in BERT variants, we hope that future work can apply this method to test linguistically motivated hypotheses about a wide range of structures, tasks and models.
Background
An RC is a subordinate clause that modifies a noun. The head of the RC needs to be interpreted twice—once in the main clause, and once inside the RC—but it is omitted from inside the RC, replaced by an unpronounced “gap”. For example, in 2.1, the RC (in bold) describes the subject of the main clause the book. Since the book is the object of the embedded clause, we say that the gap is in the object position of the RC (indicated by underscores).
. The books that my cousin likes were interesting. (Object RC)
RCs can structurally differ from the Object RC in 2.1 in several ways: the overt complementizer that can be excluded, as in 2.1; the gap can be in the subject instead of object position of the embedded clause, as in 2.1; and so on. The five types of RCs we consider in this paper are outlined in Table 1.
. The books my cousin likes were interesting. (Reduced Object RC)
. My cousin that likes the books was interesting. (Subject RC)
These differences do not affect the strategy that a system that follows the grammar of English should use to determine the number of the verb: regardless of the internal structure of the RC, a verb outside the RC should agree with the subject of the main clause, whereas a verb inside the RC should agree with the subject of the RC. Thus, a model that does not properly identify the boundaries of the RC will often predict a singular verb where a plural one is required, or vice versa.
2 Iterative Null Space Projection (INLP)
The feature subspace—the space spanned by all the learned directions ()—is a subspace of the original representation space that contains information useful to linearly decode with high accuracy. The orthogonal complement of (the null space; ) is a subspace in which it is not possible to predict with high accuracy.
AlterRep: Generating Counterfactual Representations
The goal of AlterRep is to generate, based on a model’s contextual representations of a set of words, a set of counterfactual representations that modify the encoding of a feature while leaving all other aspects of the representations intact.We aim to propose a concrete instantiation that approximates the counterfactual. If swapping these counterfactual representations for the model’s original representations changes the model’s probability distribution over predicted words in a way that aligns with the feature’s linguistic functions, we say that the model uses for word prediction in a manner that is consistent with the grammar of English.
For our case study, we use a feature with two possible values: ‘’ if the word is inside an RC and ‘’ if it is not.In the experiments below, we will only apply this procedure to sets of representations of words that are all in a particular type of RC (for example, Object RCs). We do, however, test whether the representations of RC boundary generalize across RC types; see Prediction 3 in §5. We generate two counterfactual representations: , which encodes that the word is inside an RC—regardless of its actual syntactic position in the sentence—and , which is similar to in all respects except it encodes that is not in an RC. Our method allows us to generate and irrespective of the feature value encoded in the original representation . If the model uses this feature appropriately, we expect and to lead to different predictions in contexts where correct word predictions depend on determining whether or not the word is inside an RC.
Recall that INLP defines a feature space where the property of interest is encoded, and a complement subspace where it is not.
We can project any word representation to the feature subspace (here, the RC subspace) or to the null space, resulting in the vectors and , respectively: maintains the information needed to predict from , while maintains all information which is not relevant for predicting . INLP can be used to generate “amnesic counterfactuals” (Elazar et al., 2021), which do not encode a given property, even if the original representation did encode that property. In the next paragraph we propose a way to use this algorithm to manipulate the value of the feature, rather than remove it.
Generating Counterfactual Representations
We obtain the counterfactual representations and as follows. As we discussed in Section 2.2, INLP identifies planes—one for each direction (row) in —each of which linearly divides the word-representation space into two parts: words that are in an RC and words that are not. From the representation of a word that is not in an RC, we can generate by pushing across the separating plane towards the representations of words that are inside an RC . Similarly, we can generate by moving further away from that plane (see Figure 2).For a word that is inside an RC, the reverse computations would be required: to generate we would move further away from the separating plane, whereas to generate we would move across the separating plane.
For any word , we expect a positive counterfactual to be classified as being inside an RC, with high confidence, according to all original RC directions — that is, . Conversely, we expect a negative counterfactual to be classified as not being in an RC, i.e., .
To enforce these desiderata, we create positive and negative counterfactuals as follows, where if and otherwise, and is a positive scalar hyperparameter that enhances or dampens the effect of the intervention.
In both cases, we subtract a direction , flipping its sign, if the sign constraints are violated, that is, if for and if for . Geometrically, flipping the sign of a direction in Equations 2 and 3 is equivalent to taking a mirror image with respect to a direction (Figure 2). This enforces the sign constraints: all classifiers predict the negative or positive class, respectively (see Appendix §A.1 for a formal proof).
Experimental Setup
Our overall goal is to assess the causal effect of RC boundary representations on our models’ agreement behavior when subject-verb dependencies span an RC (that is, where an RC intervenes between the head of the subject and the corresponding verb). We test whether we can modify the representation of the masked verb outside the RC such that, compared to the original representations, the model assigns higher probability to either the correct form (after negative intervention) or to the incorrect one (after positive intervention). We first describe the models we use (§4.1), then the dataset we use to obtain RC subspaces and generate counterfactual representations (§4.2), and finally the dataset we use to measure the models’ agreement prediction accuracy in sentences containing RCs, before and after the counterfactual intervention (§4.3).
We use BERT-base (12 layers,768 hidden units) and BERT-large (24 layers, 1024 hidden units) Devlin et al. (2019), as well as the smaller BERT models released by Turc et al. (2019): BERT-medium (8 layers 512 hidden units), BERT-small (4 layers, 512 hidden units), BERT-mini (4 256), and BERT-tiny (2 128). In all experiments, we intervene on a single layer at a time, and continue the forward pass of the original model through the following layers.
2 Generating Counterfactual Representations
To create the training data for the INLP classifiers, we used the templates of Prasad et al. (2019) to generate five lexically matched sets of semantically plausible sentences, one for each type of RC outlined in Table 1, as well as two additional sets of sentences without RCs; these included sentences with nearly the same word order and lexical content as the sentences in the other sets. Each set contained sentences. All verbs in the training sentences were in the past tense; this ensured that the subspaces we identified did not contain information about overt number agreement, making it unlikely that AlterRep will alter agreement-related information that does not concern RCs.
Identifying and Altering RC Subspaces
To identify RC subspaces, we used INLP with SVM classifiers as implemented in scikit-learn. We identified different subspaces for each of the five types of RCs listed in Table 1. For example, in 4.2, the bolded words were considered to be in the RC. \ex. My cousin that liked the book hated movies.
For the negative examples, we took representations of words outside of the RC, either from outside the bolded region of the same sentence, or from inside or outside the bolded region of the coordination control sentence.
. My cousin liked the book and hated movies.
We selected the negative examples in this manner for two reasons: first, to ensure that the same word served as a positive example in some context and as a negative example in others (e.g., book in 4.2 and 4.2); and second, to ensure that the same RC sentence included both positive and negative examples (e.g., book and cousin in 4.2).
Hyperparameters
INLP has a hyperparameter which sets the dimensionality of the RC subspace; this parameter trades off exhaustivity against selectivity.In particular, running INLP for 768 iterations—the dimensionality of BERT representations—yields the original space, which is exhaustive but not useful in distilling RC information. We set ; In Appendix §A.3 we demonstrate that the trends we observe are not substantially affected by this parameter.
AlterRep has an hyperparameter, , that determines the magnitude of the counterfactual intervention (§4.2). We use ; In Appendix §A.4 we show that the trends we observe are similar for other values of .
3 Measuring the Effect of the Intervention on Agreement Accuracy
We measure the models’ agreement prediction accuracy using a subset of the Marvin and Linzen (2018) dataset in which the subject is modified by an RC. The noun inside the RC either matched 4.3 or mismatched 4.3 the subject of the matrix clause in number:
. The skater that the officer loves is/are happy.
. The skater that the officers love is/are happy.
The Marvin and Linzen dataset contains sentences where the intervening RC is either a subject RC or a (reduced or unreduced) object RC. We augmented this dataset with lexically matched sentences with (reduced or unreduced) passive RC interveners, using attribute varying grammars (Mueller et al., 2020). Finally, we only considered sentences with copular main verbs (is and are) to ensure that both the singular and plural forms of the verb are highly frequent. We used 1750 sentences per construction.
Computing Agreement Accuracy
We performed masked language modeling (MLM) on the dataset described earlier in this section. In each sentence, we masked the copula, started the forward pass, performed the intervention on the representation of the masked copula in the layer of interest, and continued with the forward pass to obtain BERT’s distribution over the vocabulary for the masked token. We repeated this process for each layer separately. We then computed the probability of error, normalized within the two copulas is and are (Arehalli and Linzen 2020):
In Appendix §A.5, we present results where the metric of success is accuracy, that is, the percentage of cases where the model assigned a higher probability to the verb with the correct number Marvin and Linzen (2018). These results are qualitatively similar.
Predictions
As discussed earlier, a system that computed agreement in accordance with the grammar of English would determine the number of the masked verb in a sentence like 5 based on the number of officers, because both officers and the [MASK] token are outside the RC; the number of the RC-internal noun skater should be ignored. \ex. The officers that love the skater [MASK] nice.
We can derive the following predictions for applying AlterRep to a system that follows this strategy:
In RC sentences where the main clause subject differs in number from the RC subject, error probability will be higher with the counterfactual , which encodes (incorrectly) that [MASK] is inside the RC, than with the original representation . Conversely, error probability will be lower with , which encodes (correctly) that [MASK] is outside the RC, than with the original .
Prediction 2: No Impact on Other Sentences.
We do not expect a difference in error probability between the original and counterfactual representations in all other sentences. This should be the case for sentences with RCs where the nouns inside and outside the RC match in number, as in 5:
. The officers that love the skaters [MASK] nice.
Since both officers and skaters are plural, most plausible agreement prediction strategies would make the same predictions regardless of whether [MASK] is analyzed as being inside the RC or outside it. Consequently, intervening on the encoding of RC boundaries is not expected to systematically change the model’s predictions.
Likewise, since the interventions are designed to modulate the encoding of RC-related properties, we do not expect the interventions to impact number prediction in sentences without RCs such as 5 and 5:If models encoded boundaries of all embedded clauses similarly we would expect a change in prediction for 5.
. The bankers knew the officer [MASK] nice.
Prediction 3: Generalization Across RC Types.
If RC boundaries are represented in an abstract way that is shared across different RC types, then the counterfactual representations will affect error probability in the same way regardless of whether the counterfactual representations were generated from subspaces estimated from sentence with the same RC type as the target sentences, or from sentences with different RC types.
Results
We begin by discussing experiments where subspaces were estimated from sentences with the same type of RC as the test sentences with agreement; we report results averaged across the five RC types. Interventions using counterfactual representations generated from the middle layers of the BERT-base (5–8 out of 12) resulted in changes in the probability of error which partially aligned with Prediction 1 (Figure 3(a)). In sentences with attractors, using the positive counterfactual resulted in an increase in the probability of error (a maximum increase of 14 percentage points in layer 7). Conversely, using the negative counterfactual generated from layers 5 and 6 resulted in a decrease in the probability of error. This decrease was much smaller (a maximum decrease of 2 percentage point in layer 6) and there was overlap in the error bars for the probability of error before and after intervention.
It is likely that the smaller effects of the negative counterfactual intervention are due to the fact that accuracy before the intervention was very high (95%) and the probability of error very low (8%), leaving very little room for change: in most cases, the original representation already correctly encoded the verb is outside the RC. In a follow-up analysis, we only considered sentences in which the model originally assigned a higher probability to the ungrammatical than the grammatical form. In these examples the decrease in probability of error was larger (a maximum decrease of 16 percentage points in layer 6; see Figure 3(c)).
While only RC interventions in the middle layers elicit the expected behavioral outcomes, probing accuracy for RC information was high for all layers (Appendix §A.2), giving further evidence to the dissociation between correlational and causal methods: probing can identify aspects of the representations that do not affect the model’s behavior.
Interventions on RC Boundary Representations Generalize Across RC Types, but not Further.
In line with Prediction 3, we observed a qualitatively similar pattern of change in error probabilities even when the counterfactuals were generated from subspaces estimated from a different RC type than the RC in the agreement test sentences. The effects were smaller, however. This suggests that while BERT’s representation of RC boundaries is partly shared across different RC types, there are also structure-specific RC boundary representations. The effect of the intervention also aligned with Prediction 2: in constructions where we do not expect RC boundaries to affect predictions—sentences without attractors and those without RCs—we did not observe significant changes in error probability (Figure 3(b)).
Intervention Based on Random Subspaces Does Not Produce Interpretable Results.
To tease apart the effect of the RC-targeted intervention from intervening on any subspace with the same dimensionality, we generated counterfactual representations from 10 random subspaces and repeated our analysis.We generated a random subspace by sampling standard Gaussian vectors instead of the INLP matrix , and then employing the same procedure described in §3. While we observed very small changes in probability of error in some cases, the pattern of changes resulting from this intervention did not align with any of our predictions (see Figure 3(d)). This suggests that the change in probability of error that resulted from intervening with RC subspaces was not merely a by-product of intervening on a large enough subspace of BERT’s original representation space.
Intervening on the Middle Layers of Other BERT Variants Yielded Qualitatively Similar Results.
We repeated the experiments on BERT-large and four smaller versions of BERT, trained on the same amount of data as the BERT-base model Turc et al. (2019). As with BERT-base, intervening on the middle layers of BERT-large (12–17 out of 24) with the RC subspaces—but not the random subspaces—resulted in predicted changes in the probability of error. Compared to BERT-base, the smaller models showed a greater change in the probability of error as a result of intervention with counterfactuals generated from random subspaces. However, when the counterfactual representations were generated from particular layers—4 and 5 (out of 8) in BERT-medium, 3 (out of 4) in BERT-mini and 2 (out of 4) in BERT-small—the change in error probability aligned with Prediction 1 over and above the changes from intervening with random subspaces. In all of these layers, intervening with the positive but not the negative counterfactual resulted in an increase in the probability of error. No such layer was observed for BERT-tiny, which has only 2 layers (see Figure 4).
Discussion
We proposed an intervention-based method, AlterRep, to test whether language models use the linguistic information encoded in their representations in a manner that is consistent with the grammar of the language they are trained on. For a given linguistic feature of interest, we generated counterfactual contextual word representations by manipulating the value of the feature in the original representations. Then, by replacing the original representations with these counterfactual variants, we characterized the change in word prediction behaviour. By comparing the resulting change in word prediction with hypotheses from linguistic theory about how specific values of the feature are expected to influence the probabilities over predicted words, we investigated whether the model uses the feature as expected.
As a case study, we applied this method to study whether altering the information encoded about RC boundaries in the contextual representations of masked verbs in different BERT variants influences the verb’s number inflection in a manner that is consistent with the grammar of English. We found that while all layers of the BERT variants encoded information about RC boundaries, only the information in the middle layers influenced the masked verb’s number inflection as predicted by English grammar. We also found that in BERT-base, counterfactual representations based on subspaces that were learned from sentences with one type of RC influenced the number inflection of the masked verb in sentences with other types of RCs; this suggests that the model encodes information about RC boundaries in an abstract manner that generalizes across the different RC types.
AlterRep interventions are based on concept subspaces identified using linear classifiers, but most neural networks components, including BERT layers, are non-linear. It is possible, then, that subsequent non-linear layers transform the counterfactual representation in a way that is not amenable to analysis using our methods. As such, while we can conclude from a positive result that the feature in question causally affects the model’s behavior, negative results should be interpreted more cautiously.
Future Work
Future work can apply this method to test linguistically motivated hypotheses about a wide range of structures and tasks. For example, linguistic theory predicts that information about semantic roles (like agent and patient) is crucial for tasks such as natural language inference (NLI) that require reasoning about sentence meaning. To test if NLI models use semantic roles as predicted by linguistic theory, we can use AlterRep to replace the original representations with counterfactual representations where the patient is encoded as the agent (and vice versa), and measure the change in performance on NLI, especially on challenge sets such as HANS McCoy et al. (2019) that evaluate sensitivity to these properties.
Related Work
Behavioral tests of neural models, such as the ability of the model to master agreement prediction (Linzen et al., 2016; Gulordava et al., 2018; Goldberg, 2019), have exposed both impressive capabilities and limitations. These paradigms focus on the model’s output, and do not link the behavioral output with the information encoded in its representations. Conversely, probing (Adi et al., 2017; Conneau et al., 2018; Hupkes et al., 2018) does not reveal whether the property recovered by the probe affects the original model’s prediction in any way (Hewitt and Liang, 2019; Tamkin et al., 2020; Ravichander et al., 2021). This has sparked interest in identifying the causal factors that underlie the model’s behavior (Vig et al., 2020; Feder et al., 2020; Voita et al., 2020; Kaushik et al., 2020; Slobodkin et al., 2021; Pryzant et al., 2021; Finlayson et al., 2021).
Counterfactuals
The relation between counterfactual reasoning and causality is extensively discussed in social science and philosophy literature Woodward (2005); Miller (2018, 2019). Attempts have been made to generate counterfactual examples Maudslay et al. (2019); Zmigrod et al. (2019); Ross et al. (2020); Kaushik et al. (2020); Hvilshøj et al. (2021) and recently to derive counterfactual representations Feder et al. (2020); Elazar et al. (2021); Jacovi et al. (2021); Shin et al. (2020); Tucker et al. (2021). Contrary to our approach, previous attempts to generate counterfactual representations were either limited to amnesic operations (i.e., focused on the removal of information and not on modifying the encoded information) or used gradient-based interventions, which are expressive and powerful, but less controllable. Our linear approach is guided by well-defined desiderata: we want all linear classifiers trained on the original representation to predict a specific class for the counterfactual representations, and we prove that is the case in Appendix §A.1.
Representations and Behavior
Previous work bridging the gap between representations and behavior includes Giulianelli et al. (2018), who demonstrated that back-propagating an agreement probe into a language model induces behavioral changes and improve predictions. Lakretz et al. (2019) identified individual neurons that causally support agreement prediction. Prasad et al. (2019) used similarity measures between different RC types extracted using behavioural methods to investigate the inner organization of information within the model. Closest to our work is Elazar et al. (2021), where the authors applied INLP to “erase” certain distinctions from the representation, and then measured the effect of the intervention on language modeling. We extend INLP to generate flexible counterfactual representations (§3) and use these to instantiate hypotheses about the linguistic factors that guide the model’s behavior.
Conclusions
We proposed an intervention-based approach to study whether a model uses a particular linguistic feature as predicted by the grammar of the language it was trained on. To do so, we generated counterfactual representations in which the linguistic property under consideration was altered but all other aspects of the representation remained intact. Then, we replaced the original word representation with the counterfactual one and characterized the change in behaviour. Applying this method to BERT, we found that the model uses information about RC boundaries that is encoded in its word representations when inflecting the number of masked verb in a manner consistent with the grammar of English. We conclude that AlterRep is an effective tool for testing hypotheses about the function of the linguistic information encoded in the internal representations of neural LMs.
Acknowledgements
This work was supported by United States–Israel Binational Science Foundation award 2018284, and has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme, grant agreement No. 802774 (iEXTRACT). We thank Robert Frank for a fruitful discussion of an early version of this work, and Marius Mosbach, Hila Gonen and Yanai Elazar for their helpful comments.
References
Appendix A Appendix
In this appendix, we prove that the method presented in §3 is guaranteed to achieve its goal: the negative counterfactual would be classified as belonging to the negative class, and the positive counterfactual would be classified as belonging to the positive class, according to all the linear classifiers trained on the original representation.
We base our derivation on the decomposition presented in §3:
Where is the nullspace of the INLP matrix , is its rowspace, and and are the orthogonal projection of a representation to those subspaces, respectively.
For the negative counterfactual defined by , it holds that would always be classified to the negative class: for every in the original INLP matrix .
Where the transition from 6 to 7 stems from being in the nullsapce of , so ; and the transition from 7 to 8 stems from the mutual orthogonality of the INLP classifiers (proved in Ravfogel et al. (2020)): since , it holds that .
Case 1: , that is, the classifier predicted the positive class on the original representation. Then, by 8,
Since is a positive scalar and by assumption , it holds that .
Case 2: , that is, the classifier predicted the negative class on the original representation. Then, by 8,
Since is a positive scalar and by assumption , it holds that .
We have proved that regardless of the originally predicted label, all INLP classifiers would predict the negative class on the negative counterfactual, which concludes the proof. ∎
A.2 Probing Accuracy
In this appendix, we provide probing results for the task on which we run INLP: detecting whether representation was taken over a word inside or outside of an RC. As INLP iteratively trains linear probes, this accuracy is equivalent to the accuracy of the first INLP classifier. In all contextualized layers, we observe probing accuracy of over 90% for all RC types (Figure 5). This contrasts with the intervention results in §6. While it is possible to linearly decode the RC boundary in all layers, only in the middle layers do we find that this concept causally influences the model’s behavior. In other words, good probing performance does not indicate main-task relevancy.
A.3 Influence of the Dimensionality of the RC Subspace
In this appendix, we analyze the influence of the dimensionality of the RC subspace. Recall that INLP is an iterative algorithm (§2.2). On the th iteration, the method identifies a single direction —the parameter vector of a linear classifier—which is predictive of the concept of interest (in our case, RC). The different directions are mutually orthogonal, and after iterations, the “concept subspace” is the subspace spanned the rows of the matrix . In the th iteration of INLP, the subspace identified so far is removed from the representation (by the operation of nullspace projection), and the next classifier is trained to predict the concept over the residual representation. As such, accuracy is expected to decrease with the number of iterations: as the number of iterations increases, the algorithm identifies directions which have a weaker association with the concept. This creates a trade-off between exhaustively – identifying all the directions which are at least somewhat predictive of the concept, and selectivity – identifying only directions which have a meaningful association with the concept.
Figure 6(a) presents positive intervention results for different RC-subspace dimensionality on sentences with agreement across RC with attractors; Figure 6(b) present negative intervention results on sentences on which the model was originally mistaken. Generally, we observe the same trends under all settings, suggesting our method is relatively robust to the dimensionality of the manipulated subspace. In figure 6(c) we present the results of intervening on subspaces of different dimensionality, for sentences where we do not expect an effect: sentences without attractors, and sentences without RCs. For all contextualized layers we do not see an effect, as expected. For and , we see an effect on the uncontextualized embedding layer. This effect may hint towards a spurious information encoded in this uncontexualized layer which is used by the model when predicting agreement, but studying this possibility is beyond the scope of this work.
A.4 Influence of α𝛼\alpha
In this appendix, we analyze the influence of the parameter in the AlterRep algorithm (Section §3) on the BERT-base model. Recall that dictates the step size one takes when calculating the counterfactual mirror image: corresponds to exact mirror image, while over-emphasizes the RC components over which we take the counterfactual mirror image.
In Figures 7(a) and 7(b) we focus on the positive intervention, which is expected to increase the probability of error, making the model act as if the masked verb is within the RC; and in Figure 7(c) we focus on the negative intervention on sentences on which the model was originally mistaken, which is expected to decrease the probability of error.
In Figure 7(b) we present the results on the control sentences: sentences without agreement across RC. Overall, the trends we observe are similar for different values of , indicating that AlterRep is relatively robust to the value of this parameter. One exception is the large values of and to a lesser degree , which result in some increase in the probability of error also in the control sentences, where we do not expect such effect (Figure 7(b)), albeit this increase is much smaller than the increase on sentences with agreement across RC. With a large-enough , the new counterfactual representation might diverge too-much from the distribution of the original representations. Notice that when compared with gradient-based methods for generating counterfactuals Tucker et al. (2021), our linear approach has the advantage of being able to control the magnitude of the intervention with a single controlled parameter, which has a clear geometric interpretation: the extent to which one pushes the representations to one direction or another when taking the mirror image.
A.5 Influence on Accuracy
In this appendix, we evaluate the impact of the intervention by its influence on the model’s accuracy, calculated as the percentage of cases where the model assigned higher probability to the correct form than to the incorrect form. We focus on the cases on which the model originally predicted incorrectly, Thus, the original accuracy on this group of sentences is 0%. We use a negative intervention, pushing the model to act as if the verb is (correctly) outside of the RC, which is expected to increase its accuracy.
In Section §4.3 we use an alternative measure: probability-of-error. The probability of error is a more sensitive measure, as it might change even when the model’s absolute preference for one form over the other has not. However, it is the absolute ranking which eventually dictates the model’s top prediction.
Figure 8 presents the results for different dimensionalities of the RC subspace. The trends are similar to the trends shown by the probability-of-error evaluation measure. Notably, in up to 30% of the cases, it is possible to flip the model’s preference from the incorrect to the correct form solely by manipulating a low-dimensional subspace within the 768-dimensional representation space.