On the Pitfalls of Analyzing Individual Neurons in Language Models

Omer Antverg, Yonatan Belinkov

Introduction

Many studies attempt to interpret language models by predicting different linguistic properties from word representations, an approach called probing classifiers (Adi et al., 2017; Conneau et al., 2018, inter alia). A growing body of work focuses on individual neurons within the representation, attempting to show in which neurons some information is encoded, and whether it is localized (concentrated in a small set of neurons) or dispersed. Such knowledge may allow us to control the model’s output (Bau et al., 2019), to reduce the number of parameters in the model (Voita et al., 2019; Sajjad et al., 2020), and to gain a general scientific knowledge of the model. The common methodology is to train a probe to predict some linguistic attribute from a representation, and to use it, in different ways, to rank the neurons of the representation according to their importance for the attribute in question. The same probe is then used to predict the attribute, but using only the kk-highest ranked neurons from the obtained ranking, and the probe’s accuracy in this scenario is considered as a measure of the ranking’s quality (Dalvi et al., 2019; Torroba Hennigen et al., 2020; Durrani et al., 2020). We see this framework as exhibiting Pitfall I: Two distinct factors are conflated—the probe’s classification quality and the quality of the ranking it produces. A good classifier may provide good results even if its ranking is bad, and an optimal ranking may cause an average classifier to provide better results than a good classifier that is given a bad ranking.

Another shortcoming of the current methodology, which we mark as Pitfall II, is the focus on encoded information, regardless of whether it is actually used by the model in its language modeling task. A few studies (Elazar et al., 2021; Feder et al., 2020) have considered this question, and shown that encoded information is not necessarily being used for language modeling, but these do not look at individual neurons. We argue that in order to evaluate a ranking, one should also examine if, and how, the kk-highest ranked neurons are used by the model for the attribute in question, meaning that modifying them would change the model’s prediction—but with respect to that attribute only. This would allow some control over the model’s output, and grant us parameter-level explanations of the model’s decisions.

In this work, we analyze three neuron ranking methods. Since the ranking space is too large (768!768! in BERT’s case), these methods provide approximations to the problem and are non-optimal. Two of these methods—Linear (Dalvi et al., 2019) and Gaussian (Torroba Hennigen et al., 2020)—rely on an external probe to obtain a ranking: the first makes use of the internal weights of a linear probe, while the second considers the performance of a decomposable generative probe. The third is a simple ranking method we propose, Probeless, which ranks neurons according to the difference in their values across labels, and thus can be derived directly from the data, with no probing involved.

We experiment with disentangling probe quality and ranking quality, by using a probe from one method with a ranking from another method, and comparing the different probe–ranking combinations. We expose the problematic nature of the current methodology (Pitfall I), by showing that in some cases, a suitable probe which is given an intentionally bad ranking, or a random one, provides higher accuracy than another which is given its allegedly optimal ranking. We find that while the Gaussian method generally provides higher accuracy, its probe’s selectivity (Hewitt & Liang, 2019) is lower, implying that it performs the probing task by memorizing, which improves probing quality but not necessarily ranking quality. We further find that Gaussian provides the best ranking for small sets of neurons, while Linear provides a better ranking for large sets.

We then turn to analyzing which ranking selects neurons that are used by the model, by applying interventions on the representation: we modify subsets of neurons from each ranking and measure—using a novel metric we introduce—the effect on language modeling w.r.t to the property in question. We highlight the need to focus on used information (Pitfall II): even though Probeless does not excel in the probing scenario, it selects neurons that are used by the model, more so than the two probing-based rankings. We find that there is an overlap between encoded information and used information, but they are not the same, and argue that more attention should be given to the latter.

We primarily experiment with the M-BERT model (Devlin et al., 2019) on 9 languages and 13 morphological attributes, from the Universal Dependencies dataset (Zeman et al., 2020). We also experiment with XLM-R (Conneau et al., 2020), and find that most of our results are similar between the models, with a few differences which we discuss. Our experiments reveal the following insights:

We show the need to separate between probing quality and ranking quality, via cases where intentionally poor rankings provide better accuracy than good rankings, due to probing weaknesses.

We present a new ranking method that is free of any probes, and tends to prefer neurons that are being used by the model, more so than existing probing-based rankings.

We show that there is an overlap between encoded information and used information, but they are not the same.

Neuron rankings and data

The ranking methods we compare include two rankings obtained from prior probing-based neuron-ranking methods, and a novel ranking we propose, based on data statistics rather than probing.

The first method, henceforth Linear (named linguistic correlation analysis in Dalvi et al. 2019),A small enhancement to the algorithm was presented in Durrani et al. (2020). trains a linear classifier on the representations to learn the task FF. Then, it uses the trained classifier’s weights to rank the neurons according to their importance for FF. Intuitively, neurons with a higher magnitude of absolute weights should be more important, or contain more relevant information, for solving the task. Dalvi et al. (2019) showed that their method identifies important neurons through probing and ablation studies, and found that while the information is distributed across neurons, the distribution is not uniform, meaning it is skewed towards the top-ranked neurons. In this work, we use a slightly modified version of the suggested approach (Appendix A.1).

The second method, henceforth Gaussian (Torroba Hennigen et al., 2020), trains a generative classifier on the task FF, based on the assumption that each dimension in {1,...,d}\{1,...,d\} is Gaussian-distributed. Then, it makes use of the decomposability of the multivariate Gaussian distribution to greedily select the most informative neuron, according to the classifier’s performance, at every iteration. This way we obtain a full neuron ranking after training only once, while applying this greedy method to Linear would require retraining the probe d!d! times, which is clearly infeasible. Torroba Hennigen et al. (2020) found that most of the tasks can be solved using a low number of neurons, but also noted that their classifier is limited due to the Gaussian distribution assumption.

The third neuron-ranking method we experiment with is based purely on the representations, with no probing involved, making it free of probing limitations (Belinkov, 2021) that might affect ranking quality. For every attribute label z∈Zz\in Z, we calculate q(z)q(z), the mean vector of all representations of words that possess the attribute and the value zz. Then, we calculate the element-wise difference between the mean vectors,

and obtain a ranking by arg-sorting rr, i.e., the first neuron in the ranking corresponds to the highest value in rr. For binary-labeled attributes, this is simply the difference in means. In the general case, Probeless assigns high values to neurons that are most sensitive to a given attribute. We note that Probeless is very fast to use, as we are only limited by averaging and sorting, as opposed to training a classifier in Linear or the expensive greedy algorithm of Gaussian.

2 Data and models

Throughout our work, we follow the experimental setting of Torroba Hennigen et al. (2020): we map the UD treebanks (Zeman et al., 2020) to the UniMorph schema (Kirov et al., 2018) using the mapping by McCarthy et al. (2018). We select a subset of the languages used by Torroba Hennigen et al. (2020): Arabic, Bulgarian, English, Finnish, French, Hindi, Russian, Spanish and Turkish, to keep linguistic diversity. The tasks we experiment with are predictions of morphological attributes from these languages. Full data details are provided in Torroba Hennigen et al. (2020) and further data preparation steps are detailed in Appendix A.2. We process each sentence in pre-trained M-BERT and XLM-R (unless stated otherwise, all results are with M-BERT), and take word representations from layers 2, 7 and 12 of each model, to see if there are different patterns in the beginning, middle and end of the models. We end up with a total of 156 different configs (language ×\times attribute ×\times layer) to test for each model. For words that are split during tokenization, we define their final representation to be the average over their sub-token representations. Thus, each word has one representation for each layer, of dimension d=768d=768. We do not mask any words throughout our work.

3 Overlaps

Before evaluating our rankings in different scenarios, we first characterize them by looking at the 100-highest ranked neurons (out of 768) from different rankings, across different configs.

Since we work with multilingual models, we expect to see overlap in the selected neurons for one attribute across different languages. Fig. 1 shows that for Probeless this is indeed the case, as some attributes share a large number of important neurons across languages. For example, number in Spanish and number in French share 70 of their 100 most important neurons, where the expected number for overlap of two random selections of neurons is only 13 (Appendix A.3). Compared to the other two rankings (Appendix A.4), Probeless is the most consistent across languages, while Gaussian rarely shows consistency, which may be a weakness.

By looking at the overlaps between important neurons selected by different rankings for the same config, we observe that for all configs, the overlap between all three rankings surpasses the expected number (which is ∼1.69\sim 1.69; Appendix A.3), meaning there are neurons that are recognized by all three rankings as important.

We further see that in most cases, the greatest overlap is between Linear and Probeless. We find it reasonable, as both of them aim to select neurons that separate classes the best—one by a classifier and the other by data statistics—while Gaussian takes a different approach, assuming a Gaussian distribution and selecting neurons only by performance. Examples from 8 configs are shown in Fig. 2.

Performing the same analysis across languages on XLM-R (Appendix A.4), the overlap size is at least as the expected one between all config pairs, and is usually greater. It may imply that XLM-R’s representations can be pruned more easily than M-BERT’s, since some neurons encode multiple attributes, and there is greater redundancy among the others.

Pitfall I: Classifiers vs. rankings

We now turn to evaluating the rankings, and present the pitfalls in doing so. Given some ranking Π(d)\Pi(d), we would like to evaluate how well it sorts the neurons for the task FF. Our first ranking-evaluation approach is the standard probing approach from previous work (Dalvi et al., 2019; Torroba Hennigen et al., 2020), where we expose the classifier to a subvector of the representation and evaluate how well it predicts the task. However, while previous work conflated rankings and classifiers—Pitfall I—we are more careful: we separate the two, and pair each ranking with two classifiers, meaning that at least one of them is completely unrelated to the ranking.

As classifiers, we experiment with both classifiers used by the first two ranking methods (Linear and Gaussian). We use the hyperparameters reported in Durrani et al. (2020) and Torroba Hennigen et al. (2020) for training the classifiers. As rankings, we experiment with the 3 ranking methods described in §2.1. For each, we use the original ranking it produces and its reversed version, referred to as top-to-bottom and bottom-to-top, respectively. To those we add a random ranking baseline, resulting in 7 different rankings overall. We compare all classifier–ranking combinations for each kk.

Since both of the first two ranking methods are inherently tied to the classifier that was used to generate them, and the third ranking is a classifier-neutral ranking, it can be used for a fair comparison between the classifiers.

First, we measure the accuracy of the probe’s predictions. Since we experiment with many different configs, we use the Wilcoxon signed-rank test (Wilcoxon, 1992) as a statistical significance test to determine whether a certain combination of a classifier and a ranking is statistically significantly better than another combination.

We also evaluate our probes by selectivity (Hewitt & Liang, 2019), defined as the difference between the classifier’s accuracy on the actual probing task and its accuracy on predicting random labels assigned to word types, called a control task. Low selectivity implies that the probe can memorize the word-type–label pair, and so high accuracy in the probing task does not necessarily entail the presence of the linguistic attribute. Thus, we prefer probes that are both accurate and selective.

2 Results

Across the 156 configs we experiment with, we observe three different accuracy patterns, demonstrated in Figs. 4(a)-4(c). In these figures, each color represents a combination of a classifier and a ranking, where a solid line is used for the top-to-bottom version of the ranking and a dotted line is for the bottom-to-top version of it, and a dashed line is used for the random ranking. Almost half of the configs follow the Standard pattern (Fig. 4(a)), in which all top-to-bottom rankings are always better than the random ranking, which is always better than all bottom-to-top rankings. The other half consists of two surprising patterns, that demonstrate the inherent flaws in this ranking-evaluation approach. In the G>L pattern (Fig. 4(b)), the Gaussian classifier performs exceptionally well, providing higher accuracy (after a certain point) using a random or even a bottom-to-top ranking, than the Linear classifier using its top-to-bottom ranking. In the L>G pattern (Fig. 4(c)), the Gaussian classifier fails quickly, and thus the Linear classifier provides higher accuracy using a random or bottom-to-top ranking than the Gaussian classifier using its top-to-bottom ranking.

Fig. 4(d) shows a t-SNE (van der Maaten & Hinton, 2008) projection after performing K-means clustering on our 156 accuracy results, where each point represents accuracy results from one config (details on clustering procedure are in Appendix A.5). It shows three clusters of configs, that correspond to the three distinct patterns. On XLM-R we see very similar results, and most configs follow the same pattern in each model (Appendix A.7). We now turn to analyze these results.

In most configs, each classifier provides better accuracy using a top-to-bottom ranking (solid lines) compared to the bottom-to-top version of the same ranking (same color, dotted line), and the random ranking (dashed lines) is in between. This is also seen in our statistical significance tests (Appendix A.6). We conclude that even if they are not optimal, all ranking methods we consider generally rank task-informative neurons higher than non-informative ones.

2.2 Which classifier is better?

In the Standard and G>L patterns (Figs. 4(a), 4(b)), Gaussian achieves better accuracy than Linear when both of them use the same ranking (including top-to-bottom Linear), especially when using small sets of neurons. Our statistical significance tests (Appendix A.6) show that Gaussian performs significantly better than Linear with 6 out of the 7 rankings we tried when using 10 neurons, with 5 rankings when using 50 neurons, and with 5 rankings when using 150 neurons. We now turn to analyze what makes Gaussian more successful, and show some exceptions.

Across all configs, Linear provides higher selectivity than Gaussian using any ranking, after a certain point (Appendix A.7). This means that Gaussian tends to memorize the word-type–label pair when solving the task. While this is apparent in all configs, we note that specifically in the part-of-speech attribute there is a large portion of function words, i.e., closed set labels (e.g., pronouns, determiners), meaning that memorization can significantly help solve the task. Thus, most configs involving part of speech belong to the G>L pattern. Since memorization is a trait of the classifier and not of the ranking, this pattern demonstrates the problematic nature of the current ranking-evaluation approach, as the results are highly dependent on the probe.

On the other hand, pattern L>G shows that there are certain configs where Gaussian is struggling to model the distribution, resulting in mediocre accuracy results—which even start decreasing at some point—sometimes even below majority baseline, as seen in Fig. 4(c). This has also been mentioned in Torroba Hennigen et al. (2020), where it was shown that in those configs, there are only a few (or no) dimensions that are informative for the attribute and are Gaussian-distributed. Thus, the Gaussian classifier tries to model these distributions with the wrong tools, and fails. Poor modeling then leads to wrong predictions and low accuracy. In general, Linear behaves similarly across configs, making it more stable.

2.3 Which ranking is better?

When looking at rankings, we would like to compare performance of the same classifier, using different rankings. We would expect that each classifier would perform best when using the ranking it has generated. However, this is not always the case. As we can see in all patterns in Fig. 4, for small sets of neurons, Linear actually achieves better accuracy when using Gaussian’s ranking (solid green) than its own ranking (solid orange). As the number of neurons increases, at some point its accuracy with its own ranking becomes higher than with Gaussian’s ranking.

We suggest two explanations for this phenomenon: First, due to its greediness, the Gaussian ranking is not guaranteed to provide the optimal subset. For a subset of size 1, it goes over all possibilities, but as the size grows there are more subsets that are not taken into consideration in the algorithm, so it is more likely to miss the best sets. Second, Gaussian assumes the embedding distribution to be Gaussian. On dimensions which are not Gaussian-distributed, it makes a less accurate evaluation of the contribution of each neuron. So, if a neuron is informative towards the attribute but is not Gaussian-distributed, its addition to the selected neurons set is unlikely to improve performance, and thus it is not selected. This is a problem with a performance-based selection criterion, where the selection of neurons depends on the performance of the probe.

To summarize, it seems that Gaussian is good at selecting specific informative neurons, but misses the rest. While Linear’s ranking is not optimal (it is definitely worse then Gaussian’s on small sets), it does seem to be more stable on different sizes. Probeless provides decent performance (and is inherently consistent), but is usually behind the other two.

Pitfall II: Encoded information vs. used information

The variance of results in our probing experiments can mostly be attributed to probing limitations (Hewitt & Liang, 2019; Belinkov, 2021), and emphasizes the need to distinguish between two properties: the probe’s classification quality, and the neuron-ranking quality. To isolate the latter, and to shed light on which ranking prefers neurons that are actually used by the model for the attribute in question (which is ignored by previous work—Pitfall II), we take a second ranking-evaluation approach: we intervene by modifying the representation in the neurons selected by the ranking, and observe if, and how, our intervention affects the language model output. This approach is more of a causal one, inspired by similar prior work (Giulianelli et al., 2018; Elazar et al., 2021; Feder et al., 2020; Lovering et al., 2021; Ravfogel et al., 2021). We note that in this section, we use only the ranking itself, detaching it from any probes, thus removing classification quality from ranking comparisons.

However, knowing that the information is being used is not enough; we would like to know to what purpose it is being used, and to verify that it only affects the specific attribute we are interested in. Thus, we perform a finer-grained analysis, and check if D(h′)D(h^{\prime}) is similar, to some extent, to D(h)D(h). For that, we define a lemmatizer L:V→VL:\mathcal{V}\to\mathcal{V}, which maps words to their lemmas, and an analyzer A:V→ZA:\mathcal{V}\to Z, which maps words to their task labels. Our goal is to intervene such that L(D(h))=L(D(h′))L(D(h))=L(D(h^{\prime})), but A(D(h))≠A(D(h′))A(D(h))\neq A(D(h^{\prime})). For example, if we intervene for tense, we would like the word “sleeps” to become “slept”. If this is the case, it implies that we have successfully identified where the task-relevant information that DD uses is encoded, and how it is being used.

We consider two methods for modifying hπ(d)kh_{\pi(d)_{k}}, and compare them.

A common modification method is trying to remove the information by ablating some neurons (Morcos et al., 2018; Bau et al., 2019; Lakretz et al., 2019), meaning we set hπ(d)k=0h_{\pi(d)_{k}}=0. By that we aim to erase the information encoded in hπ(d)kh_{\pi(d)_{k}}.

For a word w∈Vw\in\mathcal{V} with attribute label z∈Zz\in Z, we attempt to translate its representation (in the geometric sense) to produce a word with attribute label z′∈Z,z≠z′z^{\prime}\in Z,z\neq z^{\prime} by taking a step in the direction of z′z^{\prime}, where bigger steps are applied to neurons that are marked as more important for the attribute. Formally, we apply the following protocol:

We calculate q(z)q(z) and q(z′)q(z^{\prime}) as in eq. (1).

Note that the rest of the neurons—those not in Π(d)[k]\Pi(d)_{[k]}—remain unaffected. Using this protocol, we give each neuron its own special treatment—an approach that was not applied before (as far we know). This can be seen as a generalization of Gonen et al. (2020).

2 Experimental setup

We handle the data the same way as in our probing experiments. However, since we analyze the model’s predictions—which may be different from the original input—we do not have gold morphology labels anymore. Thus, for morphologically analyzing the model’s predictions (LL and AA), we use spaCy (Honnibal et al., 2020). Out of the languages we used in our probing experiments, in this section we use only those that are supported by spaCy (English, Spanish and French). We calculate q(z)q(z) based on the entire training set, and perform our interventions on the test set. We compare the same 7 rankings we used in our probing experiments (§ 3.1).

3 Metrics

For our intervention experiments, we first measure the error rate of the language model. We want error rate to be high, since high error rate means we modified parts of the representation that have been used by the model in its prediction.

While inspecting predictions that are wrong after intervening (D(h′)≠wD(h^{\prime})\neq w, where ww is the true word), we categorize them by L(D(h′))L(D(h^{\prime})) and A(D(h′))A(D(h^{\prime})). If our intervention were successful, meaning we changed only the word’s specific attribute, and not other information, then we expect to see L(D(h))=L(D(h′))L(D(h))=L(D(h^{\prime})) and A(D(h))≠A(D(h′))A(D(h))\neq A(D(h^{\prime})); that is, correct lemma but wrong value (CLWV). For example, if the word “makes” becomes “made” when intervening for tense, then it is considered as a correct type of error, but if it becomes “make” or “prepared” it does not. Thus, we define CLWV as the portion of those errors out of all predictions.

4 Results

Across most configs, about 400 neurons from layer 2 and 200–300 neurons from layers 7 and 12 can be ablated without any implications on the output, meaning error rate remains the same; an example is shown in Appendix A.8. Moreover, when error rate does grow, CLWV is very low. By qualitatively analyzing those errors we saw that most predicted words are common words, e.g., “and”, “if” in English. After ablating 600–700 (80%–90%) neurons from the representations, we observe a lot of errors, but most of them are because the word is predicted as nonsensical punctuation. Another major concern is that in some configs, ablating by a bottom-to-top ranking provides better results than by the top-to-bottom version of the same ranking. In general, there are no distinct differences between the rankings. Thus, from here on we focus on translation rather than ablation.

4.2 Translation is effective

Across all translation experiments (Fig. 6 shows one example, more are in Appendix A.9), CLWV increases until a certain saturation point, after which it remains constant or drops a little.We define “saturation point” as the first point from which there are two consecutive points where the value increase is by a factor lower than 1.051.05. This means that we reached neurons that are not relevant for the attribute, and modifying them can result in loss of other information—error rate grows while CLWV does not. Thus, we are interested in the CLWV value at the saturation point (higher is better), and in the number of neurons modified at the saturation point (lower is better). We would also like the difference between the error rate and CLWV at the saturation point to be as small as possible. All terms considered, we perform a sweep search on the values of β\beta in the range $onadevset.Wefindthatlowon a dev set. We find that low\betavaluesprovidelowCLWV,whilehighvaluesprovidehigherCLWVbutalsowidenthegapbetweenerrorrateandCLWV.Wefindvalues provide low CLWV, while high values provide higher CLWV but also widen the gap between error rate and CLWV. We find\beta=8tobeabalancedpoint,andthusreporttestresultswithto be a balanced point, and thus report test results with\beta=8$ in three configs in Table 4.4.2, and the rest of the configs in Appendix A.9. The results for XLM-R are given in Appendix A.10.

Compared to ablation, translating a relatively small number of neurons results in a higher error rate, and these errors are closer to what we would expect. For example, translating only 50 neurons selected by Probeless in Spanish gender layer 2 results in 37%37\% CLWV and 49%49\% error rate, while ablating 50 neurons from the same config and ranking gives 0%0\% CLWV errors and only 1%1\% of error rate. We further note that unlike in ablation experiments, here our rankings are inherently consistent: across all configs, all top-to-bottom rankings perform better than the rest, while random rankings’ error rate sometimes increases a little, and bottom-to-top rankings do not manage to affect the model’s output at all (Fig. 6 is one example, more are in Appendix A.9).

4.3 Probeless is the most effective ranking for interventions

A clear trend from our M-BERT results (Table 4.4.2, Fig. 6 and Appendix A.9) is that in most cases, Probeless achieves higher CLWV values, and does so using a smaller number of neurons, than the other two rankings. Furthermore, its error rate is significantly higher than the other two. This implies that Probeless tends to select neurons that are being used by the model, more so than the other rankings. However, while it does select neurons that are relevant for the attribute in question (CLWV is relatively high), it also tends to select neurons that are used by the model for other kinds of attributes (the difference between error rate and CLWV is relatively high). Among Linear and Gaussian, Linear seems to have the upper hand, with higher CLWV values in most configs. This provides another evidence that the superiority of Gaussian in the probing experiments may be due to the quality of its classifier, and specifically its memorization ability, rather than the quality of the ranking it produces, as here only the ranking affects results. In XLM-R, it seems that Linear and Probeless both have the lead, with Gaussian falling behind (Appendix A.10). We perform additional experiments on monolingual models (Appendix A.11), where Probeless is again superior.

Discussion and conclusion

In this work, we show two pitfalls with the common approach for ranking neurons according to their importance for a morphological attribute, and compare different ranking methods that follow this approach. We show that to evaluate a ranking in a probing scenario, one should separate between the ranking itself and the quality of the classifier that is using the ranking—Pitfall I. While previous work concentrated on encoded information—Pitfall II—we show that it is not the same as information used by a model, by showing that Gaussian is inferior in the interventions scenario, in contrast to our probing results. This implies that high probing accuracy does not necessarily entail that the information is actually important for the model. This conclusion is also present in prior work (Elazar et al., 2021; Feder et al., 2020; Ravfogel et al., 2021), but it has been largely neglected in studies of individual neurons via probes. We propose a new, fast-to-use ranking method that relies solely on the data, without training any auxiliary classifier, and show that it is valid, and prefers neurons that are being used by the model, more so than other ranking methods. We also propose a method for intervening within the model’s representations such that it transforms the output in a desired way.

In our intervention experiments, modifying too many neurons results in more errors that are not related to the true word. This proves the importance of looking into individual neurons, especially when trying to intervene in the inner workings of the model. For example, Gonen et al. (2020) try to change the language of a word by intervening with the representation, using the same translation method we use, but with the same coefficient for every neuron, and on the entire representation. Our results imply that they may get better results by using our finer-grained method.

Our work contributes to the effort of improving the interpretability of language models, and more generally of neural networks. Better explanations and controls over a model’s outputs can ameliorate its fairness, for example in the case of gender bias: our intervention method can guide users on how to reduce such biases, by pointing to model components (neurons) responsible for gender and offering intervention methods to control model behavior w.r.t a particular property. On the other hand, malicious actors could use such capability to increase discrimination. Exposing the capabilities may also help develop defense mechanisms.

All our results are reproducible using the code repository we will release. All experimental details, including hyperparameters, are reported in §3.1, 4.1 and 4.2. As language models, we used the implementation of the transformers library (Wolf et al., 2020). We performed our experiments on NVIDIA RTX 2080 Ti GPU. All data preparation details are reported in §2.2 and Appendix A.2.

Acknowledgments

We thank Lucas Torroba Hennigen for his helpful comments. This research was supported by the ISRAEL SCIENCE FOUNDATION (grant No. 448/20) and by an Azrieli Foundation Early Career Faculty Fellowship. YB is supported by the Viterbi Fellowship in the Center for Computer Engineering at the Technion.

References

Appendix A Appendix

In Dalvi et al. (2019), after the probe has been trained, its weights are fed into a neuron ranking algorithm. However, we observed that the original algorithm distributes the neurons equally among labels, meaning that each label would contribute the same number of neurons at each portion of the ranking, regardless of the amount of neurons that are actually important for this label. For example, if for label A there are 10 important neurons and for label B there are only 2, then the first 10 neurons in the ranking would consist of 5 neurons for A and 5 for B, meaning that 3 non-important neurons are ranked higher than 5 important ones. Thus, we chose a different way to obtain the ranking: for each neuron, we compute the mean absolute value of the ∣Z∣|Z| weights associated with it, and sort the neurons by this value, from highest to lowest. In early experiments we found that this method empirically provides better results, and is more adapted to large label sets.

A.2 Data preparation

We remove any sentences that would have a sub-token length greater than 512512, the maximum allowed for M-BERT, the language model we use for generating representations. As in Torroba Hennigen et al. (2020), we remove attribute labels that are associated with fewer than 100100 word types in any of the data splits. This mostly removes function words, and we found it makes it harder for probes to use memorization for solving the task. The morphological attributes we experiment with include (in UniMorph annotations): Animacy, Aspect, Case, Definiteness, Gender and Noun Class, Mood, Number, Part of Speech, Person, Polarity, Possession, Tense and Voice.

A.3 Expected overlap between random rankings

Given ii rankings, we calculate the expected size of overlap between the first MM neurons across all rankings:

after initializing C=1C=1. Then, we calculate the expected number of overlapping neurons by:

since C[n,m,k](nm)i\frac{C[n,m,k]}{\binom{n}{m}^{i}} is the probability to have exactly kk overlapping neurons. We thus get E2(768,100)≈13.02E_{2}(768,100)\approx 13.02 and E3(768,100)≈1.69E_{3}(768,100)\approx 1.69.

A.4 Overlaps

Fig. 7 shows overlaps between the 100 most important neurons chosen by Linear and Gaussian for different configs. Both of them provide less overlaps then Probeless, with Gaussian having almost no overlaps at all, showing its inconsistency across languages.

Fig. 8 presents the same analysis, but for XLM-R, with Probeless ranking (equivalent to Fig. 1 for M-BERT). We see far more overlaps in XLM-R, and no red squares (describing a lower overlap size than the expected one), implying that the information is more condensed in XLM-R than in M-BERT.

A.5 Clustering probing results

For each config out of the 156 we experimented with, we have results of 14 classifier–ranking combinations, each of length 150, the max kk (number of neurons) we used. For clustering these results, we first remove all combinations involving a bottom-to-top ranking, as these add a lot of noise to the clustering algorithm, making it focus on irrelevant signals. Thus, our results matrix is of shape $.Wethenreshapethematrixtoshape. We then reshape the matrix to shape[156,8\times 150]andrunK−meansoveritwithand run K-means over it withK=3$. Projecting the K-means output with t-SNE gives us Figs. 4(d) and 10(a).

A.6 Statistical significance tests

Table 2 shows the results of our statistical significance tests. The three rows in each cell correspond to using 10, 50 and 150 neurons. If there is an * in the [i,j][i,j] cell, is means that the p-value under the null hypothesis that probe jj is better than probe ii is lower than 0.050.05, when using the matching number of neurons. For example, we see that there is an * in the first and second rows in the cell,meaningwecanconfidentlyrejectthehypothesisthatLinearbyLinearisbetterthanGaussianbyGaussianwhenusing10or50neurons,butwecannotdosofor150neurons.Infact,lookingatthecell, meaning we can confidently reject the hypothesis that Linear by Linear is better than Gaussian by Gaussian when using 10 or 50 neurons, but we cannot do so for 150 neurons. In fact, looking at the cell shows us that when using 150 neurons, Gaussian by Gaussian is not better than Linear by Linear.

While we do not show random and bottom-to-top rankings in Table 2 for clarity, we asserted that each classifier is statistically significantly better when using a top-to-bottom ranking compared to a random ranking, and when using a random ranking compared to a bottom-to-top ranking.

A.7 Probing: Additional Results

Fig. 9 complements Fig. 4, including graph lines that are missing in Fig. 4 due to its readability.

Fig. 10 shows XLM-R probing results. XLM-R provides very similar results to M-BERT, apparent in Figs. 10(a) and 10(b) compared to Figs. 4(d) and 4(c), respectively.

Selectivity examples from both models are provided in Fig. 11. In all configs, both in M-BERT and XLM-R, Linear is significantly more selective than Gaussian using any ranking.

A.8 Ablation results

One ablation example is shown in Fig. 12. No matter the ranking, ∼400\sim 400 neurons can be ablated with little impact on the output, and CLWV remains low. This behaviour is generally consistent across all configs we experimented with.

A.9 Translation results

All of M-BERT’s translation results (complementing Table 4.4.2) are found in Table 3, and examples from two configs are in Fig. 13. As described in §4.4.3, across most configs, Probeless achieves higher CLWV at the saturation point, and gets there earlier (using less neurons), than the other two rankings—in contrast to probing results. Among the probing-based rankings, Linear generally provides better results than Gaussian.

We also note that there are certain attributes that seem harder to control for, e.g., English number and French tense.

A.10 XLM-R translation results

XLM-R translation results (equivalent to Table 3 in M-BERT) are shown in Table 4. The superiority of Probeless is not so clear in XLM-R compared to M-BERT, with Linear providing good competition. Gaussian on the other hand, still falls behind.

We note that in XLM-R the CLWV values are somewhat lower compared to M-BERT. A possible explanation to that could be the difference in tokenization between the models.

A.11 Translation on Monolingual Models

As Linear and Probeless are somewhat equal on XLM-R, we perform experiments on additional models, to break the tie. We experiment with three monolingual models: bert-base-cased for English, dccuchile/bert-base-spanish-wwm-cased for Spanish, and camembert-base for French. The results are reported in Table 5. Probeless is superior in most of these experiments, both in terms of CLWV value and at number of modified neurons when reaching the saturation point, with Linear coming second.

It is worth noting that in these models, saturation point values are generally higher, and achieved using fewer neurons, than in M-BERT and XLM-R. We believe it is due to their relative simplicity compared to multilingual models, and their smaller vocabulary.