What Is One Grain of Sand in the Desert? Analyzing Individual Neurons in Deep NLP Models

Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, James Glass

Introduction

While neural networks have achieved state-of-the-art performance in NLP and other spheres of Artificial Intelligence (AI), their opaqueness remains a cause of concern (?). Interpreting the behavior of neural networks is considered important for increasing trust in AI systems, providing additional information to decision makers, and assisting ethical decision making (?).

Recent work attempted to analyze what linguistic information is captured in such models when they are trained on a downstream task like neural machine translation (NMT). A typical framework is to generate vector representations for some linguistic unit and predict a property of interest such as morphological features. This approach has also been applied for analyzing word and sentence embeddings (?; ?), and hidden states in NMT models (?; ?). The analyses reveal that neural vector representations often contain substantial amount of linguistic information. Most of this work, however, targets the whole vector representation, neglecting the individual dimensions in the embeddings. In contrast, much work in computer vision investigates properties encoded in individual neurons or filters (?; ?).

We address this gap by studying individual dimensions (neurons) in the vector representations learned by end-to-end neural models. We aim to increase model transparency by identifying specific dimensions that are responsible for particular properties. We thus strive for post-hoc decomposibility, in the sense of (?). That is, we analyze models after they have been trained, in order to uncover the importance of their individual parameters. This kind of analysis is important for improving understanding of the inner workings of neural networks. It also has potential applications in model distillation (e.g., by removing unimportant neurons), neural architecture search (by guiding the search with important neurons), and mitigating model bias (by identifying neurons responsible for sensitive attributes like gender, race or politenessE.g., controlling the system to generate outputs with the right honorifics (“Sie” vs. “du”) in German.). In this work we lay out a methodology for identifying and analyzing individual neurons, and open the call to explore such use cases to the research community.

To this end, we propose two methods to facilitate neuron analysis. First, we perform an extrinsic correlation analysis through supervised classification on a number of linguistic properties that are deemed important for the task (for example, learning word morphology lies at the heart of modeling various NLP problems). Our classifier extracts important individual (or groups of) neurons that capture certain properties. We call this method Linguistic Correlation Analysis. Second, we propose an alternative methodology to search for neurons that share similar patterns in independently trained networks, based on the assumption that important properties are captured in multiple networks by individual neurons. We call this method Cross-model Correlation Analysis. Such an analysis is more intrinsic and helpful for highlighting important neurons for the model itself, and in the case when annotated data (supervision) may not be available. Both machine translation and language modeling are fundamental AI tasks that have seen tremendous improvements with neural networks in recent years. We evaluated our methods for analyzing neurons on these two tasks.

We provide quantitative evidence that our rankings are correct by performing several ablation experiments: from masking out important neurons to removing them completely from the training. We then conduct a comprehensive analysis of the ranked neurons. Our analysis reveals interesting findings such as i) open class categories such as verb (part-of-speech tag) and location (semantic entity) are much more distributed across the network compared to closed class categories such as coordinating conjunction (e.g., “but/and”) or a determiner (e.g., “the”), ii) the model recognizes a hierarchy of linguistic properties and distributes neurons based on it, and iii) important neurons extracted from the Cross-model Correlation method overlap with those extracted from the Linguistic Correlation method; for example, both methods identified the same neurons capturing position as salient. In summary, we make the following contributions:

A general methodology for identifying linguistically-meaningful neurons in deep NLP models.

An unsupervised method for finding important neurons in neural networks, and a quantitative evaluation of the retrieved neurons.

Application to various test cases, investigating core language properties through part-of-speech (POS), morphological, and semantic tagging.

An analysis of distributed vs. focused information in NMT and NLM models.

Related Work

Much of the previous work has looked into neural models from the perspective of what they learn about various language properties. This includes analyzing word and sentence embeddings (?; ?; ?), recurrent neural network (RNN) states (?; ?), and NMT representations (?; ?; ?). The language properties mainly analyzed are morphological (?; ?), semantic (?) and syntactic (?; ?; ?).

Most of this work used an extrinsic supervised task and target entire vector representations. We study the individual neurons in the vector representation and propose a simple supervised method to analyze individual/groups of neurons with respect to various properties and linguistic tasks. As an alternative to supervision which is limited to labeled data, we propose an unsupervised method based on correlation between several networks to identify salient neurons.

Some recent work on neural language models and machine translation analyzes specific neurons of length (?; ?) and sentiment (?). However, not much work has been done along these lines. We present both intrinsic and extrinsic methods to analyze models at the neuron level to gain a deeper insight.

In computer vision, there has been much work on visualizing and analyzing individual units such as filters in convolutional neural networks (?; ?, among others). Even though some doubts were cast on the importance of individual units (?), recent work stressed their contribution to predicting specific object classes via ablation studies similar to the ones we conduct (?).

Methodology

We then train a logistic regression classifier on the {zi,li}\{\mathbf{z}_{i},\mathbf{l}_{i}\} pairs using the cross-entropy loss. We opt to train a linear model because of its explanability; the learned weights can be queried directly to get a measure of the importance of each neuron in zi\mathbf{z}_{i}. From a performance point of view, earlier work has also shown that non-linear models present similar trends as of linear models in analyzing representations of neural models (?; ?). In order to increase interpretability and to encourage feature ranking in the classification process, we use elastic net regularization (?) as an additional loss term. Formally, the model is trained by minimizing the following loss function:

Elastic net regularization enjoys the sparsity effect as in Lasso regularization, which helps identify important individual neurons. At the same time, it takes groups of highly correlated features into account similar to Ridge regularization, avoiding the selection of only one feature as in Lasso regularization. This strikes a good balance between localization and distributivity. This is particularly useful in the case of analyzing neural networks where we hypothesize that the network consists of both individual focused neurons and a group of distributed neurons, depending on the property being learned. The regularization terms are controlled by hyper-parameters λ1\lambda_{1} and λ2\lambda_{2}. We search for the best hyper-parameter values that maintain good accuracy while accomplishing the desired goal of selecting the salient neurons for a property, as described in the evaluation section.

Cross-model Correlation Analysis

Evaluation using Neuron Ablation

Given a trained classification model, we keep N% top or bottom neurons and set the activation values of all other neurons to zero in the test set. We then reevaluate the performance of the already trained classifier. We expect to see low performance (prediction accuracy) when using only the bottom neurons versus using only the top neurons. We also retrain the classifier with only the selected N% neurons. This serves multiple purposes: i) it confirms the results from the zeroing-out method, ii) it shows that much of the performance can be regained using the selected neurons, and iii) it facilitates the analysis of how distributed a particular property is across the network.

Experimental Settings

We experimented with two architectures: NMT based on sequence-to-sequence learning with attention (?) and an LSTM based NLM (?).We focus on standard architectures for these tasks and leave exploration of recent variants such as the Transformer (?) or QRNN (?) for future work. We trained a 2-layer bidirectional NMT model with 500-dimensional word embeddings and LSTM states. The system is trained for 20 epochs, and the model with the best development loss is used for the experiments. We follow similar settings to train a unidirectional NLM model.

We experimented with English↔\leftrightarrowFrench (EN↔\leftrightarrowFR) and German→\rightarrowEnglish (DE→\rightarrowEN) language pairs. We used a subset of 2 million sentences from the United Nations multi-parallel corpus (?) for EN↔\leftrightarrowFR and from the data made available for the IWSLT campaign (?) for DE→\rightarrowEN. We split the parallel data for each language pair into three equal subsets to train three different models. For language models, we used the source side of the parallel corpora.

We evaluated our linguistic correlation method by selecting standard tasks of part-of-speech (POS), morphological and semantic tagging. The former two capture word structure in a language and the latter captures its nuanced meaning. Additionally we considered some general properties, such as the position of words in a sentence and predicting a months of year tag.

We used 20k source-side sentences, randomly extracted from the MT training data, for training the classifier, and 4k sentences in the official test sets for testing. We tagged these sentences with standard taggers for the different properties; the details of these taggers can be found in the supplementary material.

Evaluation

In this section, we present the evaluation of our techniques:

We first evaluate the classifier performance to ensure that the learned weights are actually meaningful for further analysis and ranking extraction. The classifiers were trained using the activations of already trained neural models (NLM and NMT encoderWe limit ourselves to encoder activations for simplicity.). Table 1 shows accuracy of the classifiers trained for different language pairs and tasks on a blind test set. The classifiers achieve higher accuracies compared to the local majority baselineSelecting the most frequent tag for each word and the most frequent global tag for the unknown words. (MAJ) in all cases, except for French (POS:NLM). The overall accuracy trend shows that the neurons possess sufficient information to predict these language properties.

Since we are using elastic net regularization, we need to tune the values for λ1\lambda_{1} and λ2\lambda_{2}. The regularization controls the final ranking of the neurons directly: an increase in the value of λ1\lambda_{1} introduces further sparsity whereas higher values of λ2\lambda_{2} encourage selection of groups of correlated neurons. Our aim is to find a balance between selecting individual neurons and a group of neurons while maintaining the original accuracy of the classifier without any regularization (λ1\lambda_{1}, λ2\lambda_{2} =0=0). Figure 2 presents the results of a grid search over various regularization values on the English POS tagging task. The accuracy difference is minimal for λ\lambda values under 1e−41e^{-4}. We selected a value of 1e−51e^{-5} for both λ1\lambda_{1} and λ2\lambda_{2} and used the same for all the experiments.

Neuron Ablation in the Classifier:

After training the classifier, we used Algorithm 1 to extract a ranked list of neurons with respect to each property set and ablated neurons in the classifier to verify rankings. We masked-out all the activations (in the test set) except for the selected N%N\% neurons and recomputed test accuracies. Table 2 summarizes the results.Similar trends were found in the morphological tagging results. Please see supplementary material if interested. Compared to ALL, the classification accuracy drops drastically for both NMT and NLM. However, the performance is distinctly better in the case of keeping the top N% neurons when compared to the bottom N% neurons, showing that the ranking produced by the classifier is correct for the task at-hand.

have been used effectively to gain qualitative insights on analyzing neural networks (?; ?). We used an in-house visualization tool (?) for qualitative evaluation of our rankings. Figure 3 visualizes the activations of the top neurons for a few properties. It shows how single neurons can focus on very specific linguistic properties like verb or article. Neuron #1902 focuses on two types of verbs (3rd person singular present-tense and past-tense) where it activates with a high positive value for the former (“Supports”) and high negative value for the latter (“misappropriated”). In the second example, the neuron is focused on German articles. Although our results are focused on linguistic tasks, the methodology is general for any property for which supervision can be created by labeling the data. For instance, we trained a classifier to predict position of the word, i.e., identify if a given word is at the beginning, middle, or end of the sentence. As shown in Figure 3(a), the top neuron identified by this classifier activates with high negative value at the beginning (red), moves to zero in the middle (white), and gets a high positive value at the end of the sentence (blue). Another way to visualize is to look at the top words that activate a given neuron. Table 3 shows a few examples of neurons with their respective top 10 words. Neuron #1925 is focused on the name of months. Neuron #1960 is learning negation and Neuron #1590 activates when a word is a number. These word lists give us quick insights into the property the neuron has learned to focus on, and allows us to interpret arbitrary neurons in a given network.

Cross-model Correlation Analysis

We incrementally ablate top/bottom neurons from the ranking and report the drop in performance of the NMT model. Figure 4 shows the effect of ablation on translation quality (BLEU). For all languages, ablating neurons from top to bottom (solid curves) causes a significant early drop in performance compared to ablating neurons in the reverse order (dotted curves). This validates the ranking identified by our method. Ablating just the top 50 neurons (2.5%) leads to drops of 15-20 BLEU points, while the bottom 50 neurons hurt the performance by only 0.5 BLEU points.

Figure 5 presents the results of ablating neurons of NLM in the order defined by the Cross-model Correlation Analysis method. The trend found in the NMT results is also observed here, i.e. the increase in perplexity (degradation in language model quality) is significantly higher when erasing the top neurons (solid lines) as compared to when ablating the bottom neurons (dotted lines).

Recall that our Cross-model method requires multiple instances of the model to extract neuron rankings. In an effort to probe whether one instance of the model can sufficiently extract similar rankings, we tried several methods that ranked neurons of an individual model based on i) variance, and ii) distance from mean (high to low), and compared these with the ranking produced by our method. We found less than 10% overlap among the top 50 neurons of the Cross-model ranking and the single model rankings. On ablating the neurons based on several ranking methods, we found the NMT models to be most sensitive to the Cross-model ranking. Less damage was done when neurons were ablated using rankings based on variance and distance from mean in both directions, high-to-low and low-to-high (See Figure 7). This supports our claim that the Cross-model ranking identifies the most salient neurons of the model.

Are the neurons discovered by the linguistic correlation method important for the actual model as well? Figure 8 shows the effect on translation when ablating neurons in ranking order determined by English POS and semantic (SEM) tagging, as well as top/bottom Cross-model orderings. As expected, the linguistic correlation rankings are limited to the auxiliary task and may not result in the most salient neurons for the actual task (machine translation in this case); ablating according to task-specific ordering hurts less than ablating by (top-to-bottom) Cross-model ordering. However, in both cases, degradation in translation quality is worse than ablating by bottom-to-top Cross-model ordering. Comparing SEM with POS, it turns out that NMT is slightly more sensitive to neurons focused on semantics than POS.

Analysis and Discussion

The rankings produced by the linguistic correlation and cross-correlation analysis methods give a sense of the most important neurons for an auxiliary task or the overall model. We now dive into neuron analysis based on these rankings.

Recall that our linguistic-correlation method provides an overall ranking w.r.t. a property set (POS/SEM tagging), and also for each individual property as described in the Methodology section. Here, we look at the number of salient neurons (extracted from the NMT models) for several different linguistic properties,We choose salient neurons for each label by selecting the top neurons that cumulatively represent 25% of the total weight mass. as shown in Figure 6. For example, in open-class categories such as nouns (NN/NOM), verbs (VB/VER.simp/VVPP) and adjectives (JJ/ADJ), the information is distributed across several dozen neurons. In comparison, categories such as end of sentence marker (SENT) or WH-Adverbs (WRB) and post-positions (APPO in German) required fewer than 10 neurons. We observed similar trend in the semantic tags: information about closed-class categories such as months of year (MOY) is localized in just a couple of neurons. In contrast, an open category like location (LOC) is very distributed.

Since some information is distributed across the network, we expect to see some neurons that are common across various properties, and others that are unique to certain properties. To investigate this, we intersect top ranked neurons coming from two different properties. Some of these comparisons are interesting. For instance, we found some common neurons across all forms of adjectives, but some neurons specifically designated to specialized adjectives (e.g., comparative (JJR) and superlative (JJS) adjectives). Similarly across tasks (POS vs. Morph), we found multiple neurons targeting different verb forms (V--F3s and V--F3p , Verb Future 3rd person singular and plural) in the fine-grained morphological tagging that are aligned with a single neuron targeting the future tense verb tag (VER:futu) in POS tagging. This demonstrates that model recognizes a hierarchy of linguistic properties and distributes neurons based on it.

In the evaluation section for our linguistic-correlation classifier, we masked-out a majority of the neurons and compared the accuracy trends to confirm our ranking. An alternative to analyze is to retrain the classifier with the top or bottom N% neurons alone. Table 4 shows the results after retraining. There are several points to note here: i) training the classifier using top neurons performs consistently better than using bottom neurons, reinforcing our previous finding. ii) The classifier is able to regain performance substantially (compared to ALL), even using only 10% neurons. iii) Using the bottom N% neurons also restores performance (although not as much as using the top neurons). This shows that the information is distributed across neurons. However, the distribution is not uniform, which results in a large difference between training using top and bottom neurons (i.e., the information distribution is skewed towards the top neurons as expected). Notably, using only 20% of the top neurons, the classifier is able to regain much of the performance drop in most of the cases. This finding entails that our method could be useful for model distillation purposes.

Analyzing the top neurons identified by our Cross-model correlation method, we found several neurons corresponding to the position of the word in a sentence. Word position has been previously found to be an important property in NMT (?). The fact that our method ranks position neurons among the top ranking neurons shows its efficacy. We also observed that the top position neurons identified by our Linguistic Correlation method are the same as identified by the Cross-model correlation method. Lastly, we found that some of the remaining top Cross-model neurons correspond to fundamental structural properties in a sentence, like relations, conjunctions, determiners and punctuations.

There is substantially a large performance difference between top and bottom neurons (Refer to Table 4). For example, averaged over all properties, the top 10% NMT neurons are 12.8% (absolute) better accuracy than the bottom 10% neurons, while the top 10% NLM neurons are 25.5% better than the bottom 10% neurons. We speculate that NMT model distributes the information more, compared to the NLM model. However, this could be an artifact of the difference in the architecture of NLM (unidirectional) and NMT (bidirectional).

Conclusion and Future Work

We proposed two methods to extract salient neurons from a neural model with respect to an extrinsic task or the model itself. We demonstrated the accuracy of our rankings by performing a series of ablation experiments. Our Cross-model Correlation method can potentially facilitate research on model distillation and neural architecture search, as it pinpoints what is especially important for the model. Our Linguistic Correlation method is primarily focused on trying to understand specific dimensions that are responsible for learning particular properties. This can be helpful for understanding and manipulating systems’ behavior. In some preliminary experiments, we were able to successfully manipulate verb tense neurons and control whether the system generates output in present or past tense. Some details are presented in (?). The source code for extraction and analysis of salient neurons is incorporated in the NeuroX toolkit (?) and is available on git.https://github.com/fdalvi/NeuroX

Acknowledgments

We thank Preslav Nakov and the anonymous reviewers for their useful suggestions on an earlier draft of this paper. This work was funded by Qatar Computing Research Institute, HBKU as part of the collaboration with the MIT Computer Science and Artificial Intelligence Laboratory (CSAIL).

Supplementary Material

We annotated the date using Tree-Tagger for French POS tags, LoPar for German POS and morphological tags, and MXPOST for English POS tags. For the semantic (SEM) tagging task, we experiment with the lexical semantic task introduced by (?).The annotated data is limited to English language only. We split the available annotated data into 42k sentences for training and 12k sentences for testing.

Table 5 shows the results for the classifier performance when masking out neurons for morphological tags. Table 6 shows the results when the classifier is retrained with N% of the neurons.

References