Editing Factual Knowledge in Language Models

Nicola De Cao, Wilker Aziz, Ivan Titov

Introduction

Using pre-trained transformer-based Language Models (LMs; Vaswani et al., 2017; Devlin et al., 2019; Radford et al., 2019; Lewis et al., 2020; Raffel et al., 2020; Brown et al., 2020) has recently become a standard practice in NLP. Factual knowledge induced during pre-training can help in downstream tasks, but it can also be incorrect or become obsolete over time (e.g., not reflecting changes of heads of states or country populations). Developing reliable and computationally efficient methods for bug-fixing models without the need for expensive re-training would be beneficial. See Figure 2 for an example of revising the memory of a model that initially misremembered Namibia’s capital.

Unlike conventional Knowledge Bases (KBs) that explicitly store factual knowledge, neural models implicitly memorize facts in their parameters. One cannot easily access and interpret their computation and memories (Ribeiro et al., 2016; Belinkov and Glass, 2019; Voita et al., 2019; De Cao et al., 2020), thus, modifying their knowledge is a challenging problem. Motivated by practical considerations, we formulate the following desiderata for a method aimed at tackling this problem (see Section 2 for a more formal treatment):

Generality: be able to modify a model that was not specifically trained to be editable (i.e., no need for special pre-training of LMs, such as using meta-learning);

Reliability: be able to successfully update a specific fact without affecting the rest of the acquired knowledge;

Consistency: the changes should be consistent across equivalent formulations of a fact (e.g., when asked to update an answer for one question, answers to its paraphrases should change accordingly).

The problem has been previously tackled in Zhu et al. (2020) and Sinitsin et al. (2020), as discussed in detail in Section 3. However, both do not ensure that the edited model will be ‘reliable’, i.e. that the rest of the knowledge would not be badly affected, and that the changes are ‘consistent’ across equivalent inputs. Additionally, Sinitsin et al.’s (2020) method requires expensive specialized training of the original network. While re-training the original network was feasible in their applications (e.g., in machine translation), it is problematic when the network is a pre-trained LM. We propose a novel method that overcomes these limitations.

We treat editing the memories of a neural model as a learning-to-update problem. We use an efficient parameterization of a hyper-network that is trained to update the LM parameters when provided with a single fact that needs to be modified. We do not require meta-learning, re-training or fine-tuning of the original network. We employ constrained optimization in training: we constrain the edited model to retain the same predictions as the original one regardless of the distance between the original and updated models in the parameter space. We show how this framework can be extended to incorporate (e.g., automatically-generated) paraphrases in training, further improving consistency. Figure 1 shows an outline of our method.

Differently from both previous methods, we do not have to select a subset of parameters to update as we let our model learn that by itself. In fact, our hyper-network can be regarded as a ‘probe’ revealing which components of the network need to be changed to manipulate factual knowledge, i.e. revealing the ‘causal mediation mechanisms’ Vig et al. (2020). We observe that the updates end up being concentrated in a restricted set of model components, even though we do not encourage any kind of sparsity. Interestingly, the most-updated components are different from the groups of parameters receiving large gradients (see Figure 4).

we define the task of knowledge editing and propose a set of evaluation metrics;

we propose KnowledgeEditor that learns to modify LMs memories efficiently and reliably while maintaining consistent predictions for semantically equivalent inputs;

we verify that our proposed method largely meets our desiderata—while other baselines based on fine-tuning fail—testing it with different LM architectures on knowledge-intensive tasks such as fact-checking and open-domain question answering;

we analyze the updates for KnowledgeEditor and the alternatives.

Task

We want to edit the memory of a neural language model such that when, presented with an input, its output reflects a revised collection of facts. Unfortunately, the knowledge of a language model is typically opaque to us, being stored non-locally across a large number of parameters and architectural components. Thus, concretely, to operationalize the task, we seek a change in the model’s parameters that affects predictions from the model only for a specific input. For a given input xx, the prediction aa made by the edited model should differ from the prediction yy made by the original model only if xx is influenced by one of the revised facts.

More formally, we have a model x↦f(x;θ)x\mapsto f(x;\theta) with trained parameters θ\theta, and a dataset of revisions ⟨x,y,a⟩∈D\langle x,y,a\rangle\in{\mathcal{D}}, i.e., xx is an input, yy is the prediction preferred by f(x;θ)f(x;\theta), and aa is an alternative prediction which we would like an edited version of the model to prefer. Concretely, we keep the model architecture ff fixed, and seek alternative parameters θ′\theta^{\prime} such that for xx, f(x;θ′)f(x;\theta^{\prime}) would prefer the prediction aa instead of yy while keeping all other predictions unchanged. In practice, we approximate the set of ‘all other predictions’ using a finite data set Ox{\mathcal{O}}^{x} of pairs ⟨x′,y′⟩\langle x^{\prime},y^{\prime}\rangle with x′≠xx^{\prime}\neq x. Moreover, predictions need not be continuous nor differentiable outputs from the model; instead, they may result from an arbitrary decision rule based on f(x;θ)f(x;\theta). For example, when f(x;θ)f(x;\theta) parameterizes a discrete distribution pY∣Xp_{Y|X} over the output space, the most standard decision rule is to output the mode of the distribution: y=arg⁡max⁡c∈Y pY∣X(c∣x,θ)y=\arg\max_{c\in{\mathcal{Y}}}~{}p_{Y|X}(c|x,\theta).Whereas in text classification solving this is straightforward (for Y{\mathcal{Y}} is small), in sequence-to-sequence we resort to beam search to approximate the mode (for Y{\mathcal{Y}} is too large or unbounded).

Optionally, for some revision ⟨x,y,a⟩∈D\langle x,y,a\rangle\in{\mathcal{D}}, we may also have a set Px{\mathcal{P}}^{x} of inputs semantically equivalent to xx (e.g., automatically-generated paraphrases). Such a set can be used in at least two ways: i) to obtain explicit supervision for changes that should be realized in tandem with ⟨x,y,a⟩\langle x,y,a\rangle; and, independently of that, ii) to evaluate whether an edited model makes consistent predictions on semantically equivalent inputs. Note that in this work we never use paraphrases at test time, only for training and evaluation of our approach; generating them at test time, while potentially helpful, would have compromised efficiency.

2 Evaluation

To test if a method gg, producing edited parameters θ′\theta^{\prime}, meets our desiderata, we measure:

success rate: how much gg successfully updates the knowledge in θ′\theta^{\prime}, measured as accuracy of revised predictions for inputs in D\mathcal{D};

retain accuracy: how well θ′\theta^{\prime} retains the original predictions of ff, measured as accuracy wrt input-output pairs in sets Ox{\mathcal{O}}^{x};

equivalence accuracy: how consistent the predictions of the revised model θ′\theta^{\prime} are for semantically equivalent inputs, measured as accuracy of the revised predictions for all Px{\mathcal{P}}^{x};

performance deterioration: how much test performance of the updated model deteriorates.1−accuracy of f(⋅;θ′)accuracy of f(⋅;θ)1-\frac{\text{accuracy of }f(\cdot;\theta^{\prime})}{\text{accuracy of }f(\cdot;\theta)}

These values are obtained by comparing predictions of f(⋅;θ)f(\cdot;\theta) and f(⋅;θ′)f(\cdot;\theta^{\prime}) for different subsets of inputs (e.g., D{\mathcal{D}}, Ox{\mathcal{O}}^{x}, Px{\mathcal{P}}^{x}) and against different targets (e.g., gold-standard, original predictions, or alternative predictions). While these metrics are straightforward to compute in principle, some can be computationally demanding. For example, retain accuracy depends on predictions for all inputs we have access to, which is potentially the entirety of the downstream task’s validation/test data.During training of gg, however, we can use sub-sampling (i.e., mini batches) to approximate the metric.

Previous work has evaluated similar versions of this task differently. Sinitsin et al. (2020) measure performance deterioration and success rate but do not measure retain accuracy nor equivalence accuracy. A small performance deterioration does not guarantee high equivalence accuracy as the former is sensitive to changes in cases where the original model makes wrong decisions. Assessing accuracy against old or revised facts, which Zhu et al. (2020) also do, does not help to measure the retain accuracy. We argue that preserving model predictions for inputs not in D\mathcal{D} is critical in production settings, where model predictions might have been extensively analyzed and tested. For x′∉Dx^{\prime}\not\in\mathcal{D}, we aim to maintain all original predictions as well as the model scores f(x′;θ′)f(x^{\prime};\theta^{\prime}) itself, effectively avoiding the need to re-calibrate the models (for example, in applications where probability estimates are used downstream).

Related work

The most straightforward strategy to edit the knowledge of a model would be to re-train it on a new dataset with additional, modified, or removed facts. This is often unfeasible as LMs require large-scale expensive training that can hardly be reproduced by the most. Sinitsin et al. (2020) propose a meta-learning approach (Finn et al., 2017) for model modification that learns parameters that are easily editable at test time (e.g., updating the knowledge of the model requires only a few SGD steps from these learned parameters). To have a reliable method, they employ a regularized objective forcing the updated model not to deviate from the original one. This technique suffers from three main limitations: i) it requires expensive and specialized pre-training, ii) it is sensitive to many hyper-parameters (e.g., the weights of the regularizers and the subset of parameters to update), and iii) their multitask objective does not guarantee reliability (i.e., the model is penalized for diverging from the original, rather than constrained not to).

Instead of penalizing an updated model for deviating from the original one, Zhu et al. (2020) use constrained optimization. They use a less computationally expensive procedure as they re-fine-tune on a specific downstream task (with altered data). Their method employs either an L2L_{2} or L∞L_{\infty} constraint between the original model’s parameters and the edited ones. However, a norm-based constraint on parameters ignores the highly non-linear nature of LMs and how parameters determine the outputs of the model. Indeed, a minimal change in parameter space may produce a completely different output for many datapoints leading to a potentially unreliable method. Additionally, they show the need to select a subset of parameters to be updated, which requires extra development effort. Zhu et al.’s (2020) method is similar to Elastic Weight Consolidation (Kirkpatrick et al., 2017), a technique developed for preventing catastrophic forgetting in neural network models.

Knowledge in Language Models

Petroni et al. (2019) show that pre-trained language models recall factual knowledge without fine-tuning, which they do by feeding specific prompts to LMs. Hand-crafted prompts have been found not to be the best option to extract knowledge from LMs, and various solutions have been proposed to understand what LMs ‘know’ (Jiang et al., 2020; Shin et al., 2020; Liu et al., 2021). Additionally, Roberts et al. (2020) show that large models can be fine-tuned to access their internal memories to answer questions in natural language without any additional context and with surprisingly high accuracy—a setting they referred to as closed-book question answering. Although performing quite well, these models cannot reach the prediction quality of alternatives that retrieve and use context. Approaches that incentivize memorization of factual knowledge show to be beneficial for many downstream tasks suggesting that research on methods that effectively edit the memory of a model is indeed important (Zhang et al., 2019; Sun et al., 2019, 2020). Some recent hybrid approaches that use both implicit and explicit memory show some benefits for question answering (Févry et al., 2020; Verga et al., 2020). Notably, language models that only rely on internal implicit memory are state-of-the-art for (multilingual-) Entity Linking (De Cao et al., 2021a, b). An effective mechanism for editing LM’s implicit memory may be applicable in all these settings.

Causal Interventions

Identification of minimal changes to neural networks needed to achieve a certain behaviour has been studied in the context of research in interpreting neural networks Lakretz et al. (2019); Vig et al. (2020); Elazar et al. (2021); Csordás et al. (2021). The components which need to be updated can be interpreted as controlling or encoding the corresponding phenomena (e.g., subject-verb agreement). Much of this research focused on modifying neuron activations rather than weights and on sparse interventions (e.g., modifying one or a handful of neurons). While far from our goals, there are interesting connections with our work. For example, our analysis of updates in Section 6.4, though very limited, may shed some light on how factual knowledge is encoded in the parameters of a model.

Method

We propose to treat the task of editing the memory of a neural model as a learning problem. Instead of defining a handcrafted algorithm to compute the new parameters θ′\theta^{\prime}, we learn a KnowledgeEditor: a model that predicts θ′\theta^{\prime} conditioned on an atomic fact that we want to modify. Concretely, KnowledgeEditor is a hyper-network (Ha et al., 2017)—i.e., a neural network that predicts the parameters of another network. Since the task requires every other prediction to stay the same—except the one we desire to change—we cast the learning task as a constrained optimization problem.

For an input xx, changing the prediction of a model f(⋅;θ)f(\cdot;\theta) to aa corresponds to minimizing the loss L(θ;x,a){\mathcal{L}}(\theta;x,a) incurred when aa is the target. Preserving the rest of the knowledge corresponds to constraining the updated parameter θ′\theta^{\prime} such that model outputs f(⋅;θ′)f(\cdot;\theta^{\prime}) do not change for x′∈Oxx^{\prime}\in{\mathcal{O}}^{x}. Our editor gg is a neural network parameterized by ϕ\phi which we choose by optimising the following objective for each data-point ⟨x,y,a⟩∈D\langle x,y,a\rangle\in{\mathcal{D}}:

The constraint pushes the updated model to predict output distributions identical to the original one for all x′≠xx^{\prime}\neq x. An alternative constraint we could employ is an LpL_{p} norm over the parameter updates such that gg is optimized to make a minimal update to the original model parameter: CLp(θ,θ′,f;Ox)=(∑i∣θi−θi′∣p)1/p{\mathcal{C}}_{L_{p}}(\theta,\theta^{\prime},f;{\mathcal{O}}^{x})=\left(\sum_{i}|\theta_{i}-\theta_{i}^{\prime}|^{p}\right)^{1/p}. This constraint was previously used by Zhu et al. (2020). However, such a constraint, expressed purely in parameter space and without regards to the model architecture ff, does not directly encourage model outputs to be close to original ones in function space (i.e., the two functions to be similar). Neural models are highly non-linear functions, so we do not expect this type of constraint to be effective. This will be empirically demonstrated in Section 6.

Tractable approximations

Non-linear constrained optimization is generally intractable, thus we employ Lagrangian relaxation (Boyd et al., 2004) instead. The constraint itself poses a computational challenge, as it requires assessing KL for all datapoints in the dataset at each training step. For tractability, we evaluate the constraint approximately via Monte Carlo (MC) sampling (see Appendix A for more details). Finally, in sequence-to-sequence models, assessing KL is intractable even for a single data point, as the sample space Y{\mathcal{Y}} is unbounded. In such cases we approximate the computation on a subset of the sample space obtained via beam search.

Architecture

Instead of predicting θ′\theta^{\prime} directly, our hyper-network predicts a shift Δθ\Delta\theta such that θ′=θ+Δθ\theta^{\prime}=\theta+\Delta\theta. A naive hyper-network implementation might be over-parameterized, as it requires a quadratic number of parameters with respect to the size of the target network. Thus, we apply a trick similar to Krueger et al. (2017) to make gg tractably predict edits for modern large deep neural networks (e.g., BERT). Namely, gg makes use of the gradient information ∇θL(θ;x,a)\nabla_{\theta}{\mathcal{L}}(\theta;x,a) as it carries rich information about how ff accesses the knowledge stored in θ\theta (i.e., which parameters to update to increase the model likelihood given aa).A version of our hyper-network that does not use gradient information converges far too slowly.

where σ\sigma is the Sigmoid function (i.e., x↦(1+exp⁡(−x))−1x\mapsto(1+\exp(-x))^{-1}), and σ^\hat{\sigma} indicates the Softmax function (i.e., x↦exp⁡(x)/∑iexp⁡(xi)x\mapsto\exp(x)/\sum_{i}\exp(x_{i})). With this formulation, the parameters for the hyper-network ϕ\phi scale linearly with the size of θ\theta. An interpretation of Equation 3 is that an update ΔW\Delta W is a gated sum of a scaled gradient of the objective and a bias term. The scale for the gradient and the bias are generated via an outer vector product as it allows for efficient parameterization of a matrix with just three vectors. The gate lets the model keep some parameters unchanged.

Margin annealing

The margin mm is a hyperparameter and therefore fixed. However, i) it is hard to choose since it is task-dependent, and ii) it should be as small as possible. If the margin is too small, however, we risk having a small feasible set, and the model may never converge. To address both issues, we pick some initial value for the margin and anneal it during training conditioned on validation performance: when the model successfully changes >90%>90\% of the predictions, we multiply the margin by 0.80.8. We stop decreasing the margin once it reaches a desirable small value. The annealing procedure prevents the model from diverging while increasingly tightening the constraint.

Experimental Setting

We aim to evaluate the effectiveness of KnowledgeEditor comparing to baselines on knowledge-intensive tasks where the importance of modifying the memory of a large LM has a broad impact. We then test our method on closed-book fact-checking and closed-book question answering with the metrics proposed in Section 2.2.

We compare against two baselines: i) fine-tuning and ii) the method proposed by Zhu et al. (2020). Fine-tuning corresponds to using standard gradient descent, minimizing the loss for the fact/prediction we want to revise. For this, we follow Sinitsin et al. (2020) and employ RMSProp (Tieleman and Hinton, 2012).We tried alternatives, RMSProp was the most effective. We set the learning rate to 10−510^{-5} and stop upon successfully changing the output of the model or having reached a maximum of 100100 gradient steps. Zhu et al.’s (2020) method extends fine-tuning with an L∞L_{\infty} constraint on parameters. We search the hyper-parameter for the penalty m∈{10−3,5×10−4,10−4,5×10−5,10−5}m\in\{10^{-3},5\times 10^{-4},10^{-4},5\times 10^{-5},10^{-5}\} selecting the best based on the sum of success rate and retain accuracy. Following both Sinitsin et al. (2020) and Zhu et al. (2020) we report these baselines fine-tuning all parameters or just a subset of them. We limit the search to selecting entire layers and base our decision on performance on a subset of the validation set. Note that selecting a subset of parameters for update requires an extensive search, which KnowledgeEditor dispenses with by automatically learning it.

2 Models and data

We evaluate on closed-book fact-checking (FC) fine-tune a BERT base model (Devlin et al., 2019) on the binary FEVER dataset (Thorne et al., 2018) from KILT (Petroni et al., 2021). We also evaluate on a task with a more complex output space: closed-book question answering (QA). For that we fine-tune a BART base model (Lewis et al., 2020) with a standard seq2seq objective on the Zero-Shot Relation Extraction (zsRE) dataset by Levy et al. (2017). We evaluate on this dataset because it is annotated with human-generated question paraphrases that we can use to measure our model’s robustness to semantically equivalent inputs. We create alternative predictions for FC simply flipping the labels, whereas for QA we pick all hypotheses enumerated via beam search except the top-1. The latter ensures high-probability outcomes under the model distribution. We generate semantically equivalent inputs with back-translation. See Appendix B for technical details on models and data collection.

Results

Table 1 reports the main results for fact-checking and question answering. Overall, KnowledgeEditor achieves high performance in all metrics. Some other methods also achieve high accuracy in some metrics but always sacrificing others (i.e., never meeting all our desiderata at once).

We compare methods along different metrics (as opposed to a single one), as there is no way to precisely determine the importance of each of these metrics. To gather more insight, we compute their stochastic convex combination with coefficients sampled from a Dirichlet distribution (with α=1\alpha=1 to ensure a very diverse set of combinations) and report in Figure 6 in Appendix C an estimate of the probability that a system outperforms another across 1,0001,000 such combinations. The probability of our full method to outperform all baselines is very high for both FC and QA (≈ ⁣97%\approx\!97\% and ≈ ⁣88%\approx\!88\%, respectively). In Figure 5 in Appendix C, we show the distributions of the combined scores (i.e., the raw data for the approximation reported in Figure 6). We then analyze different aspects of our method and the baselines.

Every method achieves an almost perfect success rate on fact-checking. All methods but ours apply updates in a loop, stopping either when the new model is successfully updated or after reaching a maximum number of iterations. The success rate for KnowledgeEditor is not 100%100\% because we do not apply more than one update even in case of failure. To this end, we also show an experiment with our method with multiple updates within a loop employing the same stopping criteria as the baselines. Note that we apply this only at test time (i.e., we do not train for multiple updates). When applying multiple updates also our method reaches a 100%100\% success rate on fact-checking and almost perfect accuracy (>99%>99\%) for QA.Even if we do not train for multiple subsequent updates, its success opens the possibility to add this at training time. We leave the exploration of this technique to future work.

Closed-book QA is a more challenging task since the output space is text and not just a binary label. In this setting, KnowledgeEditor achieves high accuracy (≈ ⁣95%\approx\!95\% or >99%>99\% with the loop). Among all methods, KnowledgeEditor gets the best success rate while also obtaining the best retain accuracy. In QA, Zhu et al.’s (2020) method does not reach a good success rate (≈ ⁣80%\approx\!80\%). We searched hyperparameters for their method also to have high retain accuracy, and indeed that is higher than regular fine-tuning. However, unlike fact-checking, regular fine-tuning for QA gets almost perfect scores but at the expense of the retain accuracy. Sequence-to-sequence models are more sensitive to a slight parameter shift. This happens because minor changes may completely alter the top-k prediction from beam search (in the case of QA). Differently, in a binary classifier (in the case of FC) the probability of a prediction can change substantially without crossing the decision boundary (usually set at 0.5 when not calibrated).

2 Retaining previous knowledge

KnowledgeEditor maintains the predictions in the validation set almost perfectly (retain accuracy is ≈ ⁣98%\approx\!98\% for both FC and QA). Conversely, as expected, our method with CL2{\mathcal{C}}_{L_{2}} has very low retain accuracy (always <50%<50\%). CL2{\mathcal{C}}_{L_{2}} suffers from catastrophic forgetting because it does not enforce the updated model to be close to the original one in function space (i.e., the two functions to be similar) but just in parameter space.

Fine-tuning all layers is successful but it affects the previously acquired knowledge negatively: retain accuracy is ≈ ⁣87%\approx\!87\% and ≈ ⁣68%\approx\!68\% for FC and QA, respectively, while performance deterioration in ≈ ⁣2%\approx\!2\% and ≈ ⁣4%\approx\!4\%. Fine-tuning a single layer is more effective as it prevents over-fitting (the best model updates the 1st layer in both FC and QA). However, in FC the updated model does not generalize on semantic equivalent inputs: the accuracy on paraphrases is much lower even than versions of our methods which do not use paraphrases in training (42%42\% vs. >81%>81\%), and even more so when compared to those which use them (>94%>94\%).

Fine-tuning with Zhu et al.’s (2020) method does not affect performance for FC much, which is not surprising since standard fine-tuning already gets almost perfect scores. Differently, in the QA setting, using their constrained optimization boosts the retain accuracy (up to +4%+4\% to normal fine-tuning) but at the cost of a low success rate (≈ ⁣80%\approx\!80\% where fine-tuning gets the perfect score).

3 Accuracy on paraphrases

We evaluate our method both with and without the additional supervision of paraphrases to improve generalization—that corresponds to have Px{\mathcal{P}}^{x} as the set of paraphrases of xx or Px={x}{\mathcal{P}}^{x}=\{x\} in Equation 1, respectively. Without this additional supervision, KnowledgeEditor is already competitive in equivalence accuracy. However, employing this additional supervision is clearly beneficial on both tasks: we get the same success rate and re-train accuracy but equivalence accuracy improves by >70%>70\% on FC and >30%>30\% on QA, respectively (for generated paraphrases). In FC, although fine-tuning of a single layer proved to be optimal in terms of success rate and retain accuracy, it performs poorly for paraphrases. That is the model successfully updates the prediction of a particular datapoint, but does not update predictions of paraphrases. This indicates that fine-tuning to edit the knowledge of a model does not generalize well, and it overfits to specific inputs. On QA, also Zhu et al. (2020) performs poorly compared to our or other methods.

When other methods perform on par or better than ours on paraphrases, they do not have good retain accuracy (e.g., see QA fine-tuning on Table 1). Fine-tuning on QA seems to generalize better than on FC, but does not preserve previous knowledge. In Table 1 we also report both the accuracy on the set of generated and human-generated paraphrases. Surprisingly, the scores on human-generated paraphrases are higher. We speculate that this happens because automatic paraphrases are sometimes not semantically equivalent or fluent.

4 Analysis of model updates

In Figure 3 we plot the distribution of logits of the original and updated model on FC for different methods. With an ideal method, all logits before and after an update have to stay the same (except the ones we want to change). From that figure, we can see distributions of different types of errors such as datapoints whose predictions were mistakenly flipped (from true to false or the other way around). These errors are mostly concentrated around the origin, where small perturbations make logits cross the decision boundary. When fine-tuning all layers, we can see a clear impact on logits, they undergo a lot of change (i.e., points do not concentrate around the diagonal). Indeed, fine-tuning makes many datapoints cross the decision boundary and their probabilities to change from the original ones. The failure of CL2{\mathcal{C}}_{L_{2}} is visible in Figure 3(b) as this method preserves almost none of the previous predictions. Instead KnowledgeEditor preserves almost all of the predicted labels as well as their probabilities (most datapoints in Figure 3(c) stay on the diagonal).

We also report visualizations of the average weight updates for the QA experiment in Figure 4. We report the setting with additional supervision from paraphrases (but the heatmaps are similar without them). There are three main observations from this plot. First, gradients are mostly concentrated on the first encoder layer and the last decoder layer. Gradients explain why the best subset of parameters to update is the first layer. Secondly, fine-tuning does not preserve gradient magnitudes and updates the whole model almost uniformly. That happens because of the optimizer’s adaptive learning rate that initially erases the gradient direction. The gradient direction plays a role only after a couple of gradient steps, but most of the time, the method only needs one step to modify its knowledge. Lastly, our updates are sparser and are not consistent with the gradient for changing the predictions. That indicates that our method learns to use the gradient in a meaningful way (i.e. ignoring some directions or manipulating its magnitude). It is surprising that the knowledge manipulation seems to be achieved by primarily modifying parameters affecting the shape of the attention distribution (WselfKW_{self}^{K} and WselfQW_{self}^{Q}) rather than, e.g., values (WselfVW^{V}_{self}). As we discussed, the hyper-network may be regarded as a probe providing insights about the mechanism used by the model to encode the knowledge Vig et al. (2020). For example, the focus on the bottom layer is already intriguing, as it contrasts with claims that memorization happens in top layers of image classification models Stephenson et al. (2021), hinting at substantial differences in the underlying memorization mechanisms in NLP and vision. Proper investigation is however outside of the scope of this study. See Appendix C for some additional analysis.

Conclusions

In this work, we explore the task of editing the factual knowledge implicitly stored in the parameters of Language Models. For this task, we formally define desiderata, the objective, and a set of metrics to measure the efficacy of different methods. We concretely evaluate that on two benchmarks based on closed-book fact-checking and question answering. We propose KnowledgeEditor, a method based on a hyper-network that learns to modify implicit knowledge stored within LM parameters efficiently and reliably. We provide comprehensive evaluations for our models against different variants of fine-tuning demonstrating the advantage of our approach. The magnitude of the updates predicted by our method may unfold the mechanisms used by the LMs to encode factual knowledge; we leave such investigation for future work.

Ethical Considerations

Technology built upon pre-trained LMs inherits some or all of their potential harms (Bender et al., 2021). Our technology for editing the knowledge of LMs does not exacerbate their potential harms and can, in fact, be used to mitigate harms, as models can be corrected once problems are discovered. However, we note that malicious uses of our knowledge editor are possible. For example, malicious agents may use the techniques presented in this work to inject incorrect knowledge into LMs.

Acknowledgments

The authors want to thank Michael Schlichtkrull, Lena Voita and Luisa Quarta for helpful discussions and support. This project is supported by SAP Innovation Center Network, ERC Starting Grant BroadSem (678254), the Dutch Organization for Scientific Research (NWO) VIDI 639.022.518, and the European Union’s Horizon 2020 research and innovation programme under grant agreement No 825299 (Gourmet).

References

Appendix A Relaxation and Approximation of Constrained Optimization

Given a objective to minimize in the form of

Equation 5 can be evaluated with automatic differentiation and optimized via gradient descent.

Appendix B Experimental setting

We evaluate on closed-book fact-checking (FC) using the binary FEVER dataset (Thorne et al., 2018) from KILT (Petroni et al., 2021). FEVER has 104,966 training and 10,444 validation instances respectively. For every input claim xx, the model predicts the probability f(x;θ)f(x;\theta) that it may be true. This is done without retrieving any evidence from a corpus, instead, just by relying on the knowledge accumulated during pre-training and encoded in its own parameters—this is similar to Lee et al. (2020) that investigate closed-book and zero-shot FC using masked-LMs. Concretely, we ask the LM to perform binary classification. We fine-tune a BERT base model (Devlin et al., 2019) with an additional linear layer on top that maps the hidden state corresponding to the BOS (beginning of a sentence) token to the probability of the positive label. Given the available supervision, we train the architecture to maximize the model likelihood penalized by entropy regularization and weight decay. The final model has an accuracy of 77.1%77.1\%.This is comparable with what reported by Petroni et al. (2021) for a larger BART model.

B.2 Question answering

We also evaluate on a task with a more complex sample space: closed-book question answering (QA). Here QA is treated as a sequence-to-sequence problem from question to answer without retrieving nor providing any evidence Roberts et al. (2020). This, as in FC, emphasises the role of the knowledge acquired in pre-training and encoded in the parameters of the model. For this task, we used the Zero-Shot Relation Extraction (zsRE) dataset by Levy et al. (2017). We prefer zsRE to other popular QA datasets such as SQuAD (Rajpurkar et al., 2016), Natural Questions (Kwiatkowski et al., 2019) or TriviaQA (Joshi et al., 2017) because it is annotated with human-generated question paraphrases that we can use to evaluate our model’s robustness to semantically equivalent inputs. zsRE is specifically constructed not to have relation overlaps between training and test (i.e. it is zero-shot). We re-split the dataset to have the same distribution in training and test splits—we are not interested in zero-shot specifically, so we avoid the additional complexity it entails. The original zsRE dataset has 147,909 training and 3,724 validation instances respectively. After re-splitting and employing all paraphrases, we have 244,173 training and 27,644 validation instances respectively. For this task, we fine-tune a BART base model (Lewis et al., 2020) with a standard seq2seq objective, i.e., maximizing the model likelihood given the observed output sequence (Sutskever et al., 2011, 2014) and regularized with dropout (Srivastava et al., 2014) and label smoothing (Szegedy et al., 2016). The final model has an accuracy (exact match between model prediction and gold standard) of 22.1%22.1\%.This is more than reported by Petroni et al. (2021) on the original split of zsRE. That is because the original split aims at zero-shot evaluation, while we have an overlap of relation types between training and validation sets.

B.3 Generating alternative predictions

Generation of alternative predictions is task-dependent as it requires producing a plausible substitute target for a given input—e.g., if we need to edit the knowledge about a head of a state, a plausible substitute label should be a person, not a random (even if well-formed) string. Fact-Checking is straightforward: we simply flip the label, as it is binary classification. For QA, we exploit high-probability outcomes under the model distribution as a proxy to plausible revisions. In particular, we pick all hypotheses enumerated via beam search except the top-1.This does not always guarantee that the alternative predictions have the same semantic type as the original one, but it is likely since the model assigns high probability to them.

B.4 Semantically equivalent inputs

We would like the updated model to be consistent for semantically equivalent inputs (see Px{\mathcal{P}}^{x} in Section 2 and 4) as opposed to just learning a new specific and isolated datapoint. This consistency is indicative of an effective editing mechanism that taps into the knowledge stored in the model. However, not all datasets come with paraphrases of its inputs (e.g., in our case FEVER does not come with paraphrases and zsRE only has paraphrases for 30%30\% for the dataset). To this end, we generate semantically equivalent inputs using round-trip translation (Sennrich et al., 2016; Wieting and Gimpel, 2018). We employ English-to-German and German-to-English Transformer models from Marian Neural Machine Translation (MarianNMT; Junczys-Dowmunt et al., 2018) provided by Huggingface Transformers (Wolf et al., 2020). We use beam search with beam size 55 to obtain 2525 paraphrases. From this set, we exclude any candidate paraphrase x^\hat{x} of xx for which the prediction y^\hat{y} supported by f(x^;θ)f(\hat{x};\theta) does not match the prediction yy supported by f(x;θ)f(x;\theta). This filtering ensures that, according to the current model, all paraphrases have the exact same prediction.

B.5 Architecture details

The original models we want to modify are a BERT base model (Devlin et al., 2019) and a BART base model (Lewis et al., 2020) for fact-checking and question answering respectively. They are both Transformer based models with 12 layers each and hidden size of 768. BERT has 12 heads, where BART has 16. They have 110M and 139M parameters respectively. BERT has a vocabulary size of 30,522 where BART has 50,265.

KnowledgeEditor has a small single-layered bidirectional-LSTM with input size 768 and hidden size of 128. The FFNN that condenses the LSTM states follows a [256, tanh⁡\tanh, 1024] architecture where the 5 FFNN have all a [1024, tanh⁡\tanh, dd] architecture where dd depends on the weight to modify. In our experiments, we do not use our model to modify biases, layer norms, word and positional embeddings of LMs. Overall, KnowledgeEditor has 54M and 67M parameters for BERT and BART respectively.

B.6 Training details

The original models which we want to modify are trained with a batch size of 256 using Adam (Kingma and Ba, 2015) (learning rate of 3e-5) with weight decay (1e-2) and a linear schedule with warm-up (50k total number of updates and 500 warm-up updates). We trained for a maximum of 20 epochs and employ model selection using accuracy on the validation set.We trained on 4 Nvidia Titian X 12GB which take approximately 10 minutes for FC and 3 hours for QA.

KnowledgeEditor models are trained with a batch size of 1024 for FC and 256 for QA using Adam (learning rate of 3e-4 for the parameters and 1e-1 for the Lagrangian multiplier) with weight decay (1e-2) and a linear schedule with a warm-up (200k total number of updates and 1k warm-up updates). We trained for a maximum of 200 epochs and employ model selection using overall accuracy (success rate and retain accuracy) on the validation set (approximated using mini-batches).We trained on 4 Nvidia Titian X 12GB which take approximately 1 day for FC and 3 days for QA. The margin for the CKL{\mathcal{C}}_{KL} is annealed between 1e-1 and 1e-3 for the fact-checking model, and between 1e-3 and 1e-5 for the BART question answering model. For the sequence-to-sequence loss, we employ a cross-entropy loss with label smoothing of 0.1.

Appendix C Additional Results

During preliminary experiments, we studied a version of our hyper-network that did not exploit gradient information (see Equation 3). Without gradient information, on FC the models converged ≈ ⁣10\approx\!10 times slower to reach the same accuracy and did not converge for QA (i.e., the model was not able to get >75%>75\% success rate and >50%>50\% retain accuracy). That suggest that the gradients are helpful and actually used by our hyper-network but should not used directly, without a modification. To better show this, in Table 2 we report correlations between different update methods and the gradient in terms of cosine similarities between updates. Naturally, fine-tuning and the gradient are highly correlated, but our method (with and without additional paraphrases supervision), poorly correlates with the others. Low cosine similarity can be due to two factors i) the model indeed projects the gradient to a different and more ‘knowledge preserving’ direction, or ii) the parameter space is so large that cosine similarity gets to zero very quickly, not revealing the genuine underlying similarity.