Do Language Models Have Beliefs? Methods for Detecting, Updating, and Visualizing Model Beliefs

Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, Srinivasan Iyer

Introduction

Language models (LMs) may not have beliefs in the same sense that people do, but there are a few reasons to analyze LMs in terms of the beliefs they may possess. The first is that this is a useful way to understand and speak about how LMs behave. When discussing whether animals have beliefs (raccoons, in particular), philosopher Daniel Dennett (1995) writes: You might as well call the state of the raccoon a belief, since if you call it a “registration” or a “data-structure” in the “environmental information store” or some other technical term, the logic you use to draw inferences about the animal’s behavior, given its internal states, will be the standard, “intentionalistic” logic of belief. Dennett bases this conclusion in the fact that we can and do draw accurate inferences about animal behavior by first understanding their beliefs. We are drawn to speak about the beliefs of LMs in the same “maximally bland (but maximally useful!)” sense. To the extent that these neural networks act intelligently in response to stimuli, we may form more accurate theories of how they work by understanding their beliefs.

The second reason for ascribing beliefs to language models is that many of the stricter definitions of belief incidentally exclude many real beliefs held by real people. Following Dennett (1995), Newen and Starzak (2020) define a belief as an informational state decoupled from any motivational state, and they outline a few additional properties of beliefs, namely that they should (1) be recombinable with motivational states and other informational states and (2) have some minimal structural organization. Further, an entity with beliefs should (1) be sensitive to new information, (2) categorize new beliefs as they develop, and (3) display some kind of logical consistency. These are all properties that come in degrees, and setting the bar too high will exclude many of the statements that people earnestly express to others in their everyday lives. Meanwhile, animals and neural networks alike store information in accordance with these properties to at least some extent.

We also note that we use the term belief rather than knowledge as in related work Zhu et al. (2020); De Cao et al. (2021) because we want to analyze beliefs of language models rather than knowledge in them. LMs may contain a great deal of knowledge to us, but in a traditional view of knowledge as Justified True Belief, it is relatively more difficult to say that an LM knows something rather than that it believes it Schwitzgebel (2019).

In the remainder of this paper, we turn our attention to three practical endeavors: detecting, updating, and visualizing beliefs in LMs. We build on work on editing models after training, which is an exciting recent direction of research with many potentially valuable use cases Sinitsin et al. (2020); Zhu et al. (2020); De Cao et al. (2021); Mitchell et al. (2021). For LMs, uses include correcting factually inaccurate outputs and preventing otherwise unwanted outputs from models (e.g. toxic generated text) without expensive data curation and retraining efforts. These are important applications given that LMs (1) struggle with future data when trained on data from the past Lazaridou et al. (2021); Dhingra et al. (2021), (2) generate morally undesirable text in many situations Gehman et al. (2020); Bender et al. (2021), and (3) simply give inaccurate outputs for tasks like question answering Lin et al. (2021). Notably, there is good evidence that scaling models to larger sizes will not fix these particular problems or may even exacerbate them in cases like imitative falsehoods in QA, so we will likely need an alternative solution Lazaridou et al. (2021); Gehman et al. (2020); Lin et al. (2021). We next outline a few key contributions of the paper. Figure 1 represents the core ideas behind these contributions.

Detecting beliefs. We measure the degree to which LMs exhibit several properties of belief-possessing systems, using models finetuned on fact verification and question answering tasks. Beyond simply checking individual model responses, we want to assess the structural properties of model outputs: Are they consistent under paraphrase? Are they logically consistent? Does changing one belief correctly change other entailed beliefs? Does it erroneously change other unrelated beliefs? Past work has focused primarily on consistency under paraphrase Elazar et al. (2021); De Cao et al. (2021); Mitchell et al. (2021). Here, we adapt data from Talmor et al. (2020) to measure consistency under entailment (including for contrapositives), and we use the Wikidata5m dataset Wang et al. (2021b) to construct logically neutral belief pairs for checking that models do treat the beliefs as independent.

Updating beliefs. We propose a Sequential, Local, and Generalizing belief update objective (SLAG) that substantially improves the performance of the KnowledgeEditor method from De Cao et al. (2021). KnowledgeEditor is a learned optimizer that edits a model’s weights in order to change its prediction on an input while satisfying other desiderata, like consistency under paraphrase. Principally, we use more difficult training data for the learned optimizer, and we also train the network to apply multiple small edits rather than just one edit. These changes markedly improve the overall update success rate and lower the rate at which other beliefs are corrupted. Moreover, we find that KnowledgeEditor almost totally fails when updating multiple beliefs in a row as opposed to a changing a single belief. In this setting, off-the-shelf optimizers are far preferable methods. However, by explicitly training the optimizer to update multiple beliefs sequentially, we are able to once again outperform off-the-shelf optimizers. Lastly, we advocate that these methods be evaluated for their ability to correct false or morally undesirable model beliefs, rather than to arbitrarily adjust model beliefs to plausible alternatives as in past work Zhu et al. (2020); De Cao et al. (2021); Mitchell et al. (2021).

Visualizing belief graphs. We explore a new form of interface with LMs, the belief graph. Given a set of beliefs, we construct belief graphs by changing each model belief and checking what other beliefs are sensitive to those changes. Each belief becomes a node, and directed edges between nodes show that updating one belief changes the other. We discuss graph metrics that help summarize the dependencies between model beliefs.

We summarize our main conclusions as follows:

∼{\sim}100M parameter models exhibit limited belief-like qualities, as paraphrase consistency scores are under 70%, and models show mixed levels of consistency under entailment (Sec. 5.1).

Off-the-shelf optimizers are surprisingly effective baselines for updating model beliefs, and they generally outperform learned optimizers when updating a single belief (Sec. 5.2).

When updating multiple beliefs in a row, method performance greatly declines (especially for learned optimizers). By using SLAG, we can improve learned optimizers’ performance beyond what baselines can reach (Sec. 5.2).

Belief graphs reveal many nonsensical dependencies between model beliefs. We find that (1) updates are mostly likely to change already incorrect model beliefs and (2) there are highly connected beliefs which influence a large fraction of all model beliefs (Sec. 6.3).

Related Work

Much past work has explored how information is stored and represented in pretrained language models Rogers et al. (2020), though few discuss what qualifies information as a model belief. Petroni et al. (2019) provide evidence that LMs store relational information between entities, and Roberts et al. (2020) show that LMs can answer open-ended questions. Subsequent work has explored how much knowledge is stored in LMs Heinzerling and Inui (2021), approaches to querying models for knowledge Hewitt and Liang (2019); Jiang et al. (2020); Voita and Titov (2020); West et al. (2021), and methods for learning more knowledge during pretraining Wang et al. (2021b, a). Most relevant to our work are studies from Talmor et al. (2020) and Elazar et al. (2021). Talmor et al. (2020) train LMs to perform True/False classification of factual claims, and they measure how a model’s belief in one fact correlates with its belief in an entailed fact. We use their LeapOfThought dataset to measure model consistency under entailment before and after updating the up-stream beliefs in models. Meanwhile, Elazar et al. (2021) measure the consistency of model predictions for sets of paraphrased inputs. We adopt their metric for paraphrase consistency as a measure of belief. In concurrent work, Kassner et al. (2021) discuss consistency under entailment and paraphrase as conditions for belief, and they measure consistency under entailment with a new dataset, BeliefBank.

Updating beliefs in language models.

Approaches to making targeted updates to model beliefs vary along a few dimensions. First is whether the methods alter model training or operate in a post-training setting. Sinitsin et al. (2020) use a meta-learning objective during training to encourage ease of editing afterwards, though the memory requirements of their approach limit its scalability beyond 100M parameter models. A larger family of methods make post-training updates to models, differing in how they formalize the update problem: Dai et al. (2021) propose a hand-crafted algorithm for updating model weights, while Zhu et al. (2020) use projected gradient descent for batches of points. De Cao et al. (2021) and Mitchell et al. (2021) frame the problem as a machine learning problem and train hypernetworks (learned optimizers) that process model gradients in order to produce a new model that (1) gives the desired output for the edited point, while (2) incorporating other objectives like minimizing the changes in predictions for other data. Here, we build directly upon the method from De Cao et al. (2021), showing where it fails and providing an improved training objective (SLAG). In particular, we find that the method struggles with updating multiple beliefs sequentially. This setting bears some commonality to the continual learning problem, though continual learning methods generally aim to learn new tasks or datasets rather than make targeted updates to specific model beliefs Parisi et al. (2019).

Not all methods edit model weights. Kassner et al. (2021) update model beliefs by adding in relevant information to the input at test time (to improve consistency under entailment). But as with retrieval-based methods, this approach does not change the model weights and hence does not influence model outputs on all other potentially relevant inputs Lewis et al. (2020); Hase and Bansal (2021).

Updating Beliefs in Language Models

Following De Cao et al. (2021), we approach the problem of updating model beliefs as a machine learning problem and train a learned optimizer to perform desired model updates. We discuss metrics for detecting beliefs in Sec. 5.1 and our approach to visualizing belief graphs in Sec. 6.3. The core ideas of our approach are outlined in Fig. 1.

We suppose we have a model fθ=pθ(y∣x)f_{\theta}=p_{\theta}(y|x) parametrized by θ\theta. For an input xix_{i} that has some undesired model output y^i=arg max⁡ypθ(y∣x)\hat{y}_{i}=\operatorname*{arg\,max}_{y}p_{\theta}(y|x), we wish to obtain a new model θ∗\theta^{*} that produces a desired output yi∗y_{i}^{*} for xix_{i}. This new model θ∗\theta^{*} should also fulfill a few other desiderata. As in past work De Cao et al. (2021); Mitchell et al. (2021), we operationalize these desiderata in the following metrics:

Update Success Rate (Main Input): The rate at which the updated model gives the desired output yi∗y_{i}^{*} for the Main Input xix_{i}.

Update Success Rate (Paraphrase): The rate at which the updated model gives the same new prediction for xix_{i} as it does for paraphrases of xix_{i}, which are inputs with the same meaning but different surface form.

Retain Rate (All Data): The rate at which the updated model’s predictions are unchanged for all other data besides the Main Input.

Δ\Delta-Acc (All Data): The change in accuracy for the updated model on all other data besides the Main Input.

In practice, Retain Rate (All Data) and Δ\Delta-Acc are computed with random subsets of a dataset, since these must be computed after every belief update. We add two metrics to those used in past work:

Update Success Rate (Entailed Data): The rate at which the updated model makes predictions that are logically entailed by the model’s prediction for the Main Input.

Retain Rate (Local Neutral): The rate at which the updated model’s predictions are unchanged for data that is similar to the Main Input but still logically neutral.

We use Update Success Rate (Entailed Data) to measure logical consistency for an updated model, since changing one belief will entail changes in logically entailed beliefs. We also split “retain accuracy" into two cases, one for randomly sampled data as in past work (All Data) and the other for specially constructed Local Neutral data. Unlike randomly sampled data, Local Neutral data is guaranteed to be logically independent of the Main Input, while still being similar (local) to it. Together, these six metrics better cover the criteria for belief outlined by Newen and Starzak (2020). We compute the metrics using data of the kind shown in Table 1. For a glossary of terms used for these metrics across papers, see Appendix Table 13.

Evaluation data.

To date, methods have been evaluated on the basis of their ability to change model predictions for all data points, including correctly and incorrectly predicted points. Moreover, the desired labels {yi∗}i=1n\{y_{i}^{*}\}_{i=1}^{n} on sequence prediction tasks have each been selected from the beam search which produced the original model prediction De Cao et al. (2021); Mitchell et al. (2021). We propose for method evaluation to focus on a more valuable use case: changing the predictions on incorrect points to be correct. In Sec. 5, we show that this is a harder task than simply changing predictions to other similar outputs, so the effectiveness of past methods has been overestimated.

Sequential updates.

The default evaluation procedure in past work on learned optimizers is to update a single model belief, evaluate the new model, then rollback the update before repeating the process for each test point. In Sec. 5, we show that it is much harder to update multiple beliefs in a row before evaluating the new model. This is notable because in practice, it is likely that model developers will want to update many beliefs of a trained model, possibly over long timescales, meaning sequential updating is a more realistic application of update methods. We obtain sequential versions of all our metrics by applying rr model updates in a row before checking the metrics, meaning there are floor(n/r)\textrm{floor}(n/r) measurements for a test set of nn points.

Belief updating method.

We use the KnowledgeEditor architecture from De Cao et al. (2021) with our training objective, SLAG. For the details of this architecture, we refer readers to Appendix A. Let it suffice for now to observe that a new model is given as a differentiable function

using the learned optimizer gϕg_{\phi}, current LM weights θ\theta, Main Input xix_{i}, current prediction y^i\hat{y}_{i}, and desired model output yi∗y_{i}^{*}. In this paper, we generalize the update step to occur in a loop. If we package the above update as θ(k+1)=θ(k)+gϕ(xi,y^i,yi∗,θ(k))\theta^{(k+1)}=\theta^{(k)}+g_{\phi}(x_{i},\hat{y}_{i},y_{i}^{*},\theta^{(k)}), then we can obtain new model parameters as

for a number of steps KK from the initial parameters θ(k)\theta^{(k)}. In fact, De Cao et al. (2021) use such a loop at test time; we incorporate the loop into training to align the train and test-time distributions.

Learned optimizer training.

The training objective for KnowledgeEditor includes differentiable terms corresponding to Update Success for the Main Input and paraphrases, as well as Retain Rate for all other data. We also include terms for Update Success on entailed data and the Local Neutral Retain Rate, when this is possible given available data. The overall objective requires several kinds of additional data for each point, which we denote by DR\mathcal{D}_{R} for other random data, DLN\mathcal{D}_{LN} for local neutral data, DE\mathcal{D}_{E} for entailed data, and DP\mathcal{D}_{P} for paraphrases of xix_{i}. For a data point xix_{i} with desired prediction yi∗y_{i}^{*}, the full objective is then:

where LTask\mathcal{L}_{\textrm{Task}} is the loss used to get gradients for fθf_{\theta}. We use the Cross Entropy loss for binary classification and sequence-to-sequence tasks.

We optimize this objective w.r.t. ϕ\phi using AdamW Loshchilov and Hutter (2019). To obtain update labels {yi∗}i=1n\{y_{i}^{*}\}_{i=1}^{n}, we always use the opposite class in binary classification. For sequence-to-sequence tasks, we use the correct label when y^i\hat{y}_{i} is incorrect, and when y^i\hat{y}_{i} is correct, we randomly select another label from the training data. This choice is in contrast to De Cao et al. (2021) and Mitchell et al. (2021), who use samples from the model beam search as update labels for all points.

SLAG objective. To better prepare the update method for evaluation in a sequential-update setting, we consider training gϕg_{\phi} to update multiple datapoints in a row. Using the per-datapoint loss in Eq. 1, we obtain our Sequential, Local, and Generalizing (SLAG) loss for a set of rr Main Inputs D={xi,y^i,yi∗}i=1r\mathcal{D}=\{x_{i},\hat{y}_{i},y_{i}^{*}\}_{i=1}^{r} as

where θt+i=Update(xi,y^i,yi∗,θt+i−1;ϕ,K)\theta_{t+i}=\textrm{Update}(x_{i},\hat{y}_{i},y_{i}^{*},\theta_{t+i-1};\phi,K) are the model parameters obtained from updating on the first ii points in D\mathcal{D} (starting from θt\theta_{t}). This objective allows us to train gϕg_{\phi} to update multiple beliefs in a row. To ensure training with this objective is still efficient, we limit how far back through the LM history we backpropagate when computing the gradient w.r.t. ϕ\phi for each term in the RHS sum of Eq. 2. Each parameter vector θt\theta_{t} depends on ϕ\phi and θt−1\theta_{t-1}. We always apply the stop-gradient function to the most recent vector θt−1\theta_{t-1} to prevent backpropagating through it (visualized in Appendix Fig. 4). This choice allows our memory use to remain constant in rr (see Appendix Fig. 5).

Experiment Setup

We run experiments with four datasets (example data shown in Appendix Table 15). (1) FEVER includes 115,409 True/False factual claims Thorne et al. (2018). We use the original test set of 10,444 points, and we randomly split the training data into 94,469 train points and 10,496 dev points. (2) zsRE includes 151,631 questions based on relational knowledge from Wikipedia, which we randomly shuffle into train/dev/test splits with 80/10/10% of the data Levy et al. (2017). 32.8% of zsRE questions in each split include paraphrases, and we measure Update Success Rate (Paraphrase) for only these points. Talmor et al. (2020) introduce (3) the LeapOfThought dataset, consisting of 33,484 factual claims that are entailed to be true or false depending on a fact and distractor statements provided as context. We drop the distractors from each input and filter the data so that the facts are unique, then shuffle the resulting 14,939 points into train/dev/test splits with 60/10/30% of the data.

We also construct (4) a sequence prediction task using data from Wikidata5m, which is a relational knowledge base with over 20 million triplets Wang et al. (2021b). We build this dataset in order to get Local Neutral data. Each input consists of an entity e1e_{1} and relation rr, and the label is another entity e2e_{2} that completes the triplet. All inputs come in pairs that share the same entity e1e_{1} but use different relations with different labels. The relations are always one of ten relations that apply to people (see Appendix Table 11). In general, the completion e2e_{2} to the Main Input triplet (e1e_{1}, r1r_{1}, e2e_{2}) has no logical consequences for its paired input, (e1e_{1}, r2r_{2}, ?). This means that changing the model belief for the Main Input should not change its belief for its neutral paired input. The paired points are also local to the Main Input, i.e. they pertain to the same entity e1e_{1} as the Main Input. We obtain four paraphrases for each Main Input using different aliases for the entity and synonyms of the relation. We construct a train set of 150k points and dev and test sets of 10k points each. See Appendix B for further details.

2 Methods Evaluated

Models. We train five models with different random seeds for each dataset, using RoBERTa-base for binary tasks and BART-base for sequence-to-sequence tasks (accuracies in Appendix Table 14). For each of the five models, we train one learned optimizer using SLAG and one with the objective from De Cao et al. (2021), which we list as KE in tables below. Our model selection criterion is the mean of: the average Update Success Rate (across data types), Retain Rate (only for Local Neutral data), and Δ\Delta-Acc for All Data. We tune the choice of SLAG objective terms for each task separately (see Appendix Table 10 for final selections; results discussed in Sec. 5.3). Other hyperparameters are given in Appendix B and memory use is shown in Appendix Fig. 5. To summarize the differences between SLAG and KnowledgeEditor: (1) we use Ktrain=KtestK_{\textrm{train}}=K_{\textrm{test}} rather than Ktrain=1K_{\textrm{train}}=1; (2) we adopt training labels using real data labels rather than alternatives from the model’s beam search; and (3) our objective terms differ following tuning.

Baselines. We use off-the-shelf optimizers as baselines. We tune the baseline hyperparameters separately for each dataset, selecting among several kinds of optimizers, learning rates, and the number of update steps. The selection criterion is the same as the criterion outlined for learned optimizers above. The resulting baselines are surprisingly strong (see Appendix Table 12 for final selections).

Hypothesis testing. We obtain 95% confidence intervals and perform hypothesis tests via block bootstrap, resampling model seeds and data points Efron and Tibshirani (1994). For ablation experiments, we run only one model seed per condition.

Experiment Results

We measure Paraphrase Consistency, Entailment Acc, and Contrapositive Acc for our finetuned task models. Paraphrase Consistency is the fraction of paraphrase pairs for which a model produces the same output Elazar et al. (2021). Entailment Acc is the accuracy of a model on data that is entailed by the Main Input. For LeapOfThought (see Table 1), “Main Input xix_{i} is true" implies “entailed input xEx_{E} has label yEy_{E}," but the inverse (¬A⇒¬B\neg A\Rightarrow\neg B) does not necessarily hold. Therefore, we compute Entailment Acc only where the Main Input prediction is correct. We do know that the contrapositive holds: “Entailed input xEx_{E} does not have label yEy_{E}" implies that “Main Input xix_{i} is false." So for Contrapositive Acc, we measure how often the model follows this rule, when the antecedent holds of its prediction.

Belief measurement results. Table 2 shows the belief metrics for each dataset. We find that ∼{\sim}100M parameter models show limited evidence of having beliefs about the world. Paraphrase consistency is 69.50% (±\pm 1.09) for zsRE and much lower for Wikidata5m (25.84%±\pm0.53). While entailment accuracy is high for LeapOfThought (85.63%±\pm1.08), the model is consistent under the contrapositive only 16.51% (±\pm 2.71) of the time. One might reasonably set the bar for qualifying as a “belief" higher than these scores. But since belief-likeness comes in degrees, we continue to refer to model beliefs for the rest of the paper. Interestingly, the metrics are much higher when the model prediction on the Main Input is correct (Table 3).

2 Can we update beliefs in LMs?

First, we compare two evaluation procedures for sequence prediction tasks: correcting model beliefs versus changing them to an alternative from the model’s beam search. We do so for zsRE using SLAG. Next, we compare belief update metrics across datasets using KnowledgeEditor, SLAG, and off-the-shelf optimizers as baselines. We report results in single-update (rtest=1r_{\textrm{test}}=1) and sequential-update (rtest=10r_{\textrm{test}}=10) settings. See Appendix Fig. 6 for an ablation across rtestr_{\textrm{test}}.

Correcting beliefs vs. changing beliefs. Given the results in Table 5, we find that correcting model outputs is harder than simply changing them to a plausible alternative. Update Success can rise by a full 2.96 (±\pm0.48; p<p{<}110−4110-4) points for Main Inputs and 2.58 (±\pm0.81; p<p{<}110−4110-4) for Paraphrases, while Δ\Delta-Acc is virtually unchanged. This suggests that that past work has overestimated the efficacy of belief update methods for actually fixing models. Henceforth we evaluate methods according to their ability to update model beliefs to be true.

Update method results (single update). Table 4 shows the results in a single-update setting. First, we find that off-the-shelf optimizers are very effective across the board. The baselines show Main Input Update Success Rates of 100% for binary tasks with positive Δ\Delta-Acc scores.Positive Δ\Delta-Acc values are possibly due to distribution shift in the test split. In FEVER, for instance, the train and dev data are 73% True, while test data is 50% True. On the dev split, AdamW achieves a negative Δ\Delta-Acc, -0.18 (±\pm0.11). On sequence prediction tasks, SGD achieves 98%+ Main Input Update Success with competitive Δ\Delta-Acc scores. When strongly tuned, these baselines outperform learned optimizers on most metrics here.

However, SLAG does surpass the baselines in a few places. All Data Retain Rate on zsRE rises by 5.77 points (±\pm1.43; p<p{<}110−4110-4), and on Wikidata5m we improve Paraphrase Update Success by 11.92 points (±\pm1.20; p<p{<}110-4$)andtheLocalNeutralRetainRateby6.40() and the Local Neutral Retain Rate by 6.40 (\pm1.41;1.41;p{<}$110−4110-4) points. The gain on Entailed Data Update Success is 3.02 points, but it is not significant (±\pm6.26; p=p{=}.345). The SLAG objective also greatly improves performance over KE for sequence prediction tasks.

Update method results (sequential updates). We give results for a sequential update setting (rtest=10r_{\textrm{test}}{=}10) in Table 6. Immediately we see this is a much more difficult setting for updating model beliefs, as update metrics are generally much lower for each dataset. Next, we observe that learned optimizers with SLAG10 (rtrain=10r_{\textrm{train}}{=}10) now outperform baselines on sequence prediction tasks. On zsRE, we improve Update Success for Main Inputs by 4.86 (±\pm0.83; p=p{=}110−4110-4) and for Paraphrases by 1.39 (±\pm0.93; p=p{=}.004), with better Δ\Delta-Acc by 0.64 (±\pm0.35; p=p{=}.0005). Improvements trend in the same direction for Wikidata5m and are all statistically significant except for the gain in Δ\Delta-Acc. The jump on Paraphrases in particular is very large (11.02±\pm1.17; p<p{<}110−4110-4). In comparison, using a non-sequential (rtrain=1r_{\textrm{train}}=1) training objective leads to drastic drops in performance.

Learned optimizers still struggle with the binary datasets compared to the off-the-shelf optimizers. The baselines achieve high update update success with much better Δ\Delta-Acc scores, by 13.12 (±\pm4.51; p=p{=}110−4110-4) on FEVER and 8.16 (±\pm1.63; p=p{=}110−4110-4) on LeapOfThought. Also on LeapOfThought, the baseline’s update success with entailed data is over 10 points higher (±\pm7.38; p=p{=}.004).

3 How does the learned optimizer objective influence performance?

Here, we discuss ablations with respect to the terms in the training objective, Eq. 1. We show the effect of KtrainK_{\textrm{train}} in Appendix Fig. 9 and the choice of optimizer training labels in Appendix Table 16.

Training objective ablation. We give objective ablation results in Appendix Table 17. Surprisingly, we do not always see that the objective terms help for the data they are intended to help with. First, we obtain mixed results for the paraphrase objective. On zsRE, the objective term seems to hinder performance, with update success dropping on Main Inputs by 0.71 (±\pm0.60; p=p{=}.021) and Δ\Delta-Acc dropping by 0.18 (±\pm0.19; p=p{=}.069), while the paraphrase Update Success Rate itself is unaffected. With Wikidata5m, however, the paraphrase term improves paraphrase update success by a large margin of 16.94 (±\pm1.03; p<p{<}110−4110-4) points. Adding the Local Neutral (LN) term with the paraphrase term greatly improves the LN Retain Rate for Wikidata5m, by 9.71 points (±\pm1.44; p<p{<}110−4110-4), though both of these terms come at a cost to Main Input Update Success, similar to zsRE. Lastly, we do not find that the entailment objective improves Entailed Data Update Success; in fact, this metric falls by 4.56 (±\pm7.22; p=.213p{=}.213) points with the objective.

Analysis

In Table 7, we show belief metrics before and after model updates using SLAG with rtest=1r_{\textrm{test}}{=}1. We observe that belief updates greatly improve paraphrase consistency and entailment accuracy for updated data. Paraphrase consistency rises by 33.14±\pm1.46 on zsRE and 59.87±\pm1.09 on Wikidata5m, while Entailment Acc rises by 17.20±\pm7.10 points. To see if these improvements depend on pre-update consistency, we plot paraphrase consistency before and after updating in Fig. 2. For zsRE, consistency rises irrespective of pre-update consistency. There is a noticeable trend for Wikidata5m paraphrases, where post-update consistency is 90.1% when pre-update consistency is maxed out, versus 77.1% for totally inconsistent pre-update beliefs. We conclude that learned optimizers can induce a fairly consistent model belief even where there is no consistent belief to begin with.

2 Which beliefs are hard to retain when updating other beliefs?

We find that the Retain Rate depends heavily on whether the predictions on that data are correct to begin with. On zsRE for instance, the retain rate on correct inputs is about 96%, while for incorrect predictions, it is about 75%. So it appears that incorrect predictions are the most sensitive to model updates, and these points merely change from one incorrect prediction to another. On FEVER, incorrect beliefs change around 4% of the time, while correct beliefs change only 2.5% of the time.

We also find that Local Neutral beliefs are much harder to avoid changing than simply random data. For Wikidata5m in Table 4, the Retain Rate on All Data is 61.51±\pm1.33, while for Local Neutral data it is a full 15.66 points lower, at 47.85±\pm0.96.

3 Belief Graphs

We now construct belief graphs for the purpose of better understanding the connections between model beliefs. We form the graphs from a set of datapoints by updating each prediction and checking what other predictions change. We represent each datapoint as its own node in a belief graph. Whenever updating a datapoint uu changes the model prediction for point vv, we draw a directed edge from uu to vv. Following our results in Sec. 5.2, we use off-the-shelf optimizers to change the model output to the opposite of its original prediction for every datapoint. The resulting graphs have up to n2−nn^{2}-n edges (no self edges). For FEVER we obtain a graph of 10,444 nodes, and for LeapOfThought we obtain a graph with 8642 nodes, which is double the original test set size because we include both Main Inputs and Entailed Data as their own nodes.

We visualize part of a belief graph in Fig. 3. This figure shows a non-random subgraph intended to give a representative view of the data (we give three random subgraphs of 20 nodes in Appendix E). On inspection, we see no reason that beliefs are connected or not connected. Whether or not changing one belief changes another appears essentially random. We come to same conclusion looking at other random subgraphs (see Appendix Figures 10, 11, 12). However, we do know of some aggregate trends from earlier results. Sec. 6.2 suggests that a model’s incorrect beliefs are most likely to change after model updates, and following Sec. 5, we have reason to believe that Local beliefs are more likely than others to change with model updates.

We highlight a few summary statistics here from Table 8 for a broader view of the graphs. First, % Edgeless is the proportion of nodes which have no in or out edges. Since this is 0 for both datasets, every belief can be changed by editing the right belief. # In Edges is the number of in edges at the 95th percentile, meaning 5% of beliefs have more in edges than this value, and the same holds of # Out Edges. These values grow to a rather large fraction of the overall datasets, suggesting that (1) some beliefs are sensitive to changes in a large fraction of all beliefs, and (2) some beliefs are influential to hundreds of other beliefs when changed. # Corrupted is the number of correct predictions changed to be incorrect following a model update. For 5% of the data, model updates cause at least 211 points to become incorrectly predicted on FEVER, and 2,752 points for LeapOfThought. Lastly, % Update-Transitivity represents the answer to the question: if updating belief A changes belief B, and updating belief B changes belief C, what proportion of the time does updating A change C? For these datasets, a logically consistent model should display 100% Update-Transitivity (see Appendix D for a caveat on this metric). We find that belief updates often yield intransitive results for both datasets.

Discussion and Conclusion

Degrees of commitment to beliefs. The data we use comes in the form of declarative statements and answers to questions. These utterances take what is called a veridical stance toward a proposition, displaying a “full commitment" to that proposition’s truthfulness Giannakidou and Mari (2020). It will be valuable for future work to explore two dimensions of uncertainty in beliefs: (1) expression of uncertainty in language, via partial or trivial commitments (like “X might be Y") and (2) expression of uncertainty mathematically, via probabilities assigned by a model to utterances or True/False values. In this paper we treat a belief as “updated" when the model output changes, but this ignores any underlying change in the distribution pθ(y∣x)p_{\theta}(y|x) that could occur even if its mode does not change.

Ethics and dual use concerns. Belief update methods may be used to either correct undesired beliefs or induce problematic beliefs in LMs, and it is not clear whether these capabilities could be separated. We propose to evaluate methods only on the basis of their ability to correct mistaken model beliefs, but the malicious use case remains. We are uncertain about how a bad belief would influence the general behavior of a model (e.g. answers to many questions), but it is possible that a belief update method could instill bad beliefs in a generally capable LM with far-reaching implications for model behavior. That said, we hope that these methods will instead be used to update undesirable moral, social, and factual beliefs in large LMs.

Conclusion. We first discuss criteria for detecting when LMs have beliefs about the world. Next, we argue for evaluating belief update methods by their ability to correct mistaken beliefs, which is harder than the evaluation done in past work. We show that strongly tuned off-the-shelf optimizers make for surprisingly good belief update methods, even surpassing specialized learned optimizers in several settings. But with a new training objective (SLAG), we are able to outperform these baselines on sequence prediction tasks when updating multiple beliefs one after another. Finally, we introduce belief graphs as a means of understanding the connections between model beliefs. We find that model beliefs are highly interconnected, with some beliefs influencing hundreds of other beliefs. While it is hard to point to concrete reasons for individual connections between beliefs, we identify several patterns in the dependencies between beliefs.

References

Appendix A Learned Optimizer Details

Architecture. KnowledgeEditor is a learned optimizer g:X×Y×Y×Θ→Θg:\mathcal{X}\times\mathcal{Y}\times\mathcal{Y}\times\Theta\rightarrow\Theta that produces new model weights by applying an adjusted gradient step to a model. For reference, we give a glossary of symbols used here in Table 9. For additional details beyond what is presented here, we refer readers to De Cao et al. (2021).

At a high level, gϕg_{\phi} first encodes an input xix_{i} and requested prediction change into a vector hh, then processes hh into two low-rank matrices AA and BB that are used to transform the model gradient on xix_{i}, ∇θL(xi,yi∗)\nabla_{\theta}\mathcal{L}(x_{i},y_{i}^{*}). For Transformer models, the method edits only attention and feed-forward weights, so all model gradients match the shape of an associated weight matrix of shape d1×d2d_{1}\times d_{2}. Formally, a new model θ∗\theta^{*} is obtained using a learned optimizer gϕg_{\phi} as follows:

where ϕ\phi consists of all LSTM and MLP parameters.

Training Algorithm. The learned optimizer objective is optimized w.r.t. ϕ\phi with AdamW through a standard procedure of randomly sampling minibatches without replacement Loshchilov and Hutter (2019). Within each batch, one datapoint is randomly selected as the Main Input, and the remaining points are used as DR\mathcal{D}_{R}. To obtain update labels {yi∗}i=1n\{y_{i}^{*}\}_{i=1}^{n}, we always use the opposite class in binary classification. For sequence-to-sequence tasks, we use the correct label when y^i\hat{y}_{i} is incorrect, and when y^i\hat{y}_{i} is correct, we randomly select another label from the training data. This choice is in contrast to De Cao et al. (2021) and Mitchell et al. (2021), who use samples from the model beam search as update labels for all points.

Appendix B Additional Training Details

Learned optimizer memory. The hypernetwork has 92m trainable parameters for RoBERTa-base (which is 125m parameters), and 105m parameters for BART-base (which is 139m parameters). To increase training efficiency, we limit how far into the task model history we backpropagate. As shown in Fig. 4, when backpropagating through task model parameters θt=θt−1+Update(xi,y^i,yi∗,θt−1;ϕ)\theta_{t}=\theta_{t-1}+\textrm{Update}(x_{i},\hat{y}_{i},y_{i}^{*},\theta_{t-1};\phi), we continue backpropagating through Update(xi,y^i,yi∗,θt−1)\textrm{Update}(x_{i},\hat{y}_{i},y_{i}^{*},\theta_{t-1}) but not θt−1\theta_{t-1}, which is also dependent on ϕ\phi. That is, we apply a stop-gradient function to θt−1\theta_{t-1}. This way, we compute the derivative ∇ϕUpdate(xi,y^i,yi∗,θt;ϕ)\nabla_{\phi}\textrm{Update}(x_{i},\hat{y}_{i},y_{i}^{*},\theta_{t};\phi). only once for each tt, rather than recomputing these gradients for all subsequent time steps. These choices allow the memory use of our training algorithm to remain constant in rr. We make the same choice for our KK looped steps in a single application of the Update function, so the gradient for the update at step kk depends only on gϕ(xi,y^i,yi∗,θ(k))g_{\phi}(x_{i},\hat{y}_{i},y_{i}^{*},\theta^{(k)}) and not θ(k−1)\theta^{(k-1)}. See Fig. 5 for a graph of memory use depending on rr and kk.

Experiment runtimes. We now give runtimes for experiments in the paper. Building the belief graphs takes 25 hours for FEVER (n=10,444n=10,444) and 17.5 hours for LeapOfThought (n=8642n=8642) on an NVIDIA RTX 2080 GPU. Computing summary statistics for graphs takes 3 hours on FEVER and 3 hours for LeapOfThought for statistics besides Update-Transitivity. We compute Update-Transitivity for LeapOfThought with a subset of 4000 points, which takes 45 hours.

All other experiments are run on a NVIDIA V100 32GB GPU. Training the task models takes 7 minutes for LeapOfThought, 45 minutes for FEVER, 4 hours for zsRE, and 10 hours for Wikidata5m. Training the learned optimizer with r=1r=1 takes 2.3 hours for LeapOfThought, 5 hours for FEVER, 9.5 hours for zsRE, and 16 hours for Wikidata5m. Training the learned optimizer with r=10r=10 takes 53 minutes for LeapOfThought, 2.9 hours for FEVER, 7 hours for zsRE, and 12.5 hours for Wikidata5m. Computing update statistics with the off-the-shelf optimizers with r=1r=1 takes 4 minutes for LeapOfThought, 30 minutes for FEVER, 2.3 hours for zsRE, and 3.9 hours for Wikidata5m. With r=10r=10, the baselines require 1 minute for LeapOfThought, 15 minutes for FEVER, 54 minutes for zsRE, and 1.8 hours for Wikidata5m. Total runtimes for each experiment should take into account multiple conditions and multiple seeds of each model being run.

B.2 Hyperparameters and Objective Terms.

Training hyperparameters. We fit our RoBERTa-base and BART-base task models to their respective datasets with the following hyperparameters: We train for 10 epochs on the binary tasks, and 20 for the sequence-to-sequence tasks. When predicting with BART-base, we use a beam search with width 5. In each case, we use AdamW from torch.optim with a LR of 1e-5 and weight decay of 1e-4. We select the best model according to the best dev set accuracy, checkpointing after each training epoch. The learned optimizers are optimized with AdamW, using a learning rate of 3e-4 and weight decay of 0. We train the learned optimizer for 5 epochs on each dataset except for LeapOfThought, which we train for 10 epochs given its smaller size. The learned optimizers are also selected based on dev set performance, with checkpointing after each training epoch. Their selection criterion is a raw average of Update Success Rate (averaged over each kind of data), Retain Rate (Local Neutral) and Δ\Delta-Acc, with terms dropped when they cannot be computed given the available data. Note that dev epochs with zsRE and Wikidata5m are fairly slow, so in order to speed up our experiments we compute dev epochs with a subset of 4000 dev points.

Learned optimizer. We give the final hyperparameter and objective terms used in each experiment in Table 10. Our objective ablation is given in 17, and we select the best performing condition for each dataset according to dev set performance, using the same selection criterion outlined previously. We keep all weight coefficients λi\lambda_{i} equal rather than tuning them. Main refers to the first term in Eq. 1, plus the KL term with random data. We use Ktrain≤5K_{\textrm{train}}\leq 5 for all experiments. For results across KK values on zsRE, see Fig. 9.

Baseline update method. We tune a baseline off-the-shelf optimizer separately for each dataset, using rtest=1r_{\textrm{test}}=1. Our performance criterion is the same as with learned optimizers, a raw average of Update Success Rate (averaged over each kind of data), Retain Rate (Local Neutral) and Δ\Delta-Acc. The grid search is over the following parameters: The off-the-shelf optimizers are from torch.optim and include {AdamW, SGD, and RMSProp} with default arguments (except for the learning rate). We consider a number of maximum steps in {5, 10, 100}. The learning rates we consider depend on the optimizer: {1e-4, 1e-5, 1e-6} for AdamW, {1e-4, 1e-5, 1e-6} for RMSProp, and {1e-1, 1e-2, 1e-3} for SGD. The LR ranges were selected after some initial manual exploration of the space. Our final hyperparameter values are shown in Table 12 for each dataset. For comparison, De Cao et al. (2021) use RMSProp with 100 update steps. The LR for zsRE and Wikidata5m may seem quite high, but this is the condition that actually does the least damage to the model’s accuracy on other data, Δ\Delta-Acc. The baseline optimizes all of the trainable parameters in the language model, unlike the learned optimizer which optimizes only attention and feedforward weights for purposes of parameter efficiency.

B.3 Wikidata5m Additional Details.

We construct four paraphrases per Main Input by selecting from a set of alternative phrasings for the entity and relation in the Main Input. The syntax for each paraphrase follows the same simple template as the Main Input, in contrast to zsRE where syntax differs between paraphrases. A couple details remain. Some relations are one-to-many, and therefore we accumulate valid completing entities from the data as possible answers; later we compute accuracy as an exact match with any possible answer. All 10 relations appear in each split of the data. Only 33.80% and 37.18% of the entities in the dev and test splits are seen in the training data, though we do not find that models perform better on entities seen in training.

B.4 LeapOfThought Additional Details

The LeapOfThought dataset consists of a fact and a claim for each datapoint, where the truth of the fact implies that the claim has label yiy_{i} (True/False). All of the facts in the data are true, while half of the claims are true and half are false. When training the learned optimizer, we treat the the facts as the Main Input when training the learned optimizer and claims as entailed data. When training the True/False classifier, we fit to the claims, for which test accuracy is 83.65 (±\pm 1.05). This seems to generalize well to the facts, as test accuracy here is 93.66 (±\pm0.87), although as the low contrapositive accuracy suggests (Table 3), the model seems to be too prone to predicting true for this data.

Since very few of the Main Inputs are predicted as false, we run into a small dilemma when fitting the learned optimizer with the use of the entailed data objective term. The entailment between fact and claim only holds when the fact is true, so we can only compute the objective when updating a point from false to true. This ends up being less than 10% of the training data. We ultimately choose to oversample points that fit this description during training of the learned optimizer, which allows the learned optimizer to fully fit to the entailed data. Also note that during learned optimizer training, we include Entailed Data from other data points besides the Main Input in the KL term in Eq. 1, and we measure Δ\Delta-Acc using both Main Inputs and Entailed Data.

Appendix C Noise in Datasets

We briefly document some shortcomings of each dataset, with reference to examples in Table 15.

FEVER. Some claims are slightly vague or ambiguous when taken on their own. For instance “Doug Ducey was the CEO of Cold Stone Creamery and offered many opportunities to new hires" is rated True, though this will depend heavily on what one thinks “many opportunities" means. Similar whether or not “L.A. Guns is a tattoo shop" depends on which “L.A. Guns" one is referring to, the tattoo shop or metal band. Of course, this is a generic issue of language, and not unique to this dataset. Some inputs seem to be a matter of person opinion: “Los Angeles is known for its food" is rated False.

LeapOfThought. Many examples use an “is a" relation, producing sentences like “A sunlight is a good health." This could be more false than true, but it’s a fairly nonsensical statement to begin with. There are also other nonsensical or vague examples in the data: ”A friar is the opposite of mineral" is labeled False. “A detective desires equal opportunity." is labeled True. It is not immediately clear what conditions would make these statements true or false.

zsRE. Some questions invoke potentially one-to-many or temporally dependent relations, though there is only one ground-truth answer per question in this dataset. For instance, a paraphrase of the question about Gifford Pinchot in Table 15 is: "What disease did Gifford Pinchot have?" A person might have had many diseases over their life which could all be valid responses. The answer is especially ambiguous for spatial relations, where a valid answer might refer to a city, region, country, province, or continent.

Wikidata. Aliases sometimes vary greatly even as they refer to the same person, or they are simply noisy. For example, as shown in Table 15, “SusunW" appears in an entity name, but this is actually a username of someone who contributed to the Wikipedia article for Margarita Nolasco Armas. Meanwhile, other aliases for J.R.R Tolkien include “Tolkienian" and “Mabel Suffield," his mother. Rephrasings of relations might also create confusing inputs, e.g. switching “child" with “has kids," “daughter", or “son." Similar to zsRE, some relations are also one-to-many and temporally dependent (like occupation), though we hope that by using many valid answers we circumvent this issue to some extent when calculating prediction correctness.

Appendix D Metric Computation and Bootstrap Details

Metric computation. The only computationally difficult metric to calculate is Δ\Delta-Acc, which requires computing the updated language model’s accuracy on other data after every single belief update. We randomly sample other data after every update for this purpose, using n=30n=30 points for zsRE and Wikidata5m and n=200n=200 points for FEVER and LeapOfThought. We ensure that all evaluation data is used at some point during this sampling by preferentially selecting data that has been infrequently selected before. We note that paraphrase consistency is easy to evaluate for a small number of paraphrases per datapoint, as we have for both zsRE and Wikidata5m. Additionally, on LeapOfThought, we compute Δ\Delta-Acc using both Main Inputs and Entailed Data.

Update-Transitivity caveat. The % Update-Transitivity metric represents the answer to the question: if updating belief A changes belief B, and updating belief B changes belief C, what proportion of the time does updating A change C? We would treat this as a normative metric that we hope to maximize, except we do not know in general whether there is a confounding belief D that determines the relationship between B and C. If changing A also changed a confounding belief D, then we might not be able to expect that C should change too. That said, when we have no reason to think there are such confounding beliefs, we would expect a logically consistent model to display 100% Update-Transitivity of their beliefs. In Fig. 3, for instance, we see no reason to suspect there are confounding beliefs for the relationship between the date Bessie Smith died and the writer of Despicable Me 2, and therefore we would expect that updating the belief about what album Hot Right Now is on would change the belief in Despicable Me 2’s authorship (which it does).

Bootstrap computation. We account for sample and seed variance by block bootstrap Efron and Tibshirani (1994). When there is a single statistic per data point, like Main Input Update Success, we form a matrix of shape n×sn\times s for nn data points and ss model seeds (where the seed was used for both task model training and learned optimizer training). We then resample rows and columns of this matrix 10,000 times, which was sufficient for convergence. When we perform hypothesis tests for the difference in statistics between conditions, we pair the data points by using the same rows of this matrix at each step of the bootstrap (i.e. we conduct paired tests). For metrics involving multiple data points per Main Input, like paraphrases or other random data, we make a simplifying assumption where we do not resample the multiple data points but just compute the average metric for those data points and treat that as the ground-truth statistics for the Main Input. We explored using a full 3-dimensional bootstrap, where we resample among these extra datapoints by constructing a matrix of shape n×s×nn\times s\times n, but it was quite slow and gave similar results to the block bootstrap.

Appendix E Additional Results

Ablation across num. sequential steps. Fig. 6 shows the results for an ablation across rtestr_{\textrm{test}} using two kinds of learned optimizers: SLAG1, where rtrain=1r_{\textrm{train}}=1, and a SLAG condition where rtrain=rtestr_{\textrm{train}}=r_{\textrm{test}}. It is critical to the success of learned optimizers to train them to update points sequentially when this is a desired application. Further, sequential updating with sequence prediction tasks is the only setting where we see learned optimizers outperform baselines across all relevant metrics.

Choosing training labels for learned optimizers. In early experiments, we found that it is beneficial to use all data points (including correctly predicted points) as Main Inputs during training, rather than restricting training to only incorrectly predicted points. We still focus on correcting wrong outputs at test time. But so we must select what label to use during optimizer training. To get a Hard Label, we use the correct label for incorrectly predicted points, and for correctly predicted points, we simply draw a label randomly from the labels in the training data. The alternative Beam Label condition uses a sample from the model’s beam search for a data point, as done in past work De Cao et al. (2021); Mitchell et al. (2021). We show update metrics for zsRE split by the desired label in Table 16. If one’s goal is to fix wrong model outputs, then it is much better to use either the correct label or a random label as the desired model output during training rather than a sample from the model’s beam search. Update success improves by 3.27 (±\pm0.65; p<p{<}110−4110-4) points for the Main Input and 2.38 (±\pm1.05; p<p{<}110−4110-4) for Paraphrases, while Δ\Delta-Acc rises by 0.15 (±\pm0.18; p=.09p{=}.09).

Which beliefs are hard to update? We hypothesize that beliefs will be easier to update when they are more belief-like to begin with. We principally measure this via the correlation between update success rate and a belief’s consistency on paraphrases before the update, for our learned optimizer in a single-update setting (r=1r=1). Surprisingly, we observe no relationship between update success and the belief consistency. The correlation between consistency and update success is near 0 for both zsRE (ρ=−.027\rho=-.027) and Wikidata5m (ρ=.013\rho=.013); see Fig. 7 for a plot of the relationship. So it appears that the learned optimizer can update model beliefs independently of how belief-like they are to begin with. We would also be interested in considering consistency under entailment, but the update success rate on LeapOfThought is already 100%, so there is no variance to explain.

Learning curve. In Fig. 8 we show the learning curve of a learned optimizer trained with SLAG on zsRE. The Main Input Update Success Rate steadily rises as a function of the training set size.

Ablation by num. update steps. Fig. 9 shows the results of an ablation across values of KK using a learned optimizer trained using SLAG with r=1r=1 on zsRE. Main Input Update Success rises by over three points by increasing KtestK_{\textrm{test}} from 1 to at least 5. Using a value of KtrainK_{\textrm{train}} that matches KtestK_{\textrm{test}} gives a further increase of about 0.5 points.