Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability

Afra Feyza Akyürek, Ekin Akyürek, Leshem Choshen, Derry Wijaya, Jacob Andreas

Introduction

There is increasing interest in using language models (LMs) as sources of information and tools for fact verification (Porter, 2023). But today’s LMs cannot robustly perform either task: they are prone to generating factually incorrect information, contradict themselves, and are difficult to update with new information (Lin et al., 2021; Liska et al., 2022; Sun et al., 2023; Gilson et al., 2023).

Even if they are imperfect judges of factuality, however, current LMs are quite reliable models of factual relations between pieces of text: they can identify logical and probabilistic relationships between statements, and generate text based on new information provided as input (Williams et al., 2017). For example, an LM that cannot answer the question How old was Charlie Chaplin when he died? may nonetheless provide a correct answer when prompted with Charlie Chaplin lived between 1889 and 1977, and recognize that this statement contradicts the claim Charlie Chaplin lived in the 21st century. How can we leverage LMs’ ability to reason about factual relations between claims to improve (and control) the text that LMs themselves generate?

Conceptually, even if a piece of information appears in an LM’s training data, updating the LM’s parameters to increase the probability of sentences containing this information may not, in general, increase the probability of sentences describing its logical consequences: fine-tuning on a definition need not cause an LM to assign higher probability to uses of the newly defined word. In general, reasoning is required to determine the deductive closure of a given training set (Armstrong, 1973)—the complete collection of inferences that can be made given the information initially available. Standard likelihood maximization approaches to LM training do not perform this reasoning by default, so some alternative procedure is needed to ensure that LMs assign high probability to a complete and consistent set of facts when they are trained and fine-tuned.

In this paper, we propose a new LM fine-tuning procedure we call Deductive Closure Training (DCT), which leverages inference-time reasoning as a source of training-time supervision. At high level, given seed text (which may be provided externally or LM-generated), DCT uses an LM to identify additional text implied by or contradicting this text, reason globally about which portions of seed and generated text are most likely to be correct given this context, and finally fine-tune on inferred-correct text. This approach builds on a large body of recent work (Mitchell et al., 2022b; Kassner et al., 2023; Hase et al., 2023) on inference-time procedures for improving models’ factual correctness, showing that these techniques may be used at training time as well.

DCT may be used in several different ways depending on the source of seed documents. If these are drawn from a trusted factual source, DCT may be used to perform supervised adaptation for factuality. If documents contain new information to be inserted into an LM, DCT provides tool for model updating (or “editing”; De Cao et al., 2021). Finally, if seed documents are generated by the model itself, DCT enables fully unsupervised fine-tuning of models for improved accuracy.

We demonstrate the effectiveness of DCT across three domains: fact verification (on the CREAK benchmark; Onoe et al., 2021), question answering with new information (on the MQUaKe benchmark; Zhong et al., 2023), and a synthetic test of edit propagation (on the “Reversal Curse” benchmark; Berglund et al., 2023). On these tasks, unsupervised application of DCT improves accuracy by up to 12%, while supervised application improves accuracy by up to 26%. These results suggest that problems such as model coherence and editing may not require specialized editing or training techniques: self-supervised objectives that optimize for coherence and completeness of LM outputs can improve their accuracy and updatability.

Related Work

DCT builds on a number of recently developed techniques for improving model accuracy via inference-time computation or training-time self-supervision.

A growing body of research adopts techniques that bootstrap language model performance at inference time. Tafjord et al. (2022); Kassner et al. (2023); Bostrom et al. (2022); Weir and Van Durme (2022) and Jung et al. (2022) build self-guided semantic chains of reasoning to support inference. Suzgun et al. (2022) propose a set of procedures that bin model-generated candidate answers by semantic equivalence and later uses aggregated probabilities to select the highest ranked predictions, analogous to self-consistency Wang et al. (2023) for textual outputs. Finally, recent work has shown promise in improving coherence by conditioning language models on relevant reference texts through retrieval augmentation Mitchell et al. (2022a); Akyürek et al. (2023). Our approach builds on this line of work by using inference-time techniques to generate supervision.

Supervised learning for factuality

LMs greatly benefit from training or post-training techniques for improving accuracy, including instruction-tuning Sanh et al. (2022), learning from feedback Ouyang et al. (2022) and loss truncation Kang and Hashimoto (2020). Closest to our approach is the work of Hase et al. (2023) which (like DCT) leverages graph-structured representations of model “beliefs” about the factual neighborhood of statements, but use these to train a hyper-network for model editing. DCT aligns with this thread in aiming to improve model training; it differs by requiring minimal or no external supervision.

Self-training

Past work has also studied leveraging LMs themselves for performance improvements Pan et al. (2023). Several studies use external tools Schick et al. (2023), binary feedback Pang et al. (2023); Liu et al. (2023) and natural language feedback Bai et al. (2022) to improve capability or reduce harms. Others propose actuality and consistency metrics, which might be used for filtering bad answers in retrospect (Honovich et al., 2021; Wang et al., 2020; Honovich et al., 2022). Related to such approaches are methods that perform multiple inference attempts and aggregate them to get a more consistent answer (Wang et al., 2022; Yoran et al., 2023). Padmanabhan et al. (2023) fine-tune LMs on self-generated text without explicit implication generation or logical inference. Of immediate relevance to the current work, Li et al. (2023) and a concurrent study by Tian et al. (2023) use LM-generated factuality labels to rank or filter LM-generated data for fine-tuning; by contrast, DCT uses LMs to explicitly extrapolate from LM-generated or externally provided information, providing a single framework for both supervised model updating and unsupervised improvement.

Method

2 Document Generation

3 Consistency Evaluation

Formally, we first associate with the seed document sis_{i} and every generated document rijr_{ij} a truth value tij∈{0,1}t_{ij}\in\{0,1\}. Given an assignment of documents to truth values, we compute the LM’s probability of this assignment as:

Next, we define a value assignment Ti={tij}T_{i}=\{t_{ij}\} to be consistent if all implications and contradictions are respected.

where tit_{i} denotes the truth value of the seed document, 1[a→b]1[a\to b] is 1 iff bb is true or aa is false, and 1[a↛b]1[a\not\to b] is 1 iff bb is false or aa is false (as in the ordinary definition of logical implication and contradiction). Finally, we select the most probable consistent assignment:

The procedure is depicted in Fig. 2b, with the highest-scoring truth value assignment shown in the blue-highlighted box. For fact-verification tasks, it is possible to derive positive supervision from statements marked as false: if the consistency evaluation step infers that Meiji was the last Japanese emperor is incorrect, then we may generate a correct example of the form Verify the following statement: Meiji was the last Japanese emperor. False. We use this strategy for our experiments on fact verification.

4 Language Model Fine-Tuning

5 Sources of Seed Data

Depending on how seed documents SS are obtained, DCT-based fine-tuning may be used to improve models in several ways:

(Semi-)supervised alignment with a trusted source: in this case, the seed set comes from an external source of supervised data. If this data is known to be reliable, we fix each seed datum’s truth value ti=1t_{i}=1 during the evaluation step. This may be combined with the unsupervised procedure.

Model updating, editing and continual learning: in this case, as with supervised updating, we treat textual descriptions of desired edits as seed documents, again fix these truth values for these seeds to 1, and fine-tune both on these documents and all their implications only.

Note that in the latter two cases (where we have fixed the truth value of seed documents to 1), the evaluation step is greatly simplified, and simply discards all generated documents that are not logically consistent with the seed. In the case of unsupervised learning, this evaluation step can (and empirically does) cause LMs to re-label sampled seed documents as well as conditionally generated ones.

Finally, we remark that the procedure described above is the basic implementation of a family of DCT-like approaches, within which many more sophisticated procedures are possible—for example: probabilistic DCT (computing marginal statement probabilities rather than hard truth assignments), contrastive DCT (replacing Eq. 3 with an objective that encourages true statements to be assigned higher probability than false ones), and multi-hop DCT (generating not just direct implications of documents, but a wider graph of related ones).

Formal Analysis of DCT

At first glance, it may seem surprising that this procedure (especially in its unsupervised form) can improve LM accuracy using only LM-generated text. In this section, we describe a set of assumptions under which DCT is guaranteed to improve accuracy on certain inputs. We focus this analysis on generation and evaluation of (question, answer) pairs, but it could be extended to the other tasks considered in this paper as well.

Questions generated by the LM with high probability are likely to be correct. (Intuitively, high-probability questions will be ones that occurred frequently in the training set, and are therefore more likely to be answered correctly; McCoy et al., 2023, though c.f. Lin et al., 2021.)

Given a question, prompting an LM with a related, correct question–answer pair increases the probability of a correct answer. (Intuitively, such prompts may steer models generally in the direction of truthfulness, as in Lin et al., 2021, and can provide concrete evidence useful for answering the new question.)

We wish to show that if these two conditions hold, DCT improves model performance.

p(a0∗∣q,q0)≥p∗p(a^{*}_{0}\mid q,q_{0})\geq p^{*}. (Conditioned on generating qq during the document generation step of DCT, the probability that the generated answer to any seed question q0q_{0} contains a correct answer is (uniformly) at least p∗p^{*}.)

Experiments

We evaluate Deductive Closure Training on a set of benchmark tasks measuring fact verification, question answering with new information, and a diagnostic model editing dataset. We use Llama-2-7B in all experiments.

We first evaluate whether AC improves models’ ability to classify factual claims as correct or incorrect. Our experiments use CREAK Onoe et al. (2021), a dataset of claims about entities. We investigate four different learning settings: unsupervised, supervised, semi-supervised, and transductive, each using a different procedure for sampling seed documents. We report results on the CREAK development set. During DCT fine-tuning, we use a linear learning rate schedule until the training loss converges—this corresponds around 30 epochs for the majority of experiments unless otherwise indicated (see Appendix A for further details on experimental settings).

Evaluation and baselines

Models are scored based on the fraction of claims they correctly label as true or false. For each data condition, we compare to a state-of-the-art baseline. For unsupervised DCT, the baseline is an ordinary few-shot prompt. For supervised DCT, the baseline fine-tunes the LM on the provided true statements. For transductive DCT, we also compare to an inference-time baseline Graph-Inference similar to those described by Mitchell et al., 2022b and Kassner et al., 2023, which generates a set of implications and contradictions for each test example, performs reasoning as in Eq. 2, then directly outputs the inferred truth value for the example (with no fine-tuning). Unlike past work, we use the base LM to generate these graphs rather than a specialized pre-trained implication generation model. All results are presented in Table 1.

Results: unsupervised DCT

For these unsupervised experiments, we perform an additional evaluation specifically aimed at measuring logical coherence as well as factual accuracy. Here we use the contrast set in CREAK, which comprises 250 pairs of lexically similar examples which have opposite truth values (for example Jason Smith was raised in Ireland and Jason Smith was raised in Scotland). In addition to accuracy, we compute the fraction of pairs that are labeled Both True (indicating incoherence), and the fraction of pairs labeled Both Correct.

Results: supervised & semi-supervised DCT

In the supervised case, we utilize a small set of externally provided claims and associated ground-truth labels to initialize DCT seed nodes. We sample 20 claims from the CREAK training set and filter those labeled as true to use as our seed documents DD. For semi-supervised learning, we pool together data generated following the unsupervised and supervised settings for fine-tuning.

All variants of DCT improve over an ordinary fine-tuning baseline; interestingly, examples generated supervisedly and self-supervisedly are complementary, such that semi-supervised learning improves over both results.

Results: transductive DCT

The previous evaluations assumed a strict train / test split. Here we study the behavior of AC in a “tranductive” setting (Gammerman et al., 1998) in which we have access to unlabeled claims from the evaluation set while updating the model. For each claim in the validation set, we generate seed text by prompting the LM to generate a set of related claims, which are then used to generate additional implications and contradictions. In addition to the inference-time baseline described above, these experiments compare to an ablated version of AC that trains only on the generated related claims.

As in other experiments, DCT outperforms the inference-time reasoning baseline as well as the related-text-only ablation.

2 Model Updating and Question Answering

Language models often hallucinate wrong information and rapidly become out-of-date after initial training. As a consequence, there has been increased interest in specialized continual learning (or “model editing”) procedures for updating LMs with new information provided in natural language without full re-training. A key desideratum is LMs should not simply assign high probability to the new fact, but all of its consequences: if we wish to update an LM encode the fact that the current U.K. prime minister is not Boris Johnson but Rishi Sunak, the LM should also produce text consistent with the fact that the current P.M.’s wife is not Carrie Johnson but Akshata Murthy. Past work has found that fine-tuning on edits, as well as many specialized editing procedures, fail to propagate information in this way.

Our experiments on this task use the counterfactual subset from MQUaKe Zhong et al. (2023) dataset, which evaluates models on their ability to answer questions about new information not provided in their training sets. To apply DCT to these model updating applications, we take as seed documents the text of the new information to be inserted into the model. During the generation phase, models are prompted to combine this information with other background knowledge related to the same topic (see Appendix B for prompting details), producing what we term Correlative Implications. Finally, because MQUaKe is a question answering dataset, we convert each generated statement into a question–answer pair using the LM, then fine-tune it on these pairs.

Evaluation and baselines

We compare DCT to ordinary fine-tuning on new information and two state-of-the-art baseline approaches for model updating: a self-training baseline by Padmanabhan et al. (2023), which fine-tunes LMs to behave out-of-context the same way they would with prompts containing the new information (see Appendix A for implementation details), and the retrieval baseline MeLLo Zhong et al. (2023), which stores new text in an external memory. We evaluate the behavior of DCT and these baselines in settings where varying numbers of new pieces of information (between 10 and 1000) are provided, and report the model’s accuracy at question answering.

Results

As shown in Table 3, DCT significantly outperforms fine-tuning, fine-tuning on continuations, and MeLLo (the previous state-of-the-art on MQUaKe). Using correlative implications systematically improves over simple implications. Combining the two sets improves on average over using either in all settings. Our qualitative analysis in Section 6 reveals that using correlative implications in statements that contain about 50% more new information than using standard implications.

3 Sanity Checks for Consistent Model Updating

In addition to naturalistic question asking tasks like MQUaKe, there has been recent interest in developing precise tests of LMs’ ability to capture simple logical implications of new facts (e.g. assigning high probability to sentences of the form B is A after training on sentences of the form A is B). We investigate whether DCT can address these issues using the “Reversal Curse” benchmark (Berglund et al., 2023). We report results on two evaluations: first, a set of celebrity parent–child pairs with training examples Jennifer Lawrence’s mother is Karen Lawrence and test examples Who is the child of Karen Lawrence?; second, a set of entity–description pairs with training examples Olaf Scholz was the ninth Chancellor of Germany and cloze-style test examples The ninth Chancellor of Germany is .

Evaluation and baselines

For these experiments, we compare to the fine-tuning baseline used in the original work of Berglund et al. (2023) as well as the fine-tuning on continuations approach by Padmanabhan et al. (2023) used in previous experiments. We use training examples as seed statements, and generate implications using the same prompt as CREAK experiments in 5.1. While we expect that an DCT-type approach specifically tailored for this benchmark could trivially re-generate all the test examples, our experiments in this section aim to evaluate whether a general-purpose prompt can improve performance on a specific class of generalizations. Following Berglund et al. (2023), we report exact-match accuracy after removing punctuation and lower-casing. In this dataset, LMs are evaluated on a mix of questions and cloze completion tasks featuring both training statements and their reversed forms.

Results

Results are shown in Table 4. On average, DCT improves accuracy on reversed statements without significantly hurting performance on original questions. Notably, however, DCT with this general-purpose prompt does not completely solve this dataset, and we leave for future work the question of whether more extensive sampling or other procedures could further improve these results.

Qualitative Analysis

To better understand how DCT improves LM performance, we manually annotated about 350 generations from various experiments to assess whether (1) double-checking improves the precision of generated implications and contradictions; (2) whether DCT incorporates model internal knowledge when making new conclusions; and (3) whether generated text includes non-trivial new inferences.

We evaluated whether the double-checking following DCT (Imp. + Cont.) improves precision. In the supervised setting for CREAK, we annotated 100 implications and contradictions generated using DCT (Imp. + Cont.). We found that 74 of these are valid. The double-checking procedure removes about 2/3 of generations, resulting in 33. Among these, 27 are valid, raising the ratio of correct statements predicted by the model from 76% to 82%.

Incorporating previous information

The MQUaKe subset used in our experiments comprises difficult multi-hop questions. Hence, generations that incorporate existing information about the entities mentioned in the edit are especially useful. We compare the set of implications generated using the DCT (Imp.) and DCT (Corr. Imp.). Respectively, only 30% and 36% of generations involve strict logical implications; however, 78% and 69% were judged to be plausible given the edit. Furthermore, 24% and 33% of the generations incorporate new information supplied by the LM. For example, given an edit Chauncey Billups is associated with the sport of pesäpallo, the LM uses background knowledge Pesäpallo is popular in Finland to generate Chauncey Billups was born in Finland.

Novelty of inferences

Lastly, we find that the implications made by the model on the “Reversal Curse” dataset vary from trivial (Jennifer Lawrence’s mother is Karen Lawrence →\to Jennifer Lawrence has a mother) to rarer ones adding world knowledge to the implication (Sadie Frost’s mother is Mary Davidson →\to Mary Davidson is the mother of a British actress, where the LM itself has supplied background knowledge about Sadie Frost). While generating implications, DCT often (but not always) generates test-set-like reversed implications on its own: the model reverses 22% of the statements of the form X’s parent is Y, 43% of statements of the form the person with property X is Y, but only 6% of statements of the form Person X has property Y. These findings suggest strong bias towards generating text that actually starts with the person as opposed to description. In general, most generated extensions are fluent, different from the source and sometimes contain new information.

Conclusion

We have described Deductive Closure Training (DCT), a supervision procedure that optimizes models toward deductive closure—encouraging them to assign high probability to a logically coherent set of factual assertions, as well as all their implications. By doing so, DCT also improves the truthfulness and updatability of models, substantially increasing accuracy on a variety of fact verification and model editing datasets in both supervised and unsupervised conditions. More generally, these results show that some factual errors in LMs stem not from limitations of their training data, but limitations of training algorithms. By using LMs themselves to reason about relationships between (and implications of) their predictions, they can be made more accurate with little or no additional supervision.

Ethical Considerations

While our experiments have focused on using DCT as a tool for bringing LMs into alignment with reliable sources, these techniques could also be used to optimize LMs toward generation of (logically consistent) false facts, increasing their effectiveness as tools for generation of misinformation.

Acknowledgments

This work was supported by the National Science foundation under grant IIS-2238240, a hardware donation from NVIDIA to MIT, as well as the Shared Computing Cluster administered by Boston University’s Research Computing Services.

References

Appendix A Experimental Details

We use the Llama-2-7B-hf checkpoint provided by HuggingFace Transformers library for all of our experiments.

We sample at temperature 0.6 and top-p 0.9 for all samples except for the set of seed documents for the unsupervised experiment in Table 1 where we used temperature 0.9 to obtain a diverse set of initial documents.

Training

For fine-tuning we use the LoRA implemention via the PEFT library Hu et al. (2022); Mangrulkar et al. (2022) and set rank to 8, alpha to 32 and dropout to 0.1. In the absence of a held-out development set, we set the learning rate to 0.0001 throughout, batch size to 4 and train for 30 epochs by default. We find that training loss typically converges after 30 epochs with the exception of the supervised experiments in Table 1 for which we train for 60 epochs. The transductive setting for CREAK results in substantially more training documents, hence we train only for 1 epoch. We use a linear learning rate scheduler with 100 warm up steps and AdamW optimizer. For fact verification training, we use weighted sampling as the class distribution is sometimes unbalanced.

Editing experiments

We use the MQUaKe-CF subset from Zhong et al. (2023) and evaluate only on the multi-hop questions. Padmanabhan et al. (2023) proposes two techniques to introduce model updates based on fine-tuning we call FT on Continuations and context distillation. We find the former approach–fine-tuning the model on the continuations when the model is conditioned on the edit sequence–to perform better on MQUaKe than distillation.

Appendix B Prompt Templates

We use a set of fixed prompts to generate our graphs, calculate model-estimated probability for the correctness of a given statement, generating a set of seed documents and automatically converting statements into questions which are available in Tables 5, 6, 7 and 8.

Appendix C Proof of 1