Falsesum: Generating Document-level NLI Examples for Recognizing Factual Inconsistency in Summarization

Prasetya Ajie Utama, Joshua Bambrick, Nafise Sadat Moosavi, Iryna Gurevych

Introduction

Recent advances in conditional text generation and the availability of large-scale datasets have given rise to models which generate highly fluent abstractive summaries Lewis et al. (2019); Zhang et al. (2019). However, studies indicate that such models are susceptible to generating factually inconsistent outputs, i.e., where the content of the summary is not semantically entailed by the input document Kryscinski et al. (2019); Goodrich et al. (2019). This motivates a new line of research for recognizing factual inconsistency in generated summaries Kryscinski et al. (2020); Pagnoni et al. (2021); Wang et al. (2020); Fabbri et al. (2021).

This factual consistency problem is closely related to the task of natural language inference (NLI) whereby a hypothesis sentence is classified as either entailed, neutral, or contradicted by a given premise sentence Condoravdi et al. (2003); Dagan et al. (2006); Bowman et al. (2015). Using an input document as the premise and a corresponding generated summary as the hypothesis, earlier solutions have adopted out-of-the-box NLI models to detect factual inconsistency, albeit with limited success Falke et al. (2019); Kryscinski et al. (2020).

This poor performance largely stems from the fact that most NLI datasets are not designed to reflect the input characteristics of downstream tasks Khot et al. (2018). Such datasets may not always capture the kinds of entailment phenomena which naturally arise from neural abstractive summarization. More importantly, there is also a discrepancy in terms of the input granularity, i.e., the premises in this consistency classification task consist of multi-sentence documents while common NLI datasets use single-sentence premises.

In this work, we introduce Falsesum, a data generation pipeline that produces NLI examples consisting of documents paired with gold summaries as positive examples and automatically generated inconsistent summaries as negative examples. We propose a novel strategy to train a text generation model to render false summaries of a given document using only supervision from an existing summarization dataset Nallapati et al. (2016). In addition, our generator supports switchable input control codes to determine the type of factual error exhibited in the generated output. This design allows Falsesum to compose diverse and naturalistic outputs which more closely resemble the inconsistent summaries generated by summarization models Maynez et al. (2020). This contrasts with previous solutions (e.g., Kryscinski et al., 2020; Yin et al., 2021), which synthesize NLI examples using rule-based transformations or language model-based replacements, limiting their diversity and ability to reflect realistic factual errors in summarization. Overall, our contributions in this paper are the following:

First, we present a novel training pipeline to create a text generation model which takes as input a pair of a document and a corresponding gold summary. It then perturbs the summary such that it is no longer factually consistent with the original document. Our strategy obviates the need for explicit examples of inconsistent summaries, using only an existing summarization dataset. We use this model to generate a large-scale NLI dataset for the task of recognizing factually inconsistent summaries. The resultant dataset consists of pairs with documents as the premise and naturalistic summaries as the hypotheses, each labeled as either entailment or non-entailment.

Second, we demonstrate the utility of our generated data for augmenting existing NLI datasets. We show that on four benchmark datasets, NLI models trained on Falsesum-augmented data outperform those trained on previous document-level NLI datasets. We conduct an analysis to show that Falsesum-generated summaries are plausible and hard to distinguish from human-written summaries. Lastly, we show that the improvement over the benchmarks is largely attributable to the diversity of factual errors that Falsesum introduces.

Related Work

This work is related to the growing body of research into factual consistency and hallucination in text generation models, particularly for summarization Cao et al. (2018). Research has found that around 30% of summaries generated by abstractive summarization models contain information which is inconsistent with the source document Kryscinski et al. (2019). This motivates the development of an automatic approach to assess factual consistency in generated summaries, in addition to the benchmark datasets to measure the progress in this task Falke et al. (2019); Kryscinski et al. (2020); Pagnoni et al. (2021); Fabbri et al. (2021).

Earlier work by Goodrich et al. (2019) proposes to use an information extraction model to extract relation tuples from the ground-truth summary text and the generated summary and then count the overlap as the measure of factuality. Eyal et al. (2019); Durmus et al. (2020); Wang et al. (2020) use a question-answering model to detect factual inconsistency by matching the predicted answers using the document and the summary as the context.

Concurrently, researchers have drawn a connection between factual consistency and natural language inference (NLI), observing that all information in a summary should be entailed by the source document. While this approach enables the summary to be directly evaluated without first extracting its intermediate semantic structure, earlier attempts were largely unsuccessful. Falke et al. (2019) use the probabilities assigned to the entailment label by NLI models to re-rank the summary candidates given by beam search but found no improvement in the consistency errors. Kryscinski et al. (2020) evaluate out-of-the-box NLI models on the task of inconsistency detection in a binary classification setting and show that the performance is only slightly better than majority voting.

In the same paper, Kryscinski et al. (2020) propose FactCC, a synthetic NLI data generation process which applies a set of transformation rules to obtain examples of inconsistent summaries (e.g., sentence negation, entity swapping). They demonstrate that the resulting NLI model performs well on realistic test cases which are obtained by manually annotating the output of several summarization models. This highlights the importance of NLI examples beyond sentence-level granularity and which more closely resemble the input characteristics of the downstream tasks Mishra et al. (2021).Contemporaneous work by Laban et al. (2022) attempts to improve the application of sentence-level NLI models to detect document-level factual inconsistencies using a learnable aggregation of sentence-level predictions. Our work is orthogonal since they can benefit from better quality training examples to train their aggregation weights.

While the FactCC model is moderately effective for detecting factual inconsistency, subsequent work indicates that it only performs well on easier test cases, where highly extractive summaries (i.e., those with high lexical overlap between a summary and the source document) tend to be factually consistent and more abstractive summaries are likely to be inconsistent Zhang et al. (2020). Furthermore, Goyal and Durrett (2021) show that the synthetic and rule-based nature of FactCC leads to lack of diversity of consistency error types and it poorly aligns with the error distribution found in more abstractive summaries.

Falsesum addresses these limitations using controlled natural language generation to construct an NLI dataset which better targets the summarization domain. Inspired by the recent work on controllable generation Keskar et al. (2019); Ross et al. (2021), we employ a generation model conditioned on an input code which controls the type of consistency errors induced. We further use the generated document-level NLI examples for augmentation and show that NLI models can benefit from the additional data without hurting their existing inference ability Min et al. (2020).

Falsesum Approach

As illustrated in Figure 1, a generation model G\mathcal{G} is trained to imitate the consistency mistakes of summarization models. Specifically, it generates perturbed summaries by either (1) incorrectly inserting pieces of information from the source document into random spans of the original summary; or (2) amending pieces of information in the summary by hallucinating new “facts” not present in the source document.

To this end, the framework identifies (♢\diamondsuiti) what information or “facts” in the source document are available to the generator; and (♢\diamondsuitii) where the incorrect information can be inserted into the gold summary, which is indicated by span masking. We obtain both by subsequently performing input preprocessing and formatting steps (§3.2 and §3.3).

Next, we define the following seq2seq task to train the model G\mathcal{G}: “Given (♢\diamondsuiti) a list of shuffled and formatted pieces of information extracted from source document and gold summary and (♢\diamondsuitii) a partially masked gold summary, fill in the blanks and generate the original gold summary.” Note that using gold summaries means that we can apply the existing summarization corpus to train G\mathcal{G} to generate more coherent and plausible sentences.

2 Input Preprocessing

Following Goodrich et al. (2019), “facts” in the source document and the gold summary are defined as an open information extraction (OpenIE) tuple, which represents the predicate and argument structures found in a sentence. We denote each relation tuple as (\textscarg0,\textscpred,…,\textscargn(\textsc{arg}_{0},\textsc{pred},\dots,\textsc{arg}_{n}), where predicate pred describes the event (what happened) and its complementing semantic arguments arg represent the who, to whom, where, or how of the event. Predicates are usually the main verb of a clause. Both predicates and their arguments consist of spans of tokens Fader et al. (2011).

We use an OpenIE implementation of PredPatt White et al. (2016); Zhang et al. (2017), a pattern-based framework for predicate-arguments extraction.We note that the quality of the OpenIE extractions may impact the overall quality of our data generation framework. As illustrated in the top half of Figure 2, we extract the relation tuples from each source document and its corresponding reference summaries. To minimize the risk of G\mathcal{G} inadvertently generating consistent summaries, we corrupt each extracted “fact” by removing one randomly chosen argument from each tuple. For instance, OpenIE may extract the following tuple from a sentence:

We then randomly choose apples\textscARG2\texttt{apples}_{\textsc{ARG}_{2}} to be removed from the tuple. We additionally lemmatize the dependency root word of each argument and predicate span, e.g., plans to give⇒plan to give\textbf{plans to give}\Rightarrow\textbf{plan to give}. This forces the model to learn to correct for grammaticality by inflecting the spans when inserting them to the masked spans. Once all such spans are extracted and processed, they are grouped and shuffled into two lists (predicates and arguments).

3 Input Formatting

In the following, we describe the key steps in the input formatting process:

Step 2: Span Reduction

Step 3: Control Code

To control the type of consistency errors generated by G\mathcal{G}, we append the string “code:” followed by either “intrinsic” or “extrinsic” into the input tokens. The code is determined randomly with equal probability of 0.50.5. Once the code is chosen, we perform the remaining formatting steps accordingly (see Table 1).

Step 4: Summary Masking

4 Training Falsesum

We run the Falsesum data generation pipeline on the train split of the CNN/DailyMail corpus Hermann et al. (2015), originally collected for question answering, but subsequently reformulated for summarization by Nallapati et al. (2016). This dataset contains English news documents paired with human-written summaries, each consisting of multiple sentences. We break the summaries down such that each Falsesum example consists of the document text and a single sentence summary. We then run the preprocessing and formatting steps on each document-summary pair. The resulting pairs of formatted input and target output are subsequently split into train and test sets which consist of 394,774 and 262,692 instances, respectively.

We use the T5-base model Raffel et al. (2020) as generator G\mathcal{G} and fine-tune it on the seq2seq task described in §3.1. The NLI examples are produced by running the fine-tuned generator on the preprocessed and formatted test split.See Appendix A for the hyperparameter details. This renders an equal number of positive and negative examples. In our experiments, we randomly sample 100,000 Falsesum examples to augment the NLI dataset.

Experimental Settings

Our experiments aim to demonstrate the effectiveness of Falsesum-generated document-level examples for NLI dataset augmentation. We evaluate the downstream performance of the NLI models by testing them against several benchmarks for determining the factual inconsistency of generated summaries. In this section, we describe the training setup of the NLI models, including the model and both the sentence- and document-level datasets.

Document-level NLI datasets

We conduct augmentation comparisons with several multi-sentence NLI datasets which obtain examples from news or summarization domains. We consider the following datasets: ANLI Nie et al. (2020), a paragraph-level NLI dataset collected via an iterative and adversarial human-in-the-loop annotation protocol. It consists of mostly Wiki data but also includes a small portion of news text; DocNLI Yin et al. (2021), a document-level NLI dataset containing multi-sentence premise and hypothesis sentences, collected by converting QA examples to NLI instances Demszky et al. (2018) and replacing words and sentences in news summaries using a language model; FactCC Kryscinski et al. (2020), a large-scale dataset specifically generated for training summary factual correctness classification models. The positive examples in FactCC are obtained by backtranslating a random sentence from a CNN/DailyMail news story, while negative examples are obtained by perturbing the sentence using predefined rules, e.g., entity swapping. For fair comparison, we sample 100,000 examples from each augmentation dataset in our experiments.

2 Benchmark Datasets

We evaluate these NLI models on four benchmark datasets to classify the factual consistency of abstractive summaries. These datasets differ in terms of the annotation protocol, the granularity of the summaries (single- or multi-sentence), the summarization corpus used, and the models used to generate the summaries that are annotated. The tasks are formulated as a binary classification with the labels “consistent” and “inconsistent”. We evaluate NLI models on these tasks by mapping the predicted label “entailment” to “consistent” and “non-entailment” to “inconsistent”. The benchmarks datasets are detailed in the following:

In addition introducing a synthetic training dataset for the task, Kryscinski et al. (2020) introduce a manually annotated test set. It contains 1,431 document and single-sentence summary pairs generated by various neural abstractive summarization models trained on CNN/DailyMail corpus.We merge the test and validation sets into a single test set.

Ranksum

Falke et al. (2019) formulate the factual consistency problem in summarization as a ranking task. They introduce a dataset consisting of 107 documents, each paired with a set of five ranked summary candidates obtained from the beam search of a summarization model. Given the manually annotated consistency label on summary candidates, the task is to re-rank the list such that the top-1 summary is factually consistent.

Summeval

Fabbri et al. (2021) introduce a comprehensive benchmark for factual consistency detection in summarization. It includes summaries generated by seven extractive models and sixteen abstractive models, which are judged by three annotators using a 5-point Likert scale.We aggregate the label as “consistent” if all annotators rated the summary as a 5 and “inconsistent” otherwise.

QAGS

The dataset collected by Wang et al. (2020) consists of 239 test set instances from XSUM Narayan et al. (2018) and 714 instances from CNN/DailyMail.This is the number of instances after we split multi-sentence summaries into separate single-sentence summary test instances, where an individual factuality judgement is available. Each instance consists of a pair of a source document and a single-sentence summary, which is labeled via majority voting on three annotators’ labels.

Results and Discussion

Performance on FactCC, QAGS, and SummEval is measured using balanced accuracy, which is suitable for class imbalanced settings, since the factually consistent label is the majority in some benchmark datasets. It is defined as the average recall of the two classes, such that majority label voting obtains only a 50% score. To measure ranking performance in Ranksum, we calculate the average Precision@1, which computes the fraction of times a factually consistent summary is ranked highest on each test instance. We perform five training runs for each setup using different random seeds and take the mean to address performance instability Reimers and Gurevych (2017).

From the results in Table 2, we observe the following: (1) Models trained on sentence-level MNLI datasets perform poorly when evaluated directly on document-level benchmarks, even after we increase the maximum input token length from 128 to 512;Average context word count is only 22 in MNLI and 546 in FactCC. (2) This limitation can be alleviated by the sentence-wise prediction strategy ([split-doc]MNLI-128),See details in Appendix B which achieves 66.63. Note, however, that this improvement comes at the expense of compute cost which is multiplied by a significant factor; (3) DocNLI and ANLI perform poorly even though they contain longer premise sentences, indicating that the length mismatch may not be the primary issue; (4) Falsesum obtains substantial improvement over the previous state-of-the-art FactCC, despite being derived from the same summarization dataset (CNN/DailyMail). This indicates that Falsesum provides higher quality examples and includes more types of entailment phenomena that occur naturally in this task.

2 Ablation Analysis on Falsesum Data

Table 3 shows the performance of the ablated models. We observe that removing contrastive pairs in the augmented training data results in a 1.06%1.06\% drop on the overall benchmarks score. We also see that removing intrinsic error examples results in the highest performance loss, −5.03%-5.03\% compared to −2.22%-2.22\% by −extrinsic-\texttt{extrinsic}. This is explained by the fact that intrinsic consistency errors are more dominant on benchmarks that are built on the CNN/DailyMail corpus Goyal and Durrett (2021). We conclude that all the above properties are important for the overall improvements obtained by Falsesum.

3 Fine-grained Evaluation

Previous work has shown that NLI models are prone to relying on fallible heuristics which associate lexical overlap with entailment labels McCoy et al. (2019). In the factual consistency task, this corresponds to models associating highly extractive summaries with the “consistent” label. This raises a question about whether Falsesum data alleviates this tendency in the resulting NLI models.

To answer this question, we partition the FactCC annotated test examples into five ordered subsets based on the lexical overlap between their summary hypothesis and the source document premise. We define an overlap score using the normalized coverage and density summary extractiveness scores introduced by Grusky et al. (2018). Both measures have the range [0.0,1.0][0.0,1.0], where \textscdensity=1.0\textsc{density}=1.0 indicates that all words in a summary are also present in the source document and \textscnormalizedcoverage=1.0\textsc{normalized coverage}=1.0 indicates that the summary is obtained by copying a continuous fragment of the source document. We then define \textscoverlap=\textscnormalizedcoverage×\textscdensity\textsc{overlap}=\textsc{normalized coverage}\times\textsc{density}.

Figure 3 shows the comparison of FactCC and Falsesum augmentation performance across varying lexical overlap scores. We see that Falsesum performs better on all subsets of the FactCC test set with the greatest performance gap appearing on the 0.90.9 overlap subset. Upon closer inspection, we see that the FactCC model makes mostly false positive classification errors on this subset, i.e., it tends to predict highly extractive summaries as “consistent”, leading to near majority voting performance of 50%50\%. Falsesum, on the other hand, better discriminates the factual consistency of examples without over-relying on lexical overlap.

4 Data Quality Analysis

Following Gururangan et al. (2018), we also evaluate the naturalness of the generated dataset. We train an NLI model using positive examples from CNN/DailyMail and Falsesum-generated negative examples. The model receives no premise so must distinguish between entailed and non-entailed hypotheses using semantic plausibility or spurious surface features, e.g., grammatical mistakes or fluency errors. The relatively low accuracy of these models on Falsesum data (shown in Table 5) suggests that, compared to FactCC and DocNLI, Falsesum-generated summaries are relatively hard to distinguish from the gold ones.

Conclusion

NLI models present a promising solution for automatic assessment of factual consistency in summarization. However, the application of existing models for this task is hindered by several challenges, such as the mismatch of characteristics between their training dataset and the target task data. This mismatch includes the difference in terms of the input granularity (sentence vs. document level premises) and the types of (non-)entailment phenomena that must be recognized.

In this work, we present Falsesum, a data generation pipeline which renders large-scale document-level NLI datasets without manual annotation. Using our training strategy, we demonstrate that it is possible to learn to generate diverse and naturalistic factually inconsistent (non-entailed) summaries using only existing (entailed) consistent summaries for training. We show that the resultant data is effective for augmenting NLI datasets to improve the state-of-the-art performance across four summary factual inconsistency benchmarks.

Acknowledgments

We would like to thank Marco Ponza, Marco Fiscato, Umut Topkara and other colleagues from Bloomberg AI for the thoughtful discussion and feedback throughout this project. We also thank Leonardo Ribeiro for comments on the earlier version of this work and the anonymous reviewers for their constructive feedback. The authors affiliated with UKP were supported by the German Research Foundation through the research training group “Adaptive Preparation of Information from Heterogeneous Sources” (AIPHES, GRK 1994/1) and by the German Federal Ministry of Education and Research and the Hessian State Ministry for Higher Education, Research and the Arts within their joint support of the National Research Center for Applied Cybersecurity ATHENE.

References

Appendix A Hyperparameters

We train a T5-base model for three epochs with batch size of 24 using the AdamW optimizer. We set the maximum source token length to 256 and the target token length to 42. We use a learning rate of 3e−53e^{-5} and fix the random seed to 1111. For decoding, we set the minimum and maximum sequence length to 10 and 60, respectively. We sample using beam search with a beam of size two. We additionally set the repetition penalty to 2.5 and the length penalty to 1.0.

Classification model

We train RoBERTa-base models on augmented and original MNLI datasets for three epochs with a batch size of 32. The learning rate is set to 1e−51e^{-5}, while the maximum input token length is set to either 128 or 512. We use the following random seeds for the five training runs: 11, 12, 13, 14, and 15.

Appendix B Aggregating Predictions

This means that it is sufficient for a summary sentence to be factually consistent given only a single entailing sentence in the source document. We then take the average scores across the summary sentences since each of them needs to be entailed by the source document. We use a similar aggregation method to evaluate augmented MNLI models on multi-sentence summaries from the Summeval and Ranksum benchmarks.

Appendix C Falsesum Details

In the preprocessing steps, we only perform the predicate and argument span extraction on the first 15 sentences for computational efficiency. For training, this is not an issue since the gold spans from the reference summary are included in the input. Additionally, we may extract multiple OpenIE relation tuples from each sentence. To avoid having overlapping spans from a single input, we randomly select two tuples from each sentence.

Appendix D Falsesum Examples

We include more examples of generated NLI instances in Table 6. We also include cases where Falsesum inadvertently generates factually consistent summaries in Table 7. Lastly, we show several examples of the formatted input and the generated output at test time in Table 8.