Get Your Vitamin C! Robust Fact Verification with Contrastive Evidence

Tal Schuster, Adam Fisch, Regina Barzilay

Introduction

Determining the truthfulness of factual claims by comparing them to textual sources of evidence has received intense research interest in recent years. An underlying, but often overlooked, challenge for this paradigm, however, is the dynamic nature of today’s written resources. An extraordinary amount of new information becomes available daily; as a result, many consequential facts are established, changed, or added to over time. We argue that the quality of fact verification systems should be measured by how well they adjust to new evidence. In this way, we seek to advance fact verification by requiring that models remain reliable and robust to the change present in practical settings.

To this end, we focus on fact verification with contrastive evidence. That is, we infuse the standard fact verification paradigm with challenging cases that require models to be sensitive to factual changes in their presented evidence (hereon referred to interchangeably as “context”). We present VitaminC,Etymology of VitaminC: Contrastive evidence keeps fact verification models robust and healthy, hence “Vitamin C.” a new large-scale fact verification dataset that is based on factual revisions to Wikipedia. The key concept is exemplified in Figure 1: there a factual revision yields a contrastive pair of contexts that are nearly identical in language and content—except that one context refutes the given claim, while the other supports it.

This type of contrastive structure exposes existing deficiencies in model behavior. To illustrate this, we train a classifier on the popular FEVER fact verification dataset (Thorne et al., 2018) and evaluate it on contrastive claim-evidence pairs. We find that the model flips its prediction from the original verdict on only 56% of the contrastive cases. When examples from VitaminC are included during training, however, the model’s sensitivity increases, flipping on 86% of contrastive cases.

Such context-sensitive inference has two main benefits. First, it ensures that the model considers the provided evidence rather than relying on built-in static knowledge, such as that obtained via language model pre-training (Petroni et al., 2019; Roberts et al., 2020). This is particularly important for scenarios in which the source of truth is mutable (e.g., the current US president, or new declarations as in Figure 1). Second, this setting discourages certain biases and idiosyncrasies—such as exploiting differences in how true vs. false claims are posed—that are common in similar crowd-sourced datasets (Poliak et al., 2018; Schuster et al., 2019). Indeed, we show that augmenting both fact verification models and NLI models with VitaminC data improves their robustness to adversarial inputs.

Furthermore, our emphasis on contrastive contexts allows us to expand on the scope of commonly considered tasks. Most of the fact verification literature focuses on resolving claims to be true or false (Popat et al., 2018; Thorne and Vlachos, 2018; Wang, 2017). The surrounding ecosystem, however, includes additional challenges, some of which we explore here: Documents such as Wikipedia articles are updated frequently; which edits represent factual changes? For a given claim and (refuting or supporting) evidence pair, which words or phrases in the evidence are most relevant? If we know that a certain claim is true, can we modify an out-dated document to be consistent with it? We show that the unique structure of our VitaminC dataset can be leveraged to provide both supervised and distantly supervised data for these new questions.

We pose a contrastive fact verification paradigm that requires sensitivity to changes in data;

We introduce VitaminC, a new large-scale dataset that supports this paradigm;

We demonstrate that training on VitaminC leads to better performance on standard tasks;

We show how VitaminC opens the door to additional research directions in fact verification.

Related Work

Fact Verification. The FEVER dataset (Thorne et al., 2018) fueled the development of many fact-checking models (e.g., see Hanselowski et al., 2018; Nie et al., 2019a, b; Yoneda et al., 2018, inter alia). The claim creation process, however, required crowd-workers to write claims related to Wikipedia articles, and was found to engender biases that allow an evidence-agnostic model to achieve unexpectedly high performance (Schuster et al., 2019). Other recent datasets cover verification against tables (Chen et al., 2020), relational databases (Jo et al., 2019), Wikipedia references (Sathe et al., 2020), multiple articles (Jiang et al., 2020), and search snippets (Augenstein et al., 2019). These resources all assume static ground truths. In contrast, VitaminC compares objective claims to a dynamic source of truth, and requires models to change their verdicts accordingly.

Annotation artifacts are common in many NLP datasets, and affect performance on adversarial and contrastive examples (Gardner et al., 2020; Ribeiro et al., 2020; Ross et al., 2020). Sentence-pair inference tasks such as fact verification (Paul Panenghat et al., 2020; Schuster et al., 2019) and NLI (Gururangan et al., 2018; McCoy et al., 2019; Poliak et al., 2018; Tsuchiya, 2018) are no exception. Alleviating this bias requires either modeling solutions (Karimi Mahabadi et al., 2020; Pratapa et al., 2020; Shah et al., 2020; Thorne and Vlachos, 2020; Utama et al., 2020b), which have limited effectiveness (Utama et al., 2020a), or adversarially removing troublesome training examples (Bras et al., 2020) or manually collecting new ones (Nie et al., 2020; Thorne et al., 2019a), which is model specific. Instead, our dataset design avoids single-sentence artifacts and provides model-agnostic challenging examples that increase the robustness of trained models.

Explainability.

Current fact verification datasets provide sentence-level rationales (DeYoung et al., 2020; Petroni et al., 2020) but do not enforce the model’s verdict to rely on them—leading to a potential discrepancy. VitaminC ensures the verdict is conditioned on the retrieved evidence. Moreover, we use the revision history as distant supervision for word-level rationales, allowing for finer-grained explanations (Camburu et al., 2018; Lei et al., 2016; Portelli et al., 2020; Thorne et al., 2019b).

Factually Consistent Generation.

Generating texts that match given facts is a known challenge (Fan et al., 2020; Kryscinski et al., 2020; Lewis et al., 2020b; Parikh et al., 2020; Shah et al., 2020; Tian et al., 2020) as language models tend to degenerate and hallucinate (Holtzman et al., 2020; Schuster et al., 2020; Zhou et al., 2020). Moreover, evaluation is non-trivial, and usually manual. VitaminC includes supervised data for training sequence-to-sequence models, and provides automatic evaluation via the fact verification classifier.

The VitaminC Dataset

VitaminC (abbreviated VitC) is based on revisions to English Wikipedia. Wikipedia has become a comprehensive online resource that is rigorously maintained by a large and active community (Benjakob and Harrison, 2019). While adversaries do try to insert disinformation, popular pages are usually quickly corrected (Kumar et al., 2016). Furthermore, Wikipedia’s policies dictate that its content should be written from a neutral perspective—or should otherwise objectively state all points of view.https://bit.ly/Wiki_Neutral_POV These properties make Wikipedia a suitable source of evidence for fact verification models. In the following section, we outline our process for mining factual revisions from Wikipedia.

We collected the 5K most-viewed English Wikipedia articleshttps://bit.ly/Wiki_popular_pages as of January 2020, along with any additional articles referred from them (on average 100 per article). We also included all articles from the FEVER dataset Thorne et al. (2018). For each article, we retrieved up to 500 of its most recent revisions. In May 2020, we added all COVID-19 related articles https://wikimediafoundation.org/covid19 and all of their 41K revisions at the time. Combined together, this resulted in a total of ∼\sim200 million revisions. For each revision, we identified all of the modified sentences and stored two versions: (1) before, and (2) after the edit.

In our task, we are only interested in edits made with an intent to introduce a factual modification—i.e., a change for which one can make a claim that is supported by one sentence, but not by the other.Many edits only reflect grammatical corrections, paraphrasing, or “Wikification” (text formatting/page linking). To expedite annotation, we trained a BERT classifier (Devlin et al., 2019) on a small labeled set of revised sentences determined to be factual (Yang et al., 2017), and used this model to select the top 305K edited sentences from the corpus for manual annotation. Trained human annotators were then presented with the sentence pairs, and were asked to mark the ones that indeed represented a factual change. Sentences lacking self-contained context were filtered (e.g., short expressions from tables or bulleted lists). Example annotations are presented in Table 1. Note that these annotations can also be recursively recycled for re-training the automated BERT classifier in the future to expand the corpus further (we also introduce this as a task, see §4.1).

2 Writing Claims

The factual Wikipedia revisions guide us in creating challenging claims for fact verification. For each revision, annotators were asked to write two symmetric claims related to the same edit:

The first should be supported by the original sentence and refuted by the revised sentence;

The second should be supported by the revised sentence and refuted by the original sentence.

When an explicit contradiction was not possible, a not enough information (NEI) relation was used. A group of 70 native English speakersWe sourced our annotators through TransPerfect. wrote and reviewed claims. During the annotation period, annotations were delivered in weekly batches, from which we examined random samples to provide feedback and request corrections. Annotators were instructed to write short and self-contained claims. Furthermore, annotators were instructed to avoid copying exact phrases and values when possible, in order to avoid a bias for substantially higher word overlap in supporting pairs over refuting pairs. For example, rather than stating, “there are xx confirmed cases of coronavirus in the US”, one can write “there are more than zz confirmed cases of coronavirus in the US”, which is supported if x>zx>z and refuted otherwise. For revisions that only add new information or that remove outdated facts without replacing them, annotators wrote a single claim.

3 Adding Synthetic Revisions

Naturally, the real Wikipedia revisions we collect mostly describe facts that frequently change over time, or that are prone to mistakes and corrections (such as quantitative values, see Appendix A.1) (Faruqui et al., 2018; Yang et al., 2017). Sensitivity to contrastive contexts, however, is desirable behavior for any claim. This can both ensure consistency with external sources of truth, and improve the model’s faithfulness via connecting the verdict with a specific evidence (Jacovi and Goldberg, 2020; Ross et al., 2020). For example, we require the model to not only classify the claim “Tom Hanks was honored by a president” as true, but to also change its verdict to false if paired with a (fictional) contrasting evidence. As a result, we can verify that the model prioritizes sentence-pair inference over memorization, which can help it generalize better. Therefore, we use the FEVER dataset to augment VitaminC with synthetic revisions to Wikipedia sentences.

We follow the setting of Schuster et al. (2019) to expand claim-evidence pairs from FEVER (Thorne et al., 2018). Specifically, given a false claim from FEVER, we ask annotators to edit the sentence that refutes it so that it will then support the originally false claim. Additionally, we ask them to write a new claim that is refuted by the new, modified sentence, but that is supported by the original version. Following this method, we obtain two claims where each can be supported or refuted by the original, or the synthetically revised, sentence. We follow the same process for constructing synthetic examples using true claims, but with flipped labels.

4 Dataset Statistics

In total, 304,671 revised Wikipedia sentences were examined by annotators, of which 107,056 (35%) were found to express a factual modification and were passed to the group of expert annotators for claim writing. As two symmetric claims with opposing facts were created (when possible) for each revision, this resulted in a total of 325,724 total claim-evidence pairs. We collected 163,180 additional pairs following the synthetic process. The data was partitioned as shown in Table 2. The assignment was done randomly by article, and is consistent with FEVER for overlapping articles. Appendix A contains additional details.

VitaminC Tasks

The unique structure of VitaminC allows us to derive annotations that provide a novel source of supervision for several fact-verification-related tasks. We describe the four main tasks we consider in this work, along with baseline models: (1) factual revision flagging, (2) fact verification, (3) word-level rationales, and (4) factually consistent generation. Figure 2 illustrates an example from VitaminC. We use the following notations:

C⁡\operatorname{\mathcal{C}} is the space of short sentences that express an arbitrary factual statement that can potentially be verified or debunked by external sources.

S⁡\operatorname{\mathcal{S}} is the space of sentences that can be found in a trusted online resource (Wikipedia in this study).

(st−1,st)(s_{t-1},s_{t}) denotes the two versions of a sentence that was revised from st−1s_{t-1} to st∈S⁡s_{t}\in\operatorname{\mathcal{S}}.

rel⁡(c,s)\operatorname{rel}(c,s) denotes the relation between the claim c∈C⁡c\in\operatorname{\mathcal{C}} and observed evidence s∈S⁡s\in\operatorname{\mathcal{S}}—which can either support cc (SUP⁡\operatorname{SUP}), refute it (REF⁡\operatorname{REF}), or not contain enough information (NEI⁡\operatorname{NEI}).

Online resources like Wikipedia are continuously changing. In order to remain a reliable and neutral source for recent information, its active community of users must constantly verify and correct the revisions of others. We define factual revision flagging as the task of identifying revisions that introduce a factual change—e.g., by either modifying a certain fact, adding a new one, or removing an existing one. Such an automated detection process can help the community moderate important articles by serving as a watchdog for factual revisions. Furthermore, tracking factual revisions to certain articles can potentially help keep reliant articles consistent (e.g., citing articles, or non-English versions).

We measure the edit distance between st−1s_{t-1} and sts_{t}, assuming that larger edits are more likely to represent substantive changes. We tune a decision threshold on the validation set.

BOW.

We use an MLP on top of a bag-of-words representation. Each sentence is encoded as e∗e_{*}, the average fastText (Bojanowski et al., 2017) word embedding of its edited words (i.e., that were removed or modified in the revision). The MLP input is then taken as [et−1;et;∣et−et−1∣;et⋅et−1].[e_{t-1};e_{t};|e_{t}-e_{t-1}|;e_{t}\cdot e_{t-1}].

ALBERT.

We train the ALBERT transformer Lan et al. (2020) using either only the edited words (diff), or the full sentence pair (full).

2 Fact Verification

Our baseline model is an ALBERT sentence-pair classifier that predicts rel⁡(c,s)\operatorname{rel}(c,s). Compared to BERT (Devlin et al., 2019), it uses fewer parameters by shrinking the embedding size and sharing layers, which we find to improve robustness.

3 Word-level Rationales

Word-level rationales provide useful explanations for predictions of neural models (Lei et al., 2016). Such explanations can be particularly useful for semi-automated fact verification, since they allow users to quickly interpret and trust the model’s verdict. Roitero et al. (2020) showed that explanations can increase the agreement between users and expert fact-checkers. In Figure 2, for example, the date of the first identified case can explain the verdict for the claim.

As first proposed by Lei et al. (2016), the standard definition of extractive rationales asks for selecting the minimal set of input tokens that is sufficient for preserving the model’s prediction. Here we use a slightly modified definition following Shah et al. (2020), where we identify the minimal set of evidence tokens where removing them will change the input’s label to NEI⁡\operatorname{NEI}.

Moreover, we want mm to be as sparse as possible. Intuitively, s⊙ms\odot m could be viewed as an incomplete revision in which the masked words that have not yet been filled in will determine the relation with the claim. We say that mm reveals the most responsible words in ss for resolving cc. Following Shah et al. (2020), we formulate an unsupervised objective as

As in Shah et al. (2020), we optimize a Lagrangian relaxation of Eq. 1, where

We keep the rel⁡\operatorname{rel} classifier (from §4.2) fixed, and train a separate ALBERT model to predict the mask mm using a Gumbel softmax Jang et al. (2017).

Distantly Supervised.

4 Factually Consistent Generation

As facts change, the sources reporting them must change as well to reflect the most recent information. In VitaminC, this is reflected via the active revisions to Wikipedia. We simulate automating this process by considering two generation tasks:

Claim Extraction.

Experiments

We present and analyze results for the models described in Section 4. Our analysis attempts to evaluate several questions: (1) How well can the current state-of-the-art models perform on the VitaminC tasks? (2) Does VitaminC increases the robustness of models against adversarial examples? (3) Can VitaminC improve interpretability by providing supervision for anchoring words?

In addition to VitaminC, we train and evaluate on several related datasets, which we briefly describe:

(Thorne et al., 2018): A popular fact verification dataset based on Wikipedia. We use the provided SUP⁡\operatorname{SUP} and REF⁡\operatorname{REF} claim-evidence pairs. For NEI⁡\operatorname{NEI} claims, we randomly sample neutral evidence from the article with the highest BM25 score.

MNLI

(Williams et al., 2018): A large and diverse dataset for natural language inference. The three-way sentence-pair entailment prediction is similar to fact verification. We use the hypothesis as the claim and the premise as the evidence and evaluate on the “mismatched” evaluation set.

Symmetric

(Schuster et al., 2019): A set of challenging symmetric, synthetic extensions to FEVER’s evaluation set that avoid claim-only bias.

Adversarial

(Thorne et al., 2019c): Adversarial examples created by participants of the FEVER 2.0 shared task. Teams were asked to create claims that break FEVER-trained models. We take all SUP⁡\operatorname{SUP} and REF⁡\operatorname{REF} claims and their gold evidence sentences.

Triggers

(Atanasova et al., 2020): A set of 186 FEVER claims paraphrased adversarially to contain universal adversarial triggers (Wallace et al., 2019). Its small size leads to high variance results.

ANLI

(Nie et al., 2020): An adversarial dataset for MNLI- and FEVER-based models. The creation was performed in three iterative rounds in which a model was trained, and then crowdworkers devised adversarial inputs, and the process repeated.

PAWS

(Zhang et al., 2019): A dataset of altered Wikipedia sentences using word swapping and back-translation. Human annotators labeled whether the modified sentence is a paraphrase or not. We evaluate whether a PAWS-trained classifier can be used for our factual revision flagging task.

2 Factual Revision Flagging

Table 3 shows the results of our baseline models on the factual revision flagging task. First, we notice that a model trained on the PAWS dataset (reaching 93.42 F1 score on PAWS test) does not transfer well to the flagging task, and performs on par with a simple edit distance heuristic. We hypothesize that this is a result of the entity scrambling technique used to synthetically revise sentences in PAWS, which is different from the edits introduced by real, factual Wikipedia revisions in practice.

Second, we see that the performance of neural models trained on the VitaminC flagging task increases with richer inputs and more advanced models—demonstrating the complexity of the task. The ALBERT (diff) model that uses only the modified word sequences from each sentence (i.e., contextual within a subspan) improves the AUC by 10 points over a BOW model that gets a similar input. The ALBERT (full) model that receives the full sentences as input (i.e., has access to even more context), further improves the AUC by 2 points. Nevertheless, the best model still only reaches 83 macro-F1, indicating the difficulty of this task.

3 Fact Verification

Table 4 summarizes the results for classifiers trained on fact verification and NLI datasets. Verifying claims against real revisions proves to be the hardest. The best model achieves 89% accuracy, lower than that on either VitaminC’s synthetic cases or the original FEVER examples. Including VitaminC examples in the training data drastically increases models’ sensitivity to contrastive examples (rightmost column)—while preserving the in-domain accuracy (only −0.42%-0.42\% for FEVER and +0.12%+0.12\% for MNLI with ALBERT-xlarge). Another evidence for the generalization properties conferred by VitaminC is its zero-shot performance to both other datsets. An ALBERT-xlarge model trained only on VitaminC reaches 76%76\% and 79%79\% accuracy on FEVER and MNLI, respectively. In contrast, the transfer accuracy for MNLI→\rightarrowFEVER is 70%70\% and for FEVER→\rightarrowMNLI is only 38%38\%.

Most importantly, models trained with VitaminC perform better on challenging adversarial datasets. On the otherhand, simply augmenting FEVER data with MNLI data has a limited effect on adversarial examples.We’ve also tried augmenting FEVER with ANLI for an ALBERT-xlarge model and find it to achieve only 73%73\%, 91%91\%, and 34%34\% on Adver., Sym., and Triggers, respectively. We conjecture that the contrastive nature of VitaminC helps models better learn the relations between the claims and evidences—and to avoid relying on certain artifacts that do not generalize well.

To further probe the value of VitaminC examples compared to FEVER ones (SUP⁡\operatorname{SUP} and REF⁡\operatorname{REF} only), we compose training sets of 100K examples using different ratios of the two datasets. As shown in Figure 3, including more VitaminC pairs continuously improves the performance on the challenging adversarial and symmetric evaluation sets.

As an additional qualitative experiment, given the recent successes of huge language models such as GPT-3 (Brown et al., 2020), we explore whether such models develop sufficient context sensitivity on their own. Appendix C shows the results of classifying several claims using a few-shot GPT-3 model. We find that GPT-3 still largely under-performs our VitaminC-trained models in terms of sensitivity—demonstrating the importance of using VitaminC’s unique structure during training.

4 Word-level Rationales

Table 5 shows the results of our baseline models for identifying word-level rationales (i.e., anchoring words in the evidence). While our unsupervised model is able to uncover some patterns, directly leveraging the structure of VitaminC to obtain distant supervision for likely anchoring words (i.e., token labels) improves both the edit prediction and the word-level rationale prediction performance.We evaluate rationales using a manually annotated test set of 300 examples (150 each from VitC real and VitC synthetic). Example predictions are provided in Appendix E.

5 Factually Consistent Generation

The revision generator aims to modify sentences so that they agree with a given claim. According to our fact verification model’s verdict, it succeeds in doing so 76% of the time. Furthermore, revisions should resemble real ones, and preserve the remaining content that is unrelated to the claim. The SARI KEEP F1 (Xu et al., 2016) of 75 shows that the model and the reference mostly agree on parts of the sentence that should be kept unchanged.

Conclusion

We presented VitaminC, a large-scale dataset for training and evaluating fact verification models using contrastive contexts. Our novel method of leveraging factual revisions to Wikipedia enabled us to create challenging examples in which a claim is paired with contexts that are lexically similar, yet factually opposing. Our results illustrated that training on VitaminC improves classifier sensitivity to subtle changes in evidence, and increases their robustness to adversarial examples.

Furthermore, we formulated several new, important tasks for fact verification that VitaminC allows us to test. We showed how the dataset’s unique “before and after” structure lends itself to training classifiers to flag factual revisions. In addition, for factual revisions, the edits reveal which words in the evidence are the most critical—which helps supervise word-level rationale models for better interpretability. Finally, we demonstrated that VitaminC can help with factually consistent text generation. We hope that this work and the range of tasks it presents will motivate and support the fact verification field in developing reliable models that can adapt to dynamically changing evidence.

Acknowledgements

We thank the TransPerfect team, Darsh J Shah and Enrico Santus for helpful discussions, as well as the members of the MIT NLP group and Andreas Vlachos for valuable feedback. This work is supported in part by the Facebook Online Safety Benchmark Award. TS is supported by in part by DSO grant DSOCL18002. AF is supported in part by a NSF Graduate Research Fellowship.

References

Appendix A VitaminC: Complementary details

We provide additional details about the VitaminC dataset.

Figure A.1 shows the distribution of claims in the VitaminC dataset by the topic of the Wikipedia article they are based on. The information was collected from DBpedia,http://dbpedia.org/ontology/ retrieving the parent class of the pages. Labels for about 25% of the articles were missing, and left blank in the diagram.

The “synthetic” part of VitaminC, which is based on the claims of the FEVER dataset, contains many claims about specific human entities. About 15% of the claims in VitaminC real are about COVID-19.

Category Distribution.

We sample 100 examples from the “real” and “synthetic” subsets of VitaminC and manually categorize their claims. Due to the creation methodology of VitaminC real, its claims mostly describe frequently updating facts, or facts that tend to be corrected. We find about half of these claims to describe changes in numerical values (e.g., number of COVID-19 cases, earnings or ratings of movies, number of awards etc.). In contrast, VitaminC synthetic mostly covers general facts about specific entities, (e.g., place of birth, date of birth, occupation, etc.). This is a result of the synthetic claims being based on the FEVER dataset, where annotators were asked to come up with claims on popular Wikipedia pages. Combined, the VitaminC dataset holds a diverse set of claims about various topics.

A.2 Inter-annotator Agreement

We ask three additional annotators to independently annotate a random set of two thousand claim-evidence pairs, evenly distributed between the development and test splits of the real and synthetic sets. The Fleiss κ\kappa score (Fleiss, 1971) between the four annotations is 0.7065, which means substantial agreement. Similar agreement scores of 0.6841 and 0.7 were reported for fact verification (Thorne et al., 2018) and NLI datasets (Bowman et al., 2015), respectively.

A.3 Claim-only Classification

Annotation artifacts are common in crowd-sourced sentence-pair inference datasets such as fact verification and NLI. Models can leverage these idiosyncrasies to achieve unexpectedly high performance when given only one sentence of the pair. For example, Schuster et al. (2019) showed that a claim-only classifier can obtain 61.7% accuracy. The VitaminC dataset avoids this bias by pairing each claim with two contrastive contexts.

All claims in the VitaminC-synthetic are paired with one refuting and one supporting evidence, making it impossible for a claim-only to perform better than random. Each claim in the VitaminC-real is paired with one refuting or neutral evidence, in addition to a supporting one. To evaluate whether models can utilize lexical cues in claims, we train a claim-only classifier on VitaminC-real and find it to achieve 50% accuracy—the same as always predicting SUP⁡\operatorname{SUP}.

A.4 Claim-evidence Word Overlap

Naturally, when pairing claims to evidence sentences, the overlapping words will be higher on average for claims with their supporting evidence. In VitaminC dataset, we want to minimize this bias in order to create challenging examples that require sentence-pair inference and cannot be solved by simple word matching techniques. Therefore, we asked annotators, when possible, to avoid copying exact phrases from the evidence to the claim (see §3.2).

Figure A.2 shows the probability density function of bigram overlaps between the claim and evidence for each relation. Similar to FEVER, the overlap ratio of supporting pairs in the VitaminC dataset is only slightly higher than the one of refuting pairs. Also, the overlap ratio of the NEI⁡\operatorname{NEI} pairs of the VitaminC real dataset is on average higher than FEVER.

Appendix B Experimental Setting

We implement all our models with the HuggingFace Transformers library (Wolf et al., 2019). When comparing across training datasets of different sizes, we train the model for the same amount of update steps, upsampling the smaller datasets. We pick the checkpoint with the highest accuracy on the development set of the training task and report performance on the test set. More details are available at https://github.com/TalSchuster/VitaminC

Appendix C GPT-3 Evaluation

The GPT-3 model has recently demonstrated impressive results in zero-shot and few-shot generation and classification tasks (Brown et al., 2020). This 175B parameters language model was trained on billions of words from online sources, including the English Wikipedia. As result, it can be applied on many tasks without any further fine-tuning—instead, one need only provide a task-specific prefix (i.e., “prompt”) with a few examples that direct the language model towards the desired output format. For example, GPT-3 achieves better than random results on ANLI with only a single example in the prompt, and over 40%40\% accuracy with 50 examples (Brown et al., 2020).

We used OpenAI’s beta API to query GPT-3. Due to our limited quota, we could not perform extensive experiments. Instead, we performed a qualitative evaluation using several examples from VitaminC test set for the claim extraction (factually consistent generation) and the fact verification tasks. Therefore, these results should be viewed as exploratory only.

We use GPT-3 to extract claims for four revisions with a sampling temperature value (T\mathcal{T}) set to either or 0.70.7. The zero value is recommended for maximizing the factual consistency as the model follows its most certain predictions. Using low temperature, however, can result in less fluent generations (Holtzman et al., 2020). Therefore, high values of T\mathcal{T} are also commonly used.

We expect GPT-3 to improve with longer prompts or fine-tuning and leave this to future research due to our limited quota.

GPT-3 for Fact Verification.

We also experiment with using GPT-3 few-shot classification capabilities for the fact verification task. We follow the ANLI few-shot format of Brown et al. (2020) and compose prompts with 6 examples (2 from each class) with random examples from VitaminC training set. We use only numerical examples to evaluate numerical claims (Figure C.3), and mixed examples for other claims (Figure C.2). We set T=0\mathcal{T}=0 as recommended for classification.

Table C.1 summarizes the results. Even with only six examples, GPT-3 seems to perform significantly better than random. Yet, its verdict is wrong in several cases that can be easily classified by humans. For example, we find it to refrain from predicting a True/False verdict even when the evidence is clear. We observe this both for a date-based (line 3.2 in Table C.1), numerical (lines 4.1-4.2), and entity-focused claims (line 5.2).

To experiment with the sensitivity of the model to the provided context, we manually modified some of the examples to provide even stronger evidence. For example, while GPT-3’s prediction for line 5.2 is acceptable as actually, Turner Broadcasting System merged with WarnerMedia in 1996, changing the evidence to another disconnected entity (The Walt Disney Company) did not change the prediction (line 5.3) as expected. Even when explicitly stating that there is no other owner GPT-3 didn’t modify its verdict (line 5.4). Similarly, when evaluating the claim about the population of Beaverton being less than 90K, GPT-3 ignores the supporting evidence and outputs a false verdict (lines 1.4-1.5). Changing the claim to state “approximately 86K” instead of “less than 90,000” modified the prediction to “Neither” (line 1.6). Only repeating the exact same number as the evidence led to a true verdict (line 1.7).

Appendix D Complementary Experiments

We report fact verification results with a fine-tuned BERT-base (Devlin et al., 2019) model in Table D.1. We find ALBERT-base to outperform BERT-base on most of the evaluated datasets. ALBERT-xlarge performed better than the two base models in all datasets except for Triggers. The Triggers dataset is very small (186 examples) and contains some unnaturally looking claims, which could explain the high variance across models.

Appendix E Example Outputs

We provide examples of predicted word-level rationales in Table E.1 and of outputs for the two generation tasks in Tables E.2 and E.3.