FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi
Introduction
Long-form text generated by large language models (LMs) has widely been used (Brown et al., 2020; Ouyang et al., 2022); nonetheless, evaluating their factual precision—whether each piece of information conveyed in a generation is factually accurate—remains challenging for two reasons. First, a generation consists of a large number of pieces of information that are a mixture of true or false, Even a single sentence consists of multiple pieces of information (e.g., 4.4 per sentence in ChatGPT, 40% of which are a mixture of supported and unsupported information). making a binary judgment inadequate (Pagnoni et al., 2021). Second, validating every piece of information is time-consuming and costly.
In this paper, we introduce FActScore (Factual precision in Atomicity Score), a new evaluation of an LM that represents the percentage of atomic facts (pieces of information) supported by a given knowledge source. Computing FActScore involves (1) breaking a generation into a series of atomic facts—short statements that each contain one piece of information (Nenkova and Passonneau, 2004; Shapira et al., 2019; Zhang and Bansal, 2021; Liu et al., 2022), and (2) assigning a binary label to each atomic fact, allowing a fine-grained evaluation of factual precision. We evaluate FActScore on the task of generating people biographies because generations consist of verifiable statements rather than debatable or subjective ones, and the scope is broad (i.e., covering diverse nationalities, professions, and levels of rarity).
We perform extensive human annotations to obtain FActScores of three state-of-the-art, commercially available LMs: InstructGPT (Ouyang et al., 2022), ChatGPT (OpenAI, 2022), and search-augmented PerplexityAI.perplexity.ai Our results indicate that commercially available LMs are riddled with errors, having FActScores of 42%, 58% and 71%, respectively. Their FActScores significantly drop as the rarity of the entities increases, e.g., for ChatGPT.
Since human evaluation is costly, we next introduce an automatic evaluation of FActScore through a model that estimates a FActScore for a given LM. Our estimator decomposes generations into atomic facts and validates each based on a given knowledge source, leveraging retrieval from the given knowledge source and strong language models. Our estimator closely approximates FActScore with an error rate of and can be applied to a range of new LMs at scale with no human effort. Our case study evaluates 6,500 generations from 13 LMs that could have cost $26K, with various findings: GPT-4 (OpenAI, 2023) and ChatGPT are far less factual than humans but are much better than public models, and there is a large variance between public models, with Vicuna (Chiang et al., 2023) and Alpaca (Taori et al., 2023) being some of the best.
In summary, our contributions are as follows.
We introduce FActScore, a new evaluation of factual precision of LMs by breaking their generations into atomic facts and validating each against a given knowledge source. Human evaluation reveals that the state-of-the-art LMs with and without search have low FActScores.
We introduce a model that approximates FActScore with an error rate of , allowing evaluation of a large set of new LMs without manual human efforts.
We open-sourced FActScore and the annotated data for public use, available via pip install factscore. We suggest future work to extend FActScore for a broader set of generations (e.g., open-ended generation) and to further improve the estimator.
Related Work
Factual precision in text generation has been an active area of research in NLP. Most prior work studies factual precision of models supervised for a specific problem such as dialogue (Shuster et al., 2021), or focuses on question answering with short answers (Kadavath et al., 2022; Kandpal et al., 2022; Mallen et al., 2023; Nori et al., 2023).
More recent work has studied factual precision of text generation beyond short answers. Lee et al. (2022) evaluates the factual precision with proxy metrics, e.g., whether named entities in a generation appear in an article of the topic. A series of concurrent work verifies the precision of the citations (attributions) provided by the model (Gao et al., 2022; Liu et al., 2023a; Yue et al., 2023; Gao et al., 2023). A concurrent work by Manakul et al. (2023) automates the identification of factual errors in LM generations without using any knowledge source; we use their method as a baseline estimator in Section 4. In contrast, our work (1) considers much longer text generationConsisting of 110–151 words (Table 1), in contrast to 18–29 in Gao et al. (2022) and 65 in Liu et al. (2023a). from a variety of state-of-the-art LMs with and without search, (2) provides their fine-grained evaluation both by human experts and through an automated evaluator that closely approaches humans, and (3) applies it to a large set of LMs at scale.
Our work is closely related to prior work on fact verification (Thorne et al., 2018; Wadden et al., 2020) where claim sentences are automatically checked against a large knowledge source like Wikipedia or scientific literature. Most literature assumes a single, atomic claim, sometimes modeled with surrounding context (Nakov et al., 2018; Mihaylova et al., 2019; Shaar et al., 2022). There also has been work that verifies a longer sentence or text through decomposition to atomic facts (Fan et al., 2020; Wright et al., 2022; Chen et al., 2022; Kamoi et al., 2023) from which we take inspiration. The primary difference between fact verification literature and our work is that we focus on long-form model-generated text rather than sentence-level human-written claims.
Prior work has used learned models to define automated evaluation scores (Zhang et al., 2020; Liu et al., 2023b). This includes model-based evaluation in summarization that considers the consistency between a summary and a source document using QA or NLI (Kryscinski et al., 2020; Wang et al., 2020; Fabbri et al., 2022; Deutsch et al., 2021; Laban et al., 2022). We take inspiration from this work, and evaluate factual precision of LM generations by considering whether pieces of information are supported by a large text corpus.
FActScore: Evaluating Factual Precision of Long-form Text Generation
We introduce FActScore, a new evaluation of an LM that considers the factual precision of atomic facts generated by the LM. We perform human evaluations to calculate FActScores of the state-of-the-art LMs (Section 3.3) and discuss results (Section 3.4). FActScore allows rigorous and fine-grained evaluation of factual precision, but is time-consuming and costly, motivating automatic evaluation in Section 4.
Key idea 1: Atomic fact as a unit. Long-form text consists of many pieces of information that can each be either true or false. Prior work has explored using a sentence as a unit; however, even a single sentence is a mix of supported and unsupported facts, e.g., in 40% of the cases with ChatGPT. Previous and concurrent work either (1) defines an additional label of partial support (Manakul et al., 2023; Liu et al., 2023a) whose definition may be subjective and can lead to low agreement, or (2) takes the strictest definition of support that requires every piece of information to be supported (Rashkin et al., 2021; Gao et al., 2022), which ignores the partial support cases, e.g., assigning 0.0 to both generations in Figure 1 even though the first generation is considerably more accurate than the second.
In this paper, we define an atomic fact as a short sentence conveying one piece of information (examples in Figure 1), similar to summarization content units (Nenkova and Passonneau, 2004). An atomic fact is a more fundamental unit than a sentence for a piece of information and provides a more fine-grained evaluation, e.g., in Figure 1, rating the first generation higher than the second.
Key Idea 2: Factual precision as a function of a given knowledge source. Prior work often considers factual precision as a single global truth (Manakul et al., 2023). In contrast, we adopt a perspective that the truthfulness of a statement should depend on a particular knowledge source that end users consider to be trustworthy and reliable. Therefore, instead of whether an atomic fact is globally true or false, we consider whether it is supported by a given source of knowledge. This has been used in the fact verification literature (Wadden et al., 2022) where conflict of information between different sources is relatively common.
Definition. Let be a language model to be evaluated, be a set of prompts, and be a knowledge source. Consider a response for and , a list of atomic facts in . A FActScore of is defined as follows.
responds means did not abstain from responding to the prompt . This definition assumes the following:
Whether or not an atomic fact is supported by is undebatable.
Every atomic fact in has an equal weight of importance, following Krishna et al. (2023).
Pieces of information in do not conflict or overlap with each other.
In the rest of the paper, we propose to use people biographies as and Wikipedia as because they satisfy these assumptions to a reasonable degree (Section 3.3). We discuss in which cases these assumptions hold or may not hold in more detail in the Limitation section.
FActScore considers precision but not recall, e.g., a model that abstains from answering too often or generates text with fewer facts may have a higher FActScore, even if these are not desired. We leave the evaluation of factual recall for future work (more discussion in the Limitation section).
2 Studied LMs
We evaluate three LMs (referred to as LM, an LM as a subject): (1) InstructGPT (text-davinci-003, updated from Ouyang et al. (2022)), (2) ChatGPT (OpenAI, 2022), and (3) PerplexityAI,\@footnotemark which incorporates a search engine with a language model.
3 Data
We perform human evaluation of factual precision based on our definition. We prompt the LM to generate people biographies and evaluate them against Wikipedia for the following reasons.
Biographies are objective (not subjective or debatable) and contain specific (not vague) information, satisfying Assumption 1 in Section 3.1.
Biographies allow evaluation across diverse nationalities, professions, and levels of rarities.
Wikipedia offers reasonable coverage of information about people and is reasonably self-consistent,See Appendix A.5 for a related analysis. satisfying Assumption 3.
We carefully design an annotation pipeline to assign a factual precision to a long-form generation through the following steps.
Step 0: Sampling people entities. We sample 183 people entities from Wikidata who have corresponding Wikipedia pages. We sample entities to annotate from a uniform distribution over categories defined in Appendix A.1.
Step 1: Obtaining generations. We feed a prompt ‘‘Tell me a bio of
Step 2: Atomic facts generation. Human annotators break a generation into a series of atomic facts. To save annotation time, we provide atomic facts broken down by InstructGPT which human annotators can take and revise. Details in Appendix A.2.
Step 3: Labeling factual precision & editing. We ask another set of human annotators to assign each atomic fact one of three labels. If the atomic fact is clearly not related to the prompt, and thus should be removed from the bio without a validation step, they assign Irrelevant. If the fact is relevant, they validate the fact based on the English Wikipedia, and label either Supported or Not-supported.
We recruit freelancers through Upwork and pay 15–25 USD per hour. Annotation requires extensive effort and time, leading to the cost of $4 per generation. We assign two freelancers for the 10% of the data and calculate the agreement rate: 96%, 90% and 88% for InstructGPT, ChatGPT and PerplexityAI, respectively. More details are provided in Appendix A.3.
4 Results
Statistics of the data and results are reported in Table 1.
InstructGPT and ChatGPT achieve FActScores of 42.5% and 58.3%, respectively. PerplexityAI, which uses a commercial search engine and thus should have a perfect FActScore if directly copying the text from the correct Wikipedia page, attains a FActScore of 71.5%. We provide a qualitative analysis of its error cases in the last paragraph of this section.
ChatGPT and PerplexityAI often abstain from answering which presumably improves their factual precision. InstructGPT rarely abstains from answering, likely because it is not trained to do so.
Irrelevant facts either (a) have dependencies on previous facts in a generation that turn out to be unsupported, or (b) are irrelevant to the prompt independent from other facts in a generation (examples in Appendix A.4). We find that (b) rarely happens with InstructGPT and ChatGPT but happens considerably with PerplexityAI, because PerplexityAI often directly copies search results even if they are largely irrelevant to the input prompt. This is in agreement with a concurrent work from Liu et al. (2023a) that shows generative search engines like PerplexityAI copy incorrect search results and generate text that is irrelevant to the input query.
Figure 2 (top) shows factual precision over varying frequency levels of topic entities (humans) in the pretraining corpora (see Appendix A.1). There is a notable decrease in FActScore as the rarity of entities increases, consistently across all LMs. This is in agreement with Kandpal et al. (2022) and Mallen et al. (2023) which show that short question answering (QA) accuracy is highly correlated with the entity frequencies in the pretraining data. However, in contrast to Kandpal et al. (2022) and Mallen et al. (2023) who report QA accuracy of models with retrieval is robust to the rarity of entities, FActScore of PerplexityAI still significantly drops as entities are rarer: a relative drop of 50% and 64% observed at the atomic-level and sentence-level, respectively.
Figure 2 (bottom) reports factual precision over relative positions in a generation. Across all LMs, the later part of the generation has significantly worse precision. This is likely because (a) information mentioned earlier is more frequently mentioned in the pretraining data (e.g., nationality, profession), and (b) error propagation affects the later part of the generation. This also implies that evaluating LMs solely based on short answers may not provide an adequate assessment of their factual precision, as it fails to account for errors that arise in the later stages of generation.
One of the surprising findings in our empricial analysis is that a FActScore of PerplexityAI (71.5%) is lower than expected despite having access to the search engine. To better understand its errors, we categorize 30 random samples whose label is Not-supported (Table 2).
Single-sentence contradiction: A single sentence from Wikipedia provides direct contradiction to the generation, either at a word level (numbers, dates, or entities) or beyond.
Page-level contradiction: Errors found after reading the entire page, often because a fact that should have been mentioned in Wikipedia if true is missing, e.g., whether the subject appears in a particular film.
Subjective: Generation is subjective, often because PerplexityAI copies subjective text from Wikipedia, e.g., directly copying a quote from a journalist without realizing it.
Fact is irrelevant: Generation is irrelevant to the subject due to a search error.
Wiki is inconsistent & wrong: In the example, Wikipedia indicates that the subject won one award from the film Kick, but also includes text that they won multiple awards from Kick, which is inaccurate and cited a news article that does not support the claim.
Annotation error: Annotators assign incorrect labels, typically because the information is not mentioned in the subject’s Wikipedia page (likely because it is insignificant).
We also find that, although PerplexityAI provides citations to the references, citations have little correlation with factual precision. 36.0% and 37.6% of supported and unsupported sentences have citations, respectively. Together with independent findings from Liu et al. (2023a), this indicates that commercial LMs that incorporate search and provide citations may not be as reliable as expected.
More analysis is provided in Appendix A.5.
Estimating FActScore for Automatic Evaluation
Human evaluation of factual precision is costly ({}_{\textsc{subj}}{}_{\textsc{subj}}$.
We describe our model (Section 4.1) and demonstrate its accuracy against human evaluation (Section 4.2). FActScore estimated by our model is then used to evaluate twelve LMs (Section 4.3).
Our estimator of FActScore first breaks a generation into a series of atomic facts and then validates each against the given knowledge source. We find taking atomic facts generated by InstructGPT (used in data collection in Section 3.3) effective and close to human, consistent with findings from prior work (Chen et al., 2022). This section thus focuses on how to validate each atomic fact against a given knowledge source.
The validation is based on zero-shot prompting of an LM referred to as an LM to distinguish from an LM. Specifically, a prompt—whose construction methods differ across four variants—is fed into an LM. The prediction is then made by comparing the conditional probability of True and False from the LM. If the logit values are unavailable (e.g., commercial LMs like ChatGPT), the prediction is made based on whether the generated text contains True or False. In Appendix B.3, we compare with an alternative prompting that generates a question and compares the answer to it and the expected answer (Kryscinski et al., 2020; Wang et al., 2020; Gao et al., 2022; Manakul et al., 2023). We empirically find that our prompting performs better due to the lack of control over the questions being generated.
The four variants we consider are as follows.
No-context LM uses
RetrieveLM retrieves passages from the given knowledge source and then prompts the LM. It first retrieves passages, constructs the prompt by concatenating retrieved passages, the given atomic fact, and ‘‘True or False?’’, and feeds it to the LM to get the prediction.
Nonparametric Probability (NP) makes a judgment based on a nonparametric likelihood. It masks out each token in the atomic fact, computes its likelihood using a nonparametric masked LM (Min et al., 2023), averages probabilities over all tokens, and makes a prediction based on thresholding.
RetrieveLM + NP is an ensemble of RetrieveLM and NP which assigns Supported only if both methods assign Supported.
We use LLAMA 7B trained on Super Natural Instructions (Inst-LLAMA, Touvron et al., 2023; Wang et al., 2022) and ChatGPT as an LM, and Generalizable T5-based Retrievers (GTR, Ni et al. (2022)) for passage retrieval. See Appendix B.1 for more implementation details.
2 Evaluation of Estimators
We report Error Rate (ER)—the difference between the ground truth and the estimated FActScore—as well as whether the estimated FActScores preserve the ranking between three LMs. Appendix B.2 discusses results with other metrics that consider individual judgments instead of aggregated judgments. We use the data in Section 3.3 as evaluation data.
Models that use retrieval are consistently better than No-context LM which either has a significantly high ER or does not preserve ranking between three LMs. This is likely because the LM has not memorized every factual information about the topic entity, thus benefiting from retrieval providing factual context. Nonetheless, just using RetrieveLM may overestimate FActScore, e.g., by up to 17% with Inst-LLAMA, when a LM is InstructGPT or ChatGPT. In this case, ensembling RetrieveLM and NP reduces an error rate by a significant margin. When a LM is PerplexityAI, single methods (either RetrieveLM or NP) give a low ER, and ensemble methods have a higher ER due to an underestimation of FActScore.
Our results show that ChatGPT is not necessarily better than Inst-LLAMA. We investigate this further in Appendix B.3. In summary, ChatGPT is better at validating each individual atomic fact. However, most errors from ChatGPT are incorrectly assigning Supported to unsupported facts, overestimating FActScore. In contrast, LLAMA+NP is not biased toward overestimation or underestimation of the factual precision, resulting in an aggregated factual precision to be closer to the ground truth. This is similar to the trade-off between system-level and segment-level correlations in summarization evaluation, which often produce different rankings (Bhandari et al., 2020; Deutsch et al., 2021).
While using retrieval is consistently better than No-context LM, the best variant of estimator depends on a LM: LLAMA+NP for InstructGPT and ChatGPT, and ChatGPT for PerplexityAI. Nevertheless, both evaluators give consistently correct ranking between three LMs, and Section 4.3 show scores from two estimators are largely correlated across 10+ LMs (0.99 Pearson’s ). We recommend users try both variants of our estimator when evaluating a new LM and report their correlation.
3 Evaluation of New LMs
Our estimator allows evaluating factual precision of a large set of new LMs at scale with no human efforts. As a case study, we evaluate ten new LMs that came out within two months at the time of conducting experiments (Table 4). These LMs were evaluated on many benchmarks but not in factual precision of long-form generation since such evaluation is costly. We aim to provide new insights on these LMs by estimating FActScore of their long-form generations.
We evaluate 10 recently-released LMs as shown in Table 4. GPT-4 (OpenAI, 2023) is a multimodal LM released by OpenAI available through an API. Alpaca (Taori et al., 2023) is based on LLAMA (Touvron et al., 2023) fine-tuned on the instructions data based on InstructGPT following the recipe from Wang et al. (2022). Vicuna (Chiang et al., 2023) is based on LLAMA fine-tuned on the outputs from ChatGPT available through ShareGPT. sharegpt.com 9dolly-v2-12b 10databricks.com 11oasst-sft-1-pythia-12b 12open-assistant.io 13StableLM-tuned-alpha-7b 14stablelm-base-alpha-7b 15mosaicml.com/blog/mpt-7b 16evol_instruct_70k Dolly9 is Pythia 12B (Biderman et al., 2023) fine-tuned on DataBricks Dolly, human-written data created by Databricks.10 Oasst-pythia11 is Pythia 12B fine-tined on human-written data collected through Open Assistant.12 StableLM-tuned-alpha13 is based on StableLM-base-alpha14 fine-tuned on the data used in the Alpaca data, DataBricks Dolly, the ShareGPT data, the GPT4All data (Anand et al., 2023) and Anthropic HH (Bai et al., 2022). MPT Chat is based on MPT 7B15 fine-tuned on the ShareGPT data, the Alpaca data, Anthropic HH, HC3 (Guo et al., 2023), and Evol-Instruct.16
We prompt each LM to generate biographies of 500 human entities as done in Section 3.3 but with no overlap in entities. We additionally include InstructGPT, ChatGPT, and human-written biographies obtained through DBPedia. Human-written biographies were unavailable for 11% of entities which we consider as abstaining from responding. See Table 5 for their statistics. In total, we evaluate 6,500 generations from 13 subjects, which would have cost $26K if they were evaluated by humans.
3.2 Results
Figure 3 shows the ranking between 13 subjects provided by the two best variants of our estimator whose scores are largely correlated, e.g., having a Pearson’s of 0.99. This evaluation allows a better understanding of these models, including:
All LMs are substantially less factual than humans. This is in contrast to prior work that claims LMs approach human performance, even for complex tasks (Ding et al., 2022; Nori et al., 2023; Lee et al., 2023) even though the task of writing biographies is fairly easy.
GPT-4 and ChatGPT are comparable in factual precision. However, as reported in Table 5, GPT-4 abstains from responding less (12% vs. 16%) and generates significantly more facts (61 vs. 37 per response).
GPT-4 and ChatGPT are significantly more factual than public models.
Within the same family of models that differ in sizes, there is a clear correlation between the model size and factual precision, e.g., Alpaca 65B > 13B > 7B, and Vicuna 13B > 7B.
Alpaca and Vicuna achieve performance that is very close to each other within the same size of models, possibly because they share the same base model and similar training data. Nonetheless, as shown in Table 5, Vicuna generates significantly more atomic facts than Alpaca does (51 vs. 17 per response). Also, Alpaca never abstains from answering while Vicuna does.
Within public models, there are large gaps in factual precision even when the model size is similar, e.g., within the 7B models, Alpaca and Vicuna () are more factual than MPT-Chat () and StableLM (). Possible factors include the choice of the base LM, the data, and the training recipe (Hoffmann et al., 2022).
We highlight that this evaluation only considers factual precision, specifically in people biographies. A holistic evaluation of LMs should include other aspects of generations such as fluency, coherence, relevance, consistency and creativity, which is out of scope of this paper.
Conclusion and Future Work
We introduced FActScore, a new evaluation of the factual precision of long-form generation from LMs that breaks a generation down into a series of atomic facts and computes a fraction of facts supported by a given knowledge source. We first performed extensive human evaluation, finding that commercial, state-the-art-art LMs—InstructGPT, ChatGPT, and search engine augmented, PerplexityAI—make a substantial amount of errors, e.g., having a FActScore of 58% in the case of ChatGPT. Since human evaluation is time-consuming and costly, we proposed a model that estimates FActScore, allowing an automatic evaluation of FActScore. We found our estimator based on retrieval over a knowledge source and competitive language models estimates FActScore close to the ground truth, and showcased its application by evaluating 12 recently-released LMs that could have cost $65K if evaluated by humans and providing insights about them.
Within four months since its initial release, FActScore has actively been used in subsequent work, evaluating factual precision of recently-proposed models Ye et al. (2023); Sun et al. (2023); Malaviya et al. (2023); Dhuliawala et al. (2023). As future work, we suggest: (1) considering other aspects of factuality such as recall (coverage of factual information); (2) further improving the estimator for a better approximation of factual precision; and (3) leveraging FActScore to correct model generations (briefly explored in Appendix C).
Limitations
All of our experiments focus on people biographies and Wikipedia, because many LMs can generate biographies with objective and specific facts (rather than subjective and vague ones) and Wikipedia has a high coverage for them. FActScore can be applied to a broader domain, e.g., text about recent events whose knowledge source can be a collection of news articles, or text about scientific findings whose knowledge source can be a collection of scientific literature. We present a proof of concept in Appendix B.5 and leave further study for future work.
Due to the assumptions made in Section 3.1, FActScore is not applicable when the facts are more nuanced, open-ended, and debatable (Chen et al., 2019; Xu et al., 2023) or with a knowledge source whose text frequently conflicts with each other (Wadden et al., 2022). Moreover, FActScore may not be suitable for the human-written text that is nuanced and includes intentional or implicit deception.
While our estimator closely approximates humans and provides consistent ranking over a large set of LMs, it is not perfect in individual judgments, and the best variant depends on the degree of how close a generation is to human-written text and its linguistic complexity. Future work can investigate how the distribution of model generation affects the performance of the estimator and further improve the estimator.
FActScore focuses on factual precision—whether each piece of information in a generation is factually supported by a reliable source of knowledge—which is only one aspect of the broader factuality problem. For instance, FActScore does not consider factual recall: the coverage of information in a generation. FActScore does not penalize a model that abstains from responding too frequently or generates fewer facts, which can be unfair since there is an inherent trade-off between precision and recall. Moreover, the boundary between precision and recall is often blurry, e.g., it is possible that, even if every piece of information in a generation is supported, it misses a significant piece of information that should have been mentioned in order to be considered as correctly responding to the input prompt (example in Table 6). We leave a more holistic evaluation of factuality for future work, and recommend reporting FActScore together with the % of abstention and the average number of atomic facts (as we did in Section 4.3).
Acknowledgement
We thank Yizhong Wang for sharing Instruction-tuned LLAMA and Alpaca models with varying sizes, and for sharing feedback on the FActScore Python package. We thank experts in Upwork for annotating the data, and Dhruba Ghosh, Jiacheng Liu and Zeqiu Wu for participating in pilot annotation and sharing feedback. We thank Akari Asai, Yanai Elazar, UW NLP members, UMass NLP members, FAIR lab members for feedback and discussion on the paper.
This research was supported by NSF IIS-2046248, NSF IIS-2202506, NSF IIS-2044660, ONR N00014-18-1-2826, ONR MURI N00014- 18-1-2670, DARPA under Contract No. FA8650-23-C-7316, an Allen Distinguished Award, and gifts from AI2. The views, opinions and/or findings expressed are those of the author and should not be interpreted as representing the official views or policies of the Department of Defense or the U.S. Government. Sewon Min is supported by a J.P. Morgan fellowship, and Kalpesh Krishna was supported by the Google PhD Fellowship.
References
Appendix A Details in Data Collection
We sample 183 human entities to be annotated as follows. We first choose entities from Wikidata whose instance of is human and have corresponding Wikipedia pages. We then categorize entities based on two dimensions: frequency and nationality, resulting in 20 categories. We then sample entities uniformly at random over all categories.
We compute freqValue as a maximum of the entity occurrence in Wikipedia provided by Kandpal et al. (2022) and the pageview count of the Wikipedia page following Mallen et al. (2023). We found using one of them could lead to an underestimate of frequency levels due to failure in entity linking or mismatch in the Wikipedia page title, and taking a maximum of them provides a reasonable solution. We then assign one of five categories: ‘Very rare’ if freqValue, ‘Rare’ if freqValue, ‘Medium’ if freqValue, ‘Frequent’ if freqValue, and ‘Very frequent’ if freqValue.
We take country of citizenship from Wikidata and assign them one of four categories: ‘North America’, ‘Europe & Middle East’, ‘Asia & Pacific’ and ‘Latin/South America & Africa’.
A.2 Details in generating atomic facts
We break out a generation automatically by splitting a generation into sentences, and feeding each sentence to InstructGPT (text-davinci-003) with a series of instructions to further break it down to a series of atomic facts. The prompt to InstructGPT is provided in Table 15. Outputs from InstructGPT are used (1) to human experts for revision (Section 3.3) and (2) for model-based evaluators (Section 4). We find human experts split and merged atomic facts from InstructGPT for 18% and 34% of the cases, respectively.
A.3 More details on annotator recruitment
We recruit freelancers through Upwork and pay 15–25 USD per hour. We recruit fact-checking experts—freelancers who mentioned fact-checking as their expertise—for Step 3. Every worker went through a qualification test of 2 hours and was tested to be highly qualified. We design one HIT to consist of three generations, one from each LM, for one prompt, because we find it saves annotation time in total. 10% of the HITs have two workers assigned to calculate the agreement rate; the rest have one worker assigned. The agreement rates are 96%, 90% and 88% for InstructGPT, ChatGPT and PerplexityAI, respectively. Appendix A.5 discusses disagreement cases in more detail. The full instructions and the interface are provided in Figure 6 and Figure 7, respectively.
A.4 Examples in annotated data
Table 7 provides examples of the human-annotated data, each atomic fact with an assigned label. Supported and Not-supported respectively indicate Wikipedia supports the fact and does not support the fact (either contradicts or does not contain any evidence). Irrelevant indicates the fact is irrelevant to the input prompt, which can further be divided into two cases: (1) the fact depends on other facts because it expands previous facts in a generation, and such other facts are Not-supported, e.g., in the first example in Table 7, and (2) the entire sentence is irrelevant to the prompt, independent from other facts in a generation, e.g., the second example in Table 7. The second case rarely happens with InstructGPT and ChatGPT, but happens considerably with PerplexityAI, i.e., 24.7% of generations of PerplexityAI have sentences marked as irrelevant without dependencies to other facts, compared to 0.5% and 1.3% in InstructGPT and ChatGPT, respectively. This is because PerplexityAI often directly copies search results even if they are largely irrelevant to the input prompt. This is in agreement with a concurrent work from Liu et al. (2023a) that shows generative search engines like PerplexityAI copy incorrect search results and generate text that is irrelevant to the input query.
A.5 Qualitative Analysis
We analyze the cases where two annotators assigned to a same generation disagree on a precision label for the same atomic fact. Categorization is provided in Table 8. The 70% is due to an inherent debatability on whether or not the fact is supported by a given source of knowledge, not satisfying Assumption 2 in Section 3.1. This is because there can be multiple interpretations of a fact, it is debatable whether or not an information can be inferred from a piece of text, or the atomic fact is subjective. For instance:
Gerhard Fischer is an inventor: Gerhard Fischer is widely known as an inventor of a metal detector, and even the title of the Wikipedia article is “Gerhard Fischer (inventor)”. However, it later turns out that he did not invent a metal detector; rather, he commercialized it.
Chadwick Boseman was a producer: Chadwick Boseman is widely known as another profession (singer) and there is no text that mentions him as a producer. However, he produced one music video.
Nonetheless, since our agreement rate is fairly high (91%), we think such cases are rare in our particular domain of people biographies. We include more discussion on other domains that such cases may be more frequent in the Limitation section.
While factual prediction is inherently a function of a knowledge source given as part of the input, a potential concern is how representative using English Wikipedia as a knowledge source for evaluating people biographies with respect to its coverage. For instance, it is possible that, especially for rare entities, the coverage of information in Wikipedia is not high enough, and LMs may be penalized by generating information that is true even if not supported by Wikipedia (i.e., supported by other sources on the web).
To quantify the effect, we randomly sample 30 unsupported facts from ChatGPT on people whose categories are either ‘rare’ or ‘very rare’, and then validate them against the entire web. We found 10% (3 out of 30 facts) are in fact supported, even though they are not supported in Wikipedia. An example is [Hibo] Wardere published her memoir titled “Cut: One Woman’s Fight Against FGM in Britain Today” which is not mentioned in Wikipedia but is found from Google Books.
Nonetheless, we found that Wikipedia has a high coverage and mentions most of the important information that we were able to find from any other sources on the web. This is in agreement with prior work that treated Wikipedia as a general knowledge source under the same reason Chen et al. (2017); Petroni et al. (2021).
Appendix B Details in Estimators
As an LM, we use the best open LM and the best commercial LM at the time of conducting experiments: LLAMA 65B (Touvron et al., 2023) and LLAMA 7B trained on Super Natural Instructions (Inst-LLAMA, Wang et al., 2022) as the former, and ChatGPT (OpenAI, 2022) as the latter. For computing nonparametric probabilities, we use a single-mask variant of NPM with BM25 as in the original paper (Min et al., 2023), and use as a thresholding hyperparameter.
For passage retrieval, we use Generalizable T5-based Retrievers (GTR, a large variant), an unsupervised dense passage retrieval system (Ni et al., 2022). We restrict retrieved passages to be from the topic entity’s page, and use . We find our estimator is not sensitive to the choice of a retrieval system (ablations provided in Appendix B.3). As a retrieval corpus, we use the English Wikipedia from 04/01/2023 which is around the time the data annotation was completed, and split each page into passages with up to 256 tokens.
We also compare with Self-check LM, a method from a concurrent work by Manakul et al. (2023). Self-check LM needs multiple samples generated from the LM. It validates the given atomic fact by prompting LM conditioning on each generated sample,Manakul et al. (2023) uses BERTScore and a supervised question answering system instead of LM prompting, however, we find LM prompting to be significantly better. making judgment (Supported or not) from each, and aggregates the results through a majority vote. This method assumes (1) the LM is available at the time of evaluation and (2) the outputs from the LM are nondeterministic, which makes it not applicable to PerplexityAI.
B.2 Segment-level vs. system-level evaluation
Besides how close the estimated FActScore is to the ground truth FActScore (Error Rate, as reported in Section 4), we also report F1. F1 evaluates how well the model validates each individual atomic fact, assuming oracle atomic facts (atomic facts by human experts) are given, and evaluates how good the estimator is in identifying facts that are not Supported (NS). Formally, let and be sets of atomic facts in a set of generations that have Not-supported as a ground truth label and as a predicted label, respectively. We define F1 as follows.
We call them micro because they consider individual decisions rather than aggregated estimation.
F1 cares about the individual decision, while ER cares about the aggregated estimation. An evaluator that has a high (better) F1 but always overestimates or underestimates factual precision may have a higher (worse) ER, e.g., Evaluator A in Figure 4. Conversely, an evaluator that has a lower (worse) F1 but is not biased toward overestimation nor underestimation may have a lower (better) ER, e.g., Evaluator B in Figure 4. Prior work in model-based evaluation mainly reports aggregated scores since the goal is a comparison between different systems being evaluated (Zhang et al., 2020; Rashkin et al., 2021; Gao et al., 2022) while we report both to see the relationship between two types of metrics. F1 and ER are also closely related to segment-level and system-level correlations to human judgments respectively, which have been extensively used in developing evaluation metrics in machine translation (Ma et al., 2019; Thompson and Post, 2020) and summarization (Bhandari et al., 2020; Deutsch et al., 2021).
Results on F1 are reported in Table 9. Self-check LM outperforms no-context LM by 4–11%, which confirms findings from Manakul et al. (2023). However, both significantly underperform methods that use retrieval. This is in contrast to Manakul et al. (2023) that reports that Self-check without retrieval achieves performance that is close to that with retrieval, likely because the data in Manakul et al. (2023) contains more frequent entities. The fact that retrieval significantly helps is consistent with findings in Section 4.2 with an ER as a metric.
Adding NP improves RetrieveLM by 2–9%, again consistent with findings in Section 4.2. This is likely because RetrieveLM often makes incorrect predictions when there is a strong bias from an LM or there are distracting passages, and considering nonparametric probabilities makes the model more robust to these factors. For instance, given an unsupported fact Samuel Oboh is Nigerian, No-context LM, Self-check LM and RetrieveLM predict Supported due to a strong name-nationality bias. NPM correctly predicts Not-supported based on a passage Samuel Oboh ... is a Canadian architect, manager, .... It is also worth noting that this is different from findings in Section 4.2 that ChatGPT is not necessarily better than LLAMA+NP based on ER.
Table 10 reports a comparison across different choices of an LM. Within the same method, Inst-LLAMA 7B outperforms LLAMA 65B, and ChatGPT outperforms both. Using retrieval is critical across all models, e.g., the best no-context model based on ChatGPT is underperformed by all models with retrieval. Using NP helps LLAMA-based models but not ChatGPT, likely because ChatGPT is less affected by incorrect prior from the LM or distracting passages.
It is worth noting that these results are somewhat different from findings in Section 4.2 that ChatGPT is not necessarily better than LLAMA+NP. This is becauase, although ChatGPT is better in validating each individual atomic fact, most errors from ChatGPT are incorrectly assigning Supported to Not-supported facts, resulting in an overestimation of FActScore. In contrast, LLAMA+NP is not biased toward overestimation or underestimation of the factual precision, resulting in an aggregated factual precision to be closer to the ground truth. This is similar to the trade-off between system-level and segment-level correlations in summarization evaluation (Bhandari et al., 2020; Deutsch et al., 2021).
B.3 Ablations
As described in Section 4.1, we use True or False as part of the prompt, so-called TF Prompting. An alternative is QA Prompting, which generates a question and the expected answer, obtains the answer for the generated question independent from the expected answer, and compares the expected answer and the predicted answer. This approach has been widely studied in the summarization literature and recent work in factual precision (Kryscinski et al., 2020; Wang et al., 2020; Gao et al., 2022; Manakul et al., 2023). Table 11 provides a comparison between two types of prompting. The TF approach significantly outperforms the QA approach, consistently over all methods. Our further analysis finds that this is due to generated questions often being overly vague or ambiguous. For instance, given a supported fact Samuel Oboh is an architect, the LM generates What is Samuel Oboh’s job? as a question and Architect as an expected answer, and the obtained answer is Vice President. Although both Architect and Vice President are correct, they are not the same, thus the model incorrectly predicts Not-supported. Such cases make the model overpredict Not-supported, leading to many incorrect predictions.
Table 12 compares RetrieveLM methods based on a few passage retrieval systems, including BM25 (Lin et al., 2021), GTR Large and GTR xLarge. Results indicate that all retrieval systems are equally good and RetrieveLM is not sensitive to the choice of the retrieval system.
Table 13 categories errors made by RetrieveLM based on ChatGPT, the evaluator with the best F1. 70% of the errors are due to retrieved passages not providing direct evidence (either support or contradiction). These are difficult even for state-of-the-art retrieval systems and language models because validating facts often requires reading the entire page rather than a single passage, e.g., an actor not appearing in a particular film. 17% of errors are made because ChatGPT is being distracted by other passages, although it assigns a correct label if only a particular, correct passage is given.
B.4 More details in evaluation of new LMs (Section 4.3)
Figure 5 reports FActScores estimated by two variants of our estimator as in Figure 3 but with 100 random subsets of the data. Specifically, we chose samples (out of ) uniformly at random across 20 categories (defined in Appendix A.1) times and report the average and the standard deviation. We use and . Results indicate that the variance is overall low, preserving ranking between 13 subjects in most cases. As expected, the variance is lower as the sample size gets larger. Finally, the estimator based on ER based on LLAMA+NP (bottom) has an overall lower variance than the estimator based on ChatGPT (top).
B.5 Feasibility in applying FActScore to other domains
As mentioned in the Limitation section, our paper mainly evaluates on people biographies using Wikipedia. Evaluating the generalizability of FActScore to other types of prompts and other domains is an avenue for future work.
As a proof of conept, we conduct small-scale studies in the NLP domain. We first manually write 10 prompts asking about NLP papers: Tell me a summary of
This suggests that FActScore can generalize beyond people biographies. However, since this is a very small-scale experiment, we strongly encourage future research to explore the generalizability of FActScore to more domains at scale.
Appendix C Editing Experiments
Our experiments in Section 4 focuses on automatically identifying factual precision errors in long-form generations by language models. Can these labels be used to actually correct errors in the long-form generations? In this section, we perform a preliminary exploration of methods to edit long-form LM generations to reflect factually correct information. We assume we have access to the human-annotated set of FActScore labels, and measure how good models are at editing incorrect sentences. In other words, we evaluate our editor models independent of the errors arising from the estimator.
We adopt a similar set of methods as Section 4.1 for our editing models. All methods below use four exemplar examples for in-context learning which were sampled from our dataset and removed for subsequent analysis. For all methods, we use OpenAI’s ChatGPT (OpenAI, 2022) as the base language model due to its generative capabilities.
No-context LM. We feed language models the prompt Input:
RetrvLM. To assist an editor model, we use a passage retrieval system to find supporting evidence from an external knowledge source (Wikipedia in our case). Our retrieval pipeline is identical to Appendix B.1, but uses 3 retrieved passages instead of 5 due to context length restrictions.
+ Atomic Facts. Additionally, we explore whether adding atomic facts and their labels assist a model with fine-grained editing. Specifically, after the input sentence we add information to the prompt of the form Fact 1 (True/False):
Non-edit baselines. Finally, we add some trivial baselines to lower-bound our editing metrics. Specifically, we measure the performance of input copying (no edits), as well as an editor with random token dropping / replacement on a random 25% subset of tokens.
C.2 Evaluation
In our data collection process (Section 3.3), along with our verification data we also collected gold-standard human written edits. Let be the input sentence and be the gold edited sentence. We evaluate the quality of the model-generated edit () using three automatic metrics,
(1) Error Localization (ErrLoc): Our first metric measures how well the editor identifies errors within the input sentence. Specifically, we first create a “token preservation string”, marking token in the input sentence as "Preserved" or "Not Preserved". We then compute the macro-averaged F1 score between the token preservation strings derived from the gold edit and the model-generated edit. We remove stopwords, punctuation and lowercase all words before performing this calculation. To equally weigh every sentence, F1 scores are independently computed for each sentence before a final averaging.
(2) Edit Correctness (EditCorr): Our second metric assesses the quality of the additional tokens added by the model-generated edit. Specifically, we check the token-level F1 score (Rajpurkar et al., 2016) comparing the new tokens added by the gold edit and the new tokens added by the model-generated edit . More concretely,
where is the set cardinality and HM denotes a harmonic mean. For this metric, we discard data points where the gold edit did not add new tokens. Similar to ErrLoc, we also remove stopwords, remove punctuation and lowercase strings before calculating EditCorr scores.
(3) SIM alignment (SimAl): Finally, due to the large output space of possible edits, we also adopt a metric which rewards paraphrases of the gold edits. We use semantic similarity embeddings from Wieting et al. (2022) which map paraphrases to a similar part of a vector space. We check the similarity between the model edit and the gold edit , normalizing it by the similarity between and the original input .We avoid taking the vector differences between the original / edited text since edit vectors (Guu et al., 2018) were not explicitly modeled in Wieting et al. (2022). Specifically,
where is the semantic similarity score (normalized to $GEGX$.
C.3 Results
We present our editing results in Table 14. Overall, we find that:
All editing models perform better than trivial lower bounds. Overall, we find that all editor models outperform lower-bound baselines like random noise. This even happens in the no-context LM setting, where ChatGPT is editing its own output (or search engine augmented Perplexity AI’s outputs), but can still perform non-trivial corrections (6.8 ErrCorr for ChatGPT correcting its own outputs vs 0.1 for a random noise editor baseline).
Retrieval significantly helps with editing performance. Across all base language models and metrics, augmenting the editor with retrieved paragraphs boosts performance (6.8 16.8 ErrCorr, 4.0 9.5 SimAl for ChatGPT correcting its own outputs). We hypothesize that the internal parametric knowledge in ChatGPT has insufficient information about the topic (as we also observed in Section 3.4) to perform fine-grained editing, and using external knowledge from Wikipedia greatly simplifies error localization and correction. This also corroborates with our findings in Section 4.2.
Atomic fact labels improve error localization and improve editing performance. Across all base language models (with or without retrieval) we observe that providing fine-grained atomic fact labels improves editing performance (16.8 28.3 ErrCorr, 9.5 19.3 SimAl for ChatGPT correcting its own outputs). Fine-grained fact correctness labels help the editor easily identify problematic tokens, as seen by the consistent improvements in ErrLoc scores (43.9 63.5 for ChatGPT correcting itself). We hypothesize atomic facts help guide the editor with its editing process (for instance, perform a more targeted search in the retrieved paragraphs), resulting in ErrCorr improvements. We also find that atomic fact labels reduces the frequency of editor copying the input verbatim or saying The input has no errors from 37.3% to 3.9%.
PerplexityAI outputs are the hardest to edit. Overall, we find the highest editing success for InstructGPT, followed by ChatGPT and the least success for Perplexity AI. We hypothesize this is because PerplexityAI already uses a search engine, so errors are much more subtle as extensively discussed in Appendix A.5.