FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios
I-Chun Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, Pengfei Liu
Introduction
Generative artificial intelligence (AI) technology, exemplified by GPT-4 OpenAI (2023) consolidates various tasks in natural language processing into a single sequence generation problem. This unified architecture enables users to complete multiple tasks (e.g., question answering Thoppilan et al. (2022), code generation Chen et al. (2021), math problem solving Lewkowycz et al. (2022), and scientific literature generation Taylor et al. (2022)) through a natural language interface Liu et al. (2023) with both unprecedented performance Bubeck et al. (2023) and interactivity.
However, at the same time, such a generative paradigm also introduces some unique challenges. Content that is automatically generated can often exhibit inaccuracies or deviations from the truth due to the limited capacity of large language models (LLMs) Ji et al. (2023); Schulman (2023). LLMs are susceptible to producing content that appears credible but may actually be factually incorrect or imprecise. This limitation restricts the application of generative AI in some high-stakes areas, such as healthcare, finance, and law. Therefore, it is crucial to identify these errors systematically to improve the usefulness and reliability of the generated content.
Current literature on detecting and mitigating factual errors generated by machine learning models focuses predominantly on a single specific task, for example, retrieval-augmented verification models for QA Lewis et al. (2020), hallucination detection models for text summarization Fabbri et al. (2022), and execution-based evaluation for code Shi et al. (2022). While these methods have proven successful within their respective areas, given the remarkable versatility of tasks and domains handled by LLMs, we argue that it is also important to have a more comprehensive factuality detection and verification framework that is similarly versatile.
Additionally, in the current literature, the task of factuality detection is usually simplified as either (i) given a claim, determining whether it is factually correct, (ii) or given evidence, determining whether the generated claim is supported. This task definition is not well suited to writing tasks that users commonly engage with when interacting with generative models (e.g., ChatGPT), where we often need to validate the factuality of a long-form generation without explicit claims and evidence.
In this paper, we propose a task and domain-agnostic framework, FacTool, which aims to detect factual errors in LLM-generated texts. We illustrate our framework in Fig. 1, where we connect the concept of “tool use” Thoppilan et al. (2022); Gao et al. (2022b); Schick et al. (2023) with “factuality detection” and demonstrate that the ability to use tools in LLMs is crucial for factuality detection. Specifically, FacTool leverages various tools, including Google Search, Google Scholar, code interpreters, Python, or even LLMs themselves, to gather evidence about the factuality of the generated content. Moreover, our framework employs the reasoning abilities of LLMs to assess the factuality of the content, given the evidence that has been gathered. We develop a benchmark and perform experiments across four tasks: knowledge-based QA, code generation, math problem solving, and scientific literature review writing.
We revisit the task of factuality detection and extend it in a way that allows for a better audit of current generative AI models.
We connect the concept of “tool use” with “factuality detection”, developing a unified and versatile framework for factuality detection across a variety of domains and tasks.
We use FacTool to evaluate the factuality of modern chatbots, and found that GPT-4 has the best factuality across almost all scenarios. Supervisely fine-tuned chatbots (Vicuna-13B) have reasonably good factuality in KB-based QA but perform poorly in more challenging scenarios, including code generation, math problem solving, and scientific literature review writing.
Related Work
Factuality detection was a topic of rigorous study even before the advent of generative AI. Existing works can be organized by their differences in terms of the “response” to be verified, the “claim” extracted from the response, and supporting “evidence”. As illustrated in Tab. 1, the creation of the FEVER dataset Thorne et al. (2018a) spawned models Zhong et al. (2020); Krishna et al. (2022) that determine whether a given fine-grained claim made based on Wikipediahttps://www.wikipedia.org/ articles is correct. In this task setting, both the claim and related evidence are given. FactCC Kryscinski et al. (2020) and QAGS-based models Wang et al. (2020) adopted different task formulations to detect factual consistency, i.e., given the evidence text, and the goal is to determine if the generated summaries or summary sentences are factually consistent with the given text. WICE-based methods Kamoi et al. (2023) decide if a fact from a Wikipedia sentence could be supported by provided evidence. RARR Gao et al. (2022a) proposed a new approach by directly prompting LLMs to generate queries, retrieve evidence and determine factuality.
Existing works typically rely on either a given claim or given evidence and target a specific use case. However, in this paper, we introduce a more challenging yet practical task setting, i.e., factuality detection without explicit claims or evidence, and propose a framework capable of addressing this challenge in a variety of scenarios.
Language models store limited knowledge within their parameters. To overcome this limitation, various tools have been introduced to assist language models in order to further expand their capabilities. For example, Press et al. (2022); Komeili et al. (2022) gathered information from the Internet to enhance question answering and dialog systems, respectively. Schick et al. (2023) trained a model capable of interacting with five tools including a calculator, a translation system, etc. Recently, Shen et al. (2023) introduced a framework that employs LLMs to connect various AI models from the machine learning communities to tackle AI tasks. Furthermore, Liang et al. (2023) proposed a new AI ecosystem that connects LLMs with millions of existing APIs to accomplish tasks. In this work, we explore tool use in LLMs for the task of factuality detection.
Revisiting Factuality in Generative AI
In most previous works, factuality has been defined as whether a claim in a text can be supported by evidence from a separate, trustworthy knowledge base, with applications in fact-checking Thorne et al. (2018b) (where the knowledge base is a large source like Wikipedia) and summarization Kryscinski et al. (2020) (where the knowledge base is an input document or documents). In this paper, we extend this definition to whether the claims made in generated signals (which could be text, code, or mathematical expressions and so on) can be supported by evidence under specific rules. Specifically, these rules can range from consistency with a knowledge base derived from Wikipedia, to a verification rule specified within a Python library, or an operational rule derived from mathematics. By adopting this broader definition, we are able to establish a unified framework for addressing factuality issues in generative AI beyond just the textual domain.
One can usually detect the factuality of a given generated signal (e.g., text) at different levels of granularity, such as sentences, and documents. A more granular assessment can be particularly valuable because it (1) not only allows users to pinpoint where inaccuracies occur Liu et al. (2021) but also (2) serves as a reward model for developers to refine their generative systems Lightman et al. (2023).
However, implementing fine-grained factuality detection is challenging due to two reasons: (1) specifying the desired granularity level without ambiguity, and (2) extracting claims in line with the predetermined granularity level. In this paper, we argue that by utilizing the powerful instruction-following ability and the natural language interface of LLMs, we can effectively address the challenge of defining and extracting fine-grained claims through claim definition-based few-shot prompting. More details can be found in §4.1.
Structurally speaking, given a prompt (e.g., a query or instruction) and the corresponding model-generated response, the fine-grained factuality detection task involves the following concepts:
Prompt () a query or instruction that users provide to the generative model.
Response () a piece of text (usually in long form) generated by the generative model.
Claim () a statement inferred from the model response, whose granularity is defined by a natural language text.
Evidence () The available information (e.g., knowledge base, pre-defined rules) that support or demonstrate the truth or validity of a claim.
2 Instantiations in Different Scenarios
Using the above task definition, we can define factuality in different application scenarios (see also in Tab.2).
Knowledge-based (KB) QA Chen et al. (2017) aims to answer questions using a given knowledge base or open-domain data source (e.g., Wikipedia). In this task, we define factuality as how well each claim in the generated answer is supported by world knowledge. In this paper, we consider a more challenging scenario: open-domain QA that requires long-form answers, rather than short ones.
The code generation task Yin and Neubig (2017) involves generating executable code based on a given query. We define factuality in code generation as how well the generated code, as a whole, can be executed correctly within a specific programming language (e.g., Python) and fulfills the provided requirements. This definition is grounded in an execution-based approach to code evaluation, which measures the correctness of generated code by executing it against some test case inputs and comparing its output to the expected output.
The math problem solving task involves the use of automated methods to address mathematical problems Cobbe et al. (2021). At the claim level, factuality in math problem solving is defined as the extent to which the generated statements adhere to the calculation rules. At the response level, factuality in math problem solving is defined as how effectively the overall mathematical solution addresses the given problem.
The scientific literature review writing task Jha et al. (2015) aims to analyze and synthesize existing research on a specific topic in a field of study. In this task, we define factuality as whether the generated scientific literature review correctly cites existing scientific literature, including the correct mention of authors and publication years.In this paper, our focus lies in examining the consistency of the relationship between the paper title, authors, and publication year. However, the task of determining the suitability of the cited paper as the most appropriate choice is left for future investigation.
Approach
We propose a tool-augmented framework for detecting factual errors that can apply a unified approach across various tasks. The motivation for using tools is twofold. On one hand, each tool embodies the domain expertise, assisting us in the effective gathering of evidence that verifies the correctness of the claim. On the other hand, the ability of LLMs to utilize multiple tools paves the way for multiple tool-augmented factuality detection. For example, by directly using ChatGPT plugins,https://openai.com/blog/chatgpt-plugins we can integrate multiple tools into a chatbot.
The framework is illustrated in Fig. 1, which consists of five main components: claim extraction, query generation, tool querying, evidence collection, and agreement verification. We elaborate each component below.
Extracting claims from responses under various task settings is challenging due to the inconsistent definitions of claims across tasks and scenarios. This inconsistency hinders the development of applications such as text summarization evaluation and factuality detection. To tackle this, we propose an approach in this paper that treats claim extraction as a process guided by LLM prompts based on the specific definition of claims. This approach offers the following advantages:
(i) Leveraging the strong instruction-following capabilities of LLMs can significantly reduce the costs associated with data annotation and model training for claim extraction.
(ii) When developing a system or constructing a dataset for an application that relies on the definition of claims, one simply needs to provide a textual definition of the claim using a large model. This enables future researchers to effectively utilize these definitions as a foundation in their work.
(iii) Our experiments demonstrate that the claim extraction module, implemented by ChatGPT, exhibits strong performance in extracting claims (atomic component units). The detailed results of these experiments are discussed in Section 6.1.
Here, we employ ChatGPT as a base LLM and apply different textual definitions of claims across four tasks. Our goal is to extract all verifiable claims within the generated text , denoted as . Detailed prompting instructions can be found in Appendix A.
The claim is defined using the concept of atomic content units (ACUs) Liu et al. (2022). Each ACU corresponds to a single atomic fact within a generated answer. In practice, we leverage ChatGPTWe have also explored other entailment-based models with BERT, and the result is no better than ChatGPT. (specifically, the “gpt-3.5-turbo” version) to extract claims based on two criteria: (i) each claim should not exceed 15 words, and (ii) it should clearly describe a fact. We also include two in-context examples from the RoSE dataset Liu et al. (2022) in our prompt to obtain more fine-grained claims. Additionally, we ask ChatGPT to resolve any coreferences or ambiguity, such as unclear pronouns and other related expressions within the claims.
We consider each generated code snippet within the response as a single claim to be verified. We extract all such code snippets that are enclosed with brackets, in other words, within a code block.
We define each claim in a step-by-step math solution as the arithmetic operation performed between known real numbers. Each of these operations contains two parts: the calculation and the calculated answer. We prompt ChatGPT to extract all such claims.
Each claim within the generated review is defined as a tuple of “(paper title, year, authors)” contained from the generated review. We then prompt ChatGPT to extract all such tuples within the generated review.
2 Query Generation
For each claim , we convert it into a list of queries that can be used to query external tools such as search engines, the Python interpreter, or Google scholar.
We prompt ChatGPT or GPT-4 to generate two search engine queries from each claim . These queries are intended to help humans in verifying the factuality of . Detailed prompting instructions can be found in Appendix A.
For each claim we generate two different types of queries: simulated test case inputs, denoted as , and potential solutions, denoted as . Both types of queries are generated by ChatGPT or GPT-4. The simulated test case inputs are function calls generated for a given code snippet, while potential solutions are repeatedly generated solutions that ChatGPT generates in response to the user prompt . In our later experiments, we generate 3 simulated test case inputs and 3 potential solutions. Detailed prompting instructions can be found in Appendix A.
We prompt ChatGPT or GPT-4 to convert all mathematical operations into executable Python code snippets. These snippets are designed to return “True” when the calculation matches the calculated answer and “False” if it doesn’t. Detailed prompting instructions can be found in Appendix A.
We use the paper title, found within the extracted claim tuple, as the query for Google Scholar. Our assumption here is that if a paper exists, it should appear as the first search result on Google Scholar when we use the paper title as the query.
3 Tool Querying & Evidence Collection
We then use the queries to query various tools to collect relevant evidence statements .
The external tool we use to help verify the factuality of the generated text is the Google Search API, which queries the internet for knowledge using the queries generated from the claims extracted from the generated text of LLM.
We use the Google Search API provided by Serperhttps://serper.dev/ to search the top pages and retrieve the most relevant search snippets included in the API’s response. We then parse the response to obtain different types of snippets such as answer boxes, knowledge graphs, and organic search results.
For each test case input and generated potential solution , we execute using as the input and collect the execution result (output) for each pair. The input-output pairs are used as test cases for verifying the chatbot generated unverified solution. The process is shown in Fig. 3.
We collect the execution results for code snippets derived from the mathematical operations. As illustrated in Fig. 2, math claims like “30 /3 = 10” are extracted and then converted into a Python executable code, for instance, “print(round(30/3, 7)==10)”.
We use the title of each paper, extracted from the text, as the query to access relevant information through the Google Scholar API provided by the Scholarlyhttps://github.com/scholarly-python-package/scholarly Python package. This allows us to retrieve key information about each paper, including the paper title, author list, and publication year.
4 Agreement Verification
In the final step, each claim, , receives a binary factuality label, , based on the level of support it receives from the collected evidence, . This labeling process is performed for every individual claim.
We prompt ChatGPT or GPT-4 to judge the factuality of the claim given the retrieved list of evidence snippets. We follow a zero-shot Chain-of-Thought Wei et al. (2023) reasoning process: initially, the model attempts to reason about whether the claim is factual or not. If an error is identified, we then ask it to explain and attempt to rectify the mistake.
We conduct a majority vote for each test case across all solutions, establishing what we refer to as the “pseudo-golden output” for that particular test case. We repeat this process for every test case. Following this, we compare the execution result of the solution that’s under verification against all the test cases with the pseudo golden output. If the results match, we classify the solution under verification as true. Otherwise, it is deemed false.
We compile the results of each code snippet execution. If any snippet returns “False”, we classify the associated generated text as false. Conversely, if all snippets yield “True”, we classify the corresponding generated text as true.
We compare the extracted claim: “(paper title, year, authors)” to the evidence: “(paper title, year, authors)” retrieved from Google Scholar API. For the paper title and year of publication, we conduct an exact, case-insensitive string match. As for the authors’ match, we prompt ChatGPT or GPT-4 to judge whether the author list in the extracted claim is a subset of the retrieved author list. All the information must be matched in order to be classified as “True”, otherwise “False”.
Dataset Construction
For KB-based QA, we evaluate our framework using RoSE Liu et al. (2022) and FactPrompts. RoSE is a text summarization dataset that provides fine-grained ACUs for each reference summary. FactPrompts is a dataset that comprises real-world prompts sourced from various platforms and datasets, such as Quora and TruthfulQA Lin et al. (2022), along with corresponding responses generated by ChatGPT. We construct the dataset using 100 reference summaries from RoSE and 50 responses from FactPrompts for our evaluation.
For code generation, we evaluate our framework using HumanEval Chen et al. (2021). HumanEval is a programming problem dataset that contains several unit tests for each problem. We use ChatGPT to generate responses based on the processed prompts of HumanEval provided in Chen et al. (2022) which solely contain the instruction of the prompt without input-output demonstrations.
For math problems, we evaluate our framework using GSM-Hard Gao et al. (2022b). GSM-Hard is a dataset constructed from GSM8K Cobbe et al. (2021) by replacing the numbers in the questions of GSM8K with larger numbers. We sampled 100 prompts from GSM-Hard that have a target solution value of positive.GSM8K involves many application questions, including calculations involving money, measurements of quantities, etc. We found that GSM-Hard examples with negative values often contained illogical situations, such as “negative 5 apples”. A positive target solution value helps prevent ChatGPT from making extra assumptions on top of the description in the problem. Then, we generate responses for these prompts using ChatGPT.
For the scientific literature review, we follow self-instruct Wang et al. (2023) to create 100 diverse prompts spanning computer science, business, law, medicine, and physics. Each prompt asks for a technical or research-oriented response that includes at least one relevant literature citation. Then, we generate responses for these prompts using ChatGPT.
2 Claim Collection
For responses from FactPrompts and GSM-Hard, we follow the idea of “claim extraction as prompting” described in §4.1, This approach allows us to reuse claim prompts as listed in Appendix A. We use ChatGPT as the model for claim extraction due to its cost efficiency and effectiveness in extracting fine-grained claims. In terms of HumanEval responses, given that the generated response to a HumanEval prompt is already in the form of a code snippet, we consider the “claim” of the response to be identical to the response itself.
3 Claim and Response Annotation
For claim annotation, the authors collectively annotate the extracted claims as either factual or non-factual. For response annotation, if one claim within the response is labeled as non-factual, then the response as a whole is considered non-factual; otherwise, the response is considered factual.
We consider the claim label to be identical to the response label since the “claim” of the response is the same as the response itself. For response annotation, we annotate ChatGPT’s responses using the execution code provided in Chen et al. (2022) against the HumanEval test cases. This allows us to distinguish between factual (those passing all tests) responses and non-factual responses.
For claim annotation, the authors collectively annotate the extracted claims as either factual or non-factual. For response annotation, we utilize the target value provided in GSM-Hard Gao et al. (2022b) to annotate the generated responses.
Experiments
We evaluate FacTool against two baselines that use LLMs to check their own inputs: Self-Check with 3-shot CoT and zero-shot CoT, which have shown to been effective on various tasks including dialogue response, math reasoning, and code generation Madaan et al. (2023); Chen et al. (2023). Both of these baselines aim to test the ability of LLM to identify its own errors without the use of any external tool. In practice, we prompt ChatGPT (gpt-3.5-turbo-0301) and GPT-4 (gpt-4-0314)We anticipate that the recently released models, gpt-3.5-turbo-0613 and gpt-4-0613, will lower the inference costs for FacTool. This expectation arises from their improved ability to produce structured responses, such as those in JSON format. While conducting our experiments on gpt-3.5-turbo-0301 and gpt-4-0314, we often ran into problems where the responses were not valid JSON, requiring us to rerun any samples with invalid response formats. The source code of FacTool will be using the latest versions of ChatGPT and GPT-4. to recognize, explain, and attempt to rectify their own errors. Following this reasoning process, the models make final judgments on the factuality of the given claim. The key difference between Self-Check (zero-shot CoT) and Self-Check (3-shot CoT) is that Self-Check (3-shot CoT) provides three demonstrations to models, while Self-Check (zero-shot CoT) does not provide any demonstrations.
We demonstrate in Tab. 4 that the claims extracted by GPT-4, ChatGPT, and Flan-T5 closely match the ACUs annotated by humans, as evaluated by ROUGE and BERTScore metrics. Note that in Exp-II, we choose ChatGPT as the claim extractor for two reasons: (1) The context length of Flan-T5 is too short (512 tokens) to effectively extract claims from lengthy responses in our dataset. (2) ChatGPT is more cost-efficient compared to GPT-4, while maintaining similar effectiveness in claim extraction.
2 Exp-II: Framework Evaluation
We evaluate FacTool and the two Self-Check baselines on the dataset constructed from each scenario. Depending on the model used for query generation and agreement verification, we have two FacTool baselines: FacTool powered by ChatGPT and FacTool powered by GPT-4. We report the accuracy, recall, precision, and F1-score at both the claim and response levels.
Tab. 5 shows the claim-level and response-level performance of FacTool and the self-check baselines. We obtain following observations.
From Tab. 5, we observe that FacTool powered by GPT-4 outperforms all other baselines across all scenarios. FacTool powered by GPT-4 achieves an claim-level F1 / response-level F1 on KB-based QA, a claim-level F1 / response-level F1 on code generation (remember that claim-level factuality is considered equivalent to response-level factuality in our experiment for code generation), a claim-level F1 / response-level F1 on math problems, and a claim-level F1 / response-level F1 on scientific literature review. Each of these figures is the highest for their respective tasks.
From Tab. 5, we show that FacTool with GPT-4 outperforms all self-check baselines across all scenarios. On FacTool powered by GPT-4 v.s. Self-Check (3) powered by GPT-4, we observe: v.s. response-level F1 on KB-based QA, v.s. response-level F1 on code generation, v.s. response-level F1 on math problems, and v.s. response-level F1 on scientific literature review.
From Tab. 5, we show that FacTool powered by GPT-4 significantly outperforms the self-check baselines in scientific literature review. On FacTool powered by GPT-4 v.s. Self-Check (3) powered by GPT-4, we observe: v.s. claim-level F1 and v.s. response-level F1. Here, Google Scholar shown to be highly robust in performing its specified task of finding citations when compared to LLM itself.
FacTool powered by GPT-4 outperforms FacTool powered by ChatGPT across all scenarios. This trend is especially significant in KB-QA, where query generation and agreement verification are harder for ChatGPT but relatively easier for GPT-4 ( v.s claim-level F1 and v.s response-level F1). On the other hand, in scenarios where query generation and agreement verification are relatively easy for both ChatGPT and GPT-4, the performance is similarly good.
From Tab. 5, we observe that self-check models have lower precision compared to FacTool. On Self-Check (3) powered by GPT-4 v.s. FacTool powered by GPT-4, we observe: v.s. response-level precision on KB-based QA, v.s. response-level precision on code generation, v.s. response-level precision on math problems, and v.s. response-level precision on scientific literature review. These figures show that self-check models tend to classify claims as “True” considerably more frequently than FacTool, suggesting a lower sensitivity for error detection.
Tab. 5 shows that Self-Check (0) powered by ChatGPT outperforms FacTool powered by ChatGPT. Through examining specific cases, we found that reasoning errors are the main reason why FacTool powered by ChatGPT performs worse than the self-check baselines. Even when provided with sufficient evidence to determine whether the claim is factual or not, the agreement verification implemented by ChatGPT can become confused. For example, for the claim “The modern-day version of fortune cookies was invented in the United States.”, the reasoning of FacTool powered by ChatGPT is self-contradictory: “The given text is not entirely factual. The modern-day version of fortune cookies was not invented in the United States. Most people nowadays believe that fortune cookies were created by a Japanese man named Makoto Hagiwara in 1914 in San Francisco. Hagiwara owned what is now called the Golden Gate Park Japanese Tea Garden, where he served tea and fortune cookies. This is supported by the provided evidences.” Detailed examples can be found in Fig. 9 of Appendix B.
2.2 Performance Analysis
We take a closer look at performance in different scenarios by examining evaluated cases.
The fact-checking capability of FacTool on KB-based QA is determined by several factors, including whether the search engine can return the most relevant snippets that could assist in determining the factuality of the given claim, the quality of the generated search engine queries, and the LLM’s ability to reason about the validity of the claim given the retrieved evidence. We found that FacTool powered by GPT-4 is especially capable under the following situations: (1) Fact-checking recent events, discoveries, or news: FacTool powered by GPT-4 successfully identify false claims such as “Argentina has not won the World Cup since 1986” and “The most valuable NFT ever sold is a digital artwork called ‘Everydays: The First 5000 Days’”. (2) Fact-checking high-precision statistics: FacTool powered by GPT-4 successfully identify false claims such as “Ireland has an obesity rate of 26.9%” and “Everydays: The First 5000 Days’ sold for million”. Detailed examples can be found in Fig. 10 of Appendix B.
The fact-checking capability of FacTool on code generation is determined by the LLM’s capability to generate high-quality test cases and potential solutions. We demonstrate that due to GPT-4’s exceptional ability to generate such high-quality test cases and potential solutions, FacTool powered by GPT-4 outperforms other baselines. For example, in “HumanEval/36”, GPT-4 is consistently generating high quality solutions, leading to its correctly identifies the mistakes in the response, while ChatGPT fails to identify the mistake. Detailed examples can be found in Fig. 11 and Fig. 12 of Appendix B.
The fact-checking capability of FacTool on math problems is determined by the LLM’s capability to generate accurate Python snippets that verify the correctness of given extracted mathematical calculations. Both FacTool powered by GPT-4 and FacTool powered by ChatGPT excel in this regard. For example, both FacTool powered by GPT-4 and FacTool powered by ChatGPT correctly identify doesn’t equal to . Detailed examples can be found in Fig. 13 of Appendix B.
The fact-checking capability of FacTool on Scientific Literature Review is determined by the LLM’s capability to identifying whether the author list generated is a subset of the actual author list. Both FacTool powered by GPT-4 and FacTool powered by ChatGPT excel in this regard. For example, both FacTool powered by GPT-4 and FacTool powered by ChatGPT correctly identify that the paper “The Impact of Artificial Intelligence on Employment” was not written by “Acemoglu and Restrepo”. Detailed examples can be found in Fig. 14 of Appendix B.
2.3 Failure Analysis
To gain a comprehensive understanding of FacTool’s performance, we conduct analysis on cases where FacTool will fail.
We summarize following sources of errors: (1) Reasoning error: Although the evidence provided is sufficient and the LLM accurately finds the most relevant information, the model fails to reason about the relationship between the claim and the provided evidence. For example, for claim “Jupiter is less dense than Saturn”, FacTool powered by GPT-4 fails to reason the relative relationship even though the evidences provided are sufficient. (2) Conflicting evidence: Conflict in evidence can cause confusion for LLM, leading to incorrect decisions. For example, for claim “Jupiter has a density of 1.33 grams per cubic centimeter”, there are conflicting evidences claiming that the density is 1.326 or 1.33g/cm3 . (3) Ambiguity in claim: Ambiguous descriptions and subjective adjectives can lead to incorrect decisions. For example, the claim “Fortune cookies are enjoyed by people all over the world.” is ambiguous and can have different answers based on different interpretations. Detailed examples can be found in Fig. 15 of Appendix B.
Errors in code generation mainly comes from: (1) Limited variety in synthetic test cases: The synthetic test cases generated by LLMs may not be fully representative or sufficiently diverse. For example, in the “HumanEval/64” sample, all the inputs of the generated synthetic test cases are composed of strings that only include lowercase letters (without uppercase letters). (2) Potential errors in code generation: The generated potential solutions could contain errors or bugs. Despite implementing a majority voting system to lessen this issue, it cannot completely eliminate the chance of bugs in the code generation process. For example, in the “HumanEval/79” sample, all the generated solutions failed to correctly “decimal_to_binary(0)” as “db0db”. Detailed examples can be found in Fig. 16 of Appendix B.
There are two major types of errors in factuality detection for math problems: (1) Round-off error: Round-off errors can occur during numerical calculations in Python. For example, FacTool powered by GPT-4 incorrectly classify the math calculation “60444034 / 12 = 5037002.83” as “False”. (2) Reasoning error: Since the claims extracted by FacTool only involve mathematical calculations, FacTool will not verify the reasoning process of the mathematical solution. For example, for the question “Kylar went to the store to buy glasses for his new apartment. One glass costs $5, but every second glass costs only 60% of the price. Kylar wants to buy 5364765 glasses. How much does he need to pay for them?”, the ChatGPT generated response contains reasoning error that incorrectly substitute the total cost as “5,364,765 * 5”. However, since FacTool only checks math calculation errors, FacTool powered by GPT-4 did not identify the reasoning error. Detailed examples can be found in Fig. 17 of Appendix B.
There are two major types of errors in factuality detection for scientific literature review: (1) Errors in title matching: Title matching can sometimes be problematic due to abbreviations in the generated citations or the retrieved title. For example, although the paper “MDMA-assisted psychotherapy for treatment of PTSD: study design and rationale for phase 3 trials based on pooled analysis of six phase 2 randomized controlled trials exists, FacTool powered by GPT-4 identify the paper title as incorrect. (2) Errors in author matching: the author matching process might sometimes not be robust. For example, although the authors of “Language Models are Unsupervised Multitask Learners" are indeed “Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever, FacTool powered by GPT-4 identify the author list as incorrect. Detailed examples can be found in Fig. 18 of Appendix B.
3 Exp-III: Using FacTool to Evaluate the Factuality of Modern Chatbots
The purpose of developing a factuality detector is to audit the actual generative chatbots to assess the reliability of the responses generated by chatbots. To this end, we evaluate the factuality of modern chatbots, including GPT-4, ChatGPT, Claude-v1, Bard, and Vicuna-13B, using FacTool powered by GPT-4. It is important to note that in Exp-III, we consider FacTool as a golden evaluator, responsible for evaluating the factual accuracy of the responses generated by different chatbots. For prompts selection, we follow the same intuition as Zhou et al. (2023): KB-QA is the most common scenario. Thus, we select 30 KB-QA prompts, 10 code prompts, 10 math prompts. and 10 scientific prompts (i.e., 3 times more KB-QA prompts compare to prompts from other scenarios) to carry out this factuality evaluation on chatbots. The KB-QA prompts are collected from Zhou et al. (2023), code prompts from HumanEval Chen et al. (2021), math prompts from Gao et al. (2022b), while the scientific prompts are generated by us. Responses for these prompts are generated by each of the evaluated chatbots.
We report both the claim-level and response-level accuracies for each chatbot, as evaluated by FacTool powered by GPT-4. Given that KB-QA responses contain significantly more claims than responses from other scenarios, we report the weighted claim-level accuracy. This weight is determined by the ratio of the number of prompts in each scenario. In other words,
Adopting the weighted-claim level accuracy evaluation helps us provide a more holistic and fair assessment of each chatbot’s factual accuracy.
Tab. 6 shows that GPT-4 has the best weighted claim-level factual accuracy and response-level accuracy compared to ChatGPT, Bard, Claude-v1, and Vicuna. Fig. 5 and 5 demonstrate fine-grained performance w.r.t each scenario (KB-QA, code, math, scientific). We observe that (a) GPT-4 has the best claim-level accuracy and response-level accuracy in most of the scenarios. (b) Supervised fine-tuned Chatbots like Vicuna-13B perform reasonably well in more common scenarios like KB-QA but less so in more challenging scenarios such as math, code, and scientific.
Conclusion
We introduce FacTool, a task- and domain-agnostic framework designed to tackle the escalating challenge of factual error detection in generative AI. We expand the conventional definition of factuality, particularly focusing on auditing the capabilities of generative AI models. Realizing that (1) the generated texts of LLMs tend to be lengthy and lack a clearly defined granularity for individual facts, and that (2) there is a scarcity of explicit evidence available during the process of fact checking, we build FacTool as a 5-step tool-augmented framework that consists of claim extraction, query generation, tool querying, evidence collection, and verification.
We demonstrate the potential of incorporating tools like Google Search, Google Scholar, code interpreters, Python, and even LLMs themselves in factual error detection through experimentation across diverse tasks such as knowledge-based QA, code generation, math problem solving, and scientific literature review writing. We believe that our holistic and adaptable framework can be easily extended to more scenarios.
Acknowledgements
We thank Yixin Liu, Zhengbao Jiang, Zhiruo Wang for the useful discussion and suggestions.
References
Appendix A Prompts
We list the claim extraction, query generation, and agreement verification prompts used in this paper. All the prompts listed are user prompts. We use the same system prompt “You are a brilliant assistant.”
Appendix B Example cases of FacTool
We list the example cases of FacTool in each scenario.