Do Large Language Models Know about Facts?
Xuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo, Lijie Wen, Philip S. Yu, Zhijiang Guo
Introduction
Large language models (LLMs) have revolutionized natural language processing (NLP) in recent years since they have significantly improved performance on various downstream tasks (Brown et al. 2020; Chowdhery et al. 2022; Ouyang et al. 2022; Touvron et al. 2023a; Touvron et al. 2023b; OpenAI 2022; OpenAI 2023). Prior efforts have shown that language models can store factual knowledge and act as knowledge bases (Petroni et al. 2019; Jiang et al. 2020b). Factual knowledge in language models acquired during pretraining can benefit knowledge-intensive downstream tasks such as question answering and fact checking (Roberts et al. 2020; Yu et al. 2023; Pan et al. 2023).
Despite advancements in LLMs, they still struggle with generating content that exhibits inaccuracies or deviations from the facts and making reasoning errors (Lin et al. 2022; Bubeck et al. 2023). These factual errors can be difficult to identify since LLMs implicitly memorize facts through their parameters rather than explicitly store factual knowledge as traditional Knowledge Bases. Accessing and interpreting the computations and memories of these models can be challenging (Ribeiro et al. 2016; Belinkov & Glass 2019), especially when APIs are the only means of interaction and many interpretation methods rely on weights and representations (Cao et al. 2020). The presence of errors in stored factual knowledge or the incorrect induction and obsolescence of certain facts over time may be contributing factors to this limitation, which in turn affects the performance of LLMs (Elazar et al. 2021; Cao et al. 2021). This limitation restricts the application of LLMs in some high-stakes areas, such as healthcare, finance, and law (Dong et al. 2022). Hence, exploring the degree to which LLMs hold factual information and their ability to reason with such knowledge is vital.
To this end, we propose the Pinocchio benchmark, a comprehensive testbed of factuality and reasoning designed for LLMs. Pinocchio contains 20K diverse factual questions that span different sources, timelines, domains, regions, and languages. Furthermore, we investigate whether LLMs are able to recognize the combination of multiple facts, reason over structured and unstructured evidence, realize facts change over time, identify subtle factual differences, and resist adversarial examples. We control for problem difficulty in each distinct reasoning task to enable fine-grained analysis.
With the Pinocchio benchmark, we evaluate whether various LLMs (Scao et al. 2022b; Zhang et al. 2022; Ouyang et al. 2022; Chung et al. 2022; Touvron et al. 2023a; Chiang et al. 2023) could store factual knowledge and perform reasoning based on it. We envision Pinocchio as a suite of benchmarks, subsets of which could be separately utilized to assess certain model abilities of interest and analyze important strengths and limitations of LLMs. For instance, in temporal tasks, we find that LLMs lack factual knowledge for up-to-date questions; in complex factual tasks that require multi-hop reasoning, LLMs still have limitations, even when various prompting strategies are employed. We hope Pinocchio can guide the researchers to understand the abilities of their models from multiple dimensions and facilitate the development of factual knowledge in LLMs.
Dataset Construction
Aiming to systematically evaluate the factual knowledge and related reasoning abilities of LLMs, we raise seven research questions, then carefully select factual statements from different sources summarized in Table 1.
Task 1: Multifacted Previous research (Petroni et al. 2019) has shown that small language models like BERT have the ability to retain relational knowledge from training data and answer “fill-in-the-blank” cloze statements. This raises the question of whether LLMs can also store and reason over multiple pieces of facts obtained during pretraining. It is not just important for LLMs to memorize individual facts accurately, but to also recognize and generate new combinations of facts from different sources. To investigate this issue, we have selected claims from the FEVER dataset (Thorne et al. 2018), which were written by human annotators based on information from Wikipedia articles. These claims are either supported or refuted by multiple facts from (the same or several) Wikipedia articles, or there is insufficient information available to verify them. To assess the performance of language models in handling various combinations of facts, we have sampled statements that require different numbers of evidence, ranging from one to many, enabling fine-grained analysis.
Task 2: Structural In addition to unstructured text, factual knowledge is also commonly stored in a structured format, such as tables, lists, or databases (Bhagavatula et al. 2013). However, current LLMs are primarily trained on unstructured text using next word prediction loss (Brown et al. 2020; Touvron et al. 2023a). In order to process structured data, it is often converted into text strings using various methods, such as linearizing tables. This raises the question of whether LLMs are capable of effectively memorizing and reasoning over facts from structured sources, similar to their performance with unstructured text. To investigate this question, we sample factual statements from the FEVEROUS dataset (Aly et al. 2021), which is constructed in a similar manner to FEVER but includes evidence in the form of tables, sentences, or both.
Task 3: Adversarial Language models are known to be vulnerable to adversarial examples that are strategically modified to deceive even advanced models with hardly noticeable changes (Shen et al. 2023). Given this knowledge, it is important to examine whether LLMs can withstand adversarial examples in the context of factuality. To investigate this, we utilize two datasets, namely Symmetric (Schuster et al. 2019) and FM2 (Eisenschlos et al. 2021). These datasets consist of adversarial examples that have been crafted using various strategies, including temporal inference and diverting to unrelated facts.
Task 4: Temporal Facts are not static but rather possess a dynamic nature. With the vast amount of new information constantly emerging, facts often undergo changes, additions, or alterations. It raises the question of whether LLMs are able to adapt to these factual changes over time. In particular, we wonder if LLMs are capable of discerning factual knowledge from different time periods, since the pretraining corpus may not be processed and organized chronologically. To explore this, we utilize the VitaminC (Schuster et al. 2021) dataset, which consists of claims based on modifications made to factual content in Wikipedia articles. Claims can be either refuted by outdated facts or supported by updated facts.
Task 5: Real-World In contrast to other tasks that assume Wikipedia has all the essential factual information, verifying viral claims on the internet often requires not only factual knowledge from various sources but also common sense and worldly knowledge. An important query we have is whether LLMs can effectively integrate diverse types and sources of knowledge acquired during training. To address this, we select claims from the FactCheck (Misra 2022) dataset, which consists of claims spread over the Internet and subsequently verified by journalists.
Task 6: Domain-Specific In addition to the tasks mentioned earlier, which primarily focus on factual knowledge in general domains, we are also interested in exploring how LLMs possess the capability to access domain-specific factual knowledge. The domain-specific setting presents unique challenges. Take the science domain as an example, LLMs need to acquire background knowledge, handle quantitative reasoning, and comprehend specialized statistical language. To investigate this further, we sample claims from PubHealth (Kotonya & Toni 2020) in the public health domain and SciFact (Wadden et al. 2022) in the science domain.
Task 7: Multi-Lingual Existing LLMs are mainly trained on English corpus because of their abundance and quality (Chowdhery et al. 2022; Touvron et al. 2023a). However, the scarcity of training data in other languages raises the question of whether LLMs can transfer the factual knowledge acquired in English to other languages. To investigate this, we collected claims from various languages including French, Chinese, and more, using the XFACT dataset (Gupta & Srikumar 2021) and the CHEF dataset (Hu et al. 2022) in a total of 27 different languages.
2 Annotation and Quality Control
We only select questions of a multi-choice format, similar to other benchmarks (Hendrycks et al. 2021b; Zhong et al. 2023), because metrics are clearly defined (i.e. accuracy), and multi-choice questions are a simple but good proxy to evaluate the potential of advanced abilities of LLMs, which we consider could be easily exploited and reflected in various downstream applications through specialized instruction tuning. Each question has four choices and only one choice is the correct answer. LLMs are intended to be used to solve these questions through prompting.
Specifically, we hired 10 undergraduate students, all with good English proficiency. We asked the students to rewrite the original claims into questions without distorting factuality while providing factuality labels for the questions. By transforming declarative statements into questions, using a Question-Answering approach can more effectively elicit factual knowledge from LLMs (Kadavath et al. 2022; Lin et al. 2022), and we also illustrate through experiments in Sec. 4.2. Note that claims in the original datasets are usually labeled based on given evidence, e.g. evidence supports or refutes the claim, but in Pinocchio, we only need to judge the factuality of the question. So we use unified labels: Yes, No, Not Sure Enough. The three labels correspond respectively to Factual, Non-Factual, and Not Enough Information for factual questions. Considering that all fact-checking datasets use a three-label system (Guo et al. 2022), we did not modify the number of labels to maintain consistency in labeling. When dealing with factuality questions in low-resource languages, for Chinese, the 5 undergraduate students we hired are native Chinese speakers. For other low-resource languages, we first use Google Translate to translate them into English and generate factuality questions, then translate the English questions back to the corresponding languages. The label distribution is shown in Table 1. We paid the annotators accordingly based on the quantity and quality of the annotations.
We ensure the quality of the annotated factuality questions in two ways. The two authors of this paper served as meta-reviewers, sampling 10 questions from each of the three categories across the seven domains in Pinocchio. The meta-reviewers judged if the factuality labels were correct. For the 210 factuality questions, the average label accuracy was 92.4%. We divided the 10 students into two groups and had each group re-annotate a random 200 questions annotated by the other group, then calculated inter-annotator agreement (IAA). The final IAA was 85.6%. Based on meta-reviewer results and IAA, the factuality labels in Pinocchio are of good quality.
Methodology
To give a comprehensive view of the status of large language models in a factual context, we evaluate 10 accessible LLMs, undergone different training stages including pretraining, instruction tuning, and reinforcement learning from human feedback (Ouyang et al. 2022), covering diverse organizations and varying in size, as shown in Table 2.
For pretraining LLMs, we adopt OPT (Zhang et al. 2022), BLOOM (Scao et al. 2022a), and LLaMA (Touvron et al. 2023a). For instruction-tuned LLMs, we adopt Alpaca (StanfordCRFM 2023), Vicuna (Chiang et al. 2023), Flan -T5 (Chung et al. 2022), and ChatGLM (Zeng et al. 2023). After undergoing pretraining, instruction tuning, and RLHF, ChatGPT (OpenAI 2022) is also taken into consideration. A detailed description of these models can be found in Appendix A.2.
2 Prompt Strategy
As illustrated in Figure 2, we employ 4 types of prompts to elicit desired responses from LLMs, namely: Zero-shot, Zero-shot with CoT (Kojima et al. 2022), Few-shot, and Few-shot with CoT (Wei et al. 2022). Specifically, we begin by providing the model with task instruction, denoted as : “You will be given a question. You should answer whether it is Yes, No, or Not Sure Enough and show your evidence”. This instruction informs the LLMs about the expected input and output. Subsequently, for any given input , we anticipate obtaining an output label from the LLMs : .
In the zero-shot setting, the LLMs are expected to provide answers based on the Question and the task instruction . We anticipate that the LLMs can directly generate the factual answer “No” when presented with : “Has gas prices gone up 99 percent since Obama became president, making it the highest gas price increase since Carter?” The zero-shot with CoT setting extends the question by adding a two-stage prompt (Kojima et al. 2022): “Let’s think step by step”, designed to encourage the LLMs to contemplate the process of determining the factual label .
Few-Shot Prompt
In the few-shot setting, we employ three prompts: Yes, No, and Not Sure Enough, as query questions () for model input. Due to space constraints, detailed examples of the prompts in Figure 2 are presented in Appendix 5. The utilization of the few-shot setup allows us to better tap into the inherent factual reasoning abilities of LLMs (Chung et al. 2022; OpenAI 2022; Wang et al. 2022; StanfordCRFM 2023). This is particularly advantageous for those models that have not been fine-tuned with specific instructions, as their factual reasoning capabilities can be showcased through effective few-shot guidance. In contrast, they may struggle to adhere to instructions in zero-shot evaluations, thereby impacting their ultimate factual performance.
In the few-shot with CoT setting, we provide potential reasoning instructions to the LLMs before presenting the factual label (). The aim is to elicit the LLMs’ innate factual reasoning abilities through examples of reasoning. As shown in Figure 2, for the : “Is there a capital called Mogadish?” Our reasoning approach entails first explaining the noun phrase in the (the subject and object), and subsequently elaborating on modifying phrases such as predicates or adjectives. Regarding the subject “Mogadish”, we begin by furnishing a detailed definition: “Mogadishu is a city in East Africa, specifically in Somalia.” Following this, we proceed to reason about the relation between “Mogadish” and “capital”: “Furthermore, the capital of Somalia is indeed Mogadishu.” Consequently, we arrive at the ultimate factual label: “Therefore, the answer is Yes.” We anticipate that the reasoning instructions provided to the LLMs will serve to stimulate its factual reasoning abilities.
Experiments
In the previous sections, we provided a detailed description of how Pinocchio was constructed and the LLMs used. In this section, we will begin by introducing the performance of various LLMs on Pinocchio across different settings and tasks, along with a detailed analysis.
In Table 2, we present the average results of 10 accessible LLMs operating under varying settings on Pinocchio, run three times each. From Table 2, we draw the following conclusions:
Regarding overall performance, we observe that, on average, LLMs without instruction tuning underperform those with instruction tuning by 16.0%. GPT family LLMs undergoing RLHF exhibit superior results, indicating that instruction tuning and RLHF optimize alignment with human knowledge, thereby improving factual question response accuracy.
Results obtained using the Few-shot setting significantly outperform those obtained when simply asking factual questions to LLMs in the Zero-shot setting, especially for models without RLHF, exhibiting an average improvement of 7.3%. This highlights the capability of some sample prompts to better extract the inherent factual knowledge of LLMs.
Using the CoT method, we observed a relative boost in performance in LLMs subjected to instruction tuning and RLHF, improving by an average of 2.1%. Notably, the factual accuracy of LLMs like OPT, BLOOM, and LLaMA was mostly stable or even decreased. A review of outputs from these untuned LLMs revealed that, post-CoT application, LLMs tend to produce related content considerations, and extensive considerations often overshadow factual discernment tasks, causing incorrect factual label outputs. In contrast, for instruction-tuned LLMs, the CoT method facilitates enhanced exploration of factual entity relations in questions, resulting in accurate factual labels. See Appendix A.5 for detailed case analyses.
The OPT model, without being tuned to instructions, struggles significantly to output correct factual labels under the settings of Zero-shot and Zero-shot CoT, often resulting in either a repetition of the original question or a refusal to output any content at all. This issue is somewhat alleviated under the settings of Few-shot and Few-shot CoT.
Additionally, we studied the hyperparameters of LLMs. Due to limited computing resources, we only explored Vicuna-7B and Vicuna-13B. We found that as model parameters increase, performance on factual questions improves correspondingly, with an average increase of 5.4%. This indicates that LLMs with more parameters can store more world knowledge and have stronger factual knowledge recognition capabilities.
In Table 3, we present the factual performance of LLMs in various tasks under the Few-shot CoT setting. This reveals the relative difficulty LLMs have in understanding and responding to factual questions in different tasks, providing insights for future training of factual knowledge in LLMs. From Table 3, it is observed that LLMs exhibit relatively poorer performance on factual questions related to the real-world, domain-specific knowledge, and multilingualism, being on average 6.4% lower compared to the other four tasks. This is attributed to the fact that the training data for LLMs typically come from general domains and are not up-to-date, which indirectly inspires the exploration of retrieval-augmented LLMs (Ram et al. 2023). We analyze the LLMs in different tasks in Sec. 4.2.
2 Analysis
In this section, we explore LLMs’ capabilities focusing on key areas like handling of multi-hop factual questions, proficiency in diverse prompt strategies, and tackling challenges like numerical reasoning and entity ambiguity. We also examine their performance on time-sensitive factual questions, against adversarial attacks, with fine-grained labels and prompts in multiple languages.
To analyze the performance of LLMs when faced with factual questions based on multiple pieces of facts that require complex logical reasoning, we categorize multifaced and structural factual questions into distinct subsets, depending on the number of “hops” necessary to validate each factual question. To maintain fairness, we randomly sampled 1,490 data pieces from each of the two datasets for verification. Figure 3(a) illustrates the data counts and Macro F1 scores of GPT-3.5-Turbo for each respective subset. The figure reveals a clear pattern: as the number of “hops” increases, the reasoning chain for deriving conclusions from existing factual knowledge extends, necessitating heightened logical reasoning capabilities from the LLMs. Consequently, the performance of the LLMs exhibits diminishing trends.
Structural Knowledge Analysis in LLMs
To investigate whether LLMs can effectively memorize factual knowledge from structured data, we divided the structural task questions into three subsets according to evidence distribution: evidence in unstructured data (Only text), structured data (Only tables), or both (Combine text and tables). Figure 3(b) shows a notable decline (Avg. -5.5%) in GPT-3.5-Turbo’s performance when evidence involves structured data, indicating LLMs’ limited ability in extracting knowledge from structured tables. The LLMs also perform less effectively when handling questions requiring the combination of both evidence types, reflecting their incapacity to integrate diverse structured evidence effectively.
Analysis of Different Factual Questions Poses Challenges
To assess the capabilities of LLMs in addressing various challenges, we partitioned each factual question within the structural task into six distinct challenges: 1) Entity disambiguation, 2) Other, 3) Multi-hop reasoning, 4) Combining tables and text, 5) Search terms not in claim, 6) Numerical reasoning, each centered around the most critical difficulty encountered during verification. Figure 3(c) illustrates GPT-3.5-Turbo’s performance and data distribution across challenges. The extensive training and large-scale parameters enhance LLMs’ performance in handling entity ambiguity. Longer reasoning chains and various forms of evidence challenge LLMs’ factual abilities. When correct inference involves unmentioned entities, LLMs may lack necessary hints from factual questions, posing significant challenges. LLMs also exhibit deficiencies in precise numerical calculations due to the inherent hallucination phenomenon, resulting in subpar performance when numerical reasoning is needed for verification.
Temporal Analysis
As time progresses, the truthfulness of certain questions may undergo changes. The temporal task encompasses such data, and we leverage this task to explore the ability of LLMs to adapt to factual changes over time. Figure 4(a) illustrates that GPT-3.5-Turbo exhibits superior performance when dealing with outdated data as compared to updated data. This discrepancy arises from the fact that LLMs are pretrained on a corpus of text prior to a specific temporal point. Consequently, LLMs lack the capability to acquire real-time, up-to-date knowledge, rendering them unable to validate questions that hinge on the most recent information for accurate assessments.
Adversarial Analysis
To evaluate the robustness of LLMs to adversarial attacks, we divide the adversarial questions into three subsets: auto-generated questions from the corpus, manually modified synthesized questions yielding adversarial ones, and artificially created adversarial questions.
Figure 3(b) presents the performance of GPT-3.5-Turbo on these three subsets. It is evident that following adversarial attacks, LLMs exhibit a substantial decrease in performance. Furthermore, factual questions that have undergone manual modifications or were artificially created prove to be more challenging compared to those that are automatically generated (Shen et al. 2023). This disparity could be attributed to the fact that automatically synthesized factual questions often contain explicit positive or negative words that hint at the outcome, and the exceptional comprehension abilities of LLMs enable them to accurately discern and provide the correct response in such cases.
Label Granularity Analysis
To assess the effect of different label granularities on LLMs’ performance, we conducted a manual re-labeling of the real-world task questions. Per the settings of Misra 2022, besides labeling as “Factual”, “Non-Factual”, and “Not Enough Information”, we also require them to annotate the dataset with six factual labels: “Factual”, “Mostly Factual”, “Mostly False”, “Non-Factual”, “Pants-Fire”, and “Not Enough Information”. We also modified the prompt for GPT-3.5-Turbo for more intricate factual responses to test its competency with nuanced labels. Results in Figure 4(c) disclosed: 1) The results show that, in general, there is a significant decrease in performance (-23.83%) when transitioning from coarse-grained justification to fine-grained justification. With finer granularity, LLMs are not only required to assess the authenticity of each question but also to judiciously employ their knowledge base to precisely gauge the credibility of each factual questions. 2) When comparing the performance of coarse-grained labels with fine-grained labels, we observe significant drops in the three categories: “Factual” by 13.3%, “Non-Factual” by 23.2%, and “Not Enough Information” by 22.3%. This indicates that finer-grained labels introduce additional options that can potentially disrupt the original judgment of the LLMs. A potential remedy could be the aggregation of multiple judgments through voting (Wang et al. 2023a).
Multilingual Task with Chinese and English Prompts
To investigate the influence of prompts in different languages on LLMs, we extracted Chinese factual questions from the multilingual tasks to create a subset. We then evaluated the LLMs’ performance when using both Chinese and English prompts, both of which are depicted in Appendix A.4. Table 4 illustrates the results, indicating that the LLMs perform better when using a Chinese prompt. This underscores the notion that employing prompts in the same language as the questions can enhance the transfer capabilities from English factual knowledge to other languages of LLMs.
Prompt Strategy Analysis
In prior research, various CoT methods have been employed to enhance the performance of LLMs. These methods include 1) augmenting the number of in-context learning examples, 2) implementing self-consistency mechanisms, which alleviates the hallucination phenomenon through majority voting after multiple judgments of LLMs (Wang et al. 2023a), 3) incorporating complex reasoning chains, which leverages the most complex CoT in prompt to steer the cognitive processes of LLMs and augment their cognitive capabilities (Fu et al. 2022), and 4) employing self-refinement strategies, which refines LLMs’ answers through continuous feedback of another LLM on responses to achieve better results (Madaan et al. 2023) and so forth. Additionally, we examined the influence of utilizing declarative claims as instances of in-context learning. We randomly sampled 200 factual questions from each task of the Pinocchio, totaling 1400 questions, to compose Pinocchio-Lite with the aim of speeding up the testing of different prompt strategies. The performance results of various CoT methods are presented in Table 5. To maintain fairness, three in-context learning examples are employed in the complex chain, self-consistency, self-refinement, and declarative claim methods. Different types of CoT prompts are shown in Appendix A.4.
It is worth noting that 1) when the number of in-context learning examples is limited, the incremental improvement in performance is marginal upon increasing the number of examples. However, beyond a specific threshold, the addition of more examples gains more performance improvement. This could be due to the inability of LLMs to fully encapsulate the correct reasoning with fewer examples. 2) Concurrently, a fascinating observation is that the LLM’s performance substantially deteriorates as the complexity of the CoT increases. This could stem from the difficulty LLMs have in extracting a generalized reasoning pattern from complex, multi-stage thinking processes with limited examples. 3) The self-consistency method markedly boosts performance by mitigating the hallucination issue in LLMs through consistency voting, enhancing their response accuracy. 4) In the self-refinement approach, the model might initially provide an incorrect response, but it can amend its mistakes through feedback and refine its answers. In the end, when no additional refinement is needed, the model often reaches the correct conclusion, achieving optimal performance. 5) Compared to the 3 shots method, the declarative claims method saw a 2.3% performance drop, illustrating that using questions as examples effectively directs LLMs in acquiring factual knowledge.
Related Work
Previous research has demonstrated that LLMs have the ability to retain and utilize factual knowledge, effectively acting as knowledge bases (Petroni et al. 2019; Petroni et al. 2020; Heinzerling & Inui 2021). This acquired factual knowledge in language models during pretraining can be advantageous for knowledge-intensive tasks like question answering and fact checking (Roberts et al. 2020; Yu et al. 2023; Pan et al. 2023). To evaluate the factual knowledge stored in language models, Petroni et al. 2019 employed cloze tests consisting of triples and prompts specifically designed to simulate missing objects. Jiang et al. 2020a explored the role of prompts in retrieving factual information from language models and devised improved prompts for probing. However, Elazar et al. 2021 demonstrated the unreliability of rank-based probing methods with paraphrased context, leading to inconsistent findings. Cao et al. 2021 contended that biased prompts and leakage of golden answers often lead to overestimations of LLMs’ knowledge storage capability. In contrast, Varshney et al. 2022 used question answering to measure models’ uncertainty regarding specific facts. Our method is more in line with Kadavath et al. 2022 and Lin et al. 2022, employing self-evaluation by querying the models to assess response accuracy regarding factual knowledge.
Benchmarks for Large Language Models
The advent of LLMs has underscored the importance of exhaustive benchmarks for effective capability assessment. Presently, there are predominantly two types of existing benchmarks. One evaluates the general knowledge and reasoning capacities of LLMs, exemplified by the MMLU benchmark (Hendrycks et al. 2021a), a multi-task evaluative measure encompassing tasks from real-world tests and literature, spanning diverse subjects like elementary math, US history, computer science, and law. Moreover, benchmarks also exist for non-English languages (Huang et al. 2023) or in a bilingual context (Zhong et al. 2023). BIG-bench (Srivastava et al. 2022) is a collaborative benchmark examining LLMs’ capabilities across 204 diverse tasks from various fields like linguistics, childhood development, software development, and more. HELM (Liang et al. 2022) employs 7 metrics over 42 tasks to assess LLMs, focusing on aspects from accuracy to robustness. Specific benchmarks like GSM8K (Cobbe et al. 2021) and MATH (Hendrycks et al. 2021a) target mathematical problem-solving, presenting elementary to competition-level problems. In program synthesis, HumanEval (Chen et al. 2021) and MBPP (Austin et al. 2021) evaluate functional correctness through program synthesis from docstrings. Additional benchmarks address instruction following (Dubois et al. 2023), tool usage (Xu et al. 2023), and decision making (Liu et al. 2023). Our benchmark mainly evaluates factual knowledge, differing from ones like TruthfulQA (Lin et al. 2022), which specifically tests truthfulness in LLMs’ generated responses, with questions structured to provoke imitative falsehoods over truthful answers.
Conclusion
In this work, we investigate whether LLMs are capable of memorizing factual knowledge and reasoning based on it, across various problem categories and prompting strategies. To this end, we curate the Pinocchio benchmark, a comprehensive test bed with 20,713 questions covering seven tasks with varying complexity. By evaluating LLMs and prompting approaches on the Pinocchio benchmark, we find that different types of LLMs employing various prompting strategies, such as multi-shots and self-consistency, still perform suboptimally on factual tasks. Improving LLMs’ factual knowledge and reasoning abilities on complex and nuanced NLP tasks remains an open research question, and we encourage future work to develop upon our proposed Pinocchio benchmark.
References
Appendix A Appendix
Pinocchio primarily serves to assess LLMs’ responses to questions concerning factual knowledge. If a model performs effectively, it would be imprudent to infer that its reliability will uniformly translate to diverse task domains (even if some degree of transfer learning is anticipated). For instance, Pinocchio does not encompass long-form generation, such as news articles, or interactive settings, such as extended dialogues with adversarial entities. Furthermore, although the questions within Pinocchio parallel real-world inquiries, they originate not from a deployed system, thus posing a potential risk of over- or under-estimating the factuality of such a system.
We postulate that Pinocchio is unlikely to prove advantageous for those intending to fabricate deceptive models with malicious intent. To effectuate deception, a model must generate erroneous responses relatively infrequently, lest humans swiftly discern its unreliability. However, acquiring a low score on Pinocchio necessitates the provision of incorrect answers to virtually all questions. To be instrumental for malevolent purposes, a model must generate highly specific false statements, such as assertions concerning a maliciously targeted victim or a particular governmental policy. Yet, Pinocchio lacks coverage of highly specific subjects, offering instead a superficial overview of general factual topics.
While Wikipedia and some news websites are exemplary collaborative resources, they inherently contain inaccuracies and noise, akin to any encyclopedia or knowledge repository. Consequently, we advise users of Pinocchio against making absolute assertions about the validated claims and discourage its utilization for the development of truth-revealing models. We refrained from collecting participants’ personal data in any form. Participants accessed our online tool exclusively using an identification number. Generated assertions must solely incorporate information deemed as general world knowledge or sourced from Wikipedia, thereby excluding any personally identifiable information or offensive content.
A.2 The detailed introduction to the LLMs
For pretraining models, OPT (Zhang et al. 2022) is an open-sourced large causal language model which perform similar in performance to GPT-3 (Brown et al. 2020). BLOOM (Scao et al. 2022a) is an open-access multilingual large language model that is suitable for non-English facts. LLaMA (Touvron et al. 2023a) is probably the best open-weight foundation model so far that achieves the highest accuracy on various English benchmarks (e.g. MMLU (Hendrycks et al. 2021a)) within open-weight models. For instruction-tuned models, Alpaca (StanfordCRFM 2023) is fine-tuned from the LLaMA model on 52K self-instructed demonstrations (Wang et al. 2023b). Alpaca behaves qualitatively similarly to OpenAI’s Text-Davinci-003 on evaluation of single-turn instruction following. Vicuna is an open-source chatbot trained by fine-tuning LLaMA on user-shared conversations collected from ShareGPT (ShareGPT 2023). Flan -T5 (Chung et al. 2022) is an enhanced version of T5 that has been instruction fine-tuned in a mixture of tasks. ChatGLM is an open bilingual language model based on the General Language Model (Zeng et al. 2023). ChatGLM is trained on Chinese and English corpus, supplemented by instruction tuning, feedback bootstrap, and reinforcement learning with human feedback (RLHF; Ouyang et al. 2022). ChatGPT (OpenAI 2022) from OpenAI that has undergone pretraining, instruction tuning, and RLHF. ChatGPT has been observed to have impressive capabilities in various aspects favoring reasoning capabilities (Qin et al. 2023).
A.3 Task Results
In this section, we present the results of all LLMs across different tasks under three different settings: Zero-shot w/o CoT, Zero-shot w/ CoT, and Few-shot w/o CoT.
A.4 Prompt Strategy
In this section, we provide the comprehensive versions of all the prompts utilized in both the main experiments and the subsequent analysis. We engaged native Chinese annotators to rephrase the English prompts while maintaining their semantic integrity, thus yielding Chinese prompts.
A.5 Case Study
We have introduced an additional scenario for investigation, which occurs frequently in the output generated by the zero-shot prompt method. We conducted an experiment involving three models: OPT, ChatGLM, and GPT-3.5-Turbo. These models are presented with the same set of questions, and their responses are shown in Figure 8. It is noteworthy that the OPT model, in both questions, reiterated the question itself without providing the corresponding answer. It is essential to mention that the actual output of the OPT model repeats the problem until it reaches the maximum output length (controlled by the "max_length" parameter), and we truncated the repeated portion.
The OPT model even declined to generate any content when presented with the zero-shot prompt, resulting in a significant number of empty responses in the statistical results. In the first question, both ChatGLM and GPT-3.5-Turbo provided correct answers. However, in the second question, when faced with more detailed information inquiries, ChatGLM failed to produce a correct response, while GPT-3.5-Turbo demonstrated proficient reasoning and provided accurate answers.