Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity

Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru Tang, Tianhang Zhang, Cheng Jiayang, Yunzhi Yao, Wenyang Gao, Xuming Hu, Zehan Qi, Yidong Wang, Linyi Yang, Jindong Wang, Xing Xie, Zheng Zhang, Yue Zhang

Introduction

The quest for mastery of knowledge has been a foundational aspiration in the development of artificial intelligence systems. Historically, seminal works by McCarthy et al. 1955 and Newell and Simon 1976 have underscored the significance of knowledge representation and reasoning in AI systems. For instance, the Cyc project embarked on an ambitious journey to codify common-sense knowledge, aiming to provide AI systems with a comprehensive understanding of the world (Lenat 1995). Concurrently, endeavors like the WordNet project by Miller et al. 1990 sought to create lexical databases that capture semantic relationships between words, thereby aiding AI systems in grasping the nuances of human language.

Amidst these pioneering efforts, the emergence of Large Language Models (LLMs), such as ChatGPT (OpenAI 2022b), GPT-4 (OpenAI 2023) and LLaMA (Touvron et al. 2023a; Touvron et al. 2023b), has been seen as a significant leap in both academics and industries, especially towards AI systems possessing vast factual knowledge (Petroni et al. 2019a; De Cao et al. 2021; OpenAI 2023). The advantages of using LLMs as knowledge bases carriers are manifold. Firstly, they reduce the overhead and costs associated with building and maintaining dedicated knowledge bases (Petroni et al. 2019c; AlKhamissi et al. 2022; Wang et al. 2023e). Additionally, LLMs offer a more flexible approach to knowledge processing and utilization, allowing for context-aware reasoning and the ability to adapt to novel information or prompts (Sun et al. 2023a; Huang and Chang 2023). Yet, with their unparalleled capabilities, concerns have arisen about the potential of LLMs to generate non-factual or misleading content (Bender et al. 2021; OpenAI 2023; Bubeck et al. 2023). In light of these advancements and challenges, this survey seeks to delve deeply into the LLMs, exploring both their potential and the concerns surrounding their factual accuracy.

Understanding the factuality of Large Language Models is more than just a technical challenge; it’s essential for the responsible use of these tools in our daily lives. As LLMs become more integrated into services like search engines (Microsoft 2023), chatbots (OpenAI 2022b; Google 2023), and content generators (Cui et al. 2023b), the information they provide directly influences decisions, beliefs, and actions of millions of people. If an LLM provides incorrect or misleading information, it can lead to misunderstandings, spread false beliefs, or even cause harm, especially for those domains that demand high factual accuracy (Ling et al. 2023b), such as health (Thirunavukarasu et al. 2023; Tang et al. 2023), law (Huang et al. 2023c), and finance (Wu et al. 2023). For instance, a physician relying on an LLM for medical guidance might inadvertently jeopardize patient health, a corporation leveraging LLM insights might make ill-informed market decisions, or an attorney misinformed by an LLM might falter in legal proceedings (Curran et al. 2023). In addition, with the advancement of LLM-based agents, the factuality of LLMs is becoming even more potent. A driver or an autonomous driving car might rely on LLM-based agents for planning or driving, where serious factual mistakes made by LLMs could cause irreversible damage. By studying the factuality of LLMs, we aim to ensure that these models are both powerful and trustworthy.

A surge of research has been directed towards evaluating LLMs’ factuality, which encompasses diverse tasks like factoid question answering and fact checking. Beyond evaluation, efforts to improve the factual knowledge of LLMs have been notable. Strategies have ranged from retrieving information from external knowledge bases to continual pretraining and supervised finetuning. Yet, despite these burgeoning efforts, a holistic overview that covers the full spectrum of factuality in LLMs remains elusive. While there are existing surveys in the field, such as those by Chang et al. 2023 and Wang et al. 2023i, that delve into the evaluation of LLMs and their factuality, they only scratch the surface of the broader landscape. There are also a bunch of recent studies focusing on hallucinations in LLMs (Rawte et al. 2023; Zhang et al. 2023c; Rawte et al. 2023; Ye et al. 2023; Huang et al. 2023d). But we differentiate between the hallucination issue and the factuality issues in Sec 2.2. Moreover, these surveys often overlook key areas we emphasize, like domain-specific factuality or the challenge of outdated information. While Ling et al. 2023a explores domain specialization in LLMs, our survey takes a more expansive look at the broader issues of factuality. To the best of our understanding, our work is the first comprehensive study on the factuality of large language models.

This survey aims to offer an exhaustive overview of the factuality studies in LLMs, delving into four key dimensions: Sec 2) The definition and impact of the factuality issue (Pranshu Verma 2023; Nori et al. 2023); Sec 3) Techniques for evaluating factuality and its quantitative assessment (Huang et al. 2023a; Min et al. 2023); Sec 4) Analyzing the underlying mechanisms of factuality in LLMs and identifying the root causes of factual errors (Liu et al. 2023b; Kotha et al. 2023); and Sec 5) Approaches to enhance the factuality of LLMs (Du et al. 2023; He et al. 2022). Notably, we categorize the use of LLMs into two primary settings: LLMs without external knowledge, such as ChatGPT (OpenAI 2022b) and Retrieval-Augmented LLMs, such as BingChat (Microsoft 2023). The complete structure of this survey is illustrated in Figure 1. Through a detailed examination of existing research, we seek to shed light on this critical aspect of LLMs, helping researchers, developers, and users harness the power of these models responsibly and effectively.

Factuality Issue

In this section, we describe the issue of factuality in large language models, as well as the impact.

There is no well-accepted and exact definition of large language models in the literature (Chang et al. 2023; Zhao et al. 2023b; Huang and Chang 2022). We mainly consider the decoder-only generative pre-trained language modes with emergent abilities, such as ChatGPT (OpenAI 2022b) and LLaMA (Touvron et al. 2023a; Touvron et al. 2023b). We also include some work that is based on models with encoder-decoder architectures, such as T5 (Raffel et al. 2020a). We do not talk about work that is only based on the Encoder-only models, such as BERT (Devlin et al. 2019) and RoBERTa (Liu et al. 2019), in this survey. To be specific, our survey includes the following LLMs:

General Domain LLMs: GPT-2 (Radford et al. 2019), GPT-3 (Brown et al. 2020), ChatGPT (OpenAI 2022b), GPT-4 (OpenAI 2023), GPT-Neo (Black et al. 2021), OPT (Zhang et al. 2022b), LLaMA (Touvron et al. 2023a), LLaMA-2 (Touvron et al. 2023b), Incite (Computer 2023; Together 2023), Claude (Claude 2023), Falcon (Almazrouei et al. 2023), MPT (Team 2023b), Vicuna (Chiang et al. 2023), FLAN-T5 (Chung et al. 2022), BLOOM (Scao et al. 2022), Baichuan & Baichuan2 (Yang et al. 2023b), PaLM (Chowdhery et al. 2022), Gopher (et al. [n. d.]), Megatron-LM (Shoeybi et al. 2019), SAIL (Luo et al. 2023a), Codex (Chen et al. 2021b), Bard (Google 2023), GLM & ChatGLM (Zeng et al. 2022), InternLM (Team 2023a), StableBeluga (Mahan et al. [n. d.]), Claude (Claude 2023), Alpaca (Taori et al. 2023), New Bing (Microsoft 2023), Ziya-LLaMA (Zhang et al. 2022a), BLOOMZ (Muennighoff et al. 2022), Chinese-LLaMA (Cui et al. 2023c), Phoenix (Chen et al. 2023c), and others.

Domain-specify LLMs: BloombergGPT (Wu et al. 2023), EcomGPT (Li et al. 2023c), BioGPT (Luo et al. 2022), LawGPT (Nguyen 2023), Lawyer LLaMA (Huang et al. 2023c), ChatLaw (Cui et al. 2023b), BioMedLM (Venigalla et al. 2022), HuatuoGPT (Zhang et al. 2023a), ChatDoctor (Li et al. 2023b), MedicalGPT (Xu 2023), Bentsao (Huatuo as its original name) (Wang et al. 2023d), Zhongjing (Yang et al. 2023c), LLM-AMT (Wang et al. 2023f), DISC-MedLLM (Bao et al. 2023), Cohortgpt (Guan et al. 2023), Deid-gpt (Liu et al. 2023e), Doctorglm (Xiong et al. 2023b), MedChatZH (Tan et al. 2023a), K2 (Deng et al. 2023), HouYi (Bai et al. 2023), GrammarGPT (Fan et al. 2023), FoodGPT (Qi et al. 2023), ChatHome (Wen et al. 2023), and others.

2. Factuality

By factuality in LLMs we refer to the capability of large language models for generating contents that follow factual information, which encompasses commonsense, world knowledge and domain facts. The factual information can be grounded to reliable sources, such as dictionaries, Wikipedia or textbooks from different domains We only consider cases where the meaning is clear and the truthfulness can be determined in this survey. Furthermore, we only take into account undisputed facts. Minor errors that may exist in reliable sources are not within the scope of our consideration. . A series of work have discussed whether LLMs can serve as knowledge bases to store factual knowledge (AlKhamissi et al. 2022; Yu et al. 2023; Pan et al. 2023a)

Existing work focus on measuring factuality in LLMs qualitatively (Lin et al. 2022b; Chern et al. 2023), discussing the mechanism for storing knowledge (Meng et al. 2022; Chen et al. 2023a) and tracing the source of knowledge issues (Gou et al. 2023; Kandpal et al. 2023). The factuality issue for LLMs receive relatively the most attention. Several instances are shown in Table 1. For instance, an LLM might be deficient in domain-specific factual knowledge, such as medicine or law domain. Additionally, the LLM might be unaware of facts that occurred post its last update. There are also instances where the LLM, despite possessing the relevant facts, fails to reason out the correct answer. In some cases, it might even forget or be unable to recall facts it has previously learned. The factuality problem is closely related to several hot topics in the field of Large Language Models, including Hallucinations (Ji et al. 2023a; Zhao et al. 2023b), Outdated Information (Nakano et al. 2022; Qin et al. 2023), and Domain-Specificity (e.g., Health (Xiong et al. 2023a; Wang et al. 2023d), Law (Cui et al. 2023b), Finance (Wu et al. 2023)). At their core, these topics address the same issue: the potential for LLMs to generate contents that contradicts certain facts, whether those contents arise out of thin air, outdated information, or a lack of domain-specific knowledge.

Therefore, we consider these three topics to fall within the scope of the factuality problem. However, it is important to note that while these topics are related, they each have a unique focus. Both hallucinations and factuality issues in LLMs pertain to the accuracy and reliability of generated content, they address distinct aspects. Hallucinations primarily revolve around LLMs generating baseless or unwarranted content. Drawing from definitions by OpenAI 2023; Ji et al. 2023a, hallucinations can be understood as the model’s inclination to "produce content that is nonsensical or untruthful in relation to certain sources." This is different from factuality concerns, which emphasize the model’s ability to learn, acquire, and utilize factual knowledge. To illustrate the distinction: If an LLM, when prompted to craft "a fairy tale about a rabbit and a wolf making friends," produces a tale about "a rabbit and a dog becoming friends," it’s exhibiting hallucination. However, this isn’t necessarily a factuality error. If the generated content contains accurate information but diverges from the prompt’s specifics, it’s a hallucination but not a factuality issue. For instance, if the LLM’s output includes more details or different elements than the prompt specifies but remains factually correct, it’s a case of hallucination. Conversely, if the LLM avoids giving a direct answer, states "I don’t know," or provides a response that’s accurate but omits some correct details, it’s addressing factuality, not hallucination. Furthermore, it’s worth noting that hallucination can sometimes produce content that, while deviating from the original input, remains factually accurate. For a more structured comparison between the factuality issue and hallucination, refer to Table 2. Outdated information, on the other hand, focuses on instances where previously accurate information has been superseded by more recent knowledge. Lastly, domain-specificity emphasize the generation of content that requires specific, specialized knowledge. Despite these differences, all three topics contribute to our understanding of the broader factuality problem in LLMs.

In this survey, our primary focus is on two specific settings: 1. Standard LLMs: Directly using LLMs for answering and chatting (OpenAI 2023; OpenAI 2022b); 2. Retrieval-Augmented LLMs: The retrieval-augmented generation (Microsoft 2023; Liu 2022). The latter is of particular interest as retrieval mechanisms are among the most prevalent methods for enhancing the factuality of LLMs. This involves not just generating accurate responses but also correctly selecting pertinent knowledge snippets from the myriad of retrieved sources.

While summarization tasks—where the goal is to produce summaries that stay true to the source input—have seen research on factuality (Maynez et al. 2020; Tam et al. 2023; Tang et al. 2022), we opted not to focus heavily on this domain in our survey. There are a few reasons for this decision. Firstly, the source inputs for summarization often contain content that is not factual. Secondly, summarization introduces unique challenges like ensuring coherence, conciseness, and relevance, which deviate from the focus of this survey. It is also worth noting that Pu et al. 2023 found LLMs to produce fewer factual errors or hallucinations compared to humans across various summarization benchmarks. However, we will still discuss some works in this area, particularly those that overlap with retrieval settings.

3. Impact

The factuality problem significantly impacts the usability of LLMs. Some of these issues have even led to losses at the societal or economic level (Pranshu Verma 2023; Sands 2023), drawing the attention of many users, developers, and researchers (Ji et al. 2023a).

Factuality issues also impacted the legal field, with a lawyer in the United States facing sanctions for submitting hallucinated case law in court. One court has mandated that lawyers indicate the portions generated by generative AI in their submitted materials (Curran et al. 2023). In addition, as part of a research study, a fellow lawyer asked ChatGPT to generate a list of legal scholars with a history of sexual harassment. ChatGPT generated a list that included a law professor. ChatGPT claimed that the professor attempted to touch a student on a class tripnd referenced an article from The Washington Post in March 2018. However, the fact is that this article does not exist, nor does the mentioned class trip (Pranshu Verma 2023). Besides, a mayor in Australia, discovered false claims made by ChatGPT stating that he was personally convicted of bribery, confessed to charges of bribery and corruption, and received a prison sentence. In response, he plans to initiate legal action against the company responsible for ChatGPT, accusing them of defamation for disseminating untrue information about him. This could potentially be the first defamation case of its kind involving an artificial intelligence chatbot (Sands 2023).

A recent study (Nori et al. 2023) provides a comprehensive evaluation of GPT-4’s performance in medical competency examinations and benchmark datasets. The evaluation utilizes the text-only version of GPT-4 and investigates its ability to address medical questions without any training or fine-tuning. The assessment is conducted using the United States Medical Licensing Examination (USMLE) (Kung et al. 2023) and the MultiMedQA benchmark (Singhal et al. 2023), comparing GPT-4’s performance against earlier models like GPT-3.5 and models specifically fine-tuned on medical knowledge. The results demonstrate that GPT-4 significantly outperforms its predecessors, achieving scores on the USMLE that exceed the passing threshold by more than 20 points and delivering the best overall performance without specialized prompt crafting or domain-specific fine-tuning.

While large language models show promise on medical datasets, the introduction of automation in the healthcare field still requires extreme caution (Shen et al. 2023). Existing metrics and benchmarks are often developed for highly focused problems. Evaluating LLM outputs in supporting real-world decision-making poses challenges (Singhal et al. 2023), including the stability and robustness of personalized recommendations and inferences in real-world contexts. Using large language models carries significant risks, including inaccurate ranking recommendations (Hirosawa et al. 2023) (e.g., differential diagnosis) and sequencing (Zhang et al. 2023b) (e.g., information gathering and testing), as well as factual errors (Shen et al. 2023), particularly important omissions and erroneous responses.

Factuality Evaluation

Evaluating the factuality of LLMs is pivotal for ensuring the reliability and trustworthiness of their generated content (Pezeshkpour 2023; Lee et al. 2022b). As LLMs become increasingly integrated into various applications, from information retrieval to content generation, the accuracy of their outputs becomes paramount. In this section, we delve into the evaluation metrics and benchmarks used for assessing the factuality of LLMs, studies that have undertaken such evaluations, and domain-specific evaluation.

In this subsection, we delve into metrics established for evaluating the factuality of LLMs. As the problem formulation is akin to natural language generation (NLG) (Celikyilmaz et al. 2021; Ji et al. 2023b), we introduce several automatic evaluation metrics typically used for NLG, as well as specifically examining the metrics for factuality.

We categorize these metrics into the following groups: (1) Rule-based evaluation metrics; (2) Neural evaluation metrics; (3) Human evaluation metrics; and (4) LLM-based evaluation metrics.

Most assessments of factuality in large language models use rule-based evaluation metrics, due to their consistency, predictability, and ease of implementation; they allow for reproducible outcomes through a systematic method. However, they can be rigid and may not account for nuances or variations in language use, context interpretation, or colloquial expressions. This means language models rated highly by these metrics may still produce content that feels unnatural or inauthentic to human readers.

An “exact match" refers to a situation where the generated text precisely matches a specific input or reference text. This means that the LLM produces output that is identical, word-for-word, to the provided input or reference text. Exact matches are often used in NLG when you want to replicate or repeat a piece of text without any variations or alterations. Exact Match is commonly used in Open-domain Question Answering (Izacard and Grave 2021b; Wang et al. 2023h).

Many factuality evaluation measurements use commonly adapted metrics, such as Accuracy, Precision, Recall, AUC, F-Measure, Calibration score, Brier score, and other common metrics used in probabilistic forecasting and machine learning, specifically in tasks involving probabilistic predictions. The common definition of these metrics involves the use of correctly predicted labels and ground-truth labels. As the input and output of LLMs are human-readable sentences, there is no unified method to convert the sentences into labels. Most evaluation will define their own way. That is, these scores are frequently not used in isolation, but rather in combination. For instance, BERTScore (Zhang* et al. 2020) uses BERT for determining the Precision and Recall, then uses F-measures to get a final weighted score. In the following, we will describe the most simple form of these scores.

The Calibration score, used in (Lin et al. 2022a; Kadavath et al. 2022) measures the agreement between predicted probabilities and observed frequencies. A perfectly calibrated model should, over a large number of instances, see the predicted probability of an outcome match the relative frequency of that outcome.

The Brier Score, used in (Kadavath et al. 2022) is a metric used in probabilistic forecasting to measure the accuracy of probabilistic predictions. It calculates the mean squared difference between the predicted probability assigned to an event and the actual outcome of the event. The Brier Score ranges from 0 to 1, where 0 indicates a perfect prediction and 1 indicates the worst possible prediction. In other words, the lower the Brier Score, the better the accuracy of the prediction. It’s worth noting that this metric is appropriate for binary and categorical outcomes, but not for ordinal outcomes. For binary outcomes, the Brier Score can be calculated as follows:

where forecastiforecast_{i} is the predicted probability, actualiactual_{i} is the actual outcome (0 or 1), and NN is the total number of forecasts made.

are widely recognized metrics in multi-choice question answering, particularly in TruthfulQA (Lin et al. 2022b). MC1: For a given question accompanied by several answer choices, the objective is to identify the sole correct answer. The model’s selection is determined by the answer choice to which it allocates the highest log probability of completion, independent of the other choices. The score is calculated as the straightforward accuracy across all questions. MC2: Presented with a question and multiple reference answers labeled as true or false, the score is derived from the normalized total probability assigned to the collection of true answers.

(Papineni et al. 2002), also known as Bilingual Evaluation Understudy metric is commonly employed in the context of factuality evaluation. This metric calculates the co-occurrence frequency of phrases in two sentences, based on a weighted average of matched n-gram phrases. This helps in quantitatively assessing the factual consistency between the generated text and its reference.

The Recall-Oriented Understudy for Gisting Evaluation (ROUGE) metric (Lin 2004) serves as a measure of the similarity between the generated text and the reference text, with the similarity grounded on recall scores. Primarily used in the field of text summarization, the ROUGE metric incorporates four distinct types. These include ROUGE-n, which assesses n-gram co-occurrence statistics, and ROUGE-l, which measures the longest common subsequence. ROUGE-w provides evaluation based on the weighted longest common subsequence, while ROUGE-s measures skip-bigram co-occurrence statistics. These diverse metrics collectively provide a comprehensive measure of the factual accuracy of generated text.

The Metric for Evaluation of Translation with Explicit Ordering (METEOR) (Banerjee and Lavie 2005) aims to address several shortcomings presented by BLEU. These include deficiencies in recall, the absence of higher order n-grams, an absence of explicit word-matching between the generated and reference text, and the use of geometric averaging of n-grams. METEOR overcomes these by introducing a comprehensive measure, calculated based on the harmonic mean of the unigram precision and recall. This offers potentially enhanced appraisals of factuality in the generated text.

(Weller et al. 2023) is an n-gram overlap measure. It quantifies the degree to which a generated passage consists of exact spans found in a text corpus. The QUIP-Score serves to evaluate the ’grounding’ ability of LLMs, specifically assessing whether model-generated answers can be directly located within the underlying text corpus. It is defined by comparing the precision of the character n-gram from the generated output to the pre-training corpus.

1.2. Neural Evaluation Metrics

These metrics operate by comparing the output of these models with a standard or reference text by learning evaluator models. This category primarily comprises three prominent metrics, namely ADEM, BLEURT, and BERTScore. Each metric approaches the evaluation slightly differently, yet all aim to assess the semantic and lexical alignment between machine-generated text and its reference counterpart, thus ensuring the factuality of the generated content.

The automatic Dialogue Evaluation Model (ADEM) metric is a cogent tool utilized to gauge the quality of responses generated by language models in a conversation. Developed by Lowe et al. 2017, this metric trains a Hierarchical RNN, in a semi-supervised fashion, to predict ratings for the machine-generated responses. The ADEM’s assessment primarily hinges on a dialogue context, designated as c\mathbf{c}, alongside the model’s response, labeled as r^\mathbf{\hat{r}}, and a reference response, specified as r\mathbf{r}. These elements, when encoded via the hierarchical RNN, inform the ADEM which then predicts a score, reflecting the proximity of the model’s response to factual accuracy and relevance.

the BERTScore metric, as introduced by Zhang et al. 2020, utilizes pre-trained embeddings from BERT to gauge the similarity between two sentences. This is done by assigning embeddings to a referenced sentence, "x", and a model-generated sentence, "x^\hat{x}", denoted as "x\mathbf{x}" and "x^\mathbf{\hat{x}}", respectively. The recall, precision, and F1-scores are then calculated to quantify the similarity between "x\mathbf{x}" and "x^\mathbf{\hat{x}}". This score provides a measure of the generated sentence’s factuality.

the Bilingual Evaluation Understudy with Representations from Transformers (BLEURT) metric (Sellam et al. 2020) employs a unique pre-training scheme, where BERT is initially trained on a significant corpus of synthetic sentence pairs. This is reinforced by multiple lexical and semantic-level supervision signals, used concurrently. Following this pre-training stage, BERT is further fine-tuned on rating data, and its objective is to estimate human rating scores accurately. The initial stage of pre-training is essential for this metric because it significantly enhances the model’s robustness, thereby effectively transforming it into a secure bulwark against quality drifts inherent in generative systems.

(Yuan et al. 2021) is a metric proposed to evaluate the quality of the generated text, such as in applications like machine translation and summarization. This metric conceives this evaluation as a problem of text generation, modeled using pre-trained sequence-to-sequence models. BARTScore uses BART, an encoder-decoder-based pre-trained model, to translate generated text to and from a reference point, earning higher scores when the text is more accurate and fluent. This metric offers several variants that can be flexibly applied in an unsupervised manner for different perspectives of text evaluation, such as fluency or factuality. Tests have shown that BARTScore can outperform existing top-scoring metrics in the majority of test settings across multiple datasets and perspectives.

1.3. Human Evaluation Metrics

Human evaluation in factuality assessment is crucial due to its sensitivity to nuanced elements of language and context that may elude automated systems. Human evaluators excel at interpreting abstract concepts and emotional subtleties that can significantly inform the accuracy of evaluation. However, they are subject to limitations such as subjectivity, inconsistency, and potential for error. On the other hand, automated evaluations offer consistent results, and efficient processing of large data sets, and are ideal for tasks needing quantitative measurements. They also provide an objective benchmark for model performance comparison. Overall, an ideal evaluation framework might blend automated evaluation’s scalability and consistency with human evaluation’s ability to interpret complex linguistic concepts.

is a metric to verify that the output of LLMs is sharing only verifiable information about the external world. As proposed by (Rashkin et al. 2023), Attributable to Identified Sources (AIS) is a human evaluation framework that adopts a binary concept of attribution. A text passage yy is deemed attributable to a set AA of evidence if, and only if, an arbitrary listener would agree with the statement "According to AA, yy" within the context of yy. The AIS framework awards a full score (1.0) if every element of content in passage yy can be linked to the evidence set AA. Conversely, it gives a score of zero (0.0) if this condition is not met.

Based on AIS, Gao et al. 2023a propose a more fine-grained, sentence-level extension of AIS called Auto-AIS, where annotators assign AIS scores to each sentence, and an average score across all sentences is reported. This procedure effectively measures the percentage of sentences that are fully attributable to the evidence. Context, such as surrounding sentences and the question the text answers, is provided to annotators for more informed judgment. A limit is also set for the number of evidence snippets in the attribution report to maintain conciseness.

During model development, an automated AIS metric is defined to approximate human AIS evaluations, using a natural language inference model, which correlates well with AIS scores. Before computing the scores, they improve accuracy by decontextualizing each sentence in context.

(Min et al. 2023) is a novel evaluation metric designed to assess the factual precision of long-form text generated by LLMs. The challenge of evaluating the factuality of such text arises from two main issues: (1) the generated content often contains a mix of supported and unsupported information, making binary judgments insufficient, and (2) human evaluation is both time-consuming and expensive. To address these challenges, FActScore breaks down a generated text into a series of atomic facts—short statements that each convey a single piece of information. Each atomic fact is then evaluated based on its support from a reliable knowledge source. The overall score represents the percentage of atomic facts that are supported by the knowledge source. The paper conducted extensive human evaluations to compute FActScores for biographies generated by several state-of-the-art commercial LLMs, including InstructGPT, ChatGPT, and the retrieval-augmented PerplexityAI. The results revealed that these LLMs often contain factual inaccuracies, with FActScores ranging from 42% to 71%. Notably, the factual precision of these models tends to decrease as the rarity of the entities in the biographies increases.

1.4. LLM-based Metrics

Using LLMs for evaluation offers efficiency, versatility, reduced reliance on human annotation, and the capability of evaluating multiple dimensions of conversation quality in a single model call, which improves scalability. However, potential issues include a lack of established validation, which can lead to bias or accuracy problems if the LLM used for evaluation is not thoroughly vetted. The decision process to identify suitable LLMs and decoding strategies can be complex and pivotal to obtaining accurate evaluations. The range of evaluation may also be limited, as the focus is often on open-domain conversations, possibly leaving out assessments in specific or narrow domains. While reducing human input can be beneficial, it can also miss out on crucial interaction quality aspects better evaluated by human judges, such as emotional resonance or nuanced understandings.

(Fu et al. 2023b) is a new evaluation framework designed to assess the quality of output from generative AI models. To provide these evaluations, GPTScore taps into the emergent capabilities of 19 different pre-trained models, such as zero-shot instruction, and uses them to judge the generated texts. These models vary in scale from 80M to 175B. Testing across four text generation tasks, 22 aspects of evaluation, and 37 related datasets, has demonstrated that GPTScore can effectively evaluate text per instructions in natural language. This attribute allows it to sidestep challenges traditionally encountered in text evaluation, like the need for sample annotations and achieving custom, multi-faceted evaluations.

(Lin et al. 2022b) is a finetuned model based on the GPT-3-6.7B, which is trained to evaluate the truthfulness of answers to questions in the TruthfulQA dataset. The training set consists of triples in the form of question-answer-label combinations, where the label can be either true or false. The model’s training set includes examples from the benchmark and answers generated by other models assessed by human evaluation. In its final form, the GPT-judge uses examples from all models to evaluate the truthfulness of responses. This training includes all questions from the dataset, with the goal being to evaluate truth, not generalize new questions.

The study conducted by the authors (Lin et al. 2022b) focuses on the application of GPT-judge in assessing Truthfulness and Informativeness using the TruthfulQA Dataset. The authors undertook the fine-tuning of two distinct GPT-3 models to evaluate two essential aspects: Truthfulness, which pertains to the accuracy and honesty of information provided by the LLM, and Informativeness, which measures how effectively the LLM conveys relevant and valuable information in its responses. From these two fundamental concepts, the authors derived a combined metric denoted as truth * info. This metric represents the product of scalar scores for both truthfulness and informativeness. It not only quantifies the extent to which questions are answered truthfully but also incorporates the assessment of informativeness for each response. This comprehensive approach prevents the model from generating generic responses like "I have no comment" and ensures that responses are not only truthful but also valuable. These metrics have found widespread deployment in evaluating the factuality of information generated by LLMs (Chuang et al. 2023; Li et al. 2023d).

(Lin and Chen 2023) is a novel evaluation methodology for open-domain dialogues with LLMs. Unlike conventional evaluation methods which rely on human annotations, ground-truth responses, or multiple LLM prompts, LLM-Eval uses a unique prompt-based evaluation process employing a unified schema to assess various elements of a conversation’s quality during a single model function. Extensive evaluations of LLM-Eval’s performance using multiple benchmark datasets indicate it is effective, efficient, and adaptable compared to traditional evaluation practices. Further, the authors stress the necessity of selecting appropriate LLMs and decoding strategies for precise evaluation outcomes, underscoring LLM-Eval’s versatility and dependability in assessing open-domain conversation systems across a variety of circumstances.

2. Benchmarks for Factuality Evaluation

In this section, we delve into the benchmarks that are prominently employed to assess the factuality of LLMs. Specific benchmarks tailored for evaluating factuality in LLMs are tabulated in Table 4.

MMLU (Hendrycks et al. 2021) and TruthfulQA (Lin et al. 2022b) stand as two pivotal benchmarks in the realm of evaluating the factuality of LLMs (OpenAI 2023; Touvron et al. 2023b; Park 2023) . The MMLU benchmark is proposed to measure a text model’s multitask accuracy across a diverse set of 57 tasks. These tasks span a wide range of subjects, from elementary mathematics to US history, computer science, law, and more. The benchmark is designed to test both the world knowledge and problem-solving ability of models. The findings from the paper suggest that while most recent models perform at near random-chance accuracy, the largest GPT-3 model showcased a significant improvement. However, even the best models still have a long way to go before achieving expert-level accuracy across all tasks (OpenAI 2023). TruthfulQA is a benchmark designed to assess the truthfulness of a language model’s generated answers. The benchmark consists of 817 questions spanning 38 categories, including health, law, finance, and politics. These questions were crafted in such a way that some humans might answer them falsely due to misconceptions or false beliefs. The goal for models is to avoid generating these false answers that they might have learned from imitating human texts. The TruthfulQA benchmark serves as a tool to highlight the potential pitfalls of relying solely on LLMs for accurate information and emphasizes the need for continued research in this area.

HaluEval (Li et al. 2023a) is a benchmark designed to understand and evaluate the propensity of LLMs like ChatGPT to generate hallucinations. A hallucination, in this context, refers to content that either conflicts with the source or cannot be verified based on factual knowledge. The HaluEval benchmark offers a vast collection of generated and human-annotated hallucinated samples, aiming to evaluate the performance of LLMs in recognizing such hallucinations. The benchmark utilizes a two-step framework, termed "sampling-then-filtering", based on ChatGPT to generate these samples. Additionally, human labelers were employed to annotate hallucinations in ChatGPT responses. The HaluEval benchmark is a comprehensive tool that not only evaluates the hallucination tendencies of LLMs but also provides insights into the types of content and the extent to which these models are prone to hallucinate.

BigBench (Srivastava et al. 2023) focuses on the capabilities and limitations of LLMs. It comprises 204 tasks from diverse domains such as linguistics, childhood development, math, common-sense reasoning, biology, physics, social bias, software development, and more. The benchmark is designed to evaluate tasks believed to be beyond the capabilities of current language models. The study evaluates the performance of various models, including OpenAI’s GPT models, on BIG-bench and compares them to human expert raters. Key findings suggest that model performance and calibration improve with scale but are still suboptimal when compared to human performance. Tasks that involve a significant knowledge or memorization component show predictable improvement, while tasks that exhibit "breakthrough" behavior at a certain scale often involve multiple steps or components.

Huang et al. 2023a propose C-Eval, the first comprehensive Chinese evaluation suite. It can be used to evaluate the advanced knowledge and reasoning abilities of foundational models within a Chinese context. The evaluation suite comprises multiple-choice questions spanning 52 diverse disciplines, with four levels of difficulty: middle school, high school, college, and professional. Additionally, C-Eval Hard is introduced for very challenging subjects within the C-Eval suite, which demands advanced reasoning abilities to solve. Their evaluations of state-of-the-art LLMs, including both English and Chinese-oriented models, have shows that there is significant room for improvement as only GPT-4 managed to achieve an average accuracy of over 60%. The authors focus on assessing LLMs’ advanced abilities within a Chinese context. The researchers assert that LLMs intended for a Chinese environment should be evaluated based on their knowledge of Chinese users’ primary interests, such as Chinese culture, history, and laws. With C-Eval, the authors aim to guide developers in understanding the abilities of their models from multiple dimensions to facilitate the development and growth of foundational models for Chinese users. Simultaneously, C-Eval has not just introduced a whole suite but subsets that can serve as individual benchmarks, thereby assessing certain model abilities and analyzing key strengths and limitations of foundational models. Experiment results have shown that although GPT-4, ChatGPT, and Claude were not exclusively tailored for Chinese data, they emerged as the top performers on C-Eval.

SelfAware (Yin et al. 2023) aims to investigate whether models recognize what they don’t know. This dataset encompasses two types of questions: unanswerable and answerable. The dataset comprises 2,858 unanswerable questions gathered from various websites and 2,337 answerable questions extracted from sources such as SQuAD, HotpotQA, and TriviaQA. Each unanswerable question is confirmed as such by three human evaluators. In the conducted experiments, GPT-4 achieves the highest F1 score of 75.5, compared to a human score of 85.0. Larger models tend to perform better and in-context learning can enhance performance.

The Pinocchio benchmark (Hu et al. 2023a) serves as an extensive evaluation platform, emphasizing factuality and reasoning for LLMs. This benchmark encompasses 20,000 varied factual queries from diverse sources, timeframes, fields, regions, and languages. It tests an LLM’s capability to discern combined facts, process both organized and scattered evidence, recognize the temporal evolution of facts, pinpoint minute factual disparities, and withstand adversarial inputs. Each reasoning challenge within the benchmark is calibrated for difficulty to allow for detailed analysis.

Kasai et al. 2022 introduce a dynamic QA platform named REALTIMEQA. This platform is unique in that it announces questions and evaluates systems on a regular basis, specifically weekly. The questions posed by REALTIMEQA pertain to current events or novel information, challenging the static nature of traditional open domain QA datasets. The platform aims to address instantaneous information needs, pushing QA systems to provide answers about recent events or developments. Their preliminary findings indicate that while GPT-3 can often update its generation results based on newly-retrieved documents, it sometimes returns outdated answers when the retrieved documents lack sufficient information.

FreshQA (Vu et al. 2023) is a dynamic benchmark designed to evaluate up-to-date world knowledge of LLMs. Its questions range from those that are never-changing to those that are fast-changing, as well as questions based on false premises. The aim is to challenge the static nature of LLMs and test their adaptability to ever-changing knowledge through human evaluations. They develop a reliable evaluation protocol that uses a two-mode system: RELAXED and STRICT for a comprehensive understanding of model performance, ensuring that answers are confident, definitive, and accurate. They also provide a strong baseline named FRESHPROMPT, which seeks to enhance LLM performance by integrating real-time data from search engines. Initial experiments reveal that, outdated training data weakens the performance of LLMs and the FRESHPROMPT method can significantly enhance it. The research underscores the need for LLMs to be refreshed with current information to ensure their relevance and accuracy in a constantly evolving world.

Several benchmarks, such as BigBench (Srivastava et al. 2023) and C-Eval (Huang et al. 2023a), encompass subsets that extend beyond the realm of factual knowledge or factuality. In this work, we specifically emphasize and focus on those subsets related to factuality.

There are benchmarks primarily designed for Pre-trained Language Models (PLMs) that can also be adapted for LLMs. Some studies use them for evaluating LLM’s factuality, but they are not so widely-used, so we have chosen to exclude them from the table for clarity. These benchmarks predominantly encompass knowledge-intensive tasks, as highlighted by (Petroni et al. 2021). Those benchmarks include NaturalQuestions (NQ) (Kwiatkowski et al. 2019), TriviaQA (TQ) (Joshi et al. 2017), OTT-QA (Chen et al. 2021a), AmbigQA (Min et al. 2020) and WebQuestion (WQ) (Berant et al. 2013) of Open-domain Question Answering (QA) task, the HotpotQA (Yang et al. 2018), 2WikiMultihopQA (Ho et al. 2020), IIRC (Ferguson et al. 2020), MuSiQue (Trivedi et al. 2022) of multi-step QA, the ELI5 (Fan et al. 2019) of the Long-form QA task, the FEVER (Thorne et al. 2018), FM2 (Eisenschlos et al. 2021), HOVER (Jiang et al. 2020), FEVEROUS (Aly et al. 2021) of the Fact Checking task, the T-REx (Elsahar et al. 2018), zsRE (Levy et al. 2017) and LAMA (Petroni et al. 2019a) to examine factual knowledge contained in pretrained language models, the WikiBio (Lebret et al. 2016) of biography generation task, the RoSE (Liu et al. 2023a), WikiAsp (Hayashi et al. 2021) of summarization task, the KILT (Petroni et al. 2021) of comprehensive knowledge intensive tasks, the MassiveText (et al. [n. d.]), Curation Corpus (Curation 2020), Wikitext103 (Merity et al. 2016), Lambada (Paperno et al. 2016), C4 (Raffel et al. 2020b), Pile (Gao et al. 2020) of langauge modeling task, the WoW (Dinan et al. 2019), DSTC7 track2 (Galley et al. 2019a), DSTC11 track5 (Zhao et al. 2023a) of dialogue task, the RealToxicityPrompts (Gehman et al. 2020) of toxicity reduction, the CommaQA (Khot et al. 2022), StrategyQA (Geva et al. 2021a), TempQuestions (Jia et al. 2018), IN-FOTABS (Gupta et al. 2020a) of diverse reasoning tasks.

Some studies (Min et al. 2023; Manakul et al. 2023) also provide a small dataset, but they mainly concentrate on the evaluation metrics or methods for factuality, we choose to discuss them in the next subsection.

3. Factuality Evaluation Studies

In this section, we delve into studies that evaluate factuality in LLMs without introducing a specific benchmark, focusing primarily on those whose main contribution lies in the evaluation methodology. We spotlight works that have pioneered evaluation techniques, metrics, or have offered distinctive insights into the factuality evaluation of LLMs.

Manakul et al. 2023 use an evaluation process that encompasses several key steps. Initially, synthetic Wikipedia articles are generated using GPT-3, focusing on individuals from the Wikibio dataset. Subsequently, manual annotation is performed at the sentence level, classifying sentences as “Major Inaccurate" , “Minor Inaccurate", or “Accurate", with “Major Inaccurate" denoting sentences unrelated to the topic. Passage-level scores are derived by averaging sentence-level labels, and identifying cases of total inaccuracies through score distribution analysis. Inter-annotator agreement is assessed using Cohen’s κ\kappa (Cohen 1960). Evaluation metrics primarily employ precision-recall curves (PR-Curves), distinguishing between “Non-Factual Sentences," “Non-Factual* Sentences" (a specific subset), and “Factual Sentences." These PR-Curves elucidate the trade-off between precision and recall for different detection methods.

BSChecker (Hu et al. 2023b) is a factual errors detection framework that detects inaccuracies on triplet-level granularity. Its pipeline comprises a claim extractor and a checker. The claim extractor extracts triplets from model’s response as claims formed with (subject, predicate, object). Meanwhile, the checker compare these claims triplets with references, categorizing each triplet to one of three hallucination labels: Entailment, Contradiction or Neutral, drawing inspiration from the Natural Language Inference task. BSChecker offers a benchmark that distinguishes three task settings: Zero Context, Noisy Context and Accurate Context covering tasks of closed-book QA, retrieval-augmented generation, summarization, closed-QA and information extraction.

Pezeshkpour 2023 propose a new metric for measuring whether a certain type of knowledge is present in LLM. This metric is based on information theory and measures knowledge by analyzing the probability distribution of predictions made by LLM before and after injecting target knowledge. The accuracy of GPT-3.5 in the Knowledge Probing task is tested on the T-REx (Elsahar et al. 2018) and LAMA (Petroni et al. 2019a) datasets.

Varshney et al. 2023 point out a common real-world occurrence where users often ask questions based on false premises. These questions are challenging for state-of-the-art models. This necessitated the creation of a new evaluation dataset. To this end, the authors have conducted a case study and compiled a set of 50 such adversarial questions, all of which the GPT-3.5 model answered incorrectly. The aim is to create a challenging experimental setup to assess the performance of models faced with such questions. In order to enhance the evaluation, corresponding true premise questions were created for each of the false premise questions. This allowe for a holistic evaluation of the model’s performances, taking into consideration both correct and incorrect premises. The authors make sure to evaluate the complete answers given by the model for correct and incorrect questions - in this context, it’s not enough for an answer to be partially correct, the entire answer needs to be accurate to be marked as correct. For example: For the false premise question “Why does Helium have an atomic number of 1?", the corresponding true premise question is “Why does Hydrogen have an atomic number of 1?".

FACTOOL (Chern et al. 2023) is a tool designed to function as a factuality detector, with the primary purpose of auditing generative chatbots and assessing the reliability of their outputs. This tool is employed to evaluate several contemporary chatbots, including GPT-4, ChatGPT, Claude, Bard (Google 2023), and Vicuna (Chiang et al. 2023). Notably, FACTOOL itself leverages the capabilities of GPT-4. For the evaluation process, the researchers have curated a diverse set of prompts: 30 from knowledge-based question answering (KB-QA), 10 each from code, math, and scientific domains. The KB-QA prompts were sourced from a prior study, code prompts were taken from HumanEval, math prompts from another distinct study, while the scientific prompts were crafted by the authors themselves. The evaluation metrics included both claim-level and response-level accuracies for each chatbot. To offer a more comprehensive and equitable evaluation, a weighted claim-level accuracy is used. The weighting is determined based on the proportion of prompts from each category. The findings are illuminating. GPT-4 emerge as the top performer in terms of both weighted claim-level factual accuracy and response-level accuracy among all the chatbots assessed. Another intriguing observation is that chatbots that underwent supervised fine-tuning, such as Vicuna-13B, exhibited commendable performance in standard scenarios like KB-QA. However, their performance dip in more intricate scenarios, including those involving math, code, and scientific queries.

Wang et al. 2023b ask several LLMs, including ChatGPT, GPT-4 (OpenAI 2023), BingChat (Microsoft 2023) to answer open questions from NaturalQuestions (Kwiatkowski et al. 2019) and TriviaQA (Joshi et al. 2017). They manually estimate the accuracy of those LLMs on open question answering, and find that though LLMs can achieve nice performance but still far away from perfect. Besides, they evaluate whether the GPT-3.5 can assess the correctness of LLM-generated responses, and find negative results, even if the golden answer is also presented. Similarity, Fu et al. 2023a ask LLMs, such as GPT-2 and GPT-4, to directly score the factuality of a summary, and find no significant correlation between LLM’s factuality indicators and human evaluations.

Kadavath et al. 2022 investigate whether language models can evaluate the accuracy of their own assertions and predict which questions they can answer correctly. It is found that larger models are well-calibrated on diverse multiple-choice and true/false questions if given in the appropriate format. The approach to self-evaluation on open-ended tasks is to ask the models to initially suggest answers, and then evaluate the probability (P[True]) that their answers are correct. This resulted in compelling performance, calibration, and scaling on a diverse range of tasks. Furthermore, self-evaluation performance improved when the models are allowed to consider many of their own suggestions before predicting the validity of a specific one.

Yu et al. 2023 explore whether the internal knowledge of LLMs can replace the retrieved documents on knowledge intensive tasks. They ask LLMs, such as InstructGPT (OpenAI 2022a), to directly generate contexts given a question rather than retrieving from database. They find the generated documents contain the golden answers more often than the top retrieved documents. Then they feed the generated docs and retrieved docs to the Fusion-in-Decoder model (Izacard and Grave 2021b) for knowledge-intensive tasks such as Open-domain QA (Kwiatkowski et al. 2019) and find the generated docs are more effective than the retrieved docs, suggesting that the LLMs contain enough knowledge for knowledge-intensive tasks.

Menick et al. 2022b propose a task named Self-supported QA to evaluate LLMs’ ability in also producing citations when generating answers. Authors ask humans to evaluate whether the responses of their proposed model GopherCite are plausible and whether they are supported by the accompanying quote evidence on datasets such as NQ, ELI5, TruthfulQA.

CONNER (Chen et al. 2023b), a framework that evaluates LLMs as generators of knowledge. It focuses on six areas: Factuality, Relevance, Coherence, Informativeness, Helpfulness, and Validity. It evaluates whether the generated information can be backed by external proof (Factuality), is relevant to the user’s query (Relevance), and is logically consistent (Coherence). It also checks if the knowledge provided is novel or surprising (Informativeness). The Extrinsic evaluation measures whether the knowledge enhances downstream tasks (Helpfulness) and its results are factually accurate (Validity).

In the realm of factuality evaluation, the Model Editing task holds a unique position, focusing on refining the internal knowledge of models. This task comes with its own set of specialized evaluation metrics. Predominantly, the Zero-Shot Relation Extraction (zsRE) (Levy et al. 2017) and CounterFact (Meng et al. 2022) serve as the primary benchmarks for assessing Model Editing techniques. When evaluating these methods, several key criteria emerge: Reliability: Post-editing, the model should consistently generate the intended output. Generalization: The model should be adept at producing the target output even when presented with paraphrased inputs. Locality: Edits should be localized, ensuring that facts not related to the specific edit remain intact. However, given the intricate web of interconnected facts, recent studies Zhong et al. 2023b; Yao et al. 2023a; Cohen et al. 2023a have advocated for a more holistic evaluation approach. They introduce broader criteria for fact updates, encompassing aspects like portability, logical generalization, among others.

RAGAS (Es et al. 2023) proposes a framework for the reference-free evaluation of Retrieval-Augmented Generation (RAG) systems in Large Language Models (LLMs). This framework primarily assesses the RAG system’s ability to identify relevant and key context paragraphs, the LLM’s fidelity in utilizing these paragraphs, and the overall quality of the generated content. The framework focuses on three aspects of quality: Faithfulness: The generated answers should be based on the given context. To evaluate faithfulness, the generated content is broken down into sentences, which are then transformed into one or more short assertions. Each assertion’s faithfulness is assessed against the retrieved content, with the faithfulness score being the ratio of faithful assertions to the total number of assertions. Answer Relevance: The generated answers should appropriately address the posed questions. This is evaluated by having the LLM generate potential questions from each piece of generated content. An embedding model then calculates the similarity between each potential question and the original question. The answer relevance score is the average of these similarities. Context Relevance: The retrieved context should be highly relevant, containing as little irrelevant information as possible. This is assessed by using the LLM to extract sentences from the retrieved content that are relevant to the question. The context relevance score is then calculated based on the proportion of these relevant sentences to the total number of sentences retrieved. The authors created a WikiEval dataset, consisting of question-context-answer triplets with human judgments. This allows for the measurement of how closely different evaluation frameworks align with human assessments of faithfulness, answer relevance, and context relevance. The paper compares two baselines: GPTScore, which requires ChatGPT to rate the three quality dimensions on a scale of 0 to 10, and GPT Ranking, which does not require ChatGPT to choose a preferred answer/context but includes definitions of the considered quality metrics in the prompt. When evaluating answer relevance, a specific prompt is used. Overall, the RAGAS framework proposed in the paper aligns more closely with human judgment compared to the two baselines.

Some work’s main contributions lie at the methods to improve the factuality of LLMs, and their evaluation part may also be informative and important when people reach the related study. So we choose to list evaluation parts in the Table 6 but not to discuss their evaluation in detail.

4. Evaluating Domain-specific Factuality

In our examination of specialized Large Language Models (LLMs), we’ve identified an extensive array of datasets and benchmarks tailored to a variety of domains. These resources not only serve as critical tools for evaluating the capabilities of LLMs but also facilitate advancements in specialized applications. We summarize them in Table 7. The distinction between this subsection and the previous two lies in its focus. This subsection delves deeper into factuality evaluation tailored to specific domains, while Sec 3.2 and 3.3 primarily concentrate on general factuality evaluation, with only a portion of their content dedicated to datasets evaluating factuality within specific domains.

Xie et al. 2023a designed a financial natural language understanding and prediction evaluation benchmark dubbed FLARE, based on their collected financial instruction tuning dataset FIT. This benchmark is used to evaluate their FinMA model. It randomly selects validation sets from FIT to choose the best model checkpoint and utilizes distinct test sets for evaluation. FLARE is a broader variant compared to the existing FLUE benchmark (Shah et al. 2022) as it also encapsulates financial prediction tasks like stock movement prediction in addition to standard NLP tasks. The FLARE dataset includes several subtasks, such as sentiment analysis (FPB, FiQA-SA), news headline classification (Headline), named entity recognition (NER), question answering (FinQA, ConvFinQA), and stock movement prediction (BigData22, ACL18, CIKM18). Performance is gauged via a variety of metrics for each task, such as the accuracy and weighted F1 Score for sentiment analysis, entity-level F1 score for named entity recognition, and accuracy and Matthews correlation coefficient for stock movement prediction. Several methods, including their own FIT-fine-tuned FinMA and other LLMs (BloombergGPT, GPT-4, ChatGPT, BLOOM, GPT-NeoX, OPT-66B, Vicuna-13B) are dedicated to their comparison. BloombergGPT’s performance is assessed in various shot scenarios, while the zero-shot performance is reported for the remaining results. Some of the baselines depend on human evaluations given that LLMs without fine-tuning fail to generate instruction-defined answers. Conversely, FinMA’s results are conducted on a zero-shot basis and can be evaluated automatically. To enable direct comparison between the performance of FinMA and BloombergGPT, despite the former not releasing their test datasets, test datasets were constructed with the same data distribution.

Li et al. 2023c propose EcomInstruct benchmark, which experiment investigates the performance of their EcomGPT language model in comparison to baseline models such as BLOOM and BLOOMZ. Categories of these baseline models include pre-trained large models with decoder-only architecture like BLOOM, and instruction-following language models like BLOOMZ and ChatGPT. The evaluation metric of the EcomInstruct involves converting all tasks to generative paradigms and using text generation evaluation metrics like ROUGE-L. Classification tasks were evaluated with precision, recall, and F1 scores. The EcomInstruct dataset, comprising 12 tasks across four major categories, is divided into training and testing sections. The EcomGPT is trained on 85,746 instances of E-commerce data. The performance of the model is assessed based on its capacity to generalize unseen tasks or datasets, with emphasis on cross-language and cross-task paradigm settings.

Wang et al. 2023a propose a localized medical benchmark called CMB, or the Comprehensive Medical Benchmark in Chinese. CMB is rooted entirely in the native Chinese linguistic and cultural framework. While traditional Chinese medicine is a significant part of this evaluation, it does not make up the entire benchmark. Both prominent LLMs, such as ChatGPT and GPT-4, and localized Chinese LLMs, including those specializing in the health domain, are evaluated using CMB. However, the benchmark is not designed as a leaderboard competition, but rather as a tool for self-assessment and understanding the progression of models in this field. CMB embodies a comprehensive, multi-layered medical benchmark in Chinese, comprised of hundreds of thousands of multiple-choice questions and complex case consultation questions. This wide range of queries covers all clinical medical specialties and various professional levels, seeking to evaluate a model’s medical knowledge and clinical consultation capabilities comprehensively.

Li et al. 2023f introduce Huatuo-26M dataset, the largest Chinese medical Question and Answer (QA) dataset to date, including over 26 million high-quality medical QA pairs. It covers a wide range of topics, such as diseases, symptoms, treatments, and drug information. The dataset is a valuable resource for anyone looking to improve AI applications in the medical field, such as chatbots and intelligent diagnostic systems. The Huatuo-26M dataset is gathered and integrated from various sources, including online medical encyclopedias, online medical knowledge bases, and online medical consultation records. Each QA pair in the dataset contains a problem description and a corresponding answer from a doctor or expert. Although a significant proportion of the Huatuo-26M dataset is constituted by the online medical consultation records, the text format data from these records is not publicly available, for unspecified reasons. This dataset is expected to be instrumental for multiple types of research and AI applications in the medical field. These areas of application extend to Natural Language Processing tasks like developing QA systems, text classification, sentiment analysis, and Machine Learning model training tasks like disease prediction and personalized treatment recommendation. Consequently, the dataset is favorable for developing AI applications in the medical field, from intelligent diagnosis systems to medical consultation chatbots.

Jin et al. 2023 use nine tasks related to NCBI resources to evaluate the proposed GeneGPT model. The dataset used is GeneTuring benchmark (Hou and Ji 2023), which contains 12 tasks with 50 question-answer pairs each. The tasks are divided into four modules - Nomenclature (Gene alias, Gene name conversion), Genomic location(Gene SNP association, Gene location, SNP location), Functional analysis(Gene disease association, Protein-coding genes), Sequence alignment (DNA to human genome, DNA to multiple species). Two settings of GeneGPT are assessed: a full setting where all prompt components are used, and a slim setting using only two components. The performance of GeneGPT is compared against various baselines including general-domain GPT-based LLMs like GPT-2, GPT-3, and ChatGPT. Additionally, GPT-2-sized biomedical domain-specific LLMs such as BioGPT and BioMedLM are evaluated. The new Bing, a retrieval-augmented LLM with access to relevant web pages, is also assessed. The result evaluation of the compared methods is based on the results reported in the original benchmark and are manually evaluated. However, the evaluation of the proposed GeneGPT method is determined through automatic evaluations.

LegalBench (Guha et al. 2023) is a benchmark for legal reasoning introduced due to the increasing use of LLMs in the legal field. It consists of 162 tasks covering six types of legal reasoning and was collaboratively constructed through an interdisciplinary process with significant contributions from legal professionals. The tasks designed aim to either demonstrate practical legal reasoning capabilities or measure reasoning skills that are of interest to lawyers. To facilitate discussions between different fields relating to LLMs in law, LegalBench tasks correspond to popular legal frameworks for describing legal reasoning, thus creating a shared language between lawyers and LLM developers. The paper not only describes LegalBench, but also presents an evaluation of 20 different open-source and commercial LLMs and highlights the types of research opportunities that LegalBench can provide.

LawBench (Fei et al. 2023) is an evaluation framework dedicated to assessing the capabilities of LLMs in relation to legal tasks. The context-specificity and high-stakes nature of the law field make it crucial to have a clear grasp of LLMs’ legal knowledge and their ability to execute legal tasks. LawBench dives deep into three cognitive evaluations of LLMs: the capability to memorize crucial legal details, understanding legal texts, and applying legal knowledge to resolve complicated legal problems. A total of 20 diverse tasks have been put together, covering five main task categories: single-label classification, multi-label classification, regression, extraction, and generation. Throughout the evaluation process, 51 LLMs were extensively tested under LawBench, compiling a spectrum of language models including 20 multilingual LLMs, 22 Chinese-oriented LLMs, and 9 law-specific LLMs. The results concluded that GPT-4 ranked as the superior LLM in the law domain, significantly surpassing its competitors. Despite the noted improvements when LLMs were fine-tuned on law-specific texts, the study acknowledged that there is still a long road ahead in achieving highly reliable LLMs for legal tasks.

Bi et al. 2023 design OceanBench to evaluate the capabilities of LLMs for oceanography tasks. OceanBench includes a total of 15 type of ocean-related tasks, such as question-answering and description tasks. The samples in OceanBench are are automatically generated from the seed dataset and have undergone manual verification by experts.

Analysis of Factuality

In the previous Section 3, we provide quantitative statistics related to evaluating factuality. In this section, we delve deeper, exploring the underlying mechanisms that influence factuality in large language models.

This subsection delves into intriguing analyses concerning the factuality of LLMs, focusing on aspects that aren’t directly tied to evaluation, or enhancement. Specifically, we explore the mechanisms through which LLMs process, interpret, and produce factual content. The subsequent sections offer an in-depth examination of different dimensions of factuality in LLMs, ranging from their knowledge storage, and awareness to their approach to managing conflicting data. While some research (Wang et al. 2021; Neeman et al. 2023) applies some similar analyses to Pretrained Language Models. This review excludes them as they are not typically considered Large Language Models.

The language model serves as a repository of knowledge, storing a multitude of information about the world within its parameters (Petroni et al. 2019b). However, the organization of this knowledge within LLMs remains largely mysterious. (Meng et al. 2022) introduce a methodology called causal tracing to measure the indirect impact of hidden states or activations. This technique was employed to illustrate that factual knowledge is primarily stored in the early-layer feed-forward networks (FFNs) of such models. Similarly, Geva et al. 2021c also suggests that a substantial portion of factual information is encoded within the FFN layers. They conceptualize the input of the FFN as a query, the first layer as keys, and the second layer as values. Consequently, the intermediate hidden dimension of the FFN can be interpreted as the number of memories within the layer, and the intermediate hidden state represents a vector comprising activation values for each memory. As a result, the final output of the FFN can be understood as the weighted sum of activated values. The authors further demonstrate that the value vectors often encapsulate human-interpretable concepts and knowledge (Geva et al. 2022; Geva et al. 2023). In addition, Chen et al. 2023a have made an intriguing finding that the language model contains language-independent neurons that express multilingual knowledge and degenerate neurons that convey redundant information by applying the integrated gradients method (Lundstrom et al. 2022). Nevertheless, it is important to note that the aforementioned studies primarily focus on the representation of individual facts, and the comprehensive understanding of how factual knowledge is precisely organized and interconnected within these models remains an ongoing challenge.

1.2. Knowledge Completeness and Awareness

This subsubsection delves into the intriguing realm of LLMs’ self-awareness, their capacity to discern their knowledge gaps, and the balance between their internally generated knowledge and externally retrieved information. We delve into the dichotomy between parametric knowledge and retrieved knowledge, exploring the promises and challenges these models bring to the forefront of knowledge-intensive tasks.

Several studies have investigated the knowledge awareness of Large Language Models, specifically assessing whether LLMs can accurately estimate the correctness of their own responses. Most of these studies treat LLMs as "black boxes," prompting the models to report their confidence levels or calculating the perplexity of the model’s output as an indicator of response likelihood. Gou et al. 2023 explore the model’s ability to validate and iteratively refine its outputs, akin to how humans interact with tools. The authors find that solely relying on self-correction without external feedback can lead to marginal improvements or even diminished performance. (Ren et al. 2023) experiment with settings either augmented or not with external document retrieval to determine whether models recognize their own knowledge boundaries. Their findings indicate that LLMs possess an inaccurate perception of their factual knowledge boundaries and tend to be overly confident about their responses. LLMs often fail to fully harness the knowledge they possess; however, retrieval enhancement can somewhat compensate for this shortcoming. Yin et al. 2023 introduce a dataset named “SelfAware" to test if models recognize what they don’t know, encompassing both answerable and unanswerable questions. The experiment suggests that models do possess some capacity to discern their own knowledge gaps, but they are still far from human levels. GPT-4 outperforms other models, instructions and In-Context-Learning (Dong et al. 2023) can enhance a model’s discriminatory ability. Kadavath et al. 2022 focus on LLM self-assessment based on Language Model calibration using multiple-choice questions. Their findings revealed that the "none of the above" option decreased accuracy, larger models showed better calibration, and RLHF hindered model calibration levels. However, simply adjusting the temperature parameter can rectify this issue. Azaria and Mitchell 2023 assess the truthfulness of statements generated by LLMs, by using the model’s internal state and hidden layer activations. The authors, employing a feedforward neural network, can classify if the model is misleading by utilizing the hidden output states.

Yu et al. 2023 explore whether the internal knowledge of LLMs can replace the retrieved documents on knowledge-intensive tasks. They ask LLMs, such as InstructGPT, to directly generate contexts given a question rather than retrieving them from the database. They find the generated documents contain the golden answers more often than the top retrieved documents. Then they feed the generated docs and retrieved docs to the Fusion-in-Decoder model (Izacard and Grave 2021b) for knowledge-intensive tasks such as Open-domain QA (Kwiatkowski et al. 2019) and find the generated docs are more effective than the retrieved docs, suggesting that the LLMs contain enough knowledge for knowledge-intensive tasks.

On the contrary, these observations have been contested in subsequent investigations. Kandpal et al. 2023 underscore the dependency of LLMs on the number of associated documents seen during pre-training. They argue that the success in answering fact-based questions is highly linked to the number of documents containing the topic of the question that were encountered in pre-training. The study further posited the necessity of scaling models extensively to achieve competitive performance for questions with minimum representation in the training data. Adding to these concerns, Sun et al. 2023d critically evaluate the factual knowledge base of LLMs, using a specifically designed Head-to-Tail benchmark comprised of 18K question-answer pairs. The results show that the understanding of factual knowledge, particularly related to torso-to-tail entities, by currently available LLMs is suboptimal.

In summary, while LLMs show promise in handling knowledge-intensive tasks, their dependency on pre-training information and limitations in factual accuracy remain significant hurdles. It underscores the need for further advancements in the field and the importance of incorporating complementary methods, such as retrieval augmentation, to enhance the learning of long-tail knowledge in LLMs.

1.3. Contextual Influence and Knowledge Conflict

This sub-subsection examines the interplay between an LLM’s inherent parametric knowledge and the provided contextual knowledge, exploring both the model’s capacity to utilize context and its behavior when confronted with conflicting information.

Some works explore the model’s capacity to utilize context, for example, Li et al. 2023e observe that larger models tend to rely on their parametric knowledge, even when faced with counterfactual contexts. This suggests that as models increase in size, they might grow more confident in their internal knowledge, potentially sidelining external context. However, the introduction of irrelevant contexts can still influence their outputs. The balance between controllability (relying on relevant context) and robustness (resisting irrelevant context) emerges as a challenge in LLM training. The study indicates that reducing context noise improves controllability, but the effect on robustness remains to be seen. In contrast, Zhou et al. 2023b propose prompt templates to guide LLMs towards more faithful text generation. Among these, opinion-based prompts prove most effective, indicating that when LLMs are queried about opinions, they adhere more closely to the context. Interestingly, the study finds that using counterfactual context enhances the model’s faithfulness, while original context, sourced from platforms like Wikipedia, might induce a simplicity bias, leading LLMs to answer questions without heavily relying on the context. Chen et al. 2023d conduct a comprehensive evaluation of LLMs’ ability to effectively utilize retrieved information. The study reveals that while retrieved documents can boost LLM performance, the presence of noise in these documents can hinder it. Yue et al. 2023 investigate the nature of LLM-generated content in relation to provided references. They categorize the generated content as attributable, contradictory, or extrapolatory to the reference. Both fine-tuned models and instruction-based LLMs struggle to accurately evaluate the alignment between generated content and references, underscoring the challenge of ensuring that LLMs produce content consistent with the provided context. Shi et al. 2023a study the distractability of LLMs on the GSM-IC dataset derived from GSM8K (Cobbe et al. 2021). They discover that all the prompting techniques are responsive to irrelevant information in the problem definition. They identify various factors of irrelevant information that impact the model’s sensitivity to irrelevant context. Moreover, they find that self-consistency prompting and incorporating irrelevant information into the exemplars can enhance models’ performance, enabling them to learn to disregard irrelevant information.

A series of studies are interested in LLMs’ behavior when confronted with conflicting information. Longpre et al. 2021 introduce the concept of knowledge conflicts, where the provided context contradicts the model’s learned information. Their findings suggest that such conflicts lead to increased prediction uncertainty, especially for in-domain examples. Observations across models, ranging from T5-60M to 11B, indicate that larger models tend to default to their parametric knowledge. Moreover, there’s an inverse relationship between retrieval quality and the tendency to rely on internal knowledge: the more irrelevant the evidence, the more the model defaults to its parametric knowledge. (Chen et al. 2022) conduct experiments on typical ODQA models, including FiD and RAG. Their results show that FiD models rarely resort to memorization (less than 3.6% for NQ) compared to RAG models. Instead, FiD primarily grounds its answers in the provided evidence. Interestingly, when confronted with conflicting retrieved passages, models tend to fall back on their parametric knowledge. (Xie et al. 2023b) explore the behavior of recent LLMs, including ChatGPT and GPT-4. Contrary to findings on smaller LMs, they discover that LLMs can be highly receptive to external evidence, even if it contradicts their parametric memory, provided the external evidence is coherent and convincing. Additionally, LLMs exhibit a strong confirmation bias, especially when presented with evidence that aligns with their parametric memory. This bias becomes even more pronounced for widely accepted knowledge. In scenarios where no relevant evidence is provided, LLMs tend to express uncertainty. However, when presented with both relevant and irrelevant evidence, they demonstrate an ability to filter out irrelevant information. Wang et al. 2023c argue that LLMs should not rely solely on either parametric or non-parametric information, but grant LLM users the agency to make informed decisions. They introduce a framework including three tasks ((1) Contextual knowledge conflict detection; (2) QA-span knowledge conflict detection; (3) Distinct answers generation) to simulate knowledge conflicts and evaluate whether LLMs’ behaviors align with the goal.

In conclusion, while studies like Li et al. 2023e and Zhou et al. 2023b emphasize the challenges and potential solutions in making LLMs more context-aware, others like Yue et al. 2023 and Xie et al. 2023b highlight the inherent biases and limitations of LLMs. The overarching theme is the need for a balanced approach, where LLMs effectively leverage both their internal knowledge and external context to produce accurate and coherent outputs.

2. Causes of Factual Errors

Understanding the root causes of these factual inaccuracies is crucial for refining these models and ensuring their reliable application in real-world scenarios. In this subsection, we delve into the multifaceted origins of these errors, categorizing them based on the stages of model operation: Model Level, Retrieval Level, Generation Level, and other miscellaneous causes. Table 1 shows examples of factuality errors caused by different factors.

This subsection delves into the intrinsic factors within large language models that contribute to factual errors, originating from their inherent knowledge and capabilities.

The model may lack comprehensive expertise in specific domains, leading to inaccuracies. Every LLM has its limitations based on the data it was trained on. If an LLM hasn’t been exposed to comprehensive data in a specific domain during its training, it’s likely to produce inaccurate or generalized outputs when queried about that domain. For instance, while an LLM might be adept at answering general science questions, it might falter when asked about niche scientific subfields (Lu et al. 2022).

The model’s dependence on older datasets can make it unaware of recent developments or changes. LLMs are trained on datasets that, at some point, become outdated. This means that any events, discoveries, or changes post-dating the last training update won’t be known to the model. For example, ChatGPT and GPT-4 are both trained on data up to 2021.09 might not be aware of events or advancements after then.

The model does not always retain knowledge from its training corpus. While it’s a misconception that LLMs “memorize" data, they do form representations of knowledge based on their training. However, they might not always recall specific, less-emphasized details from their training parameters, especially if such details were rare or not reinforced through multiple examples. For example, ChatGPT has pretrained with Wikipedia, but it still fail in answering some questions from NaturalQuestions (Kwiatkowski et al. 2019) and TriviaQA (Joshi et al. 2017), which are constructed from Wikipedia (Wang et al. 2023b).

The model might not retain knowledge from its training phase or could forget prior knowledge as it undergoes further training. As models are further fine-tuned or trained on new data, there’s a risk of “catastrophic forgetting" (Goodfellow et al. 2015; Wang et al. 2022; Chen et al. 2020; Zhai et al. 2023) where they might lose certain knowledge they were previously know. This is a well-known challenge in neural network training, where networks forget previously learned information when exposed to new data, which also happens in large language models (Luo et al. 2023b).

While the model might possess relevant knowledge, it can sometimes fail to reason with it effectively to answer queries. Even if an LLM has the requisite knowledge to answer a question, it might fail to connect the dots or reason logically. For instance, ambiguity in the input (Liu et al. 2023d) can potentially lead to a failure in understanding by LLMs, consequently resulting in reasoning errors. In addition, Berglund et al. 2023 find the LLM traps in the reversal curse, for example, it knows that A is B’s mother but fail to answer who is B’s son. This is especially evident in complex multi-step reasoning tasks or when the model needs to infer based on a combination of facts (Tan et al. 2023b; Kotha et al. 2023).

2.2. Retrieval-level Causes

The retrieval process plays a pivotal role in determining the accuracy of LLMs’ responses, especially in retrieval-augmented settings. Several factors at this level can lead to factual errors:

If the retrieved data doesn’t provide enough context or details, the LLM might struggle to generate a factual response. This can result in generic or even incorrect outputs due to the lack of comprehensive evidence.

LLMs can sometimes accept and propagate misinformation present in the retrieved data. This is especially concerning when the model encounters knowledge conflicts, where the retrieved information contradicts its pre-trained knowledge, or multiple retrieved documents contradict each other (Li et al. 2015). For instance, (Longpre et al. 2021) observed that the more irrelevant the evidence, the more likely the model is to rely on its intrinsic knowledge. Recent studies, such as (Pan et al. 2023b), have also shown that LLMs are susceptible to misinformation attacks within the retrieval process.

LLMs can be misled by irrelevant or distracting information in the retrieved data. For example, if the evidence mentions a "Russian movie" and "Directors", the LLM might incorrectly infer that "The director is Russian". (Luo et al. 2023a) highlighted this vulnerability, noting that LLMs can be significantly impacted by distracting retrieval results. They further proposed instruction tuning as a potential solution to enhance the model’s ability to sift through and leverage retrieval results more effectively.

Additionally, when dealing with long retrieval inputs, the models tend to show their best performance when processing information given at the beginning or end of the input context, according to Liu et al. 2023b. In contrast, the models are likely to experience a significant decrease in performance when they are required to extract relevant data from the middle of these extensive contexts.

Even when the retrieved information is closely related to the query, LLMs can sometimes misunderstand or misinterpret it. While this might be less frequent when the retrieval process is optimized, it remains a potential source of error. For instance, in the ReAct study (Yao et al. 2023b), the rate of errors dropped significantly when the retrieval process was improved.

2.3. Inference-level Causes

During the generation process, a minor error or deviation at the beginning can compound as the model continues to generate content. For instance, if an LLM misinterprets a prompt or starts with an inaccurate premise, the subsequent content can veer further from the truth (Zhang et al. 2023g; Varshney et al. 2023).

The decoding phase is crucial for translating the model’s internal representations into human-readable content (Chuang et al. 2023; Massarelli et al. 2020). Mistakes during this phase, whether due to issues like beam search errors or suboptimal sampling strategies, can lead to outputs that misrepresent the model’s actual "knowledge" or intention. This can manifest as inaccuracies, contradictions, or even nonsensical statements.

LLMs are a product of their training data. If they’ve been exposed more frequently to certain types of content or phrasing, they might have a bias toward generating similar content, even when it’s not the most factual or relevant. This bias can be especially pronounced if the training data has imbalances or if certain factual scenarios are underrepresented (Felkner et al. 2023; Gallegos et al. 2023). The model’s outputs, in such cases, reflect its training exposure rather than objective factuality. For example, research by (Hossain et al. 2023) suggests that LLMs can correctly identify the gender of individuals who conform to the binary gender system, but they perform poorly when determining non-binary or neutral genders.

Enhancement

This section discusses methods to enhance factuality in LLMs across different phases, including LLM generation, retrieval-augmented generation, inference-phase enhancements, and domain-specific factuality improvements, outlined in Figure 2.

Table 8 provides a summary of enhancement methods and their respective improvements over baseline LLMs. It’s essential to recognize that various research papers may employ distinct experimental settings, such as zero-shot, few-shot, or full settings. Consequently, when examining this table, it’s important to note that performance metrics for different methods, even when evaluating the same metric on the same dataset, may not be directly comparable.

When focusing on standalone LLM generation, enhancement strategies can be broadly grouped into three main categories:

(1) Improving Factual Knowledge from Unsupervised Corpora (Sec 5.1.1): This involves refining the training data during pretraining, such as through deduplication and emphasizing informative words (Lee et al. 2022a). Techniques like TOPICPREFIX (Lee et al. 2022b) and sentence completion loss are also explored to enhance this approach.

(2) Enhancing Factual Knowledge from Supervised Data (Sec 5.1.2): Examples in this category include supervised fine-tuning strategies (Chung et al. 2022; Zhou et al. 2023a) focus on finetuning with labelled data or integrating structured knowledge such as knowledge graphs (KGs) or making precise adjustments to model parameters (Li et al. 2023d).

(3) Optimally Eliciting Factual Knowledge from the Model (Sec 5.1.3, 5.1.4, 5.1.5): This category encompasses methods like Multi-agent collaboration (Du et al. 2023) and innovative prompts (Yu et al. 2023). Additionally, novel decoding methodologies, such as factual-nucleus sampling, are introduced to further improve factuality (Lee et al. 2022b; Chuang et al. 2023).

Some work (Liu et al. 2023c; Sun et al. 2023b) aim for improve the factuality of large multi-modal models, we choose to not emphasize them.

Pretraining plays a pivotal role in equipping the model with the factual knowledge derived from the corpus. By emphasizing strategies during this phase, the model’s inherent factuality can be significantly enhanced. This approach is particularly crucial for addressing challenges like immemorization and forgetting.

Methods employed during the foundational pretraining of the model.

Lee et al. 2022a develop two distinct tools for deduplicating training data, addressing the issue of redundant texts and long repeated substrings present in the training sets of current LLMs. These tools effectively reduce the recall of memorized texts in model outputs. Remarkably, they achieve similar or even superior accuracy with fewer training steps.

Sadeq et al. 2023 introduce a modification to the Masked Language Model (MLM) training objective used in the pretraining of LLMs. They discover that high-frequency words do not consistently contribute to the model’s ability to learn factual knowledge. To address this, they devise a strategy that encourages the language model to prioritize informative words during unsupervised training. This is achieved by masking tokens more frequently based on their informative relevance. To quantify this relevance, they utilize Pointwise Mutual Information (PMI) (Fano and Hawkins 1961), positing that words with elevated PMI values, in relation to their adjacent words, are likely to be more informatively pertinent. Experimental results indicate that this innovative approach significantly bolsters the efficacy of pretrained language models across various tasks, including factual recall, question answering, sentiment analysis, and natural language inference, in a closed-book setting.

Iterative pretraining processes that allow the model to progressively refine and update its knowledge base.

Lee et al. 2022b introduce TOPICPREFIX as a pre-processing method and the sentence completion loss as training objective. Some factual sentences may be unclear when the LM training corpus is chunked, especially when these sentences contain pronouns (e.g., she, he, it). So they prepend TOPICPREFIX (e.g., Wikipedia document name) to sentences in the factual documents to transform each sentence into an independent factual statement. They also introduce the sentence completion loss, with the aim of enabling the model to capture facts from entire sentences rather than just focusing on the associations between subwords. For implementation, they establish a pivot tt for each sentence and require zero-masking for all token prediction losses before tt. This pivot is only necessary during the training phase. Experiments show that such methods can further reduce the factual errors than standard factual-domain adaptive training.

1.2. Supervised Finetuning

Supervised fine-tuning leverages labeled datasets to refine the model’s performance. This approach serves a dual purpose: it imparts specific task or knowledge base-oriented information to the model and addresses challenges like immemorization and forgetting. Several studies, such as Chung et al. 2022; Zhou et al. 2023a, emphasize the pivotal role of supervised fine-tuning in eliciting the inherent knowledge of the base model, subsequently enhancing its reasoning capabilities.

A cyclic fine-tuning approach, where the model undergoes consistent refinement using sequential sets of labeled data.

Moiseev et al. 2022 investigate how to inject structured knowledge from a knowledge graph into LLMs. The approach involves directly training T5 using triplets containing relationship knowledge. Previous methods often describe the triplets using prompts and then train LLMs with masked language model task, but some triplets are not easy to describe. They compare three knowledge-enhancement fine-tuning methods, which are MLM training on the C4 corpus (Raffel et al. 2020b), masking the subject or object in KG triplets, and masking the subject or object in the KELM corpus (Agarwal et al. 2021). Experiments show that effectiveness of the latter two methods achieved better exact match scores in closed-book QA tasks (Jiang et al. 2019; Welbl et al. 2018; Joshi et al. 2017; Kwiatkowski et al. 2019). This demonstrates that training directly based on KG triplets is an effective way to inject knowledge into the model

Sun et al. 2023c introduce negative samples and contrastive learning to supervised finetuning process with MLE loss. The sources of these negative samples are either from the Knowledge Graph or are generated by large language models. While traditional contrastive learning can only function at the token or sentence level, the intrinsic value of span information cannot be overlooked, they employ Named Entity Recognition to extract this critical span data. The training process incorporates a blend of MLE loss, standard contrastive learning loss, and span-based contrastive learning loss, with the parameters finely tuned to optimize results. Experiments indicate that the method delivers performance comparable with SOTA KB-based methods but offers significant benefits in efficiency and scalability.

Yang et al. 2023a presents a development framework for Knowledge Graph-enhanced Large Language Models (KGLLMs), drawing from established technologies. They detail various enhancement strategies, including a before-training enhancement that refines input quality by integrating factual data; a during-training enhancement that synergizes textual and structural knowledge, utilizing tools like graph neural networks and attention mechanisms; multi-task learning which focuses on knowledge-guided pre-training tasks to bolster the factual knowledge acquisition of LLMs; and a post-training enhancement that fine-tunes LLMs for domain-specific tasks using knowledge-rich data. Additionally, the significance of prompt learning in LLMs is highlighted, emphasizing the importance of choosing appropriate prompt templates. They also suggest knowledge graphs as a valuable resource for crafting these templates to harness domain-specific knowledge.

FactTune (Ovadia et al. 2023) aims to use Reinforcement Learning to improve the factuality of LLM, so scoring the factuality of the content generated by the model is needed. However, human labels are too costly, so authors use two methods to check the factuality of the content generated by LLM, namely 1) Reference-based, which uses external knowledge base retrieval and Factscore (Min et al. 2023) method to score; 2) Reference-free, which uses Another LLM’s confidence to score. Then, they utilized the Directly Preference Optimization (Rafailov et al. 2023) algorithm to optimize the model. 50% and 20+% error rate reductions were observed in Llama’s bio generation tasks and medical QA respectively.

Instead of directly finetuning the model, model edit is a more precise approach to enhance the model’s factuality. By editing specific areas that are related to the fact, the model can correctly express that fact without compromising other unrelated knowledge. Current editing methods (Yao et al. 2023a) can be categorized into weight-preserved and weight-modified paradigms. When modifying the model’s weight, KN (Dai et al. 2022) and ROME (Meng et al. 2022) first analyze representations to locate those underlying factual errors and then directly update the relevant weights. Meanwhile, KE (De Cao et al. 2021) and MEND (Mitchell et al. 2022a) employ a hypernetwork to learn the necessary weight changes. While effective, the robustness and generalization of directly updating weights remains an open question.

Li et al. 2023d present a technique known as Inference-Time Intervention (ITI) designed to boost the factual accuracy of large language models. ITI adjusts model activations during inference to enhance the truthfulness of responses. This method, categorized under activation editing, is both adjustable and less intrusive, setting it apart from weight editing techniques. Drawing inspiration from prior research, they utilize steering vectors, which are proven effective for style transfer, to guide model activations. They detail a multi-step model selection process involving the calibration of intervention strength hyperparameters, pinpointing truth-telling related heads, and determining their truth-telling directions. The TruthfulQA dataset serves as the foundation for training and validation, with a rigorous 2-fold cross-validation employed to prevent data leakage. Experiments on the TruthfulQA benchmark demonstrate that ITI improves Alpaca’s (Taori et al. 2023) truthfulness from 32.5% to 65.1%.

1.3. Multi-Agent

Engaging multiple models in a collaborative or competitive manner, enhancing factuality through their collective prowess, helps in the immemorization and reasoning failure problems.

Du et al. 2023 propose an approach to enhance the performance of language models by treating different LLMs as intelligent agents engaged in multi-agent debates. In this method, multiple instances of language models present and debate their respective answers and reasoning processes, ultimately reaching a consensus on the final answer after multiple rounds of debate. If the debate’s answers fail to converge, prompts are modified to reduce the stubbornness of the two agents. This approach has been demonstrated to significantly improve mathematical and reasoning abilities across various tasks while enhancing the factual accuracy of generated content. Moreover, this method can be directly applied to existing black-box models, making it applicable to all research tasks using the same prompts.

Cohen et al. 2023b develop a fact-checking mechanism. Drawing parallels to a scenario where a witness is interrogated for the veracity of their claims, they utilize a LLM to gather statements from a QA dataset, which are either factually correct or incorrect. During the statement generation, the model is provided with a golden answer, prompting it to produce both accurate and inaccurate statements, inherently labeling each statement. For every QA pair, another LLM, acting as an interrogator, generates a series of questions. A separate LLM, playing the role of the respondent, answers these queries. This iterative questioning and answering continues until the interrogator is satisfied, culminating in a conclusion. The authors conduct experiments on datasets like LAMA (Petroni et al. 2019a), PopQA (Mallen et al. 2023), NaturalQA (Kwiatkowski et al. 2019), and TriviaQA (Joshi et al. 2017). The precision is the portion of incorrect claims, out of the claims rejected by the examiner and the recall is the portion of incorrect claims rejected by the examiner, out of all the incorrect claims. Measured by the F1 score, the LM vs LM method consistently outperforms the baseline by a significant margin, ranging from ten to over twenty points across all datasets.

1.4. Novel Prompt

Introducing innovative or tailored prompts to extract more factual and precise responses from the LLM, can better assist the model elicit the knowledge in its parameters and improve the reasoning ability.

Yu et al. 2023 introduce a novel approach called Generate-then-Read (GENREAD). Document retrievers are replaced by LLM generators. In this article, the LLM is prompted to generate multiple contextual documents on a given question. The authors clustered these document embeddings, and sampled documents from different clusters to ensure diversity of contextual documents. With these generated in-context demonstrations, LLMs achieved better results on knowledge-intensive tasks than retrieving from external corpus such as Wikipedia.

Weller et al. 2023 introduce a metric called QUIP-Score to measure the dependency on pre-trained data. They index Wikipedia to swiftly determine the dependency of an LLM’s response. By using specific prompts, like "Based on evidence from Wikipedia:" they aim to evoke the LLM’s recall of content from its training dataset. In addition to grounding prompts, they also introduce anti-grounding prompts to encourage the LLM to respond without referencing its training data, for instance, "Respond without using any information from Wikipedia." The motivation behind this approach is the belief that guiding the LLM to reference more of the knowledge it acquires during pre-training can reduce the generation of incorrect information. To quantify this grounding, they propose the QUIP-Score metric to gauge the similarity between the model’s generated content and the most relevant content in Wikipedia. In their experiments conducted on datasets like TQ, NQ, HotpotQA, and ELI5, the results show that while adding the said prompt doesn’t significantly improve traditional QA metrics, it notably boosts the scores on the QUIP metric.

Khot et al. 2023 present Decomposed Prompting. Complex tasks are broken down into multiple simpler tasks via prompting and then it can be addressed by task-specific LLMs. For instance, for tasks involving extremely long input sequences, this technique systematically decomposes the input into shorter sequences for individual processing. Notably, the authors observed that when combined with a retrieval module, this approach significantly enhances performance on open domain multi-hop QA tasks.

Dhuliawala et al. 2023 present Chain-of-Verification (CoVe) to reduce factual errors.The CoVe strategy involves the model initially crafting a response, subsequently formulating verification queries to assess its initial draft, independently responding to these queries to maintain unbiased answers, and ultimately producing a validated reply. The CoVe method encompasses four pivotal stages: 1) Drafting an initial reply based on a query using the LLM, 2) Generating verification questions from the query and initial answer to pinpoint potential errors, 3) Responding to each verification question and comparing these answers with the initial response to detect discrepancies, and 4) If discrepancies are found, producing a revised answer that integrates the verification outcomes. The entire procedure is executed by prompting the same LLM differently to achieve the intended results. Experiments show that Chain-of-Verification can reduce errors in diverse tasks, including list-based questions from Wikidata (Vrandečić and Krötzsch 2014), closed book MultiSpanQA (Li et al. 2022c) and longform text generation (Min et al. 2023).

1.5. Decoding

Decoding methodologies, such as beam search and nucleus sampling, play a crucial role in directing the model to produce outputs that are both factual and coherent. By refining the decoding process, challenges like snowballing errors or erroneous decoding, as detailed in Sec 4.2.3, can be effectively addressed.

Lee et al. 2022b propose a new decoding sampling algorithm called factual-nucleus sampling that achieves a better trade-off between generation quality and factuality when compared to prevailing decoding algorithms. They postulate that the randomness of sampling is more harmful to factuality when applied to the generation of the latter portion of a sentence as opposed to its initial segment. So the factual-nucleus sampling algorithm, an adaptation of nucleus sampling, can dynamically adjust the ’nucleus’ probability throughout the generation of each sentence, progressively reducing randomness with each successive generation step. The ω\omega-bound parameter is provided to prevent the p-value from becoming too small and hurting diversity. Experimental results show that a factual-nucleus sampling algorithm can improve the factuality of generation while maintaining generation quality, e.g., diversity and repetition.

Chuang et al. 2023 propose “Decoding by Contrasting Layers" to mitigate these hallucinations. This approach leverages the differences in logits obtained from projecting the later layers versus earlier layers to the vocabulary space, taking advantage of the known localization of factual knowledge in LLMs. The results of this study demonstrate that DoLa consistently enhances the truthfulness of LLM-generated content across various tasks, such as multiple-choice and open-ended generation tasks, showcasing its potential to significantly improve the reliability of LLMs in generating accurate and truthful facts.

2. On Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) has emerged as a widely adopted approach to address certain limitations inherent to standalone LLMs, such as outdated information and the inability to memorize (Liu 2022; Chase 2022). These challenges are elaborated upon in Sec 4.2.1. Yet, while RAG offers solutions to some issues, it introduces its own set of challenges, including the potential for insufficient information and the misinterpretation of related data, as detailed in Sec 4.2.2. This subsection delves into various strategies devised to mitigate these challenges. Within the realm of retrieval-augmented generation, enhancement techniques can be broadly categorized into several pivotal areas:

(1) The Normal Setting of Utilizing Retrieved Text for Generations (Sec 5.2.1):

(2) Interactive Retrieval and Generation (Sec 5.2.2): Examples here include the integration of Chain-of-Thoughts steps into query retrieval (He et al. 2022) and the use of an LLM-based agent framework that taps into external knowledge APIs (Yao et al. 2023b).

(3) Adapting LLMs to the RAG Setting (Sec 5.2.3): This involves methods like the one proposed by Peng et al. 2023, which combines a fixed LLM with a plug-and-play retrieval module. Another notable approach is REPLUG (Shi et al. 2023b), a retrieval-augmented framework that treats the LLM as a black box and fine-tunes retrieval models using language modeling scores.

(4) Retrieving from Additional Knowledge Bases (Sec 5.2.5 and Sec 5.2.4): This category includes methods that retrieve from external parametric memories (Chen et al. 2023e) or knowledge graphs (Zhang et al. 2023f) to enhance the model’s knowledge base.

A normal RAG setting works by retrieving external data and passing it to a LLM during the generation phase. We follow the framework proposed by LlamaIndex (Liu 2022) and LangChain (Chase 2022) to decouple the process into the following modules and steps:

(1) Document loaders are used to load documents from various sources. These loaders can fetch different types of documents (HTML, PDF, code) from various locations (private s3 buckets, public websites).

(2) Document transformers are employed to extract relevant parts of the documents. This may involve splitting or chunking large documents into smaller chunks. Different algorithms for this task are employed, optimized for specific document types like code or markdown.

(3) Text embedding models are employed to capture the semantic meaning of the text.

(4) Vector stores are employed to efficiently store and search the embeddings.

(5) Retrievers, such as the Parent Document Retriever, Self Query Retriever, and Ensemble Retriever, are used to retrieve the data from the database.

Here, the Parent Document Retriever allows for the creation of multiple embeddings per parent document to retrieve smaller chunks while maintaining a larger context. The Self Query Retriever separates the semantic part of a query from other metadata filters, allowing for more accurate retrieval. The Ensemble Retriever enables the retrieval of documents from multiple sources or using different algorithms.

Borgeaud et al. 2022 suggest scaling the size of the text database for retrieval as a complementary path to scaling language models. There is a pre-collected text database with a total of over 5 trillion tokens. These chunks are stored in the form of key-value pairs, with each chunk as a unit, and similarity retrieval is performed on kk-nearest neighbours from key-value database using the L2L_{2} distance on Bert embeddings. The input sequence is splited into chunks, Retrieval Transformer (Retro) model retrieves text similar to the previous chunk to improve the predictions in the current chunk. The model calculate cross-attention between the input text and the retrieved text chunks to generate better answers. With only 25× fewer parameters of GPT-3, its performance on Pile is quite comparable.

Lazaridou et al. 2022 present a method that capitalizes on the few-shot capabilities of large-scale language models to enhance their grounding in factual and current information. Drawing from semi-parametric language models, the approach conditions LMs using few-shot prompts based on data sourced from Google Search. For any given query, the method retrieves pertinent documents from the web, extracts the top 20 URLs, and processes them to obtain clear text. These documents are segmented into paragraphs, and the most relevant ones are chosen using TF-IDF based on their similarity to the query. The LMs are then conditioned using few-shot prompts that incorporate the retrieved paragraphs. This k-shot prompting technique is augmented with an evidence paragraph, creating a prompt structure that encompasses evidence, query, and response. The method also involves generating multiple answers from the model and reranking them using different probabilistic factorizations. The experimental results indicate that by conditioning on retrieved evidence, the 7B Gopher LM (et al. [n. d.]) surpassed the performance of the 280B Gopher LM, with relative improvements reaching up to 30% on the NQ (Kwiatkowski et al. 2019) dataset.

2.2. Interactive Retrieval

While retrieval systems are designed to source relevant information, they may occasionally fail to retrieve accurate or comprehensive data. Additionally, LLMs might struggle to recognize, or even be misled by, the retrieved content, as detailed in Sec 4.2.2. Implementing an interactive retrieval mechanism can address these challenges, aiding in sourcing more appropriate information and guiding the LLM towards improved content generation. In this subsubsection, we explore methods that employ the Chain-of-Thoughts and Agents mechanisms to achieve effective interactive retrieval.

In recent studies, there is a growing interest in integrating Chain-of-Thoughts (Wei et al. 2022) steps into query retrieval. He et al. 2022 introduce a method that generates multiple reasoning paths and their corresponding predictions for each query. This process involves retrieving relevant knowledge from external sources like Wikidata (Vrandečić and Krötzsch 2014), WordNet (Miller 1992), and ConceptNet (Speer et al. 2017). The faithfulness of each reasoning path determines based on entailment scores, contradiction scores, and MPNet similarities (Song et al. 2020) with the retrieved knowledge. The prediction with the highest faithfulness score is chosen as the final result. This method demonstrates superior performance in tasks such as commonsense reasoning (Geva et al. 2021b), temporal reasoning (Jia et al. 2018), and tabular reasoning (Gupta et al. 2020b), outperforming the baseline CoT reasoning and self-consistency methods (Wang et al. 2023g). On a related note, Trivedi et al. 2023 propose IRCoT, an innovative retrieval technique that interweaves the CoT process. In this approach, each generated sentence during the CoT combines with the question to form a retrieval query. The subsequent reasoning step then produces by the Language Model using both the retrieval results and the prior reasoning. This interleaving method finds to enhance the performance of both retrieval and CoT in Open-domain QA. Experiments proves that this is beneficial for models of different sizes, including GPT-3(175B) (Brown et al. 2020) and Flan-T5 (Chung et al. 2022) families.

FLARE (Jiang et al. 2023a) is a dynamic solution to address the limitations of previous RAG works which either retrieve only once at the onset of generation or do so based on fixed intervals. Single retrievals are insufficient for long-form generation due to the evolving information needs during the process, and fixed intervals for previous token queries can be inappropriate, FLARE determines "when and what to retrieve" dynamically. The decision of "when" is based on whether the current sentence contains a token with a generation probability below a set threshold. If it doesn’t, the sentence is accepted and moves to the next generation step; otherwise, retrieval augmented generation occurs. For the "what", the current sentence is used as a query. To address the challenge of low-probability tokens affecting retrieval accuracy, two solutions are proposed: masking low-probability tokens, and using LLM for generation of these tokens as queries. Testing on tasks like Multihop QA (Ho et al. 2020), Commonsense Reasoning (Geva et al. 2021a), Long-form QA (Stelmakh et al. 2022), and Open-domain Summarization (Hayashi et al. 2021) using GPT3.5, results show FLARE outperforms baselines, with both query generation methods showing comparable performance.

Using an LLM-based agent framework that leverages external knowledge APIs as tools or requesting such APIs as actions.

Yao et al. 2023b present a new framework named ReAct that integrates Chain-of-Thoughts reasoning with action. Through in-context learning, the LLM’s CoT output is transformed into descriptions of reasoning processes and action behaviors. Subsequently, these action descriptions are standardized and executed, with the results being incorporated into the next prompt. In terms of results, for tasks like Fact checking and QA, while CoT has 14% of its correct answers containing incorrect reasoning steps or facts, ReAct only has 6%. In the incorrect answers, 56% of CoT’s errors were due to errors in reasoning steps or facts, whereas ReAct has no factual errors. However, ReAct have 23% of its errors resulting from search errors.

Shinn et al. 2023 propose a prompt engineering framework named Reflextion to enable LLMs to reflect on and correct previous errors. They use linguistic feedback to strengthen an agent’s actions instead of adjusting model weights. Specifically, the Reflextion Agent, an LLM, first interacts with its environment by generating an "Action" via ReAct (Yao et al. 2023b) or Chain-og-Thoughts (Wei et al. 2022) in few-shot scenarios, which results in an "Observation". This "Observation", whether a reward, error message, or natural language feedback, provides insights on the Agent’s current "Action". When the Agent receives a failure signal, it triggers a self-reflextion mechanism, utilizing the LLM, to summarize the reasons for the failure into its Memory module, creating a "long-term memory". On subsequent generations, the Agent can review all past reflextion memories to prevent mistakes. Experimental findings indicate that Reflextion achieves a 10-20% performance increase over the baseline methods ReACT and CoT on datasets like AlfWorld (Shridhar et al. 2021), HotPotQA (Yang et al. 2018), and HumanEval (Chen et al. 2021b).

Varshney et al. 2023 introduce a comprehensive framework aimed at reducing factual inaccuracies. They use models to recognize entities and generate questions, we recognize these models as tools for the LLM-based agent. During the generation process, pivotal concepts encompassing names, geographical locales, and temporal references are ascertained within the contextual sentence employing entity extraction, keyword distillation, or directives to the LLM. The logit output values corresponding to these discerned concepts act as surrogates for confidence estimates. Should these values fall beneath a predetermined threshold, the mechanism then fetches a pertinent document to corroborate the generated information. The query methodology employed for such a retrieval hinges on posing a binary (Yes/No) query to the LLM. In scenarios where the validation is unsuccessful, the framework directs the model to rectify the erroneous output, either by omission or by substitution, drawing upon the knowledge from the consulted document. Empirical evaluations, specifically in the domain of article generation, underscore the efficacy of the delineated approach. Notably, factual error rates exhibited by GPT-3 witnessed a substantial decline, from 47.5% to a mere 14.5%, when subjected to this methodology. The diagnostic facet of their approach manifests an 80% recall, and the rectification mechanism adeptly rectifies 57.6% of the factually incorrect outputs that were accurately pinpointed.

Self-RAG (Asai et al. 2023) builds upon the Retrieval-Augmented Generation (RAG) framework by incorporating a self-reflective approach. This includes on-demand retrieval, prompting the LLM to consider whether the retrieved documents are relevant and supportive of the argument, thereby enhancing the factuality of the LLM’s output. Self-RAG achieves this by having the model produce special tokens (i.e., reflection tokens) during the output process. The authors employ an end-to-end training method that enables the model to generate these special tokens. At the inference level, Self-RAG decodes a retrieval token for each input and each segment of the previously generated content. If the token is ’no,’ the model proceeds with the normal output; if ’yes,’ it performs a retrieval. The input, previous output, and each retrieved document are then fed into another session of the model, which assesses the relevance of each document. If a document is relevant, the model further evaluates whether it is supportive and should be used in generation. Based on this document, the model produces a segment and applies soft and hard constraints using the generated reflection token to select the most appropriate document. Self-RAG then incorporates the most suitable document into the continued generation of content, providing citations as needed. This process repeats until the entire content is generated. Self-RAG requires training data, which is annotated by GPT-4. Experiments based on the Llama-2 model on Short-form QA (PopQA (Mallen et al. 2023), TQA (Joshi et al. 2017)), Closed-set QA (Pub (Zhang et al. 2023e), ARC (Clark et al. 2018)), and Long-form NLG with Citations (Bio Generation (Min et al. 2023), ALCE-ASQA (Gao et al. 2023b)) have demonstrated its effectiveness.

2.3. Retrieval Adaptation

Recent research (Wang et al. 2023b; Ren et al. 2023) has highlighted that merely using the retrieved information in LLMs doesn’t always enhance their ability to answer factual questions. This underscores the importance of enabling LLMs to better adapt to the retrieved data to produce more accurate content. In this section, we delve into various strategies that facilitate this adaptation. Specifically, we explore three methodological approaches: prompt-based methods, SFT-based methods, and RLHF-based methods.

Leveraging prompts to navigate the retrieval process, ensuring the extraction of pertinent and factual data.

Peng et al. 2023 introduce the LLM-Augmenter system, a system using a fixed LLM combined with a plug-and-play retrieval module to help the LLM perform better in tasks that are particularly sensitive to factual errors. This system enhances its performance by enabling the LLM to use a series of modules (e.g. allowing the LLM to interact with external knowledge) to assist the LLM in generating results grounded in evidence. And they use automated feedback generated by utility functions (e.g., the factuality score of a LLM-generated response) to modify LLM’s candidate response options. The author evaluates the system’s performance on information-seeking dialog (Galley et al. 2019b) and Wiki QA (Chen et al. 2021a) and experiments show that the system can significantly reduce ChatGPT’s errors without sacrificing the fluency and informativeness of the generated content.

Optimizing the LLM or retrieval system through training to enhance the alignment between generation tasks and the retrieved content.

Izacard et al. 2022 introduce a comprehensive architecture named ATLAS which is composed of the Contriever (Izacard et al. 2021) retriever of the dual-encoder architecture and the T5 (Raffel et al. 2020a) language model with Fusion-in-Decoder (Izacard and Grave 2021b). The training objectives for the retriever consist of four components: Attention Distillation (Izacard and Grave 2021a), where the retriever is trained on the average attention scores from the language model for each article; End-to-end training of Multi-Document Reader and Retriever (EMDR2) (Singh et al. 2021), which involves using the query and the top-K retrieved articles from the current retriever as input and loss computation against the standard answers to train the retriever; Perplexity Distillation (PDist), where the retriever is trained to predict how much the perplexity of standard answers would improve for each document; Leave-one-out Perplexity Distillation (LOOP), which trains the retriever to predict how much worse the prediction of the language model gets when removing a document from the top-K results. The LM’s training objectives consist of three parts: prefix language modeling, masked language modeling, and title to section generation. Besides, they optimize and accelerate retriever training using techniques such as full index update, Re-ranking, and Query-side fine-tuning. Atlas achieves notable accuracy on Natural Questions using only 64 examples, outperforming a 540B model with 50x fewer parameters.

Shi et al. 2023b introduce REPLUG, a retrieval-augmented framework that considers the LLM as a black box, freezes its parameters and tunes retrieval models with supervision signals using language modeling scores. In this framework, the input context and the document are encoded through the dual encoder architecture, cosine similarity is then calculated to retrieve related documents. The likelihood for each retrieved document and the language model scores are computed. Then they can update the retrieval model parameters by minimizing the KL divergence between retrieved document likelihood and the language model"s score distribution. Ablation experiments demonstrate that this method significantly improves the performance of the original language models and the improvements are not coming from ensembling random documents.

Luo et al. 2023a focus on utilizing instruction-tuning to denoise the retrieval results. They gather retrieval outcomes from various search APIs and domains, leading to the creation of a new search-grounded dataset. This dataset encompasses instructions, grounding information, and responses. Notably, it includes both pertinent results and those that are unrelated or disputed. The model needs to learn to ground on useful search results. After fine-tuning the LLaMa-7B model on this dataset, the resulting model, named SAIL-7B, exhibits superior performance in transparency-sensitive tasks such as open-ended QA and fact-checking.

Menick et al. 2022a use reinforcement learning from human preferences (RLHP) to train a 280 billion parameter model named GopherCite that generate answers along with high quality supporting evidence. They firstly collect data from existing models and have it rated by humans. The data is used for fine-tuning and reward model training. A supervised fine-tuning model is trained to produce accurate quotes with proper syntax. A reward model is created to rank model outputs based on overall quality. Finally, a reinforcement learning policy is optimized to align model behavior with human preferences, improving quoting performance. The model may decline to answer when the reward model score is too low. According to human evaluation, the model achieves better supported and plausible rating on the subset of Natural Questions dataset (Kwiatkowski et al. 2019) than the previous SOTA (FiD-DPR) (Izacard and Grave 2021b).

2.4. Retrieval on External Memory

Currently, most LLMs enhance their factuality by retrieving knowledge in the form of text snippets from external storage and incorporating them into the context. Some researchers are exploring the storage of knowledge in non-textual forms and integrating this knowledge into models through specialized methods.

Li et al. 2022a store knowledge in the form of key-value pairs in memory. The key is obtained by encoding knowledge using a Doc Retrieval Embedder, while the value is encoded using a Transformer encoder. Similar to traditional retrieval-based LLMs, the model encodes the input using a Query retrieval embedder and retrieves knowledge from memory. The retrieved value is then integrated into the model’s multi-head attention layer through cross-attention to enhance factuality.

G-MAP (Wan et al. 2022) does not explicitly store knowledge in storage but uses a general domain PLM as external memory. To mitigate catastrophic forgetting during adaptive pretraining, G-MAP introduces a frozen-parameter general domain PLM (PLM-G) during the fine-tuning of the domain-specific PLM (PLM-D). During fine-tuning, the input is provided to both PLM-G and PLM-D. The hidden states from each layer of PLM-G are stored in a cache, and a Memory-Augmented Strategy is used to extract hidden states from certain layers, which are then concatenated and integrated into PLM-D’s Memory-Augmented Layer. The study also compared four Memory-Augmented Strategies, with Chunk-based Gated Memory Transfer performing the best.

Fine-tuning methods and some model editing methods store new knowledge in new model parameters through continual pretraining. The difference is that fine-tuning methods store many pieces of knowledge in a matrix parameter, while model editing establishes a new neuron for each piece of knowledge.

Houlsby et al. 2019 propose to add Adapter modules to fine-tune pre-trained deep learning models on a new task with minimal changes to the original model by inserting small, trainable modules between existing layers. During fine-tuning, the main body of the pre-trained model is frozen, and the Adapter module learns knowledge specific to downstream tasks. The Adapter method reduces the computational requirements for model fine-tuning while enhancing the model’s factuality for specific domains. However, the addition of the Adapter module also increases the overall parameter count of the model, somewhat reducing the model’s inference performance.

To enhance the model’s understanding of entities and thereby improve factuality, KALA (Kang et al. 2022), EaE (Févry et al. 2020), and Mention Memory (de Jong et al. 2022) store encoded entities in external memory. During generation, the retrieved entity embeddings are integrated into the model layers.

Kang et al. 2022 introduced KALA to reduce overhead and catastrophic forgetting during Adaptive Training. KALA not only establishes a memory for entities and their encodings but also uses a KG to store relationships between these entities. For a given mention in the input, the corresponding entity is first determined. Based on the KG, the encoding of this entity and its neighboring entities is retrieved from memory. Through GNN weighted aggregation, the encoding for this mention’s corresponding entity is obtained. Finally, the Knowledge-conditioned Feature Modulation (KFM) is introduced in the model layer, integrating the encoding result into the representation of all tokens involved in the mention.

Févry et al. 2020 introduce mention detection, entity linking, and MLM during model training. The model queries the top 100 entity embeddings from the entity storage that is closest to the current mention and integrates them using attention.

de Jong et al. 2022’s TOME model is an improvement over the EaE model. Instead of storing entity embeddings in memory, TOME stores the embeddings of entity mentions. For marked entity mentions in the input, TOME retrieves all related entity mention embeddings from memory and integrates them into the model through a memory attention layer.

Similar to the three methods mentioned above, knowledge plugin (Zhang et al. 2023h) also introduce entity-related knowledge. However, instead of integrating the knowledge directly into the model layers, they utilize a pre-trained mapping network. This network maps the entity embeddings to the token embedding space of the Pre-trained Language Model (PLM). Ultimately, the mapped entity embeddings are injected at the input embedding level, facilitating the knowledge insertion process.

Xue et al. 2023 address the challenge of improving factual consistency in knowledge-grounded dialogue systems by introducing extended feed-forward networks (FFNs) and leveraging reinforcement learning, resulting in more accurate and reliable responses, as demonstrated on the WoW (Dinan et al. 2019) and CMU_DoG (Zhou et al. 2018) datasets.

To further tackle the factual correction task (Thorne and Vlachos 2021), Gao et al. 2023a explore the integration of LLMs, such as GPT-3, with search engines to improve their precision and memory. The goal is to use search engines to search for evidence and correct sentences generated by LLMs. The proposed method RARR is that, for each input sentence, a set of questions is generated, and web pages are searched to verify the consistency of information with the input sentence. The paper evaluates the modifications based on attribution and preservation criteria, with both manual and automatic verification methods. The primary evaluation metric is F1, considering both attribution and preservation aspects to assess the effectiveness of the approach in enhancing LLM-generated sentences.

Chen et al. 2023e is a follow-up work to (Gao et al. 2023a) and (Thorne and Vlachos 2021). Similar to EFEC, this paper fine-tunes a T5 model to serve as an editor, but it introduces negative samples during fine-tuning. The model PURR is trained to take user questions, perform Google searches to retrieve the top 5 web page summaries (used as positive samples), and generate noise by replacing some content in these positive samples using the language model. A sequence-to-sequence model is then trained to correct the noisy sentences back to their correct versions. This approach differs from EFEC, which used a mask-and-fill approach. PURR represents an improvement over EFEC, focusing on directly training a language model to edit incorrect sentences into correct ones using Google search for generating positive samples, ultimately leading to increased F1 scores.

2.5. Retrieval on Structured Knowledge Source

We discuss studies that retrieves on structured repositories, such as knowledge graphs and databse, to source factual data during generation in this subsubsection.

Zhang et al. 2023f utilize knowledge graphs (KG) for retrieval to tackle factual errors. They observe that there can be inconsistencies between a user’s request and the content in the KG. For instance, when a user mentions a full name, the KG might only have its abbreviation, leading to imperfect retrieval results. To rectify this, they propose a method to rephrase the user’s request. Their approach involves generating an SQL query based on the user’s input and the database metadata using a LLM. And then they query the database and ask the LLM to identify which entity in the sentence corresponds to an entity in the database, thereby creating a mapping. Using the entity names from the database, the LLM is prompted to reformulate the question. If a database query results in multiple rows for a selected column and item, a new question is generated using a greedy approach, prompting the user for more specific details until a conclusive answer is reached. Experiments show that this method exhibits notable enhancements in comparison to contemporary state-of-the-art techniques in mitigating language model inaccuracies.

StructGPT (Jiang et al. 2023b) is a general prompt framework to support LLMs reasoning on structured data (e.g., KG, Table, and Database). In general, the core of this framework is that they construct the specialized interfaces to collect relevant evidence from structured data (i.e., reading), and let LLMs concentrate on the reasoning task based on the collected information (i.e., reasoning). Specially, they propose an invoking-linearization-generation procedure to support LLMs in reasoning on the structured data with the help of the interfaces. By iterating this procedure with provided interfaces, our approach can gradually approach the target answers to a given query. Experiments conducted on three types of structured data, including KGQA, TableQA, and Text-to-SQL, show that StructGPT greatly improves the performance of LLMs, under the few-shot or zero-shot settings.

Baek et al. 2023 propose to inject the factual knowledge from knowledge graphs into (large) language models (up to GPT-3.5), by retrieving the relevant facts from knowledge graphs based on their textual similarities with the input question and then injecting them as the prompt of language models. This approach improves the performance of language models on knowledge graph question answering tasks (Sen et al. 2022) by up to 48% on average, compared to baselines without knowledge graphs.

3. Domain Factuality Enhanced LLMs

Domain Knowledge Deficit is not only an important reason for limiting the application of LLM in specific fields, but also a subject of great concern to both academia and industry. In this subsection, we discuss how those Domain-Specific LLMs enhance their domain factuality.

Table 9 lists the domain-factuality enhanced LLMs. Here, we include several domains, including healthcare/medicine (H), finance (F), law/legal (L), geoscience/environment (G), education (E), food testing (FT), and home renovation (HR).

Based on the actual scenarios of Domain-Specific LLMs and our previous categorization of enhancement methods, we have summarized several commonly used enhancement techniques for Domain-Specific LLMs:

(1) Continual Pretraining: A method that involves continuously updating and fine-tuning a pre-trained language model using domain-specific data. This process ensures that the model stays up-to-date and relevant within a specific domain or field. It starts with an initial pre-trained model, often a general-purpose language model, and then fine-tunes it using domain-specific text or data. As new information becomes available, the model can be further fine-tuned to adapt to the evolving knowledge landscape. Continual pretraining is a powerful approach for maintaining the accuracy and relevance of AI models in rapidly changing domains, such as technology or medicine (Zhang et al. 2023a; Yang et al. 2023c).

(2) Continual SFT: Another strategy for enhancing the factuality of AI models. In this approach, the model is fine-tuned using labeled or annotated data specific to the domain of interest. This fine-tuning process allows the model to learn and adapt to the nuances and specifics of the domain, improving its ability to provide accurate and contextually relevant information. It can be particularly useful in applications where access to domain-specific labeled data is available over time, such as in the case of legal databases, medical records, or financial reports (Bao et al. 2023; Li et al. 2023b).

(3) Train From Scratch: It involves starting the learning process with minimal prior knowledge or pretraining. This approach can be likened to teaching a machine learning model with a blank slate. While it may not have the advantage of leveraging pre-existing knowledge, training from scratch can be advantageous when dealing with completely new domains or tasks where there is limited relevant data available. It allows the model to build its understanding from the ground up, although it may require substantial computational resources and time (Ross et al. 2022; Venigalla et al. 2022).

(4) External knowledge: involves augmenting a language model’s internal knowledge with information from external sources. This method allows the model to access databases, websites, or other structured data repositories to verify facts or gather additional information when responding to user queries. By integrating external knowledge, the model can enhance its fact-checking capabilities and provide more accurate and contextually relevant answers, especially when dealing with dynamic or rapidly changing information. Below, we introduce these methods (Wang et al. 2023f; Fan et al. 2023).

For each Domain-specific LLM, we list its respective enhancement methods, which are presented in Table 9.

4. Healthcare domain-enhanced LLMs

These LLMs have emerged as powerful tools in the medical field, offering a diverse range of capabilities. These models, such as CohortGPT(Guan et al. 2023), ChatDoctor(Li et al. 2023b), DeID-GPTLiu et al. 2023e, BioMedLMVenigalla et al. 2022, DoctorGLMXiong et al. 2023a, MedChatZHTan et al. 2023a, BioGPT(Luo et al. 2022), GeneGPT(Jin et al. 2023), Almanac(Zakka et al. 2023), and MolXPT(Liu et al. 2023f), harness the potential of LLMs to revolutionize healthcare. They are equipped with features like classifying unstructured medical text into disease labels, improving performance with knowledge graphs and sample selection strategies, fine-tuning on large datasets of patient-doctor dialogues, enabling automatic medical text de-identification, excelling in medical question-answering tasks, handling traditional Chinese medical question-answering, and outperforming in various biomedical NLP tasks. Some models interact with web APIs for genomics questions, while others specialize in clinical guidelines and treatment recommendations. These LLMs not only demonstrate state-of-the-art performance on healthcare domain but also emphasize the importance of domain-specific training and evaluation, showcasing their potential in transforming healthcare and clinical decision-making.

Zhang et al. 2023a present HuatuoGPT, a medical language model that uses data from ChatGPT and doctors, resulting in state-of-the-art performance in medical consultations. It is based on Baichuan-7B and Ziya-LLaMA-13B-Pretrain-v1, continually pre-trained on both distilled data (from ChatGPT) and real-world data (from Doctors).

Likewise, Yang et al. 2023c introduce Zhongjing, the first Chinese medical language model based on LLaMA, which utilizes a comprehensive training pipeline and a multi-turn medical dialogue dataset. Specifically, it is enhanced with a multi-turn medical dialogue dataset called CMtMedQA, consisting of 70,000 authentic doctor-patient dialogues, enabling complex dialogue and proactive inquiry. The backbone model used is Ziya-LLaMA-13B-v1, and the evaluation dataset is CMtMedQA and huatuo-26M (Li et al. 2023f).

Wang et al. 2023f unveil a system, LLM-AMT, that improves large-scale language models like GPT-3.5-Turbo and LLaMA-2-13B with medical textbooks, notably enhancing open-domain medical question-answering tasks. At the same time, the external knowledge source is a Hybrid Textbook Retriever comprising 51 textbooks from the MedQA dataset and Wikipedia.

Bao et al. 2023 present DISC-MedLLM, a solution that uses LLMs to provide accurate medical responses in conversational healthcare services, utilizing strategies like medical knowledge graphs, real-world dialogue reconstruction, and human-guided preference rephrasing to create high-quality SFT datasets, applied on Baichuan-13B-Base. The paper uses various datasets for fine-tuning, including Re-constructed AI Doctor-Patient Dialogue, MedDialog, cMedQA, Knowledge Graph QA pairs (CMeKG), Behavioral Preference Dataset (Manual selection), MedMCQA, MOSS6, and Alpaca-GPT.

Similarly, Guan et al. 2023 introduce CohortGPT, a model that uses LLMs for participant recruitment in clinical research by classifying complex medical text into disease labels. CohortGPT enhances ChatGPT performance with the use of a knowledge graph as auxiliary information and a CoT sample selection strategy. The tasks involve IU-RR (Preparing a collection of radiology examinations for distribution and retrieval) and MIMIC-CXR (Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports). Li et al. 2023b introduce ChatDoctor, a refined LLaMA-based model. It is fine-tuned using a dataset of 100,000 patient-doctor dialogues, and equipped with a self-directed information retrieval mechanism.

Liu et al. 2023e introduce DeID-GPT, a framework leveraging GPT-4 for automatic medical text de-identification. Additionally, the paper mentions the use of HIPAA identifiers as an extra knowledge source to enhance the de-identification process.

Venigalla et al. 2022 present BioMedLM, a domain-specific LLM trained on PubMed data, for medical QA tasks. Xiong et al. 2023b introduce DoctorGLM, a Chinese-focused language model fine-tuned for healthcare-specific tasks. Both Tan et al. 2023a and Luo et al. 2022 introduce dialogue and generative Transformer language models for traditional Chinese medical question-answering and biomedical NLP tasks, respectively.

Jin et al. 2023 present GeneGPT, a method for teaching LLMs to answer genomics questions using NCBI Web APIs. Zakka et al. 2023 introduce Almanac, a LLM with retrieval capabilities for medical guideline recommendations. Lastly, Liu et al. 2023f introduce MolXPT, a unified language model adept in molecular property prediction and molecular generation.

5. Legal domain enhanced LLMs

These LLMs, such as LawGPT (Nguyen 2023), and ChatLaw (Cui et al. 2023b), have been fine-tuned to provide comprehensive legal assistance, from answering intricate legal queries and generating legal documents to offering expert legal advice. Leveraging extensive corpora of legal text, these models ensure context-aware and accurate responses. Moreover, their continual development involves injecting domain knowledge, designing supervised fine-tuning tasks, and incorporating retrieval modules to address issues like hallucination and ensure high-quality legal assistance. These innovations not only pave the way for more accessible and reliable legal services but also open up new avenues for research and exploration within the legal domain.

Nguyen 2023 introduce LawGPT 1.0, a fine-tuned GPT-3 language model for the legal domain to provide conversational legal assistance, including answering legal questions, generating legal documents, and offering legal advices. The paper mentions the use of a large corpus of legal text for fine-tuning the model to adapt it to the legal domain.

Savelka et al. 2023 evaluate the performance of GPT-4 in generating explanations of legal terms in legislation, comparing a baseline approach to an augmented approach that uses a legal information retrieval module to provide context from case law, revealing improvements in quality and addressing issues of factual accuracy and hallucination.

Huang et al. 2023c address the challenge of enhancing LLMs like LLaMA for domain-specific tasks, particularly in the legal domain, by injecting domain knowledge during continual training, designing appropriate supervised finetune tasks and incorporating a retrieval module to improve factuality during text generation. They release their data and model for further research in Chinese legal tasks.

Cui et al. 2023a introduce ChatLaw, an open-source legal LLM, designed for the Chinese legal domain. The paper introduces a method to improve model factuality during data screening, and a self-attention method for error handling. The paper uses various datasets for fine-tuning ChatLaw, including a collection of original legal data, data constructed based on legal regulations and judicial interpretations, and crawled real legal consultation data. The primary model used in this paper is Ziya-LLaMA-13B, which serves as the backbone for ChatLaw, tailored for the Chinese legal domain and optimized to handle legal questions and tasks. Additionally, the paper uses of a vector database retrieval method, keyword retrieval, and a self-attention method to enhance the model’s performance in the legal domain.

6. Finance Domain-enhanced LLMs

These LLMs combine sophisticated language models designed specifically for commercial and financial tasks to deliver robust processing capabilities. They focus on creating tailor-made solutions optimized for both financial text analysis and e-commerce settings, trained on datasets containing myriad business-related tasks and copious financial tokens. They are designed to perform a plethora of functions, ranging from understanding and generating instructions for various E-commerce assignments to identifying sentiment, recognizing named entities, and answering questions in financial contexts. The models are further fine-tuned for zero-shot generalization on diverse tasks and benchmarks.

Li et al. 2023c introduce EcomGPT, a language model tailored for E-commerce scenarios, trained on the newly created EcomInstruct dataset, which consists of 2.5 million instruction data spanning various E-commerce tasks and data types. The dataset covers product information, user reviews, and more. It defines atomic tasks and Chain-of-Task tasks to enable comprehensive training for E-commerce scenarios. The backbone model used is BLOOMZ, which is fine-tuned with the EcomInstruct dataset. The evaluation dataset includes 12 tasks, encompassing classification, generation, extraction, and other E-commerce-related tasks.

Wu et al. 2023 introduce BloombergGPT, a specialized 50 billion-parameter language model for the financial domain, trained on a massive 363 billion token dataset, which combines Bloomberg’s extensive financial data sources with general-purpose datasets. The dataset used in this paper is an extensive 363 billion token dataset, which includes a significant portion of financial data from Bloomberg’s sources (51.27% of the training data). BloombergGPT is based on a decoder-only causal language model architecture known as BLOOM. The evaluation includes various financial NLP tasks such as sentiment analysis, named entity recognition, binary classification, and question answering.

7. Other Domain-Enhanced LLMs

are expertly designed, leveraging vast corpora to provide precise and robust results pertaining to geoscience and renewable energy. K2, a trailblazer in geoscience LLM, was trained on a massive geoscience text corpus and further refined using the GeoSignal dataset. Meanwhile, the HouYi model, another pioneering LLM focusing on renewable energy, harnessed the Renewable Energy Academic Paper dataset, containing over a million academic literature sources. These LLMs are fine-tuned to deliver adept performance in their respective fields, showing substantial capabilities in aligning their responses with user queries and renewable energy academic literature.

Deng et al. 2023 introduce K2, the first LLM designed specifically for geoscience, which is a LLaMA-7B continuously trained on a 5.5 billion token geoscience text corpus and fine-tuned using the GeoSignal dataset. The paper also presents resources like GeoSignal, a geoscience instruction tuning dataset, and GeoBench, the first geoscience benchmark for evaluating LLMs in the context of geoscience.

Bai et al. 2023 present the development of the HouYi model, the first LLM specifically designed for renewable energy, utilizing the newly created Renewable Energy Academic Paper (REAP) dataset, which contains over 1.1 million academic literature sources related to renewable energy, and the HouYi model is fine-tuned based on general LLMs such as ChatGLM-6B.

Bi et al. 2023 present OceanGPT, the first-ever LLM in the ocean domain, which is expert in various ocean science tasks. They also propose a novel framework called DoInstruct to automatically obtain a large volume of ocean domain data. OceanGPT is evaluated in OceanBench and shows a higher level of knowledge expertise for oceans science tasks.

are used for assisting education scenarios. An example is GrammarGPT (Fan et al. 2023), which provides an innovative approach to language learning, particularly focusing on error correction in Chinese grammar. It is an open-source LLM designed for native Chinese grammatical error correction, which leverages a hybrid dataset of ChatGPT-generated and human-annotated data, along with heuristic methods to guide the model in generating ungrammatical sentences. The backbone model used is phoenix-inst-chat-7b.

are language models specifically designed to meet the distinct requirements of food testing protocols. For example, Qi et al. 2023 introduce FoodGPT, a LLM for food testing that incorporates structured knowledge and scanned documents using an incremental pre-training approach, with a focus on addressing machine hallucination by constructing a knowledge graph as an external knowledge base, utilizing the Chinese-LLaMA2-13B as the backbone model and collecting food-related data for training.

are domain-specific language models tailored for home renovation tasks. For example, Wen et al. 2023 introduce ChatHome, which uses a dual-pronged methodology involving domain-adaptive pretraining and instruction-tuning on an extensive dataset comprising professional articles, standard documents, and web content relevant to home renovation. The backbone model is Baichuan-13B, and the evaluation datasets include C-Eval, CMMLU, and the newly created "EvalHome" domain dataset, while the fine-tuning data sources encompass National Standards, Domain Books, Domain Websites, and WuDaoCorpora.

Conclusion

Throughout this survey, we have systematically explored the intricate landscape of factuality issues within large language models (LLMs). We began by defining the concept of factuality (Sec 2.2) and proceeded to discuss its broader implications (Sec 2.3). Our journey took us through the multifaceted realm of factuality evaluation, encompassing benchmarks (Sec 3.2), metrics (Sec 3.1), specific evaluation studies (Sec 3.3), and domain-specific evaluations (Sec 3.4). We then delved deeper, probing the intrinsic mechanisms that underpin factuality in LLMs (Sec 4). Our exploration culminated in the discussion of enhancement techniques, both for standalone LLMs (Sec 5.1) and retrieval-augmented LLMs (Sec 5.2), with a special focus on domain-specific LLM enhancements (Sec 5.3).

Despite the advancements detailed in this survey, several challenges loom large. The evaluation of factuality remains an intricate puzzle, complicated by the inherent variability and nuances of natural languages. The core processes governing how LLMs store, update, and produce facts are yet not fully revealed. And while certain techniques, like continual training and retrieval, show promise, they are not without limitations. Looking ahead, the quest for fully factual LLMs presents both challenges and opportunities. Future research might delve deeper into understanding the neural architectures of LLMs, develop more robust evaluation metrics, and innovate on enhancement techniques. As LLMs become increasingly integrated into our digital ecosystem, ensuring their factual reliability will remain paramount, with implications that impact across the AI community and beyond.

References