A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation
Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, Dong Yu
Introduction
Recently developed large language models such as GPT-3 Brown et al. (2020), InstructGPT Ouyang et al. (2022), PaLM Chowdhery et al. (2022), LLaMA Touvron et al. (2023), and several others Taori et al. (2023); Scao et al. (2022); Wei et al. (2022a); Wang et al. (2022) have achieved remarkable performance on a wide range of language understanding tasks. Furthermore, they have been shown to possess an impressive ability to generate fluent and coherent text. Despite all these abilities, their tendency to ‘hallucinate’ critically hampers their reliability and limits their widespread adoption in real-world applications.
Hallucination in the context of language models refers to the generation of text or responses that seem syntactically sound, fluent, and natural but are factually incorrect, nonsensical, or unfaithful to the provided source input Maynez et al. (2020); Holtzman et al. (2020); Ji et al. (2023); Koehn and Knowles (2017). These hallucinations can lead to serious consequences such as spreading of misinformation and violation of privacy. Thus, in this work, we focus on the crucial problem of ‘addressing’ large language models’ hallucinations.
We propose to actively ‘detect’ and ‘mitigate’ hallucinations during the generation process. This is crucial as we show that when a generated sentence is hallucinated, the chances of hallucination in the subsequently generated sentences increase. Thus, actively detecting and mitigating hallucinations is also important to prevent the propagation of hallucinations in the subsequently generated sentences. We divide our approach into two stages, Detection and Mitigation.
In the hallucination detection stage, we first identify the candidates of potential hallucination, i.e., the key ‘concepts’ of the generated sentence. Next, leveraging the logit output values of the model, we calculate model’s ‘uncertainty’ on the identified concepts. We demonstrate that this uncertainty provides a signal for hallucination. However, we note that this is an additional signal and not a necessary requirement for our approach. Then, we check the correctness of the ‘uncertain’ concepts through a validation procedure where we: (a) create a query that tests the correctness of the information pertaining to the concept, (b) retrieve knowledge relevant to the validation question, (c) answer the validation question leveraging the retrieved knowledge, and verify the corresponding information in the generated sentence to detect hallucinations.
This is followed by the hallucination mitigation stage in which we ‘repair’ the potentially hallucinated sentence using the retrieved knowledge as evidence. Figure 2 illustrates the key steps of our approach. Furthermore, we conduct a systematic and wide study exploring multiple techniques to achieve the objective of each of the steps. Importantly, we show that simply instructing the model achieves the corresponding objectives of these steps.
We design an experimental setup where we prompt the model to write about topics from diverse domains such as sports, politics, music, literature, etc. Then, we annotate the correctness of the first five generated sentences for each topic. We first demonstrate the individual efficacy of our detection and mitigation techniques. Specifically, the detection technique achieves a recall of and the mitigation technique successfully mitigates of the correctly detected hallucinations. Importantly, our mitigation technique does not introduce new hallucinations even in the case of incorrectly detected hallucinations, i.e., false positives. Then, we show that the proposed active detection and mitigation approach successfully reduces the hallucinations of the GPT-3.5 (text-davinci-003) model from to on average (Figure 1). We further demonstrate the effectiveness and wide applicability of our approach in addressing hallucinations through three additional studies: (1) Using another LLM from a different model family (Vicuna-13B), (2) Adapting the approach to answer multi-hop questions, and (3) Assessing it on the ‘false premise questions’.
Approach
We propose to actively detect hallucinations and mitigate them during the generation process. This is crucial as we show that a generated sentence is hallucinated more often when the model has already hallucinated in its previously generated sentences for the input (Section 3.1.1). Similarly, a generated sentence is relatively less often hallucinated when the model has not hallucinated in its previously generated sentences. Thus, actively detecting hallucinations and mitigating them is also important to prevent the propagation of further hallucinations in subsequently generated sentences. To this end, we iteratively generate sentences through the model and actively detect and mitigate hallucinations. Figure 2 illustrates the key steps of our approach.
In section 2.2, we detail the steps of our hallucination detection approach, i.e., identifying the important ‘concepts’ of the generated sentence, i.e., the candidates of potential hallucination (2.2.1), calculating model’s uncertainty on the concepts using the logit output values (2.2.2), and checking the correctness by creating validation query (2.2.3), finding relevant knowledge (2.2.4), and verifying information leveraging the retrieved knowledge (2.2.5). We describe various techniques to achieve the objective of each of these steps and also elaborate on several important points such as using a ‘self-inquiry’ method to answer validation questions without using an external knowledge source and trade-off between executing the validation procedure in parallel for all the concepts and in sequential order based on their ‘uncertainty’. For each step, we also indicate the most preferred technique with (*) and provide our justification.
In section 2.3, we detail our hallucination mitigation approach. Specifically, we ‘repair’ the hallucinated sentence by removing or substituting the hallucinated information leveraging the retrieved knowledge as evidence, and can also utilize the retrieved knowledge as context (prepended to the input) to generate the next sentence.
2 Hallucination Detection
In the first step, we identify the important concepts from the generated sentence. We identify these concepts because validating the correctness of the entire sentence at once is infeasible; this is because a sentence may contain a number of different facets all of which can not be validated at once. On the other hand, individually validating the correctness corresponding to the concepts provides opportunities for accurately detecting hallucinations. Thus, the objective of this step is to identify the candidates of potential hallucination. We note that a concept or keyphrase is essentially a span of text consisting of one or more words. We study the following techniques to identify the concepts:
Entities are usually an important part of a sentence, thus, we use an off-the-shelf entity extraction model to identify the concepts. A limitation of this method is that a concept need not necessarily be an entity and can be a non-entity span also. We address this limitation with a keyword extraction model.
To also identify the non-entity concepts, we explore an off-the-shelf keyword extraction modelhttps://huggingface.co/ml6team/keyphrase-extraction-kbir-kpcrowd. This model uses Keyphrase Boundary Infilling with Replacement (KBIR) as its base model and fine-tunes it on the KPCrowd dataset Kulkarni et al. (2021).
Since state-of-the-art language models perform remarkably well on a wide range of tasks, in this technique, we directly instruct the model to identify the important concepts from the generated sentence. An important characteristic of this technique is that it doesn’t require calling a task-specific tool (entity or keyword extraction model) for this task.
Table 7 (in Appendix A.1) illustrates examples of concepts identified using the three techniques. It shows that the entity extraction model misses many important concepts while the keyword extraction model identifies a lot of insignificant concepts also. In contrast, instruction technique successfully identifies all the important concepts. Moreover, it doesn’t require calling a task-specific tool. Thus, we represent this technique with (*), our preferred technique for this step.
2.2 Calculate Model’s Uncertainty
GPT-3 Brown et al. (2020) and several other publicly available models also provide logit output values in their prediction response. Thus, we study if these logit output values can be utilized to detect hallucinations. However, we note that this is an additional source of information and not a necessary requirement for our hallucination detection method as some models that are available only via API calls do not provide these logit output values.
Recall that a concept can consist of more than one token also (note that the model provides logit output values at the level of tokens); thus, we study three different techniques for calculating a probability score for a concept. Consider a concept consisting of tokens and having the maximum softmax probabilities as for the token positions respectively. We obtain these probabilities by applying the softmax function over the logit values for each token position. We study the following techniques:
: In this technique, we simply take the average of the probabilities of the tokens corresponding to the concept:
: Here, we take a normalized product of the probabilities of the tokens:
: Here, we take the minimum of probabilities as the score.
This is our preferred technique for this step as the other techniques average out the effect of model’s uncertainty on the tokens while low probability in even one token of the concept provides a strong evidence of the model being uncertain. For example, if the model is uncertain on the name of the USA president then its uncertainty on the first token (‘Joe’) would be high but on the next token (‘Biden’) would be very low as the token ‘Joe’ is frequently followed by the token ‘Biden’. Thus, averaging or normalizing the probabilities will have a limited capability to capture this signal.
Through our experiments (Section 3.1.2), we show that this score (especially ‘MIN’) indeed provides a signal for hallucination, i.e., the more uncertain a model is on a concept (low probability score), the more likely it is to be hallucinating about that concept. However, we note that this score is just a signal for hallucination and in no way provides a guarantee for presence of hallucinations. We utilize this signal and check for hallucinations with respect to the uncertain concepts using our validation procedure (2.2.3-2.2.5).
For models that do not provide the logit output values, all or some heuristically selected concepts (depending on the computational and latency budget of the system) can be passed to the validation stage for detecting hallucinations.
2.3 Create Validation Question
We start the validation procedure for a concept by creating a question that tests the correctness of the information (in the generated sentence) pertaining to the concept. We create Yes/No Questions, i.e., questions for which the answer is either a ‘Yes’ or a ‘No’. Table 8 shows examples of validation questions. For creating these questions, we explore the following two techniques:
Here, we use an off-the-shelf answer-aware question generation model.
Here, we directly instruct the model to create a validation question checking the correctness of the information about the selected concept. For the same reason as in the concept identification step, this is our preferred technique as it does not require calling a task-specific tool.
We note that instead of Yes/No questions, Wh-questions can also be used for validation. We prefer Yes/No questions as it is relatively easier to check the answer for these questions. We explore Wh-questions for a case study for answering multi-hop questions (Section 4.2).
2.4 Find Relevant Knowledge
In order to answer the validation question, we retrieve knowledge relevant to it which serves as additional context. For generality and wide coverage, we use web search (via Bing search API) for retrieving this knowledge. However, we note that any other search API or knowledge corpus can also be utilized for this purpose.
We also explore a self-inquiry technique where we directly prompt the model to answer the validation question. In this technique, the model relies on its parametric knowledge to answer the validation question. This technique has several drawbacks as compared to web search such as lack of a reliable strategy to extract the parametric knowledge from the model and staleness of the parametric knowledge.
Note that the proposed knowledge retrieval step in our approach has several benefits, such as (a) it does not retrieve knowledge when it is not required, i.e., when the model is already sufficiently confident (since we show it is less likely to hallucinate in such scenarios), (b) it individually retrieves knowledge pertinent to the concept(s) on which the calculated probability score is low thus providing it sufficient and relevant context for accurate validation / mitigation.
2.5 Answer Validation Question
In this step, we prompt the model to answer the validation question (leveraging the retrieved knowledge as context) and verify its response. If the validation procedure succeeds for all the uncertain concepts then we continue generating the next sentence; otherwise, we interrupt the generation process, mitigate the potential hallucination in the sentence, and then continue generation.
Validation of different concepts can be done in a sequence (in ascending order of their calculated probability score) or in parallel. However, running this in parallel would require starting multiple threads which may not be supported by all machines. Thus, in this work, we study only the sequential validation strategy but note that it can be made more efficient by running it in parallel. We regard this sequential validation as a greedy exiting strategy as we proceed to the mitigation stage on detection of the first potential hallucination.
3 Hallucination Mitigation
For mitigating the hallucination in the generated sentence, we instruct the model to repair the generated sentence by either removing or substituting the hallucinated information using the retrieved knowledge as evidence. Table 6 shows the instructional prompts for different steps of our approach.
Note: We note that the result of the validation procedure is contingent on the retrieved knowledge and the model’s ability to leverage that knowledge in answering the validation question. Thus, a case is plausible in which the validation procedure reports hallucination even though the sentence is actually not hallucinated. However, in Section 3.2, we show that our approach performs fairly well on this task. Moreover, it achieves a very high recall demonstrating its efficacy at detecting hallucinations. Moreover, in Section 3.3, we show that our mitigation approach does not introduce new hallucinations even in the case of incorrectly detected hallucinations, i.e., false positives.
4 Design Decisions
We note that dealing with the hallucination problem is a complex task and prior work has shown that breaking down a complex task into simpler sub-tasks helps the model in solving the task Wei et al. (2022b); Zhou et al. (2023); Khot et al. (2023). Thus, we break down this task into individual sub-tasks which are considerably easier for the model. For the same reason, we also break down the validation procedure into several steps.
Our preferred technique for retrieving knowledge is web search because the web is more likely to contain the updated knowledge in comparison to a knowledge corpus whose information can become stale, outdated, and obsolete.
We note that our detection and mitigation techniques can also be applied in a “posthoc” manner after complete response generation. However, it has several limitations which are addressed by our “active” approach. The “active” approach prevents the propagation of hallucinations in the subsequently generated sentences, i.e., if hallucination is detected in the initially generated sentences then it would be mitigated and course correction would be done for the subsequently generated sentences. However, the “post-hoc” approach does not provide such an opportunity of course correction. In other words, in the “active” approach, the model sees the mitigated / corrected sentences while generating the subsequent sentences; thus, its output will be more correct, coherent, and fluent. In contrast, in the “posthoc” approach, the generated sentences are based on the initially generated previous sentences and thus the mitigated sentence will not be able to influence the generation of subsequent sentences; thus, the output would not be as coherent and fluent as the active approach.
Our approach results in improvements in the form of reduced hallucinations and thus makes the model more reliable; however, it comes at the expense of increased inference cost. However, we believe that at current time, to enable the widespread adoption of LLMs, it is more important to address their reliability and trustworthiness concerns because computational advancements are ongoing at a rapid pace. Moreover, even larger models with multi-fold times more parameters such as PaLM (540B) Chowdhery et al. (2022), Gopher (280B) Rae et al. (2021), and MT-NLG (530B) Smith et al. (2022) are also being developed which have even higher inference cost showcasing a larger focus of the community on developing better performing systems. However, we note that our approach can be made more efficient by various techniques discussed before such as validating concepts in parallel and executing these intermediate steps using a smaller low-cost model.
Experiments and Results
In this section, we first demonstrate the two findings that motivate our approach (3.1.1 and 3.1.2). Then, we show the individual efficacy of our hallucination detection and mitigation techniques in 3.2 and 3.3, respectively. Finally, in 3.4, we show the effectiveness of the proposed active detection and mitigation approach in addressing hallucinations.
In our experimental setup, we prompt the large language model (GPT-3.5: text-davinci-003) to write about various topics. Specifically, we use a total of topics from diverse domains. Figure 3 shows the distribution of different domains in our topic set. In each domain, we include different kinds of topics; for instance, Sports domain consists of sports persons, administrators, teams, and games, Music consists of musicians, songs, music labels, and bands, Politics includes politicians, political parties, and elections, Film & TV includes actors, TV personalities, shows, and movies, History includes historians and events, etc. For selecting the names of people, we use randomly sampled names from the top 20% of longest articles in WikiBio dataset Lebret et al. (2016) as done in Manakul et al. (2023). Similarly, for the other topics, we randomly sample from the longest Wikipedia articles. This is done to ensure that no obscure or ambiguous concept is selected.
Equipped with the list of topics, we give the following input prompt to the model: ‘‘Write an article about
1 Motivating Findings
Recall that we consider the first five sentences generated by the model for each topic and annotate their correctness. Since the sentences are sequentially generated, we investigate the relationship between ‘hallucination in a generated sentence’ and ‘hallucination in the previously generated sentences’ for an input. Since there are two binary variables, there exist four possibilities in this relationship, i.e., a sentence is hallucinated and there was hallucination in the previously generated sentences (A), the sentence is not hallucinated and there was hallucination in the previously generated sentences (B), the sentence is hallucinated and there was no hallucination in the previously generated sentences (C), the sentence is not hallucinated and there was no hallucination in the previously generated sentences (D). For illustration, consider a sample case for sentence 3, the two binary variables are whether sentence 3 is hallucinated and whether there was hallucination in the previously generated sentences (i.e. in sentence 1 OR sentence 2). Figure 4 demonstrates this relationship for sentences 2, 3, 4 and 5 aggregated over all the topics in our data. We do not show this for sentence 1 as there is no previously generated sentence for it.
From this figure, we draw the following inferences:
(a) A > B: Cases A and B correspond to the scenario when there is hallucination in the previously generated sentences. It can be observed that A is considerably greater than B which implies that when there is hallucination in the previously generated sentences, a sentence is hallucinated more often. Moreover, the gap keeps increasing as the sentence number increases.
(b) A > C: Cases A and C correspond to the scenario when a generated sentence is hallucinated. It can be observed that A is greater than C which implies that a generated sentence is hallucinated more when there is hallucination in the previously generated sentences as compared to when there is no previous hallucination.
(c) D > C: Cases C and D correspond to the scenario when there is no hallucination in the previously generated sentences. Here, D is greater than C which implies that when there is no hallucination in the previously generated sentences, a generated sentence is more often not hallucinated.
(d) D > B: Cases B and D correspond to the scenario when a generated sentence is not hallucinated. D is greater than B which implies that a generated sentence is not hallucinated more when there is no previous hallucination as compared to when there is previous hallucination.
This shows that hallucination in a sentence often results in further hallucinations in the subsequently generated sentences and thus actively detecting and mitigating hallucinations can not only fix the current hallucination but can also prevent its propagation in the subsequently generated sentences.
Next, we demonstrate the utility of logit output values in detecting hallucinations.
1.2 Logit Output Values Provide a Signal for Hallucination
In this subsection, we first show the trend of hallucination with the probability score. Note that this score is calculated using the logit output values. Then, we demonstrate the benefit of identifying concepts from the generated sentence in detecting hallucinations. Finally, we compare the efficacy of different probability calculation techniques in detecting hallucinations.
In order to study the relationship between logit output values and hallucination, we annotate correctness at concept-level also (in addition to sentence-level annotations described earlier). Specifically, for each identified concept, we mark whether the information about it in the generated sentence is hallucinated or not. This can be different from sentence-level annotation as it focuses only on the correctness of the information about the concept in the sentence. Table 9 shows examples of both sentence-level and concept-level annotations.
Figure 5 shows the trend of hallucination with our calculated probability scores at both sentence and concept levels. For a sentence, we use the minimum across tokens of all its identified concepts as the probability score and for a concept, we use the minimum across all its tokens as the probability score. It can be observed that as the probability score increases (or uncertainty decreases), tendency to hallucinate decreases. This shows that these probability values can be utilized as a signal for hallucination, i.e., the low probability concepts in a generated sentence can be considered as candidates of potential hallucination and their correctness in the generated sentence can be validated for detecting hallucinations. On average, we observe an absolute difference of between the probabilities of concepts when the model is hallucinating vs when it is not hallucinating.
Now, we demonstrate the benefit of identifying concepts from a sentence and leveraging the logit output values corresponding to their tokens for detecting hallucinations. To this end, we plot precision-recall curves for the hallucination detection task corresponding to two methods that use the probabilities calculated from the logit output values. The blue curve corresponds to the technique in which we use the minimum probability across all tokens of the sentence and the orange curve is for the technique in which we use the minimum over only the tokens of the identified concepts. Figure 6 shows the two curves. The orange curve achieves higher area under the precision-recall curve implying that utilizing the probabilities of the concept tokens provides a stronger signal for hallucination as compared to the probabilities corresponding to all the tokens.
Figure 7 shows the Precision-Recall curves for the hallucination detection task (at concept-level) using the three probability calculation techniques, i.e., Minimum, Average, and Normalized (described in 2.2.2). The ‘Minimum’ technique achieves the highest area under the curve and hence is better at the hallucination detection task.
2 Hallucination Detection Performance
In this subsection, we demonstrate the hallucination detection performance of various techniques at both sentence and concept-levels.
Table 1(a) and 1(b) show the hallucination detection performance of the self-inquiry and web search techniques at sentence-level and concept-level, respectively. For sentence-level results, we predict the sentence to be hallucinated if the validation procedure fails on any identified concept. Note that in these results, we do not leverage the uncertainty score to select concepts for validation, instead we validate all the identified concepts. We study the relationship of recall with probability thresholds in Figure 12 (in Appendix). From the tables, it can be observed that the web-search technique achieve considerably high recall in detecting hallucinations.
Here, we emphasize on the high ‘recall’ of web-search technique as we show that our mitigation approach does not introduce any new hallucinations even in the case of incorrectly detected hallucinations, i.e., false positives (3.3). Figure 12 shows the recall of hallucination detection vs Probability threshold plot for Self Inquiry and web search techniques at both sentence-level and concept-level. Web-search is consistently and considerably better than self-inquiry.
3 Hallucination Mitigation Performance
On sentences where our validation procedure (using Web search) reports hallucinations, we apply our mitigation technique. We note that a sentence which is reported as hallucination can either be actually hallucinated or not hallucinated, i.e., it could also be a false positive. Table 2 shows the result of our method. It successfully mitigates the hallucination on of the correctly detected hallucinations (True Positives); we refer to this metric as ‘success’. Furthermore, it achieves this at minimal ‘deterioration’ (), i.e., it incorrectly converts a minimal of the non-hallucinated instances to hallucinated. Table 10 (in Appendix) shows examples where our mitigation technique successfully mitigates the hallucinations.
Table 11 (in Appendix) shows examples where our mitigation technique fails to mitigate the hallucinations. We observe that in many of the failure cases, our technique fixes some hallucinated content of the sentences but fails to fix ALL the hallucinated content from them. For instance, example 1 and 2 in Table 11 correspond to type of failure.
Furthermore, in some of the failure cases, our technique results in a sentence which is no longer hallucinated but it not completely related to the topic. For instance, the fourth example in Table 11 about the topic ‘Harry S. Kennedy’; the model generates “Harry S. Kennedy was an American politician who served as the 35th President of the United States from 1961 to 1963.” which is wrong and our mitigation technique modifies it to “John F. Kennedy was an American politician who served as the 35th President of the United States from 1961 to 1963.” which is factually correct but not related to the topic ‘Harry S. Kennedy’. This happens because the output of the mitigation step is contingent on the information in the retrieved knowledge.
4 Active Detection and Mitigation
The two findings in Section 3.1 motivate our approach of addressing hallucinations in which we actively detect hallucinations leveraging the logit output values and mitigate them during the generation process to prevent their propagation. Specifically, using the calculated probability scores, we identify the uncertain concepts and check their correctness using our validation procedure. We generate one sentence at a time and when our detection method reports hallucination, we fix it using our mitigation approach and continue generating the next sentence. We demonstrated separate detection and mitigation efficacy in 3.2 and 3.3, respectively. Figure 1 compares the percentage of hallucination in the output of GPT-3.5 model and our active detection and mitigation approach. Our approach reduces the percentage of hallucinations from to . In Figure 8, we demonstrate this comparison for different categories of hallucination. It shows that our approach reduces hallucinations for all categories.
To further demonstrate the effectiveness and wide applicability of our approach, we present three interesting additional studies. In the first study (Section 4.1), we experiment with another large language model, Vicuna-13B Chiang et al. (2023) and show that our approach performs well with this model also and considerably reduces the hallucinations it its output (Figure 9). We select this model since it is widely popular and publicly available to use. In the second and third studies, we adapt our approach to two different types of questions and show its effectiveness. Specifically, in Section 4.2, we adapt it to answer multi-hop questions and in Section 4.3, we show experiment with the false premise questions.
Additional Experiments
Figure 9 compares the percentage of hallucinations (on the ‘article generation task’) in the output of Vicuna-13B and our proposed active detection and mitigation approach. It shows that our approach considerably reduces the hallucinations similar to the case with GPT-3.5 model. This study is conducted on randomly sampled topics (i.e, generated sentences) from the topic set described in Section 3. We note that similar to the setup with the GPT-3.5 model where we used instructional prompts with GPT-3.5 for all the steps of the approach (such as identifying key concepts and creating validation questions), here, in this setup, we use the Vicuna-13B model itself for all those steps. This result demonstrates the generality and applicability of our approach for other models also.
2 Multi-hop Questions
In this study, we show that our approach can be adapted to improve the performance on multi-hop questions. Table 12 shows examples of these questions. Recall that our approach works by mitigating hallucination / incorrectness in the sentences generated by the model. Thus, if we can enable the model to answer these multi-hop questions step by step, then our active detection and mitigation approach can be applied to these steps, leading to correct predictions. To this end, we prompt the model and provide in-context examples demonstrating it to answer a given multi-hop question step by step. Table 13 (in Appendix) shows the prompt with in-context examples used for this purpose. Specifically, for a new question, the model generates the answer in multiple steps (one step at a time) and for each step, we apply our technique in which we first identify the low probability concepts from the sentence, validate their correctness using web search results, mitigate the hallucination (if detected), and then proceed to generate the next step. In our case study, we sample multi-hop bridge questions from the validation set of HotpotQA Yang et al. (2018) and evaluate the performance.
Table 3 shows examples of responses generated using our approach. Figure 10 shows the performance achieved by different methods on this task. First, it shows the performance of the GPT-3.5 model; the model answers of the questions incorrectly. Then, it shows the performance of the GPT-3.5 model with in-context examples; it results in a slight improvement over the zero-shot performance. Then, it shows the performance of the model on leveraging the knowledge retrieved from the web (using the question as the search query) as context to answer the question. As expected, the model’s performance improves, i.e., it results in fewer incorrect predictions. Finally, we show the performance of our active detection and mitigation approach which results in considerably fewer hallucinations (26%), i.e., higher percentage of correct answers. This demonstrates the effectiveness of our approach in improving the performance on multi-hop questions.
3 False Premise Questions
We perform this experiment because LLMs have already been shown to perform remarkably well on a wide range of ‘correct’ questions, i.e., questions that are factually correct and make the right assumptions Khashabi et al. (2020); Brown et al. (2020); Zhang et al. (2022); Lourie et al. (2021); Chowdhery et al. (2022); Rae et al. (2021). However, users in real-world applications often ask questions that are based on false premises / pre-suppositions such as “Why energy is absorbed in exothermic reactions?” and “Why do floppy disks have higher storage capacity than USB drives?”.
We observe that state-of-the-art models often struggle to appropriately respond to such questions; thus, such questions serve as another challenging evaluation setting. To this end, we conduct a case study and compile a set of such adversarial questions, i.e., we compile questions on which GPT-3.5 model generates an incorrect response. This is done to create a challenging experimental setup for evaluation as the model generates incorrect output for such questions. Furthermore, we also create a true premise question corresponding to each false premise question. Table 14 (Appendix) shows examples of these question pairs. For this task, we evaluate correctness at the complete answer level as this is a question answering task and the entire answer needs to be correct for the answer to be considered as correct.
We note that an ideal response from a system for such questions depends on the application. For instance, some applications may require identifying such questions and then abstaining on them like the selective prediction systems Varshney and Baral (2023); Kamath et al. (2020); Xin et al. (2021); Varshney et al. (2022). Some applications may additionally require suggesting a ‘rectified’ question and providing response to that rectified question such as the search engines. Our approach supports these requirements by using the validation and mitigation step on the given input question.
Specifically, we first retrieve the relevant knowledge (via Bing Search using the question as query). Then, conditioned on the retrieved knowledge, we prompt the model to respond ‘Yes’ if the question makes factually correct assumptions, otherwise respond ‘No’. If the response to this prompt is No, then we proceed to modify the question using the mitigation step. Table 15 shows both the instructional prompts used for identifying and rectifying a potentially false premise question. This step enables identifying false premise questions and also rectify them to facilitate the system in providing an appropriate response. Importantly, we also show that our approach does not incorrectly modify a true premise question. This is crucial because if the user’s question is correct then the system’s response must be pertinent to that and not to its modified variant.
Recall that the false premise questions in our evaluation set are adversarially collected, i.e., GPT-3.5 gives an incorrect response to all of these questions. First, we evaluate the performance of GPT-3.5 model when relevant knowledge (retrieved via bing search using the question as the search query) is given as context to answer the question. We find that even with the retrieved knowledge, GPT-3.5 manages to answer only false premise questions correctly, i.e., hallucinates on the remaining questions. In contrast, our approach answers questions correctly and hallucinates only on . Figure 11 shows this comparison. Furthermore, we note that even in some of these hallucinated responses, some of the individual sentences in the responses are correct. However, since we focus on complete answer correctness, we mark them as incorrect. Table 16 shows responses on a few false premise questions generated by the GPT-3.5 model, GPT-3.5 model leveraging the retrieved knowledge as context, and our approach.
We analyze the performance of our approach in rectifying the questions; it successfully repairs false premise questions while not incorrectly modifying any true premise question. Though this step makes modifications in a small number of true premise questions ( instances), it does not change the semantics of those questions as shown in the Table 4. We note that not incorrectly modifying a true premise question is an important characteristic of our approach.
4 Other Applications
Our approach has utility in a variety of other applications also such as Abstractive Summarization and Claim Verification. In abstractive summarization where the generated summary has been shown to be often hallucinated Cao et al. (2022); Zhao et al. (2020); Chen et al. (2021) can be improved using our approach. Note that in the validation procedure of our approach, the relevant knowledge for this task will be retrieved from the original document instead of the web. However, for open-summarization, knowledge can be additionally retrieved from the web also. Our approach can be adapted for the claim verification task also as we can first identify the key sub-claims and then verify each sub-claim using the validation procedure. Here, the mitigation step will be useful for providing explanation behind the model’s decision. We leave exploring these other usecases of our approach for future work.
Related Work
Advancements in the field of natural language processing led to the development of models that possess an impressive ability to generate fluent and coherent text. However, these models are vulnerable to a phenomenon called text hallucination. Prior work Maynez et al. (2020); Huang et al. (2021); Ji et al. (2023) has categorized text hallucinations into two classes: Intrinsic (when the generated output contradicts the source content) and Extrinsic (when the generated output cannot be verified from the source content, i.e., it that can neither be supported nor contradicted by the source).
One thread of research pertaining to hallucinations has focused on studying different causes of this phenomenon such as training data quality Wang (2019); Lee et al. (2022a), source-target divergence Dhingra et al. (2019), ill-suited modeling Aralikatte et al. (2021); Feng et al. (2020); Li et al. (2018), and randomness during inference Dziri et al. (2021); Tian et al. (2019); Lee et al. (2022b).
The other thread focuses on addressing the hallucination problem Manakul et al. (2023); Azaria and Mitchell (2023); Lee et al. (2022b); Du et al. (2023); Zhang et al. (2023). Manakul et al. (2023) propose a sampling-based hallucination detection approach in which they first sample multiple responses from the model and then measure the information consistency between the different responses. They posit that when a language model knows a given concept well, the sampled responses are likely to be similar and contain consistent facts; on the other hand, for hallucinated facts, stochastically sampled responses are likely to diverge and may completely contradict one another.
Another recent work Azaria and Mitchell (2023) leverages LLM’s internal state to identify the truthfulness of a statement. Using an annotated dataset, they train a separate classifier that takes the LLM’s activation values as input and predicts its truthfulness. Lee et al. (2022b) hypothesize that the randomness of sampling is more harmful to factuality when it is used to generate the latter part of a sentence than the beginning of a sentence and propose a new sampling algorithm named factual-nucleus sampling that dynamically adapts the ‘nucleus’ p along the generation of each sentence. Du et al. (2023) propose an approach motivated by The Society of Mind and multi-agent settings in which multiple models individually propose and jointly debate their responses and reasoning processes to arrive at a common answer. In our approach, we leverage the logit output values, web search, and actively detect and mitigate hallucinations. We demonstrate the effectiveness of our approach on a variety of tasks, including article generation, multi-hop question answering, and false premise question answering.
Conclusion
In this work, we proposed an approach that actively ‘detects’ and ‘mitigates’ hallucinations of the large language models. Through systematic and extensive experiments with the article generation task, we showed that our approach successfully reduces the hallucinations of the GPT-3.5 (text-davinci-003) from to on average. We also demonstrated the individual efficacy of our detection and mitigation techniques. Specifically, our detection technique achieves a high recall and the mitigation technique successfully mitigates a large fraction of the correctly detected hallucinations. Notably, the mitigation technique does not introduce new hallucinations even in the case of incorrectly detected hallucinations, i.e., false positives. We further demonstrated the effectiveness and wide applicability of our approach and presented several interesting studies including evaluation with another LLM (Vicuna) and answering multi-hop and false premise questions. Overall, our work addresses the LLMs’ hallucination problem and thus contributes to improving their reliability and trustworthiness, a crucial step en route to enabling their widespread adoption in real-world applications.
References
Appendix
Appendix A Fruther Details of the Approach
Table 6 shows the instructional prompts used for different steps of our approach. We note that these techniques are the preferred techniques as they do not require calling an external task-specific tool to achieve the corresponding objectives.
Table 7 shows examples of concepts identified using the three methods, i.e., Entity Extraction, Keyword Extraction, and Instructing the Model. It shows that the entity extraction model misses many important concepts while the keyword extraction model identifies a lot of insignificant concepts also. In contract, instruction technique successfully identifies majority of the important concepts.
A.2 Create Validation Question
Table 8 shows examples of validation questions corresponding to each concept created via instructing the model technique. It shows examples of both the question types, i.e., Yes/No and Wh questions. We prefer Yes/No questions as it is relatively easier to check the answer for these questions. We leave exploring Wh-questions for validation for future work.
Appendix B Evaluation Data
Table 5 shows the statistics of the sentences generated by the GPT-3.5 (text-davinci-003 with temperature 0) model. A sentence has words on average and each sentence has key concepts that are identified by our instruction technique.
Table 9 shows examples of sentence-level and concept-level hallucination annotations.
Appendix C Recall of Hallucination Detection vs Probability Threshold
Figure 12 compares recall of hallucination detection for self-inquiry and web search techniques at different probability thresholds. Web search considerably outperforms self-inquiry at all thresholds.
Appendix D Hallucination Mitigation Examples
Tables 10 and 11 show examples where our mitigation technique successfully mitigates the hallucinations and where it fails, respectively. We observe that in many of the failure cases, our technique fixes some hallucinated content of the sentences but fails to fix ALL the hallucinated content from them. Furthermore, in some of the failure cases, our technique results in a sentence which is no longer hallucinated but it not completely related to the topic.
Appendix E Multi-hop QA Experiment
Table 12 shows examples of multi-hop bridge questions from HotpotQA dataset. Table 13 shows the prompt with in-context examples used for prompting the model to answer multi-hop questions step by step. Table 3 shows examples of responses generated using our approach for multi-hop bridge questions.
Appendix F False Premise QA Experiment
Table 14 shows examples of false premise and true premise question pairs. Table 16 shows responses generated on a few false premise questions by the GPT-3.5 (text-davinci-003) model, GPT-3.5 (text-davinci-003) using the retrieved knowledge as context, and our approach.