How Language Model Hallucinations Can Snowball
Muru Zhang, Ofir Press, William Merrill, Alisa Liu, Noah A. Smith
Introduction
Language models are increasingly being deployed to interface with humans in open-ended information-seeking and problem-solving settings. Despite their diverse capabilities and extreme fluency, a major open challenge is that LMs still hallucinate by making up facts or citing sources that do not exist (Maynez et al., 2020; Liu et al., 2023, i.a.), often while sounding extremely plausible.
Hallucination is commonly attributed to knowledge gaps in LMs Zheng et al. (2023), motivating mitigation strategies through retrieval over knowledge bases Lewis et al. (2020); Shuster et al. (2021); Peng et al. (2023) But, do LMs only hallucinate when they do not “know” a fact? We present a setting where LMs often generate hallucinations that they immediately recognize as wrong when presented in isolation. Specifically, after an LM answers a question incorrectly, it usually justifies that answer by making incorrect assertions that it separately acknowledges as incorrect (Figure 1).
To study this behavior empirically, we automatically construct three question-answering (QA) datasets. These datasets span different domains: determining whether a number is prime, whether there is a U.S. senator satisfying two given constraints, and whether two cities are connected given a set of flights between cities. Empirically, we find that ChatGPT OpenAI (2022) and GPT-4 OpenAI (2023) commit to an answer within the first token (Yes/No) over 95% of the time; these answers are often incorrect, and then followed by an incorrect explanation. Yet, when presented with the incorrect explanation alone, we find that the LM is likely able to recognize it as incorrect.
We refer to this phenomenon as hallucination snowballing. We hypothesize that LMs produce snowballed hallucinations for consistency with earlier hallucinations (rather than due to a “knowledge gap” in the model), as they recognize the snowballed hallucination is incorrect when presented in isolation (i.e., in a separate interaction session).
While prompting strategies that encourage the LM to reason before stating an answer improve accuracy on the task, our work points to the broader issue that conditioning on faulty context leads LMs to produce extremely simple mistakes that they wouldn’t otherwise make. Indeed, when prompting with “Let’s think step by step” Kojima et al. (2023), snowballed hallucinations still occur in 95% of cases where the model fails to answer correctly. We observe that sometimes even when “Let’s think step by step” does lead to the right answer, it uses invalid reasoning chains.
In this paper, we demonstrate the phenomenon of hallucination snowballing by leveraging recent LMs’ tendency to state and justify their answers. Rather than over-committing to its previously generated context, we believe that LMs should acknowledge their initial mistake, and then revise their answer. We have indeed observed GPT-4 doing this in a limited number of cases; amplifying this behavior would be beneficial, as well as developing new methods in which LMs can backtrack.
Why do we expect hallucination snowballing?
In this section, we explain why we hypothesize that LMs are susceptible to hallucination snowballing. We predict that snowballing will occur on questions with two key properties:
Initial committal: The prompt leads the LM to first state an answer (before outputting the explanation). This applies to many yes/no questions.
Inherently sequential: Transformers cannot find the answer within one timestep because of their limited reasoning abilities within one timestep.
We now discuss how these properties may lead to snowballed hallucination.
In English and many other languages, speakers often say the final Yes/No answers to questions before explaining their answer. We therefore hypothesize that LMs and especially instruction-tuned LMs Wei et al. (2021); Sanh et al. (2021); Ouyang et al. (2022); Wang et al. (2022) will reflect this answer format where the answer comes before the explanation. Indeed, on our datasets (presented in §3.1), we observe that GPT-4 and ChatGPT immediately commit to an answer to the question: the first token is Yes or No 95.67% and 98.40% of the time for GPT-4 and ChatGPT respectively. In the remaining cases, the model often commits to an answer within the first few tokens of the response (e.g., “There is no record of a U.S. Senator…”). Crucially, once the LM generates Yes or No, that token remains in the context, and coherence would require commitment to that choice through the subsequent justification. Thus, the model produces an answer to a complex question in a single timestep, and it then continues by generating an explanation for that answer, which inevitably will be incorrect.
Inherently sequential.
Furthermore, transformers cannot solve inherently sequential reasoning problems like primality testing or graph connectivity within a single timestep,Technically, this holds only for inputs above a certain hardness level, i.e., the size of the prime number for primality testing, or the size of the graph for graph connectivity. as documented in recent theoretical results (Merrill and Sabharwal, 2023).Merrill and Sabharwal (2023) show that, with a single generation step, bounded-precision transformers cannot solve any problem outside the complexity class , which corresponds to a highly parallelizable subclass of both (log-space) and (polynomial-time). Graph connectivity is an -complete problem, which means it cannot be in unless , i.e., all of can be parallelized to a surprisingly high degree. Primality testing was shown to be in (Agrawal et al., 2004) but cannot be in unless it is also in ; i.e., any can be factored with bits of overhead. In summary, unless standard complexity-theoretic conjectures are false, graph connectivity and primality testing are outside and thus are too inherentially sequential for transformers to solve in a single generation (cf. Merrill and Sabharwal, 2023). Our graph connectivity and primality datasets are concrete instantiations of these problems. Because the transformer must use one step to answer a question that requires multiple timesteps to answer correctly, it will necessarily sometimes commit to an incorrect answer. We hypothesize that this leads the LM to hallucinate supporting incorrect facts that it otherwise would not generate.
Experiments
We design three QA datasets with the properties described in §2 to probe hallucination snowballing, and evaluate ChatGPT and GPT-4. We first check whether the LM returns the correct answer to the given question, and we show that when the model returns the wrong answer, it frequently provides an incorrect explanation for that wrong answer. We automatically extract the incorrect claim in the explanation and ask the same LM to check whether its claim is correct. See Table 1 for a representative example from each dataset.
We design three QA datasets, each containing 500 yes/no questions that we expect are not answerable by transformers in one timestep. To aid evaluation, the questions are designed so that an incorrect answer would be justified with easily verifiable claims.
For each dataset, we fix one specific label for all examples, so that if the model chooses the incorrect answer (e.g., that 9677 is not prime), it would produce a specific claim to support it (e.g., an incorrect factorization). This enables us to systematically examine model-written justifications for incorrect answers.
For this dataset, we query the primality of 500 randomly chosen primes between 1,000 and 20,000; the correct answer is always Yes. When the model answers incorrectly, we expect it to justify its answer with an incorrect factorization.
Senator search
This dataset consists of 500 questions of the form “Was there ever a US senator that represented the state of and whose alma mater was ?” where is a U.S. state and is a U.S. college. For these questions, the correct answer is always No. When the model answers incorrectly, we expect it to falsely claim that a particular senator both represented and attended .
To create the dataset we consider all U.S. states and a manually constructed list of twelve popular U.S. colleges (see §A for the full list); for each possible pair, we generate a question following the template, and manually remove pairs where the answer is Yes.
Graph connectivity
For each of the 500 questions in this dataset, we present 12 flights among 14 cities, and ask if there is a sequence of flights from a particular city to another. The problem always corresponds to the same underlying directed graph structure (see §A.1), where flights are edges and cities are nodes. For each instance in the dataset, we randomly assign letters from the English alphabet to name the nodes. To formulate the query, we sample a source city and destination city in different subgraphs, with the additional constraint that corresponds to a source node, and a leaf node, so that 1-step heuristics cannot be used to solve the problem.
We formulate the problem as a flight-finding question in natural language so that it sounds more natural: in the prompt, we list the twelve flights (“There is a flight from city F to city K; there is a flight from city G to city N, …”), followed by the question “Is there a series of flights… from to ?”. Note the correct answer is always No. When the model answers incorrectly, we expect it to justify its answer with a flight that does not exist.
2 Inference Setup
Language models. We run all experiments on ChatGPT (gpt-3.5-turbo) and GPT-4 with greedy decoding.
Our experiments are zero-shot (i.e., we do not show the model any example QA pairs in the prompt). We focus on the model behavior under the direct prompt (see §A for full examples), which is the most common way users interact with LMs. See §4 for experiments with the zero-shot chain-of-thought style prompting method.
For each dataset, we perform a two-stage evaluation. First, we evaluate the model’s accuracy (i.e., how many of the questions it answers correctly). When either models is incorrect, empirically it always generates a justification. In the second stage, we assess whether the model can identify the incorrect step in the explanation.
For a given question, we evaluate the model’s response by examining whether the output begins with either Yes or No. In cases where the response does not fall into these categories, we manually determine the answer conveyed by the model.
3 LM Recognition of Snowballed Hallucinations
We probe whether LMs recognize their snowballed hallucinations by verifying the model’s incorrect claims in the output against the model itself. Note that our recognition procedure relies on heuristics gained from manual examination of the model output, and these heuristics might not work on other models (e.g., a different model might not provide factors when supporting the claim that a number is not prime).
For each sample where the model thinks there is a series of connecting flights (where answer starts with Yes), we manually extract the list of flights from the model’s output and identify the invalid or discontinuous flights.
We then, in a new session, ask the model to verify whether the extracted flights are valid based on the flight information, and if consecutive flights are indeed connected. We manually assess the verification output to check if the model correctly detects the error. See Appendix Table 3 for how we prompt the model and an example of successful verification.
Primality Testing
For each sample where the model answers that the number is not prime, we extract the factors the model uses to justify it. The extraction is done by putting the output in the context and asking “What are the factors proposed in the above text? List them out.” We use ChatGPT for extraction with one-shot demonstration (for its fast inference speed); we manually checked 30 examples and found that it can always extract the correct factors.
We then, in a new session, ask the model to verify each extracted factor individually. See Appendix Table 4 for an example of successful verification.
Senator Search
For each sample where the model thinks there is such senator, we extract the name of the senator the model uses to justify the existence, by putting the output in the context and asking “What is the senator mentioned in the above text? Just give the name”. Again, we use ChatGPT and manually observed perfect extraction on 30 examples.
We then, in a new session, ask the model if that senator’s alma mater is the college in the question and has represented the state in the question. See Appendix Table 5 for an example of successful detection.
4 Results
Figure 2 shows that both ChatGPT and GPT-4 experience very low accuracy across the board. With the exception of ChatGPT on the Senator Search dataset, all models achieve less than 50% accuracy.(See Appendix Table 6 for a breakdown of the error rate by dataset.) We observe that GPT-4 performs worse than ChatGPT across all datasets despite popularly being considered superior to ChatGPT OpenAI (2023). While ChatGPT has an average accuracy of , GPT-4 has only .
Hallucination detection
Here, we check whether the model can identify that the incorrect claim is wrong when it is presented alone. As shown in Figure 2, ChatGPT detects 67.37% of incorrect claims in explanations (i.e., snowballed hallucinations), and GPT-4 detects 87.03%. Notice that when the model fails the verification (an example in Appendix Table 12), we do not consider it a snowballed hallucination.
Overall, we find that ChatGPT and GPT-4 are both extremely susceptible to hallucination snowballing, leading to extremely simple mistakes.
Can we prevent snowball hallucinations?
We hypothesize that hallucination snowballing occurs because LMs are trained to model continuations consistent with their current context (the given prompt and prior outputs). Although a fix to the fundamental problem might require more than just inference-time modification, in this section we study the effectiveness of two inference strategies in alleviating hallucination snowballing: prompting (§4.1) and decoding or training methods (§4.2).
In this section, we examine the effectiveness of better prompts on preventing snowballed hallucination by using a different zero-shot prompt that encourages the model to generate the reasoning chain before the answer. Since the outputs generated under these prompts are less structured, we manually inspect them to determine correctness and the presence of snowballed hallucinations.
For each task, we append “Let’s think step-by-step” at the end of the original question (shown in Table 1). As shown in Figure 3, the model can solve the Senator Search task perfectly, achieve 10% error rate on Primality Testing, and 30% on Graph Connectivity. Despite the large improvement in accuracy, we identify a potential issue: the model sometimes hallucinate while outputting the reasoning chain, which causes snowballed hallucination in future steps. For example, in the below output,
Step 3: From city E, we have three options: a flight to city N, a flight to city B, or a flight to city C.
Step 4: The only option that could potentially lead us to city M is the flight from city E to city C.
ChatGPT incorrectly states that there are three options in the step 3 (there are only two), inducing the snowballed hallucination “or a flight to city C” (ChatGPT can verify that E C is not a valid flight in a separate session). As shown in Figure 3, GPT-4 still has a high overall snowballed hallucination rate at 94.90% averaged across tasks, and ChatGPT also obtains a similarly high snowballed hallucination rate.
Finally, while our experiments have focused on simple multi-step problems that are suitable for breaking down step-by-step, we hypothesize that hallucination snowballing appears in open-ended text generation more broadly, where one mistake in the generation triggers more Arora et al. (2022). In these cases, better prompting would neither be able to anticipate nor fix these mistakes.
2 Algorithmic Corrections
During decoding, the temperature controls the sharpness of the output distribution, with higher spreading probability mass away from the model’s most likely prediction for each next word. Our experiments in §3 used greedy decoding, which is equivalent to . At and , both error rates and snowballed hallucination rate remain similarly high, in both GPT-4 and ChatGPT (Figure 4).
Top-k and nucleus sampling
Using sampling methods such as top- sampling or nucleus sampling Holtzman et al. (2020) would not help since they only narrow the range of tokens to be considered, and thus can only increase the probability that the model will immediately commit to an answer.
Beam search
The argument for hallucination snowballs in §2 relies on the fact that, once a model generates some tokens committing to an answer, they remain in the context and influence later generations. One potential way around this is beam search, i.e., maintaining a beam of high-probability sequences at each timestep rather than a single sequence. In principle, if some sequences in the beam after the initial token do not commit to an answer (or commit to the right answer), their continuations may eventually have higher probability than those that initially commit incorrectly and later produce incorrect reasoning as a result. If so, beam search would solve the snowball hallucination problem. Unfortunately, we cannot test the effect of beam search on hallucination snowballs because the OpenAI API does not support beam search.
Learning strategies
A more general way to further reduce snowballing might be to change aspects of the pretraining or instruction tuning phases. In particular, a greater emphasis on having the model produce a reasoning chain before generating an answer could be a good way to accommodate its computational limitations and avoid committing to wrong answers that force hallucinations.
In addition, we hypothesize that finetuning on data with backtracking might improve a model’s performance on the tasks we present. This could be accomplished by, for example, giving a question, followed by a wrong solution, and then issuing a phrase like “Sorry, that was incorrect” before giving the correct solution. This solution is related to the “Review your previous answer and find problems with your answer.” prompt from Kim et al. (2023).
Related Work
Hallucination in text generation is a well-studied problem (Rohrbach et al., 2018; Maynez et al., 2020; Raunak et al., 2021, i.a.) that has recently become more prominent due to ChatGPT’s tendency to produce plausible-sounding falsehoods. Hallucinations are often attributed to knowledge gaps in LMs Zheng et al. (2023), and several works have shown the promise of using retrieval over knowledge bases to mitigate them Lewis et al. (2020); Shuster et al. (2021); Peng et al. (2023). Our work demonstrates hallucination can be induced from context, thus motivating further mitigation techniques.
Hallucination snowballing is likely the result of exposure bias: LMs were only exposed to gold history during training, but during inference, conditions on possibly erroneous previous predictions. Prior work linked this to compounding hallucinations in machine translation Wang and Sennrich (2020) and open-ended text generation Arora et al. (2022). We go beyond demonstrating error propagation by showing that the propagated errors (which we call snowballed hallucinations) are recognized by the LM itself.
Our observations are related to previous findings that LMs hallucinate when given questions that contain false presuppositions (e.g., “Which linguist invented the lightbulb?”; Kim et al., 2021, 2022) or that are otherwise misleading (e.g., “Who really caused 9/11?”; Lin et al., 2022), in that faulty context misguides the LM. However, our work differs in that our questions are not intentionally misleading, showing that this failure mode may be triggered even on innocent information-seeking queries to the LM.
LM (in)consistency
Our work adds to a growing body of work demonstrating the extent to which LMs are inconsistent across different prompts on the same issue. For instance, allowing an LM to generate intermediate steps Nye et al. (2021); Wei et al. (2022); Press et al. (2022) enables it to reach a different answer than it otherwise would. Other work has shown that simply prepending “Professor Smith was given the following instructions” to a prompt can improve performance, despite providing no valuable information about the problem itself Lin et al. (2022).
Conclusion
We define the phenomenon of hallucination snowballing and demonstrate its prevalence in generations from state-of-the-art models, leading to hallucinations on simple facts that wouldn’t otherwise occur. Our findings point to the risk of training language models that prioritize fluency and coherence indiscriminatively at the expense of factuality, and we encourage future work to study remedial actions at all levels of model development.
Limitations
We focus on hallucination snowballing in the context of question answering in English, and we do not explore it on other tasks, such as summarization or code generation.
In addition, we only conduct experiments on two proprietary models, namely ChatGPT and GPT-4, due to their state-of-the-art performance on many benchmarks OpenAI (2023). Due to the limitations of the APIs for these models, we do not have access to the probability distributions they output and do not have the ability to finetune them. This restricts our ability to explore potential mitigation strategies. Having access to the output distributions would allow us to investigate mitigating the snowballing hallucination issue using alternative sampling methods such as beam search. Having the ability to finetune the model would allow us to explore whether instruction tuning with different annotations could lead to better handling of the questions we use to instigate hallucination snowballing.
Acknowledgements
We thank Sofia Serrano, Yizhong Wang, Yanai Elazar, Michael Hu and Richard Yuanzhe Pang for their valuable feedback and fruitful discussions. While writing this paper, Ofir Press was a visitor at New York University’s Center for Data Science, hosted by Kyunghyun Cho.
References
Appendix A Dataset Details
In this dataset, the list of flights can be represented by a directed graph. We generated the flight information to ensure all the graphs share a specific connection pattern, with the node names randomly chosen among the 26 letters in the English alphabet. For an illustration of the underlying graph structure, see Figure 5.
A.2 Senator search
The twelve colleges used in the datasets are: MIT, University of Chicago, Johns Hopkins University, California Institute of Technology, Duke University, Northwestern University, Dartmouth College, Brown University, Vanderbilt University, Rice University, University of Washington. We constructed this list by taking a list of top universities in the U.S. and excluding from it universities which also appeared on The U.S. News & World Report’s list of Top 10 Colleges for Members of Congress.
Appendix B Additional Results
We provide the detail breakdown of the question-answering accuracy in Table 6 and the hallucination detection accuracy in Table 7.