Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models
Matthew Dahl, Varun Magesh, Mirac Suzgun, Daniel E. Ho
Introduction
The legal industry is on the cusp of a technological metamorphosis, driven by recent advancements in AI—particularly in the form of large language models (LLMs). As these AI models have shown increasing proficiency in law-related tasks, such as first-year law school exams (Choi et al., 2022), the uniform bar exam (OpenAI, 2023b), statutory reasoning (Blair-Stanek et al., 2023), and issue-rule-application-conclusion (IRAC) analysis (Guha et al., 2023), many have started asking whether AI tools and systems might soon displace human lawyers altogether (Bommasani et al., 2022; Perlman, 2023, inter alia). Despite their potential to reform how legal work is performed, however, there remains a critical challenge to LLMs’ widespread adoption: the issue of hallucinations. LLMs like ChatGPT can sometimes generate text that is inconsistent with current legal doctrine and case law, and in the legal field, where adherence to the source text is paramount, unfaithful or imprecise interpretations of law can lead to nonsensical—or worse, harmful and inaccurate—legal advice or decisions.
In this work, we present the first evidence documenting the nature and frequency of hallucinations in the legal domain. While anecdotal discussion of this problem has surfaced in national media—for instance, the New York Times reported on a lawyer who faced sanctions for using ChatGPT-generated fictional case citations in a brief (Weiser, 2023), and SCOTUSBlog highlighted ChatGPT’s misinformation regarding a supposed dissent by Justice Ruth Bader Ginsburg in the landmark gay rights case Obergefell v. Hodges (Romoser, 2023)—a systematic analysis of this issue has been lacking. Our study seeks to fill this gap, providing insights essential for evaluating LLMs’ effectiveness in general legal settings.
Hallucinations have previously been studied in settings such as neural machine translation and abstractive summarization, as well as general-knowledge question-answering (QA) and dialogue generation (Ji et al., 2023; Mündler et al., 2023; Manakul et al., 2023; Dhuliawala et al., 2023). We advance this literature by focusing on open-domain hallucination in the legal context, where factual and textual accuracy are crucial. Traditionally, a key challenge in studying this mode of hallucination has been obtaining the appropriate source content against which non-factual model output can be judged (Agrawal et al., 2023), but we overcome this problem by leveraging the structured nature of American case law. Because American cases follow a standard legal schema (with parties, issues, citations, dispositions, and so forth), we are able to use the tabular metadata that is recorded in legal databases on these dimensions to create knowledge queries that simulate basic legal research tasks for each case. We apply these tasks to a random sample of cases across each level of the federal judiciary—the U.S. District Courts (USDC), the U.S. Courts of Appeals (USCOA), and the U.S. Supreme Court (SCOTUS)—and evaluate them using three popular LLMs: OpenAI’s ChatGPT 3.5, Google’s PaLM 2, and Meta’s Llama 2.
Our findings reveal the widespread occurrence of legal hallucinations: When asked a direct, verifiable question about a randomly selected federal court case, LLMs hallucinate between 69% (ChatGPT 3.5) and 88% (Llama 2) of the time (Figure 1). However, we also find that LLMs perform better on cases that are newer, more salient, and from more prominent legal jurisdictions, suggesting that LLMs suffer from a kind of legal “monoculture” (Kleinberg and Raghavan, 2021) that leads them to recite an unduly homogenized notion of the law. We then investigate two additional potential failure points for LLMs, beyond their raw hallucination rates: (1) their susceptibility to contra-factual bias, i.e., their ability to respond to queries anchored in erroneous legal premises (Sharma et al., 2023; Wei et al., 2023), and (2) their certainty in their responses, i.e., their self-awareness of their propensity to hallucinate (Kadavath et al., 2022; Xiong et al., 2023; Tian et al., 2023b; Yin et al., 2023; Azaria and Mitchell, 2023). Our results indicate that not only do LLMs often provide seemingly legitimate but incorrect answers to contra-factual legal questions, they also struggle to accurately gauge their own level of certainty without post-hoc recalibration. Overall, we conclude that while these LLMs appear to offer a way to make legal information and services more accessible and affordable to all, their present shortcomings—particularly in terms of generating accurate and reliable statements of the law—significantly hinder this objective.
Preliminaries and Background
We first provide a brief overview of language models (LMs) for readers who may not necessarily have a deep technical background. LMs are functions that map text to text: When a user provides a text input (known as a “prompt”), the model produces a text output (referred to as a “response”). If the prompt takes the form of a question, the response can be understood as an answer to that question. An LM generates its response by selecting the most probable sequence of tokens that follow the prompt’s tokens; therefore, it essentially functions as a probability distribution over these tokens. In this work, we focus on “large” language models (LLMs). These are models that contain billions of parameters and are trained on vast corpora reflective of the scope of the Internet.Hence, the largeness of a language model is a dual reference to its parameter count and the scope of its training corpus.
Formally, we define an LLM as a function f_{\tau}:{\color[rgb]{0,0,1}\definecolor[named]{pgfstrokecolor}{rgb}{0,0,1}\text{prompt}}\mapsto{\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}\text{response}}, where operates by sampling language tokens from a conditional probability distribution that is learned by optimizing over a training corpus hopefully reflective of facts about the world. We equate with the model’s response and with the input prompt. is a temperature parameter that controls the shape of the probability distribution under sampling and is configurable by the user. When , the distribution becomes degenerate and the model’s response is theoretically deterministic—thus, the model must always return the most likely next token.In practice, non-determinism may persist due to a model’s implementation details, e.g., the “mixture of experts” architecture (Puigcerver et al., 2023; Chann, 2023). This deterministic response is typically referred to as the “greedy” response. When is high, however, the distribution becomes more uniform and the model’s response becomes more stochastic—this is the default behavior of chatbots like ChatGPT, which are often used to generate diverse and creative responses.
2 Hallucination in Language Models
LLMs have shown promise on a number of legal research and analysis tasks (Blair-Stanek et al., 2023; Choi et al., 2022; Fei et al., 2023; Guha et al., 2023; Nay et al., 2023; OpenAI, 2023a; Trozze et al., 2023), but the problem of legal hallucination has so far only been studied in closed-domain applications, e.g., where a model is used to summarize the content of a given judicial opinion (Deroy et al., 2023; Feijo and Moreira, 2023) or to synthesize provided legal text (Savelka et al., 2023). In this paper, by contrast, we examine hallucination in an open-domain setting, i.e., where a model is tasked with providing an accurate answer to an open-ended legal query. This setting approximates the situation of a lawyer or a pro se litigantA pro se litigant is a litigant who is representing themselves, possibly without formal legal training. seeking advice from a legal chatbot.
In the context of such question-answering (QA) scenarios, the study of hallucinations in LMs is still in its infancy, even outside the legal field. There is no universally accepted definition or classification of LM hallucinations (Ji et al., 2023). However, as Kalai and Vempala (2023) show, LMs that assign a positive probability to every response token must hallucinate at least some of the time. A key contribution of our work is thus to propose a typology of legal hallucinations, which we summarize in Table 1. It is essential to recognize that there are various ways in which an LLM might generate hallucinated information, and not all types are equally concerning for legal professionals.
First, a model might hallucinate by producing a response that is either unfaithful to or in conflict with the input prompt, a phenomenon referred to as closed-domain hallucination. This is a major concern in tasks requiring a high degree of accuracy between the response and a long-form input, such as machine translation (Xu et al., 2023) or summarization (Cao et al., 2018). In legal contexts, such inaccuracies are particularly problematic in activities like summarizing judicial opinions, synthesizing client intake information, drafting legal documents, or extracting key points from an opposing counsel’s brief.
Second, LLMs might also hallucinate by producing a response that either contradicts or does not directly derive from its training corpus. Following Agrawal et al. (2023), we conceptualize this kind of hallucination as one form of open-domain hallucination. In general, the output of a language model should be logically derivable from the content of its training corpus, regardless of whether the content of the corpus is factually or objectively true.For example, if a training corpus consisted of J. K. Rowling’s Harry Potter series, we would expect an LLM to produce the sentence “Tom Marvolo Riddle” in response to a query about Voldemort’s real name. However, if the training corpus consisted solely of Jane Austen’s Pride and Prejudice (for instance), we would consider this LLM output to be a hallucination—because there would be no basis in the training data for making such a claim about Voldemort. In legal settings, this kind of hallucination poses a special challenge to those aiming to fine-tune the kind of general purpose foundation models that we study in this paper with proprietary, in-house work product.For example, this kind of firm-specific fine-tuning is the business model of a prominent legal tech startup, Harvey.ai (Ambrogi, 2023). For example, firms might have a catalog of internal research memos, style guides, and so forth, that they want to ensure is reflected in their bespoke LLM’s output. At the same time, however, insofar as creativity is valued, certain legal tasks—such as persuasive argumentation—might actually benefit from some lack of strict fidelity to the training corpus; after all, a model that simply parrots exactly the text that it has been trained on could itself be undesirable. Defining the contours of an unwanted “hallucination” in this context requires value judgements about the balance between fidelity and spontaneity.
Finally, the third way that an LLM can hallucinate is by producing a response that lacks fidelity to the facts of the world, irrespective of how the LLM is trained or prompted (Maynez et al., 2020). We consider this to be another type of open-domain hallucination, with the key concern being “factuality” in relation to the facts of the world (cf. Wittgenstein, 1998 ). In our context, this is perhaps the most alarming type of hallucination, as it can undermine the accuracy required in any legal context where a correct statement of the law is necessary.
3 Hallucination Trade-offs
In this paper, we investigate only the last kind of hallucination. As mentioned, the first two modes of hallucination are not always problematic in the legal setting: these kinds of hallucinations could actually be somewhat desirable to lawyers if they resulted in generated language that, for example, removed unnecessary information from a given argument (at the expense of being faithful to it) or invented a novel analogy never yet proposed (at the expense of being grounded in the lexicon) (Cao et al., 2022). However, what a lawyer cannot tolerate is the third kind of hallucination, or factual infidelity between an LLM’s response and the controlling legal landscape. In a common law system, where stare decisis requires attachment to the “chain” of historical case law (Dworkin, 1986), any misstatement of the binding content of that law would make an LLM quickly lose any professional or analytical utility.
Focusing on non-factual hallucinations alone, however, comes with certain trade-offs. One of the advantages of our typology is that it makes clear that it may not always be possible to minimize all modes of hallucination simultaneously (Kalai and Vempala, 2023); indeed, reducing hallucinations of one kind may increase hallucinations of another. For example, if a given prompt contains information that does not conform to facts about the world, then ensuring response fidelity with respect to the former would by definition produce infidelity—i.e., hallucination—with respect to the latter. More generally, although fidelity to the prompt is necessary for avoiding closed-domain hallucination, there is an important sense in which prioritizing such behavior might actually induce the kind of open-domain hallucination that we center in this paper.
These trade-offs present unavoidable challenges for prospective users of legal LLMs. When responding to a query, should an LLM be skeptical of its prompt or sycophantic to it? If it has been trained on case law from one jurisdiction, should it enforce adherence to that training corpus even when responding about the law in another jurisdiction? If facts about the world conflict with each other—as legal rules often do—should the LLM preserve that nuance or refrain from introducing information outside the scope of a query? Questions like these are ultimately questions about which kinds of legal hallucinations are more and less preferable, and they are questions whose answers require both empirical evidence and normative arguments. We supply some of the empirics in this paper (see Sections 5.1.6 and 5.2), but stress that the normative considerations are crucial and should be a topic of future legal hallucination research.
Profiling Hallucinations Using Legal Research Tasks
To empirically assess non-factual hallucination, we adopt a QA framework where the goal is to test an LLM’s ability to produce accurate information in response to different kinds of legal queries. We develop fourteen tasks that are representative of such queries, which we group into three categories in order of increasing complexity and list in Figure 2.
In the low complexity category, we ask for information that we consider relatively easy for an LLM to reproduce. The information in this category does not derive from the actual content of a case itself, so it does not require higher-order legal reasoning skills to internalize. Instead, this information is readily available in a case’s caption or its syllabus—standard textual locations whose patterns even non-specialized LLMs should be able to recover. We therefore expect LLMs to perform best on these tasks:
Existence: Given the name and citation of a case, ascertain whether the case actually exists. This basic evaluation provides preliminary insights into an LLM’s knowledge of actual legal cases: if it cannot distinguish real cases from non-existent ones, it probably cannot offer detailed case insights. We use only real cases in our prompts, so affirming their existence is the correct answer.In Appendix E.1, we experiment with using fake cases as well.
Court: Given the name and citation of a case, provide the name of the court that ruled on it. This task assesses an LLM’s knowledge about legal jurisdictions, an important building block of a case’s precedential value. We perform this task across three different levels of the federal judiciary: the Supreme Court, whose opinions bind all lower courts; the Courts of Appeals, which operate one level down and set the law for courts within their geographic region; and the District Courts, whose opinions do not carry general lawmaking or binding authority over other courts.
Each level of the court system has a different reporter, or the series of volumes that opinions are published in. This is relevant because the reporter is included in the citation that we provide to the LLM, essentially revealing the level of the hierarchy that an opinion is from. All and only SCOTUS cases are published in the U.S. Reports. Opinions from the thirteen Courts of Appeals are published in the Federal Reporter, and District Court cases are published in the Federal Supplement. Because of this, we expect this task to be more difficult as we descend the hierarchy of courts. There is only one court associated with the U.S. reporter, but 13 associated with the Federal Reporter and 94 associated with the Federal Supplement. For USCOA cases, we require the name of the specific Court of Appeals, and for USDC cases, we require the name of the specific District Court.
Citation: Given a case name, supply the Bluebook citation of the case. This query tests an LLM’s ability to associate a given dispute with its official record in a reporter volume at a particular page, which is the key way in which different opinions reference and link to each other. For USCOA cases, we further specify that we want the citation for the Court of Appeals opinion, and for USDC cases, we further specify that we want the citation for the District Court opinion. We test for citation equality using eyecite (Cushman et al., 2021).
Author: Given the name and citation of a case, supply the name of the opinion author. This query tests an LLM’s ability to associate a given case with a particular judge, which is important for contextualizing a case in the broader jurisprudential landscape. For SCOTUS and USCOA cases, we further specify that we want the name of the majority opinion author. We accept a fuzzy match of the opinion author’s name as accurate.
2 Moderate Complexity Tasks
Next, in the moderate complexity category, we start to require an LLM to evince knowledge of actual legal opinions themselves. To answer the queries in this category, an LLM must know something about a case’s substantive content; these queries seek information that must be collated from idiosyncratic portions of its text. Of course, a database-augmented LLM might still be able to retrieve some of this information without ever actually internalizing the content of a case, but we expect this text-based knowledge to be less available than the information described in the low complexity category. Specifically, we ask for the following information:
Disposition: Given a case name and its citation, state whether the court affirmed or reversed the lower court. This query tests an LLM’s knowledge of how the court resolved the instant appeal confronting the parties in the case, which is the first step for determining the holding that is created by the case. This is essentially a binary classification task where we accept correct “affirm” or “reverse” labels as accurate. We filter out all ambiguous dispositions (e.g., reversals in part) and we do not ask this query of USDC cases because District Courts are courts of original jurisdiction.While it is possible for some administrative agency decisions to be appealed to a district court, this occurs infrequently enough that we choose not to ask for case disposition at the district court level.
Quotation: Given a case name and its citation, supply any quotation from the opinion. This query tests an LLM’s ability to produce some portion of an opinion’s text verbatim, which is an important feature for lawyers seeking to use a case to stand for a specific proposition. Normally, such memorization is considered an undesirable property of LLMs (Carlini et al., 2022), but in this legal application it is actually desirable behavior. We accept any fuzzy string of characters appearing in the majority opinion as accurate.
Authority: Given a case name and its citation, supply a case that is cited in the opinion. This query probes an LLM’s understanding of the chain of precedential authority that supports a given opinion. We do not distinguish between positive and negative citations for this task; we accept any precedent cited in any way in the text of the majority opinion as accurate. We extract and match citations on their volumes, reporters, and pages using eyecite (Cushman et al., 2021).
Overruling year: Given a case name and its citation, supply the year that it was overruled. This query tests an LLM’s ability to recognize when a given case has been subsequently altered, which is crucial information for lawyer seeking to determine whether a given precedent is still good law or not. This task is the most complicated in this category because it requires the LLM to draw connections between multiple areas of the case space. We accept only the exact year of overruling as accurate, and we limit this task to only those SCOTUS cases that have been explicitly overruled (n=279).In Section 5.2, we experiment with prompting with cases that have never been overruled as well.
3 High Complexity Tasks
Finally, in the high complexity category, we seek answers to tasks that both presuppose legal reasoning skills (unlike the low complexity tasks) and are not readily available in existing legal databases like WestLaw or Lexis (unlike the moderate complexity tasks). These tasks all require an LLM to synthesize core legal information out of unstructured legal prose—information that is frequently the topic of deeper legal research. In Section 4.3, we explain how we test LLMs’ knowledge of some of these more complex facts without necessarily having access to the ground-truth answers ourselves:
Doctrinal agreement: Given two case names and their citations, state whether they agree or disagree with each other. This query requires an LLM to show knowledge of the precedential relationship between two different cases, information that is essential for higher-order legal reasoning. We use Shepard’s treatment codes as a basis for constructing this task, filtering out all ambiguous citation treatments (e.g., neutral treatments) and coarsening the unambiguous codes into “agree” and “disagree” labels that we accept as accurate. For this task, we use a relatively balanced dataset of 2,839 citing-cited case pairs coded as “agree,” and 2,161 citing-cited case pairs coded as “disagree.” This task is limited to SCOTUS cases as our underlying dataset only contains thorough Shepard’s data for citations to the Supreme Court.
Factual background: Given a case name and its citation, state its factual background. This query tests an LLM’s understanding of the concrete fact pattern underlying a case, which is helpful in assessing the relevance of the case to current research and in drawing parallels with other cases.
Procedural posture: Given a case name and its citation, state its procedural posture. This query tests an LLM’s understanding of how and why a case has arrived at a particular court, which aids in understanding the precise question presented and standard of review applicable.
Subsequent history: Given a case name and its citation, state its subsequent procedural history, if any. This query tests an LLM’s knowledge of any other related proceedings that concern the given case after a particular decision, which is information that can change or clarify the legal significance of the case.
Core legal question: Given a case name and its citation, state the core legal question at issue. This query tests an LLM’s ability to pinpoint the main issue or issues that a court is addressing in a case, which is the most important factor in assessing whether a case is apposite or not.
Central holding: Given a case name and its citation, state its central holding. This query tests an LLM’s knowledge of the legal principle that a given case stands for, i.e., the precedent that future cases will rely upon or distinguish from. Articulating the holding of a case is crucial for legal analysis and argumentation and is the most complex task that we evaluate.
Experimental Design
We aim to profile hallucination rates across several legally salient dimensions, including hierarchy, jurisdiction, time, and case prominence. Thus, we construct our test data with an eye toward making statistical inferences on these covariates.
We begin with the universe of case law from each level of the federal judicial hierarchy—namely, SCOTUS, USCOA, and USDC—that has been published in the volumes of the U.S. Reports, the Federal Reporter, and the Federal Supplement. To ensure balance over time and place, we then perform stratified random sampling using year strata for the SCOTUS cases, circuit-year strata for the USCOA cases, and state-year strata for the USDC cases. We draw 5,000 cases from each level of the judiciary. Finally, we merge these units with metadata obtained from the Caselaw Access Project (2023), the Supreme Court Database (Spaeth et al., 2022), the Appeals Courts Database Project (Songer, 2008; Kuersten and Haire, 2011), the Library of Congress (Congress.gov, 2023), and Shepard’s Citations (Fowler et al., 2007; Black and Spriggs, 2013).More information about how we use these metadata to construct each query is available in Appendix A.
2 Reference-based Querying
One effective way to study hallucinations in the open-domain setting is to use a test oracle—or an external reference—to detect and adjudge non-factual responses (Lin et al., 2022; Lee et al., 2023; Li et al., 2023). Such oracles are usually difficult and costly to construct (Krishna et al., 2021), but we exploit the tabular metadata described in Section 4.1 to develop ours. Our assumption is that while LLMs have access to and are trained on the raw text of American case law, which is in the public domain (Henderson et al., 2022), they have not yet explicitly memorized these cases’ attendant metadata, which exist separately from the cases’ textual content and which we have aggregated from disparate sources.
These metadata enable us to construct reference-based queries for the first nine of our tasks (Figure 2). These queries take the form of question-and-answer triples , where is the question, is the LLM’s greedy answer retrieved from calling , and is the known ground-truth answer. Our estimand of interest for each task is the population-level hallucination rate , which we estimate by averaging over the binary outcomes of our randomly sampled cases:
3 Reference-free Querying
Reference-based querying lets us directly recover our population parameter of interest, but two problems limit the effectiveness of the approach. First, we are restricted to asking questions for which digestible metadata exist and a clear answer has been recorded, which rules out many more complex inquires. Second, precisely because these queries can be answered with tabular data, legal database-augmented LLMs (Cui et al., 2023; Savelka et al., 2023) are likely to soon solve or at least mask hallucinated responses to these queries (Shuster et al., 2021; Peng et al., 2023).
To test the tasks that cannot be easily verified against an external legal database, we employ reference-free querying instead, which detects hallucinations by exploiting the stochastic behavior of LLMs at higher temperatures (Agrawal et al., 2023; Manakul et al., 2023; Min et al., 2023). This approach is rooted in the theory that hallucinations are more likely to originate in flat probability distributions with higher next-token uncertainties, whereas factual answers should always have a high probability of being the generated response given a prompt. Thus, by repeatedly querying an LLM at a non-greedy temperature, we can estimate the model’s hallucination rate by examining its self-consistency—factual responses should not change, but hallucinated ones will.
Most reference-free approaches implicitly assume that the LLM is calibrated, i.e., that there is indeed some correlation between its self-consistency and its propensity to hallucinate. For reasons that we discuss in Section 5.3, we are unwilling to make this assumption in our legal setting. We therefore adopt a slightly different implementation that is still reference-free, but only requires contradiction, not consistency (Mündler et al., 2023). Specifically, for our final five tasks (Figure 2), we construct reference-free queries in the form of question-and-answer triples , where is the question, is one LLM answer retrieved by calling once, and is another LLM answer retrieved by calling again. Detecting a hallucination then amounts to detecting a logical contradiction between the two stochastic answers: any such contradiction guarantees non-factuality, because two contradictory answers cannot both be correct (Mündler et al., 2023).
To identify these contradictions at scale, we feed both answers into GPT 4 and ask it for its assessment. This technique does not assume anything about ’s calibration—it just requires that GPT 4 possess logical reasoning skills sufficient to compare ’s two responses and accurately label them as contradictory as not. To justify this reliance on GPT 4, we manually label a portion of the reference-free responses ourselves and conduct an intercoder reliability analysis to ensure that GPT 4 is indeed able to perform this task. Full information about our procedure and a validity check is provided in Appendix C. (We find that GPT 4’s reliability is comparable to human labeling of contradictions.)
An important caveat of this approach is that it only allows us to establish a lower bound on the hallucination rate for our reference-free queries:
Although self-contradiction guarantees hallucination, the inverse does not hold: two answers may be logically consonant but still lack fidelity to the law. Because we are unwilling to assume calibration, we accept this inferential limitation, but, as we show below, even the lower bounds on hallucination rates are quite high and informative.
4 Models
We perform our experiments using three popular LLMs: OpenAI’s ChatGPT 3.5 (gpt-3.5-turbo-0613) (OpenAI, 2023b), Google’s PaLM 2 (text-bison-001) (Anil et al., 2023), and Meta’s Llama 2 (Llama-2-13b-chat-hf) (Touvron et al., 2023).
We run each query under both zero- and three-shot prompting setups. We provide the full text of the prompts we use for each query, along with the few-shot examples, in Appendix B. All of our raw queries and responses, including timestamp logs, are available in our GitHub repository.
Results
We begin by presenting our main results profiling the LLMs’ hallucination rates, which cut to the core of popular concerns over LLMs’ suitability for various legal research tasks (Section 5.1). Then, after showing that hallucinations are generally widespread, and highlighting some salient variation in the LLMs’ tendencies to hallucinate, we turn to two additional challenges that threaten LLMs’ utility for legal research: (1) their susceptibility to contra-factual bias, i.e., their ability to handle queries based on mistaken legal premises (Section 5.2), and (2) their certainty in their responses, i.e., their self-awareness of their propensity to hallucinate (Section 5.3).
Tables 2, 3, and 4 report our estimated hallucination rates and their standard errors for each category of our tasks. We find that hallucinations vary with they substantive complexity of the task (Section 5.1.1), the hierarchical level of the court (Section 5.1.2), the jurisdictional location of the court (Section 5.1.3), the prominence of the case (Section 5.1.4), the year the case was decided (Section 5.1.5), and the LLM queried (Section 5.1.6). We do not find substantial differences between zero-shot and few-shot prompting, so we focus our discussion on the few-shot results alone.Throughout, we also drop from our analyses any instances of LLMs refusing to respond to our queries. We document these abstentions in Appendix F, but they are generally rare and do not affect our findings.
As we hypothesized in Section 3, we first observe that hallucinations increase with the complexity of the legal research task at issue, which we visualize in Figure 3. Starting with the low complexity category (Table 2), the LLMs perform best on the simple Existence task, though this is in part driven by their tendency to always answer “yes” when asked about the existence of any case. (In Appendix E.1 we demonstrate this problem by asking about the existence of fake cases instead.) The models begin to struggle more when prompted for information about a case’s Court, Citation, or Author. Hallucinations then surge among the moderate complexity tasks (Table 3), all of which require the LLMs to evince knowledge of the actual content of a legal opinion. We note that these results are not just a product of different evaluation metrics: although the Quotation task, for example, requires near-word reproduction of particular sentences and phrases to be judged correctly, the Disposition task simply asks for binary responses from the model. Yet, the LLMs hallucinate widely in both setups.
The results for the high complexity tasks (Table 4) confirm this general pattern of poor performance, but must be interpreted slightly differently. First, the Doctrinal agreement task is another binary classification task, so the LLMs’ hallucination rates on this task—near 0.5—represent little improvement over random guessing. This suggests that LLMs are not yet able to perform the kind of legal reasoning that attorneys perform when they assess the precedential relationship between cases—a core purpose of legal research.
Second, regarding the remaining tasks in the high complexity category, it is important to keep in mind that the hallucination rates that we report for these tasks are only lower bounds on the true rates, as these tasks are evaluated using our reference-free method (Section 4.3). To provide some context for these bounds, we note that in a similar self-contradiction setup, Mündler et al. (2023) found that GPT 3.5 hallucinated about 14.3% of the time on general QA queries. On our legal QA queries, GPT 3.5 and our other LLMs far surpass this baseline rate—and it is possible that the true hallucination rate is even higher.
For example, we find that even on one of the easier reference-free tasks—Procedural posture—our LLMs hallucinate at least 57% of the time. Performance degrades further on the most complex Core legal question and Central holding tasks, with hallucinations arising in response to at least 75% of our queries. Hallucinations are lowest among GPT 3.5 responses to the Subsequent history task at the SCOTUS level, but this is because the model simply tends to state that the litigation concluded with the Supreme Court decision. This may not actually be correct—many Supreme Court cases result in a remand and have additional procedural history in lower courts. However, we are unable to capture this kind of mistake, as our methodology only permits us to identify hallucinations where the model contradicts itself. We are not able to capture repeated incorrect answers as instances of hallucination, meaning that our estimate of hallucination in the SCOTUS Subsequent history task is likely to understate the rate of hallucination by a larger margin that other tasks.
Taken together, these results invite skepticism about LLMs’ abilities to conduct complex forms of legal research. Our reference-free tasks, in particular, raise serious doubts about LLMs’ knowledge of substantive aspects of American case law—the very knowledge that attorneys must often synthesize themselves, instead of merely looking up in a database.
1.2 Hallucinations Vary by Court
We next examine trends by hierarchy, exploring LLMs’ abilities to restate the case law of the three different levels of the federal judiciary.
We find that across all tasks and all LLMs, hallucinations are lowest in the highest levels of the judiciary, and vice-versa (Figure 4). Thus, our LLMs perform best on tasks at the SCOTUS level, worse on tasks at the USCOA level, and worst on tasks at the USDC level. These results are encouraging insofar as it is important for LLMs to be knowledgeable about the most authoritative and wide-ranging precedents, but discouraging insofar as they suggest that LLMs are not well attuned to localized legal knowledge. After all, the vast majority of litigants do not appear before the Supreme Court, and may benefit more from knowledge that is tailored to their home District Court—their court of first appearance.
1.3 Hallucinations Vary by Jurisdiction
To better understand the relationship between different courts and hallucinations, we next zoom in on the middle level of the judicial hierarchy—the Courts of Appeals—and examine horizontal heterogeneity across the circuits.Because not all Courts of Appeals were created at the same time, for parity in comparison here we exclude from our results cases decided before 1982, the year the youngest circuit—the Federal Circuit—was created. We report the full, non-truncated results in Appendix Section E.2, which are largely consistent with these post-1981 results. Figure 5 depicts these results geographically, showing lower hallucination rates in lighter colors and higher rates in darker colors. Pooling our tasks together, we see the best performance in the Ninth Circuit (comprising California and adjacent states in yellow), the Second Circuit (comprising New York and adjacent states in soft green), and the Federal Circuit (a special court headquartered in Washington, D.C.). By contrast, performance tends to be worst in the circuits in the geographic center of the country.
These results confirm popular intuitions about the influential role that the Second, Ninth, and Federal Circuits play in the American legal system. Because it encompasses New York City, the Second Circuit has traditionally had a significant impact on financial and corporate law, and many landmark decisions in securities law, antitrust, and business litigation have come from this court. The Ninth Circuit, on the other hand, handles more cases than any other federal appellate court, and often issues rulings that advance progressive positions that lead to disproportionate review by the Supreme Court. Finally, the Federal Circuit exercises exclusive appellate jurisdiction over patent and trademark cases, inter alia, enjoying unique influence in those areas of the law.
Perhaps surprisingly, however, our results do not confirm popular intuitions about the D.C. Circuit, which is generally thought to be the most influential appellate division. In our tasks, our LLMs perform only at an average level within this circuit—not worst, but not best either.
1.4 Hallucinations Vary by Case Prominence
To probe the role of legal prominence more directly, we move to SCOTUS-level results next, examining the relationship between case importance and hallucinations. To measure case prominence within this single level of the judiciary, we use the Caselaw Access Project’s PageRank percentile scores, a metric of citation network centrality that captures the general legal and political prominence of a case.
We find that case prominence is negatively correlated with hallucination, reaffirming our results from above (Figure 6). However, we also note that a sharp slope change occurs around the 90th prominence percentile in the GPT 3.5 and PaLM 2 models. This suggests that the bias of these LLMs—but not Llama 2—may be skewed even more toward the most well-known decisions of the American legal system, even within the SCOTUS level.
1.5 Hallucinations Vary by Case Year
Because case law develops in virtue of new decisions building on old ones over time, the age of a case may be another useful predictor of hallucination. Examining this relationship at the SCOTUS level in Figure 7, we find a non-linear correlation between hallucination and age: hallucinations are most common among the Supreme Court’s oldest and newest cases, and least common among its post-war Warren Court cases (1953-1969). This result suggests another important limitation on LLMs’ legal knowledge that users should be aware of: LLMs’ peak performance may lag several years behind the current state of the doctrine, and LLMs may fail to internalize case law that is very old but still applicable and relevant law.
1.6 Hallucinations Vary by LLM
Finally, we also partition our results by the LLM itself and compare across models. We find that not all LLMs are equal: GPT 3.5 performs best overall, followed by PaLM 2, followed by Llama 2 (see Figure 1 at beginning of manuscript).
We also discover tendencies towards different inductive biases, or the predisposition of an LLM to generate certain outputs more frequently than others. In Figure 8, we highlight one of these biases for our SCOTUS-level Author task, which asks the LLM to supply the name of the justice who authored the majority opinion in the given case. Each LLM we test has slightly different inductive preferences; some err towards the most recognizable justices, but others are a little more difficult to explain. For example, Llama 2 disproportionately favors Justice Story—an influential jurist who authored the famous Amistad opinion, among others—whereas PaLM 2 prefers Justice McLean—also an important jurist, but one more known for his dissents than his majority opinions, such as his dissent in the infamous Dred Scott case. Across the board, all our LLMs tend to overstate the true prevalence of justices at a higher magnitude than they understate them, as indicated by the greater dispersion of the points above the line in Figure 8.
These biases demonstrate one way that LLMs inevitably encounter the kind of hallucination trade-off that we discuss in Section 2.3. If the inductive bias that an LLM learns from its training corpus is not well-aligned with the true distribution of facts about the world, then the LLM is likely to make systematic errors when queried about those facts. Moreover, the persistence of inductive biases also increases the risk of LLMs instantiating a kind of legal “monoculture” (Kleinberg and Raghavan, 2021). Instead of accurately restating the full variation of the law, LLMs may simply regurgitate information from a few prominent members of the response set that they have been trained on, flattening legal nuance and producing a falsely homogenous sense of the legal landscape.
2 Contra-factual Bias
We now turn to the first of two potential failure points that we seek to highlight for LLMs performing legal research tasks, beyond their sheer propensity to hallucinate: their bias toward accepting legal premises that are not anchored in reality and answering queries accordingly. We view this behavior as a particular kind of model sycophancy (the tendency of an LLM to agree with a user’s preferences or beliefs, even when the LLM would reject the belief as wrong without the user’s prompting; Sharma et al., 2023; Wei et al., 2023) or general cognitive error (Tversky and Kahneman, 1974; Jones and Steinhardt, 2022; Suri et al., 2023).
This bias poses a subtle but pernicious challenge to those aiming to use LLMs for legal research. When a researcher is learning about a topic, they are not only unsure about the answer, they are also often unsure about the question they are asking as well. Worse, they might not even be aware of any defects in their query; research by its nature ventures into the realm of “unknown unknowns” (Luft and Ingham, 1955). This is especially true for unsophisticated pro se litigants, or those without much legal training to begin with. Relying on an LLM for legal research, they might inadvertently submit a question premised on non-factual legal information or folk wisdom about the law. As discussed in Section 2.3, this then forces a trade-off: if the LLM is too intent on minimizing prompt hallucinations, it runs the risk of simply recapitulating the user’s misconception and producing a non-factual hallucination instead.
To test whether this risk is real in the legal setting, we evaluate two modified versions of our reference-based queries, but with premises that are false by construction. Specifically, we ask the LLMs to (1) provide information about an author’s dissenting opinion in an appellate case in which they did not in fact dissent and (2) furnish the year that a SCOTUS case that has never been overruled was overruled. In both cases, we consider failing to provide the requested information an acceptable answer; any uncritical answering of the prompt is treated as a hallucination.
Table 5 reports the results of this experiment and Figure 9 summarizes them by LLM. In general, LLMs seem to suffer from contra-factual bias on these legal information tasks. As in the raw hallucination tasks, contra-factual bias hallucinations are higher in lower levels of the judiciary. Substantively, they are also greatest for the question with a false overruling premise, possibly reflecting the increased complexity of the question asked.
Llama 2 performs exceptionally well, demonstrating little contra-factual hallucination. However, this success is linked to a different kind of hallucination—in many false dissent examples, for instance, Llama 2 often states that the case or justice does not exist at all. (In reality, all of our false dissent examples were created with real cases and real justices—just justices who did not author a dissent for the case.) Under our metrics for contra-factual hallucination, we choose to record these examples as successful rejections of the premise. The kind of error that Llama 2 makes here is already measured in its poor performance on other tasks, especially Existence.
3 Model Calibration
The second potential failure point that we investigate is model calibration, or the ability of LLMs to “know what they know.” Ideally, a well-calibrated model would be confident in its factual responses, and not confident in its hallucinated ones (Kadavath et al., 2022; Xiong et al., 2023; Tian et al., 2023b; Yin et al., 2023; Azaria and Mitchell, 2023). If this property held for legal queries, researchers would be able to adjust their expectations accordingly and could theoretically learn to trust the LLM when it is confident, and learn to be more skeptical when it is not (Zhang et al., 2020). Even more importantly, if an LLM knew when it was likely to be hallucinating, the hallucination problem could be in principle solvable through some form of reinforcement learning from human feedback (RLHF) or fine-tuning, with unconfident answers simply being suppressed (Tian et al., 2023a).
To study our LLMs’ calibration on legal queries, we estimate the expected calibration error (ECE) for each of our tasks. Appendix D describes our estimation strategy in full, but, intuitively, it entails extracting a confidence score for each LLM answer that we obtain and comparing it to the empirical hallucination rate that we observe. Table 6 reports the results of this analysis at the task level, and Figure 10 pools our findings at the LLM level by plotting those two metrics—confidences and empirical non-hallucination frequencies—against each other, binned into 10 equally-sized bins (represented by the dots). In a perfectly calibrated model, the confidences and empirical frequencies would be perfectly correlated along the diagonal.
Overall, we note that PaLM 2 () and GPT 3.5 () are significantly better calibrated than Llama 2 (). Diving into the task-level results, we see that across all LLMs, calibration is poorer on our more complex tasks and on tasks directed toward lower levels of the judicial hierarchy. ECE is also higher on our partially open-ended tasks such as Court and Author. In these tasks, the LLM has a large but finite universe of responses, and the high ECE for these tasks reflects the LLMs’ tendencies to over-report on the most prominent or widely known members of the response set.
In all cases, the calibration error is in the positive direction: our LLMs systematically overestimate their confidence relative to their actual rate of hallucination.In Appendix D, we explore whether this bias can be corrected with an ex post scaling adjustment, but conclude that challenges remain. This finding, too, suggests that users should exercise caution when interpreting LLMs’ responses to legal queries, especially those of Llama 2. Not only may they receive a hallucinated response, but they may receive one that the LLM is overconfident in and liable to repeat again.
Discussion
We began this paper with a question that has surged in salience over the last twelve months: Will AI systems like ChatGPT soon reshape the practice of law and democratize access to justice? Although recent research suggests that LLMs are performing increasingly well on a number of legal benchmarking tasks (Blair-Stanek et al., 2023; Choi et al., 2022; Fei et al., 2023; Guha et al., 2023; Nay et al., 2023; OpenAI, 2023a; Trozze et al., 2023), we highlight the problem of legal hallucinations, which remains a serious obstacle to the adoption of these models. Confirming popular perceptions (Weiser, 2023; Romoser, 2023), we show that factual legal hallucinations are widespread in the language models that we study—OpenAI’s ChatGPT 3.5, Google’s PaLM 2, and Meta’s Llama 2—on the bulk of the legal research tasks that we profile (Section 5.1).
We also push beyond conventional wisdom by surfacing two additional behaviors that threaten LLMs’ utility for legal research: (1) their susceptibility to contra-factual bias, i.e., their inability to handle queries containing an erroneous or mistaken starting point (Section 5.2), and (2) their certainty in their responses, i.e., their inability to always “know what they know” (Section 5.3). Unfortunately, we find that LLMs frequently provide seemingly genuine answers to legal questions whose premises are false by construction, and that under their default configurations they are imperfect predictors of their own tendency to confidently hallucinate legal falsehoods.
These findings temper recent enthusiasm for the ability of off-the-shelf, publicly available LLMs to accelerate access to justice (Tito, 2017; Perlman, 2023; Tan et al., 2023). Our results suggest that the risks of using these generic foundation models are especially high for litigants who are:
Filing in courts lower in the judicial hierarchy or those located in less prominent jurisdictions
Seeking more complex forms of legal information
Formulating questions with mistaken premises
Unsure of how much to “trust” the LLMs’ responses
In short, we find that the risks are highest for those who would benefit from LLMs most—indigent or pro se litigants. Not only do LLMs hallucinate widely, their current implementations lack the behavioral features that such users would require. Ideally, LLMs would do best at localized legal information (rather than SCOTUS-level information), be able to correct users when they ask misguided questions (rather than accepting their premises at face value), and be able to moderate their responses with the appropriate level of confidence (rather than hallucinating with conviction). We therefore echo concerns that the proliferation of LLMs may ultimately exacerbate, rather than eradicate, existing inequalities in access to legal services (Simshaw, 2022; Draper and Gillibrand, 2023). Indeed, increased reliance on LLMs may actually produce a kind of legal “monoculture” (Kleinberg and Raghavan, 2021), with users being fed information from only a limited subset of judicial sources that elide many of the deeper nuances of the law.
We also emphasize that the challenges presented by legal hallucinations are not only empirical, but also normative. Although data-rich and moneyed players certainly stand at an advantage when it comes to building hallucination-free legal LLMs for their own private use, it is not clear that even infinite resources can entirely solve the hallucination problem we diagnose. As we discuss in Section 2.3, model fidelity to the training corpus, model fidelity to the user’s prompt, and model fidelity to the facts of the world—i.e., the law—are normative commitments that stand in tension with each other, despite all being independently desirable technical properties of an LLM. Ultimately, since hallucinations of some kind are generally inevitable (Kalai and Vempala, 2023), developers of legal LLMs will need to make choices about which type(s) of hallucinations to minimize, and they should make these choices transparent to their downstream users.For example, Casetext, a prominent legal AI provider, claims that its CoCounsel tool works by “eliminating the serious limitations—like hallucinations—that curb the professional utility of [its] model” (Casetext, 2023, emphasis added). However, without further qualification of the types of hallucinations that have been eliminated, blanket statements of hallucination elimination may actually further obscure the true risks of using LLMs from the user. Only then can individual litigants decide for themselves whether the legal information they seek to obtain from LLMs is trustworthy or not.
In the meantime, more experienced legal practitioners may find some value in consulting LLMs for certain tasks, but even these users should remain vigilant in their use, taking care to verify the accuracy of their prompts and the quality of their chosen LLM’s responses.
Acknowledgements
We thank Neel Guha, Sandy Handan-Nader, Adam T. Kalai, Peter Maldonado, Chris Manning, Joel Niklaus, Kit Rodolfa, and Andrea Vallebueno for helpful discussions and feedback.
References
Appendix
Appendix A Data Sources
In Section 4.1, we describe the data and sampling strategy that we use to construct our queries. For clarity, Table 7 links these data to each task. Our exact procedure for sampling, merging, and aggregating across these datasets is available in our GitHub repository.
Appendix B Prompt Templates
The full zero-shot and few-shot prompt templates for all of our queries are shown in Figures 14 to 28. The few-shot examples presented are those used for the SCOTUS queries; appropriate cases from the other levels of the judiciary are used in the USCOA and USDC versions.
Appendix C Contradiction Detection Approach
In Section 4.3, we describe our strategy for assessing hallucinations for our reference-free tasks. Briefly, because we do not have access to ground-truth labels for these tasks, we exploit the stochastic behavior of LLMs at higher temperatures and check for contradictions in their responses to repeated queries. In Figure 29, we share the template for the contradiction elicitation query that we send to GPT 4 to perform this contradiction labeling at scale. We frame our contradiction check as an entailment task (Wang et al., 2021) as we find that this prompting strategy is the most performant in our setting.
To confirm that GPT 4 is able to reliably detect the contradictions that we envision, we crosscheck GPT 4’s conclusions against our own expert knowledge on a subsample of our data. Specifically, for 100 of our queries, two of us independently recoded the LLMs’ responses for contradictions. We report Cohen’s coefficient (Cohen, 1960) for our agreement with GPT-4 on these queries in Table 8. In general, a value greater than 0.80 is considered “almost perfect” agreement, and one greater than 0.60 is considered “substantial” agreement (Landis and Koch, 1977). Our values suggest that we are well-justified to use GPT 4 for contradiction detection; indeed, one of us agrees with GPT 4 more than with the other human coder.
To give the reader a sense of GPT 4’s reasoning abilities in this setting, in Figures 30 and 31 we share some examples of its reasoning process and contradiction conclusions.
Appendix D Expected Calibration Error (ECE)
In Section 5.3, we estimate the expected calibration error (ECE) of our LLMs for each of our tasks. Studying model calibration directly is not possible in our setup because we do not always observe our LLMs’ conditional probability distributions, which we require in order to determine whether they are in fact correlated with their response accuracies. Specifically, two of the models that we evaluate—OpenAI’s ChatGPT 3.5 and Google’s PaLM 2—are closed-source and do not expose this information to the user. To overcome this hurdle, we instead estimate the distributions by drawing samples from the model at temperature 1 and comparing those responses to its greedy response as follows:
Following Kumar et al. (2020), we use a plug-in estimator that bins the data into 10 equally-sized slices and subtracts off an approximated bias term.
D.2 Temperature Scaling
As we report in Section 5.3, our results suggest that out of the box, the LLMs we evaluate are not well-calibrated on legal queries. However, a model that is uncalibrated under its unscaled probability distribution is not necessarily uncalibrated tout court; the purpose of a LLM’s temperature parameter, after all, is to allow an sophisticated user to adjust the distribution as needed. To explore whether such temperature scaling could indeed affect our results, here we perform ex post Platt scaling (Guo et al., 2017) on the raw distribution and check for improvements in the measured ECE.
Figure 11 visualizes the rescaled ECE for each of our LLMs, and Table 9 quantifies the numerical gains. As expected, rescaling generally improves the ECE. However, the rescaling procedure is not perfect: GPT 3.5 and PaLM 2 remain relatively uncalibrated in the confidence interval. The pooled ECE of Llama 2 improves substantially, but this is due to the entire distribution being compressed into the interval—the rescaled LLama 2 model is simply not confident in any of its responses. Overall, these results confirm our conclusion in Section 5.3 that LLMs face calibration challenges on legal knowledge queries.
Appendix E Supplementary Analyses
In Table 2 we report results for the Existence task, where our LLMs are asked whether or not a given case exists. Performance on this task is generally strong; however, because all our prompted cases are in fact real, it is unclear whether these results are due to the LLMs’ genuine knowledge of a case’s existence or simply their tendency to always answer “yes” to this type of question. Accordingly, here we conduct a supplemental analysis where we repeat the existence query, but using fake cases instead of real cases. Each fake case citation is constructed with plausible party names and the reporter that is appropriate for the court at issue, e.g., SolarFlare Technologies v. Armstrong, 656 F.3d 262 for a fake USCOA case.
Appendix Table 10 reports the results of this Fake case existence experiment. We see that GPT 3.5 and PaLM 2 are both prone to simply asserting the existence of any case—real or fake—though PaLM 2 is more discriminating at the SCTOUS level. Llama 2, on the other hand, appears immune from this behavior, but recall from Table 2 that this is just bias in the opposite direction: it simply denies the existence of any case, real or fake. Altogether, these results belie the LLMs’ seemingly excellent performance on the Existence task: even here, they may not possess any actual knowledge of the true details of many cases.
E.2 Hallucination Rates at the USCOA Level (No Time Cutoff)
In Figure 5 in Section 5.1.3 above, we show the relationship between hallucinations and USCOA jurisdiction. However, in that figure, we exclude cases decided prior to 1982 in order to fairly compare rates across older and younger jurisdictions. Here, we share the results for that analysis with no cases excluded. As Figure 12 suggests, the non-truncated results substantively mirror the truncated ones: the Ninth, Eleventh, and Federal Circuit continue to perform best. However, in this figure, the Eleventh Circuit has a somewhat stronger showing as well. We believe that this is best explained by its relative infancy: it split from the Fifth Circuit in 1981 so it is composed of newer cases only.
E.3 Hallucination Rates by State
In Section 5.1.3, we show within-USCOA sources of hallucination heterogeneity on a geographic map. Figure 13 depicts a similar analysis for our USDC tasks, aggregated to the state level. Confirming our results above, we observe the lowest level of hallucinations in the state of New York, which is comprised of the Northern District, Southern District, Western District, and Eastern District of New York.
Appendix F Abstention Rates
Occasionally, our LLMs abstain from providing an answers to our queries. For example, they may plead ignorance or simply claim that they are unable to answer. When this occurs, we drop these responses from our analyses. Table 11 reports the LLMs’ abstention rates for each task, which are generally low.