LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models
Neel Guha, Julian Nyarko, Daniel E. Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N. Rockmore, Diego Zambrano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory M. Dickinson, Haggai Porat, Jason Hegland, Jessica Wu, Joe Nudell, Joel Niklaus, John Nay, Jonathan H. Choi, Kevin Tobia, Margaret Hagan, Megan Ma, Michael Livermore, Nikon Rasumov-Rahe, Nils Holzenberger, Noam Kolt, Peter Henderson, Sean Rehaag, Sharad Goel, Shang Gao, Spencer Williams, Sunny Gandhi, Tom Zur, Varun Iyer, Zehua Li
Introduction
Advances in large language models (LLMs) are leading American lawyers and administrators to reexamine the practice of law .In using “LLMs”, we are referring to language models which evince in-context learning capabilities (also referred to as “foundation models” ). This behavior has traditionally been observed in models with at least a billion parameters. Proponents have argued that LLMs could alter how lawyers approach tasks ranging from brief writing to corporate compliance . By making legal services more accessible, they could eventually help alleviate the United States’ long standing access-to-justice crisis . This perspective is informed by the observation that LLMs possess special properties which, it is argued, make them more suited for legal tasks. The models’ capacity to learn new tasks from limited labeled data would reduce the manual data annotation costs that ordinarily burden the development of legal language models . Their apparent proficiency at sophisticated reasoning tasks would also make them ideal for the rigor of law, which requires parsing obtuse texts with heavy jargon, and inferential processes which combine different modalities of reasoning .
This excitement, however, is tempered by the fact that legal applications often involve significant risk . Existing work has shown that LLMs are capable of generating content that is offensive, misleading, and factually incorrect . Such behaviors—if replicated in legal applications —could result in substantial harms , with much of the potential burden imposed on traditionally marginalized and under-resourced populations . The safety implications thus create a pressing need to develop infrastructure and processes for benchmarking LLMs in legal contexts.
However, significant challenges face practitioners seeking to assess whether LLMs can perform legal reasoning. The first challenge is the limited ecosystem of legal benchmarks . The majority of existing benchmarks, for example, focus on tasks which models learn by finetuning or training on task-specific data . These benchmarks do not measure the aspects of LLMs which generate excitement for law—namely, their ability to perform many different tasks using only few-shot prompts. Relatedly, benchmarking efforts have focused on professional certification exams like the Uniform Bar Exam , but these are not always representative of the actual use-cases for LLMs. The second challenge is the incongruity between the ways in which existing benchmarks and lawyers frame “legal reasoning.” Existing benchmarks coarsely generalize all tasks involving legal data or laws as measuring “legal reasoning.” In contrast, lawyers recognize that legal reasoning is a broad umbrella term encompassing many distinct types of reasoning . Different legal tasks require different skills and bodies of knowledge. Because existing legal benchmarks fail to draw these distinctions, it is difficult for legal professionals to contextualize the performance of modern LLMs within their own understanding of legal competency. In short: legal benchmarks do not use the same vocabulary or conceptual frameworks as the legal profession.
In light of these limitations, we believe that rigorously evaluating the legal reasoning capabilities of LLMs will require the legal community to take a more proactive role in the process of benchmarking. To that end, we present LegalBench: the first steps towards constructing an interdisciplinary collaborative legal reasoning benchmark for the English language.https://github.com/HazyResearch/legalbench/ Over the past year, the authors of this paper—drawing from their diverse legal and computer science backgrounds—came together to assemble 162 tasks (from 36 different data sources), each of which measures a specific type of legal reasoning. LegalBench is thus, to the best of our knowledge, the first open-source legal benchmarking effort. We believe that this style of benchmark construction—where domain experts take an active and participatory role in the crafting of evaluation tasks—illustrates one approach to interdisciplinary collaboration in LLM research. Importantly, we believe it also shows that legal professionals have an essential role to play in the assessment and development of LLMs for law.
As a research project, we highlight three components of LegalBench:
LegalBench was constructed from a mix of existing legal datasets (restructured for the few-shot LLM paradigm), and hand-crafted datasets created and contributed by legal professionals (included as authors on this work). The legal professionals involved in this collaboration were asked to contribute datasets that they believed to either measure an interesting legal reasoning skill, or to capture a practically useful application for LLMs in the law. High performance on LegalBench tasks thus provides useful information, allowing lawyers to validate their assessment of an LLM’s legal competency, or identify an LLM that could be used in their workflow.
LegalBench tasks are organized into an extensive typology which describes the types of legal reasoning required to perform the task. Because this typology is drawn from frameworks familiar to the legal community, it enables legal professionals to meaningfully engage in discussions of LLM performance, using a terminology and conceptual framework familiar to them .
Finally, LegalBench is intended as a platform to support further research. For AI researchers who lack legal expertise, LegalBench comes with significant support for understanding how to prompt and evaluate different tasks. And as more of the legal community begins to engage with the potential impact and role of LLMs, we hope to grow LegalBench by continuing to solict and incorporate tasks from legal professionals.Cognizant of LegalBench’s current skew towards American law, we hope that additional contributions incorporate tasks from other jurisdictions.
In this paper, we make the following contributions:
First, we present a typology for organizing and describing legal tasks in terms of the types of reasoning they require. This typology is drawn from frameworks lawyers use to describe legal reasoning .
Second, we provide an overview of the tasks in LegalBench, describing the process by which they were constructed, important dimensions of heterogeneity, and limitations. A full description of each task is provided in the Appendix.
Finally, we use LegalBench to evaluate 20 LLMs from 11 different families, across a range of size points. We make observations regarding the performance of different models and present an initial study into different prompt-engineering strategies. Ultimately, these results are intended to highlight different directions of future work that LegalBench may enable.
We hope that this benchmark will be interesting to a diverse set of communities. Practitioners may use these tasks to determine whether and where LLMs can be integrated into existing workflows to improve outcomes for clients. Legal academics may benefit from observing the types of annotation that LLMs are capable of , and different forms of empirical scholarly work they may enable. Computer scientists may benefit from studying the performance of these models in a domain like law, where distinct lexical properties and unique tasks may surface new insights.
Before we progress further, we note that the purpose of this work isn’t to evaluate whether computational systems should replace lawyers and legal officers, or to understand the positive and negative impacts of that replacement . Rather, our goal is to construct artifacts that enable the relevant stakeholders and affected communities to better understand, empirically, the capacity for LLMs to perform different types of legal tasks. Given the proliferation of computational legal tools, we believe that answering this question is vital for ensuring their safe and ethical usage.
Related work
Understanding the extent to which NLP models can perform tasks or skills traditionally associated with lawyers—or be useful in legal analysis—has been the focus of significant work . Researchers have approached this question in a variety of ways . First, prior work has identified manually arduous tasks currently performed by lawyers—like forms of document review or case summarization —and developed benchmarks to assess the performance of current state-of-the-art techniques. Here, research has focused on the aspects of legal text which are often challenging for NLP methods, like the length of documents or the presence of jargon . A second line of work has focused on developing tasks to evaluate forms of inferential reasoning common to law . This includes, for instance, tasks which require a model to identify the best supporting statement for an argument , or perform statutory reasoning . Other work has focused on creating datasets for pretraining models , non-English/multilingual tasks , legal judgement prediction , legal role labeling , and different forms of retrieval .
Importantly, the majority of previous benchmarking efforts have focused on language models which learn by supervised training or finetuning (e.g., BERT variants ), and researchers have consequently studied questions related to the role of domain specific datasets . More recently, researchers have begun to ask whether large language models (LLMs) like GPT-3/4 can perform legal reasoning , citing to evidence of these models’ capacity to perform sophisticated reasoning tasks in domains like math or programming . Unlike BERT-based models, LLMs are evaluated on their ability to learn tasks in-context, primarily through prompting. While a few works have experimented with LLMs on existing benchmarks , most evaluations focus on standardized tests or other exam equivalents . Studies have explored the role of prompt-engineering , potential applications , questions regarding human-LLM interaction , and comparisons to older finetuned-models .
LegalBench builds on prior work in several ways. First, LegalBench enhances opportunities to study legal reasoning in LLMs, by making available 162 evaluation tasks. LegalBench systematizes and standardizes these tasks for LLM evaluation, specifying potential prompts, in-context demonstrations, and metrics. Second, LegalBench presents a framework for organizing and comparing tasks, allowing researchers to identify trends in performance across groupings of tasks. This enables researchers, for instance, to distinguish between task types for which current LLMs are highly performant, and task types for which further work is needed.
A notable consequence of focusing on few-shot LLMs is that LegalBench can contribute a much more diverse set of legal reasoning tasks. Traditional NLP methods require a large training set and a smaller evaluation set. The cost of legal annotations means that constructing benchmarks has required extraordinary financial investment or a “natural” source of existing annotations . Because the few-shot prompting regime requires only a few labeled demonstrations, creating large training sets isn’t necessary, and the effort they otherwise would have consumed can be allocated towards developing new tasks.
2 Connections to other LLM benchmarking efforts
We highlight connections to two broader research efforts. First, we draw inspiration from existing efforts within NLP and machine learning to define fine-grained measures of performance, which allow researchers to discuss model capabilities with precision and specificity. Examples include the diagnostic set of the GLUE Benchmark , the “reasoning patterns” studied in , the task organization used in HELM , and the BigBench effort . Fine-grained measurements are valuable because they allow researchers to identify how particular modifications to model architectures or training regimes affect performance. They hold particular value for the field of legal NLP, in which researchers continue to debate how best to specialize language models to the domain
We additionally draw inspiration from other large-scale collaborative efforts in AI, including the BigBench project , and studies in medicine . In particular, we believe that LegalBench illustrates a new model of open-source and interdisciplinary collaboration between the legal and AI communities. To the extent that LLMs gain adoption for legal tasks, legal professionals will be primarily charged with supervising them and selecting application use-cases. Involving the legal community in the design and construction of evaluation tasks allows for the construction of benchmarks which are more responsive to their interests and information needs.
The LegalBench typology
LegalBench identifies six types of legal reasoning that LLMs can be evaluated for: (1) issue-spotting, (2) rule-recall, (3) rule-application, (4) rule-conclusion, (5) interpretation, and (6) rhetorical-understanding. We first justify the selection of these types by providing background on how the legal profession frames “legal reasoning,” and the connections to our typology. We then illustrate how task datasets may be used to evaluate LLMs for each type, using examples from LegalBench.
Though this framework draws heavily on American legal thought, we find it can be easily extended to characterize LegalBench tasks that implicate non-American bodies of law. We also note that our types are non-exhaustive, and in future work hope to consider additions to these types.
American legal scholars often describe “legal reasoning” as the process of determining the legal conditions that arise from a set of events or occurrences, with reference to both prior cases and codified laws . A common framework for executing this type of legal reasoning is the Issue, Rule, Application and Conclusion (IRAC) framework . In this framework, legal reasoning decomposes into four sequential steps.
First, lawyers identify the legal issue in a given set of facts (issue-spotting). An issue is often either (1) a specific unanswered legal question posed by the facts, or (2) an area of law implicated in the facts. Depending on the setting, a lawyer may be told the issue, or be required to infer a possible issue.
Second, lawyers identify the relevant legal rules for this issue (rule-recall). A rule is a statement of law which dictates the conditions that are necessary (or sufficient) for some legal outcome to be achieved. In the United States, rules can come from a variety of sources: the Constitution, federal and state statutes, regulations, and court opinions (case law). Importantly, rules often differ between jurisdictions. Hence, the relevant rule in California might be different than the relevant rule in New York.
Third, lawyers apply these rules to the facts at hand (rule-application). Application, or the analysis of rule applicability, consists of identifying those facts which are most relevant to the rule, and determining how those facts influence the outcome under the rule. Application can also involve referencing prior cases involving similar rules (i.e. precedent), and using the similarities or differences to those cases to determine the outcome of the current dispute.
Finally, lawyers reach a conclusion with regards to their application of law to facts, and determine what the legal outcome of those facts are (rule-conclusion).
We illustrate this framework with a simple example. Suppose that BusinessMart—a large manufacturing corporation—is being sued by Amy in federal court on diversity jurisdiction.Diversity jurisdiction gives federal courts the ability to hear cases between parties that are “citizens” of different states. BusinessMart sells the majority of its goods in Texas, has its headquarters (where its CEO and board members sit and work) in California, and maintains a factory in Florida. A court is trying to determine—for the purposes of diversity jurisdiction—where BusinessMart’s “principal place of business is.”
Issue-spotting: Here, a narrow issue is offered—where is BusinessMart’s principal place of business?
Rule-recall: A lawyer would recognize that the most relevant rule here comes from the case Hertz Corp. v. Friend,Hertz Corp. v. Friend, 559 U.S. 77 (2010). in which the Supreme Court determined “that the phrase ‘principal place of business’ refers to the place where the corporation’s high level officers direct, control, and coordinate the corporation’s activities.”
Rule-application: Applying this rule to the facts above yields two observations. First, a corporation’s CEO and board members are examples of high level officers referred to in Hertz that control and conduct a company. Second, the place where BusinessMart’s high level officers control the company is California, as that is where the CEO and board sit and work.
Rule-conclusion: Based on the chain of inference spelled out in the application stage, a lawyer would thus conclude that California is BusinessMart’s principal place of business.
The extent to which the outcome of the application and conclusion steps follow each other is dictated by the level of ambiguity in the fact patterns. When the law on a particular question is clear and there is little ambiguity in the facts (as the case in the above example), then the application and conclusion steps point towards the same outcome. Sometimes however, the facts may be unclear or contested, and reasonable minds may differ as the conclusion step. For now, LegalBench focuses entirely on the former setting (unambiguous answers), and all tasks are considered to have objectively “correct” answers.
Though IRAC is the most formal framework for legal reasoning, lawyers recognize a variety of skills which are useful to practice of law . For instance, lawyers are often required to exercise interpretive skills, in order to identify the rights, obligations, or limitations of certain legal language (e.g., what a contractual clause may or may not enable). They must also exhibit rhetorical skills, and understand the types of arguments that are made. Though these tasks require the knowledge base and skill set of lawyers, they, arguably, do not always fit neatly within the IRAC framework. Hence, we consider these to be distinct from the examples offered in the previous section.
2 Evaluating legal reasoning in large language models
LegalBench identifies six categories of legal reasoning. For each category, we describe how a LLM task may evaluate the typified legal reasoning, using examples from LegalBench.
LegalBench evaluates issue-spotting through tasks in which an LLM must determine if a set of facts raise a particular set of legal questions, implicate an area of the law, or are relevant to a specific party. Issue tasks evaluate a LLM’s ability to reason over the legal implications of different activities, events, and occurrences. An example of an issue-spotting task is the learned_hands_benefits task, which requires an LLM to determine (Yes/No) whether a post on a public legal aid forum raises issues related to welfare law (i.e., public benefits or social services). The box below shows how a LLM might be prompted for this task.
LegalBench evaluates rule-recall through tasks which require the LLM to generate the correct legal rule on an issue in a jurisdiction (e.g., the rule for hearsay in US federal court). A rule task can be an open-ended generation task—in which the LLM must generate the text of the rule for a jurisdiction—or a classification task—in which the LLM must determine whether the rule exists in that jurisdiction. Anchoring to jurisdiction is important, as legal rules differ across different jurisdictions. Rule tasks are particularly useful for measuring hallucinations . An example of a rule-recall task is rule_qa, a question-answer task where questions include asking the model to state the formulations for different legal rules, identify where laws are codified, and general questions about doctrine.
LegalBench evaluates rule-conclusion through tasks which require an LLM to determine the legal outcome of a set of facts under a specified rule. LLMs are evaluated purely on whether their predicted outcome is correct. For example, the ucc_v_common_law task asks a LLM to determine whether a contract is governed by the Uniform Commercial Code (UCC) or the common law of contracts. The LLM is always provided with the relevant rule, via the prompt (see below).
LegalBench evaluates rule-application through the same tasks used to measure rule-conclusion. When evaluating rule-application however, we prompt the LLM to provide an explanation of how the rule applies to a set of facts, and evaluate the quality of the generated explanation along two dimensions: (1) whether the explanation is correct, and (2) whether it contains analysis. Each metric captures a different dimension upon which a particular rule-application may be good.
Correctness corresponds to the criteria that explanations should not contain errors. We focus on five types of errors: misstatements of the legal rule, misstatements of the fact pattern, incorrectly asserting the legal outcome, logic errors, and arithmetic errors. Analysis corresponds to the criteria that explanations should contain inferences from the facts that are relevant under the rule, and illustrate how a conclusion is reached. Consider, for example, an explanation which restates the rule, the fact pattern, and the predicted legal outcome. If the predicted legal outcome is correct, than the explanation in its entirety would be correct, because it contains no error. However, as prior works have noted , examples like this are conclusory, and often unsatisfactory in the context of legal work.
To standardize evaluation and enable future work, we have released an “answer guide” for each task used for rule-application, which contains the inferences required for each sample, and describes common modes of errors. All evaluations in LegalBench for rule-application have been performed with respect to this answer-guide.
Table 1 presents an examples of how three generations (corresponding to the Alice/Bob example above) would be evaluated under the above metrics. The first generation is incorrect, because it misstates the rule. The second generation is correct because it contains no falsehoods, but performs no analysis because it does not articulate inferences. The third generation is both correct and contains analysis, because it has no errors, and explicitly mentions an essential inference (e.g., that a bike is a “good”).
LegalBench evaluates interpretation through tasks which require the LLM to parse and understand a legal text. Interpretive tasks provide the LLM with a text, and ask the LLM to either extract a relevant piece of information, answer a question, or categorize the text by some property. Interpretive tasks are among the most studied and practically relevant tasks in LegalBench, and many have been taken from actual use-cases. An example of an interpretive task is cuad_audit_right, which asks the LLM to determine if a contractual clause contains an “audit right.” An example is shown below:
LegalBench evaluates rhetorical-understanding through tasks which require an LLM to reason about legal argumentation and analysis. In these tasks, an LLM is provided with a legal argument (usually excerpted from a judicial opinion), and asked to determine whether it performs a certain function or has a certain property. An example is the definition_classification task, in which an LLM must determine if a sentence from a judicial opinion provides a definition of a term.
Rhetorical-understanding example: definition_classification Does the sentence define a term? Sentence:“To animadvert carried the broader implication of “turn[ing] the attention officially or judicially, tak[ing] legal cognizance of anything deserving of chastisement or censure; hence, to proceed by way of punishment or censure.” 1 Oxford English Dictionary 474 (2d ed.1989).” Answer: Yes We emphasize one aspect of LegalBench: IRAC in this work is used as an organizing principle for grouping tasks. On a law exam, a student would be expected to generate an answer which structurally resembles IRAC, where each step builds on the inferences of the previous step . LegalBench tasks, in contrast, each evaluate a single type of legal reasoning. Hence, a task like learned_hands_benefits can only be used to evaluate issue-spotting, and not rule-recall. In future work we hope to add tasks which evaluate multiple steps jointly.
LegalBench tasks
Appendix F discusses each task in detail, providing a description of the reasoning that each task evaluates, how task data was constructed, task examples, and evaluation protocols. This section provides an overview of LegalBench.
LegalBench tasks are drawn from three sources. The first source of tasks are existing available datasets and corpora. Most of these were originally released for non-LLM evaluation settings. In creating tasks for LegalBench from these sources, we often significantly reformatted data and restructured the prediction objective. For instance, the original CUAD dataset contains annotations on long-documents and is intended for evaluating extraction with span-prediction models. We restructure this corpora to generate a binary classification task for each type of contractual clause. While the original corpus emphasized the long-document aspects of contracts, our restructured tasks emphasize whether LLMs can identify the distinguishing features of different types of clauses. The second source of tasks are datasets that were previously constructed by legal professionals but never released. This primarily includes datasets hand-coded by legal scholars as part of prior empirical legal projects (e.g., ). The last category of tasks are those that were developed specifically for LegalBench, by the authors of this paper. Overall, tasks are drawn from 36 distinct corpora.
In August 2022, we published a call for tasks, describing the goals of the project and its structure . We publicized the project through mailing lists and legal computational conferences. Submitted tasks were vetted for legal correctness and task validity. Task contributors are drawn from diverse professional backgrounds within the law (e.g., academics, practitioners, computational legal researchers) and constitute the authors of this paper.
LegalBench comes with support designed to enable non-law AI researchers to use and study LegalBench tasks. First, each LegalBench task is accompanied by extensive documentation describing how the task is performed, its legal significance, and the construction procedure. The objective of this documentation is to provide AI researchers with a working understanding of the mechanical processes behind each task, for the purposes of better understanding LLM performance. Second, each task is accompanied by a “base” prompt, which contains task instructions and demonstrations. The base prompt is provided to promote replicability and standardization. We anticipate that future research efforts building off of LegalBench will identify higher performing prompts/prompt formats. We intended to update the LegalBench GitHub repository with these prompts as they are discovered.
We note several limitations of the current LegalBench tasks (additional limitations are noted in Appendix B). First, when this project began, most LLM context-windows were constrained to a few pages of text. As a result, the initial round of LegalBench tasks does not involve longer documents. We hope to include such tasks in future work, particularly as recent technical developments have resulted in significantly longer context windows . Second, LegalBench’s tasks focus on legal reasoning questions with objectively correct answers. LegalBench is thus not helpful for evaluating legal reasoning involving degrees of correctness or tasks where “reasonable minds may differ.” Third, LegalBench only considers English language tasks, is skewed towards certain jurisdictions (American law), and certain areas of the law (contracts). Thus, the current iteration of the benchmark limits inferences regarding how LLMs may generalize to legal tasks involving other jurisdictions. As we continue to solicit and incorporate contributions to LegalBench, we hope to add tasks addressing these limitations. Finally, LegalBench evaluates IRAC abilities independently, while law exams and other legal work requires lawyers to generate outputs which follow IRAC in a multi-hop matter (i.e., each aspect is applied to the same fact pattern).
2 Dimensions of variation
All LegalBench tasks contain at least 50 samples, with an average task size of 563 samples (Appendix D.4). These tasks are comparable in size to those used in benchmarking efforts like BigBench , HELM or RAFT . LegalBench tasks also span different formats: multiple-choice questions (35 tasks), open-generation (7 tasks), binary classification (112 tasks), and multi-class/multi-label classification (8 tasks).
LegalBench provides tasks for each of the reasoning categories discussed above: rule-recall (5 tasks), issue-spotting (16 tasks), rule-application (16 tasks), rule-conclusion (16 tasks), interpretation (119 tasks), and rhetorical-understanding (10 tasks). Tasks are predominantly drawn from areas of law implicating civil matters, including contracts (58 tasks), civil procedure (8 tasks), evidence law (1 task), and corporate law (58 tasks). The skew towards interpretation tasks and tasks from contract law can be explained by the ubiquity of legal documents from these areas (e.g., contracts, terms-of-service agreements, disclosures, and etc.) and their immediate commercial implications .
Legal language is highly heterogeneous, varying in sentence structure, vocabulary, and rhetorical style across different legal areas and document types . This poses a distinct challenge for LLMs, which are extremely sensitive to structure of input text and the vocabulary used . LegalBench tasks are drawn from a diverse set of legal language types, thus enabling researchers to study performance variation across different categories of legal text. Specifically, LegalBench encompasses tasks with language drawn from plain English (32 tasks), legal opinions (11 tasks), merger agreements (34 tasks), contracts (55 tasks), statutory text (3), and other sources.
3 Tasks
We offer a brief summary of the tasks present in each reasoning category.
There are 17 issue-spotting tasks. 16 tasks are derived from the “Learned Hands” Dataset (Section F.14). Each of these tasks is a binary classification task, in which the LLM must determine if a post from /r/legaladvice implicates a particular domain of law (e.g., immigration). The last task is the corporate_lobbying task (Section F.7), which requires determining if a legislative bill has legal implications for a described company.
There are 5 rule-recall tasks. Two tasks require an LLM to either generate the citation for a particular legal quote, or identify if a candidate citation is correct (Section F.3). The remaining three tasks are:
rule_qa, in which the LLM must generate the text of different legal tests and identify where they’re codified (Section F.25).
international_citizenship_questions, in which the LLM must answer yes/no questions about citizenship requirements in different countries (Section F.13).
nys_judicial_ethics, in which the LLM must answer yes/no questions corresponding to different ethical rules under the guidance provided by the New York State Advisory Committee on Judicial ethics (Section F.17).
There are 12 tasks used for both rule-application and rule-conclusion.
Six tasks evaluate an LLM’s ability to apply the diversity jurisdiction test to information about plaintiff and defendant citizenships and the amount-in-controversy for different claims (Section F.9). This requires both arithmetic and logical reasoning. The simplest (diversity_1) involves one plaintiff, one defendant, and one legal claim. The most complex (diversity_6) involves two plaintiffs, two defendants, and two claims against each defendant.
abecrombie evaluates an LLM’s ability to apply the Abercrombie test to classify how distinctive a product/service name is for a particular product/service (Section F.1).
hearsay evaluates an LLM’s ability to identify—given a particular piece of evidence and an issue being litigated—whether the evidence would count as hearsay for that issue (i.e., an out-of-court statement introduced to prove the truth of the matter asserted) (Section F.11).
personal_jurisdiction evaluates an LLM’s ability to identify when a court in a particular forum may excercise personal jurisdiction over a defendant, given basic facts about the defendant’s place of domicile, their interactions with the state, and the claims brought against them by plaintiffs (Section F.21).
successor_liability evaluates an LLM’s ability to identify the potential successor liability exceptions present in fact patterns describing a sale of assets from one company to another (Section F.29).
telemarketing_sales_rule evaluates an LLM’s ability to identify whether the representations made by a company covered under the Telemarketing Sales Rule violate either 16 C.F.R. § 310.3(a)(1) and 16 C.F.R. § 310.3(a)(2), which outline a series of specific telemarketing sales practices defined as “deceptive” (Section F.31).
ucc_v_common_law evaluates an LLM’s ability to determine whether a particular contract is covered by the Uniform Commercial Code (UCC) or the common law, given information about the contract (Section F.33).
consumer_contracts_qa, which evaluates an LLM’s ability to determine the rights/obligations imposed by terms of service clauses from popular websites (Section F.5).
contract_qa, which evaluates an LLM’s ability to identify different types of contractual provisions.
14 tasks designed from the ContractNLI dataset . Each task evaluates an LLM’s ability to identify whether a candidate contract excerpt adheres to a task-specific assertion (Section F.6).
38 binary-classification tasks designed from the CUAD dataset . Each task evaluates an LLM’s ability to identify whether a candidate contractual clause is of a certain type (Section F.4.1).
insurance_policy_interpretation, which evaluates an LLM’s ability to determine whether a particular claim is covered by an insurance policy (Section F.12).
jcrew_blocker, which evaluates an LLM’s ability to identify whether a particular loan clause is a J.Crew Blocker provision (Section F.4.2).
34 tasks from the MAUD dataset , which evaluates an LLM’s ability to answer multiple-choice questions about the content of excerpts from merger-agreements (Section F.16). Each task corresponds to a different question.
9 tasks from the OPP-115 dataset , each of which evaluates an LLM’s ability to determine whether a privacy policy clause discusses a particular issue (Section F.18). Each task is a binary classification task corresponding to a different issue.
privacy_policy_entailment , which evaluates an LLM’s ability to answer entailment questions from privacy policies (Section F.22).
privacy_policy_qa , which evaluates an LLM’s ability to determine if a clause from a privacy policy contains the answer to a particular question (Section F.23).
2 tasks designed from the SARA dataset , which evaluate an LLM’s ability to interpret and apply sections of the tax-code (Section F.26).
10 tasks which evaluate an LLM’s ability to identify when a supply chain disclosure discusses or describes a particular type of information (Section F.30). Each task corresponds to a different disclosure objective.
unfair_tos , which evaluates an LLM’s ability to classify clauses from terms of service agreements into one of muliple categories (Section F.4.3).
There are 10 tasks which evaluate rhetorical-understanding.
canada_tax_court_outcomes evaluates an LLM’s ability to identify the outcome of a tax court decision, based on the text of the decision (Section 15).
2 tasks evaluate an LLM’s ability to (1) identify sentences from US Supreme Court opinions which define a term, and (2) extract that term (Section F.8).
function_of_decision_section evaluates an LLM’s ability to identify the function that an excerpt of a legal opinion has (e.g., statement of rule) (Section F.10).
legal_reasoning_causality evaluates an LLM’s ability to identify when an excerpt of a court’s opinion relies on statistical evidence (Section F.15).
oral_argument_question_purpose evaluates an LLM’s ability to identify the purpose that a particular question (from Supreme Court oral arguments) plays (Section F.19).
overruling evaluates an LLM’s ability to identify when a sentence from a judicial opinion overrules a previous case (Section F.20).
scalr evaluates an LLM’s ability to assess which holding statement (amongst several options) best answers a provided legal question.
2 tasks evaluate an LLM’s ability to identify whether excerpts of judicial reasoning rely on certain textualist tools (Section F.32). Each task corresponds to a different tool.
Results
We use LegalBench to conduct a three-part study.
In the first part (Section 5.2), we conduct a sweeping evaluation of 20 LLMs from 11 different families, at four different size points. We use this study to make initial observations on performance differences across families, the role of model size, and the gap between open-source and commercial LLMs.
In the second part (Section 5.3), we show how LegalBench can be used to conduct in-depth evaluations of models. To illustrate, we use LegalBench to highlight similarities and differences in the performance of three popular commercial models: GPT-4, GPT-3.5, and Claude-1.
In the final part (Section 5.4), we show how LegalBench can support the development of law-specific LLM methods. We focus on prompting, and conduct a series of experiments that begin to surface tradeoffs and challenges with regards to guiding LLMs towards certain tasks.
Ultimately, our study here serves to illustrate the types of analyses that LegalBench enables, and highlight potential directions for future work.
We study three commercial API-access models. From the OpenAI GPT family, we study GPT-3.5 (text-davinci-003) and GPT-4 . Results from these models were retrieved between May and August of 2023. From the Anthropic family, we study Claude-1 (v1.3) . Results from this model were retrieved in July of 2023. These models are believed to be large (hundreds of billions of parameters), though exact details on their architecture and training process are unknown. It is thus possible that some LegalBench tasks leaked into pretraining data. Details on the extent to which different LegalBench tasks have been previously made available online can be found in Appendix D.
We study 17 open-source models at three different size points: 3B parameters, 7B parameters, and 13B parameters. All inference was performed on two-GPU GCP 40GB A100s, using the Manifest library . HuggingFace links for each model are provided in Appendix G.
From Together, we study three models: Incite-Instruct-7B, Incite-Base-7B, and Incite-Instruct-3B .
From Meta’s OPT family, we study three models: OPT-2.7B, OPT-6.7B, and OPT-13B .
From TII’s Falcon family, we study Falcon-7B-Instruct .
From MosaicML’s MPT family, we study MPT-7B-8k-Instruct .
From LMSYS’ Vicuna family, we study Vicuna-7B-16k and Vicuna-13B-16k .
From Google’s FLAN-T5 family, we study Flan-T5-XL (3B parameters) and Flan-T5-XXL (11B parameters) .
From Meta’s LLama-2 family, we study LLaMA-2-7B, and LLaMA-2-13B .
From the Wizard family, we study WizardLM-13B .
From the BigScience BLOOM family, we study BLOOM-3b and BLOOM-7B .
Our selected LLMs represent only a sample of the models available. For instance, we do not evaluate LLMs larger than 13B parameters, which have been observed to perform well . Studied LLMs are also “general domain,” in that we don’t find evidence that any were specifically customized to perform well on legal text.We note that as of July 2023, we were unable to identify public law-specific English large language models to evaluate. In future work we hope to expand our evaluation to a broader set of LLMs.
1.2 Prompts
We designed a prompt for each task by manually writing instructions for the task, and selecting between zero and eight samples from the available train split to use as in-context demonstration. The number of samples selected depended on the availability of data and the sequence length of samples (Appendix G.2). For instance, the inputs to the Supply Chain Disclosure tasks are disclosure statements between 1-2 pages long, making the inclusion of multiple demonstrations infeasible. For application evaluation, we augmented the prompt with an instruction for the LLM to explain its reasoning.
We used the same prompts across all LLMs with one exception. In contrast to the OpenAI and open-source LLMs, Anthropic recommends specific prompting formats when using Claude.https://docs.anthropic.com/claude/docs/introduction-to-prompt-design This includes surrounding in-context samples with
LLM outputs were generated using next-token generation at a temperature of . For classification/extraction tasks, we terminated at a new-line token. For rule_qa and all application tasks except diversity_jurisdiction_6 we generated 150 tokens. For diversity_jurisdiction_6 we generated 300 tokens.
We believe there is significant scope for improving and refining prompts on LegalBench. Hence, our results here provide a lower-bound on performance, as better prompts may elicit higher scores. Our prompts correspond to what we believe would be reasonable, based on experience with prompt engineering in other settings, and the guidance provided by model developers. We make all prompts available as a starting point for future work on LegalBench.
1.3 Evaluation
Classification tasks are evaluated using "exact-match" (following HELM ). Because some tasks contain significant label imbalances, we use balanced-accuracy as a metric. For extraction tasks, we perform normalization on generated outputs to account for differences in tense/casing/punctuation. A few tasks (e.g., successor_liability and ssla_individual_defendants) requires the LLM to produce multiple classes or extracted terms per instance. For these, we evaluate using F1. Appendix E provides more details.
Rule-application tasks were evaluated manually by a law-trained individual, who analyzed LLM responses for both correctness and analysis.For the six diversity jurisdiction tasks, we sampled 30 instances from each task. For all other rule-application tasks, we manually evaluated the entirety of the dataset. This type of manual evaluation is consistent with previous works evaluating LLM generations in the legal domain . As rule-application requires LLMs to generate “explanations” detailing legal reasoning—a capability primarily exhibited by larger models—we only evaluated GPT-4, GPT-3.5, and Claude-1. rule_qa was also manually evaluated by a law-trained individual. Appendix E provides more details on our approach to manual grading. All manual evaluation was performed with reference to a grading guide, which we additionally make available.
2 Performance trends
Table 2 provides the average task performance for all 20 models in five reasoning categories (issue-spotting, rule-recall, rule-conclusion, interpretation, and rhetorical-understanding). The first block of rows corresponds to large commercial models, the second block corresponds to models in the 11B-13B range, the third block corresponds to models in the 6B-7B range, and the final block corresponds to models in the 2B-3B range. Table 3 provides the average task performance for the three large models on rule-application. Appendix G provides full results for each model on each task.
Overall, we find significant variation in performance across tasks, suggesting that LegalBench captures a diverse spectrum of difficulty (Appendix G). These results emphasize that assessments of LLM capabilities for legal applications must be made on a task-by-task basis, and informed by the nuances of specific tasks. While certain types of tasks appear beyond the scope of current-day LLMs, others seem more within reach. In this section, we offer preliminary observations on performance trends across model size, family, and reasoning categories.
Within LLM families, we observe that larger models usually outperform smaller models. For instance, Flan-T5-XXL (11B parameters) outperforms Flan-T5-XL (3B parameters) on average across all five reasoning categories, and LLaMA-2-13B outperforms LLaMA-2-7B on average across four reasoning categories. Notably, the margin of the gap varies across LLM families and reasoning categories. For instance, on rule-recall, the 7B Incite-Instruct model outperforms the 3B Incite-Instruct model by almost 10pts, while the 6.7B OPT model outperforms the 2.7B OPT model by less than 1pt. We additionally note that the largest LLM (GPT-4) outperforms virtually all other models.
Even for LLMs of the same size, we find considerable differences in performance. For instance, we observe significant gaps in performance between Flan-T5-XXL (11B parameters) and Vicuna-13B-16k (13B parameters), across all reasoning categories. This suggests, unsurprisingly, that the choice of pretraining data, regime of instruction-tuning, and architecture play an important role in determining performance, and that certain configurations may be better aligned for LegalBench tasks. Interestingly, we observe that such choices may affect which types of reasoning categories LLMs appear to perform well at. For instance, we observe that WizardLM-13B performs worse than all peers on issue-spotting tasks, best on rule-recall tasks, and nearly matches the performance of the best-performing peer on rule-conclusion tasks. Comparing Incite-7B-Instruct to Incite-7B-Base also provides insight into the effect of instruction-tuning across different categories, at one size point (7B parameters). We observe that instruction-tuning improves performance on four categories (issue-spotting, rule-conclusion, interpretation, and rhetorical-understanding), and worsens performance on rule-recall.
We additionally find that family-specific trends appear to hold across different size points. For instance, the Flan-T5 models outperform all others at both the 3B and 13B scale, while the Vicuna models appear to underperform competitors at both the 7B and 13B scale. We attribute the Vicuna models’ low performance to their frequency tendency to generate poorly-formed outputs, which did not map to the expected verbalizer tokens (e.g., blank spaces, random characters, etc.).In further experimentation, we found that writing prompts using the “### Human:” and “‘### Assistant:”’ templates did not appear to help. This could possibly be attributed to the type of data used to fine the model (e.g., user-conversation), although more in-depth experimentation is necessary.
Finally, we find evidence that open-source models are capable of performance that matches or exceeds certain commercial models. For instance, Flan-T5-XXL outperforms GPT-3.5 and Claude-1 on two categories (issue-spotting and rhetorical-understanding), despite the relative gap in parameter count. Notably, the gap between closed and open-source models is largest for the rule-conclusion category. Amongst LegalBench tasks, rule-conclusion tasks most like the other types of multi-step/common-sense reasoning tasks where commercial LLMs have been found to perform well.
3 Comparing GPT models
This section provides a more in-depth study of performance, focusing on the three commercial models (GPT-4, GPT-3.5, and Claude-1). The purpose of this section is to illustrate how LegalBench enables fine-grained analysis of LLM performance. In particular we highlight how LegalBench can provide more rigorous empirical support for anecdotal observations arising out of the legal community’s use of these models, and explain performance differences between models.
We first consider average model performance across all issue-spotting tasks. We observe that GPT-4 outperforms GPT-3.5 and Claude-1 (both at ).Statistical significance is computed using a paired -test over the tasks in the category. In absolute terms, issue tasks present the largest gap in performance between GPT-4 and other closed-API models, with an absolute margin of points. GPT-3.5 and Claude-1, in contrast, appear to match each other in performance, separated by an average gap of only 2 points. We additionally find that the open-source models perform poorly here. On 9 tasks, Incite-Base collapses to predicting a single class for all samples.
We note one limitation to our results: because 16/17 of our issue-spotting tasks are drawn from one source (Learned Hands data), average issue performance is skewed by properties of the Learned Hands data distribution (i.e.,, user-generated questions). For instance, though GPT-3.5 outperforms Claude-1 on 12/16 Learned Hands tasks, Claude-1 outperforms GPT-3.5 on the one non-Learned Hands task (corporate_lobbying). Despite the skew, we observe that these tasks appear to vary in difficulty. While GPT-4’s balanced-accuracy on learned_hands_torts is only 70.6%, on three tasks—learned_hands_immigration, learned_hands_traffic, and learned_hands_estate—it scores 95%.
3.2 Rule-recall
We first consider average model performance across all rule-recall tasks. While GPT-4 outperforms GPT-3.5 (), we surprisingly find that Claude-1 also outperforms GPT-3.5 (), and appears almost on par with GPT-4. Moreover, Claude-1 outperforms GPT-4 on three tasks: rule_qa, international_citizenship_questions, and nys_judicial_ethics. This is the only task category where Claude-1 provides performance comparable to GPT-4. Because little is known regarding the architecture and training processes for these models however, it is difficult to explain why this is the case.
Because rules/laws can be analogized to law-specific “facts,” rule-recall tasks are similar to general domain LLM tasks designed to measure “hallucination.” There, an extensive literature has documented the propensity for LLMs to both generate factually incorrect information, and answer fact-based questions incorrectly . Our results align with the primary findings of that literature. For example, we observe that the small open source models perform considerably worse than the larger models, consistent with the observation that model size plays an important role in fact-retention. Overall, performance on the rule-recall tasks also lend additional empirical support to more anecdotal reports—from the legal community—regarding how LLMs often misstate the law or cases .
3.3 Rule-application
Application tasks evaluate whether LLMs can explain how a legal rule applies to a set of facts, and verbalize the necessary inferences. With respect to correctness, we observe that GPT-4 outperforms both GPT-3.5 () and Claude-1 (). Across LLMs, we find that variation in performance across tasks is consistent with subjective impressions of task difficulty. For instance, performance on diversity_jurisdiction_1 (an easy task requiring a model to determine if an amount is greater than $75k and if the plaintiff and defendant are from different states) is much higher than performance on successor_liability (a harder task requiring a model to identify multiple successor liability exceptions in a fact pattern describing a complex transaction).
We observe that LLM generations may be incorrect in many different ways. On the Diversity Tasks, LLMs sometimes perform incorrect arithmetic operations or mathematical comparisons (i.e., stating that 75,000). On telemarketing_sales_rule in contrast, LLMs will cite to an incorrect portion of the rule. For instance, a generation may explain that certain conduct by a telemarketer runs afoul of the rule because the telemarketer failed to make a mandatory disclosure (16 CFR § 310.3(a)(1)), but cite to the portion of the rule prohibiting misrepresentations (16 CFR § 310.3(a)(2)). Examples of other types of incorrect generations can be found in Table 4.
With respect to analysis, we observe that GPT-4 again outperforms both GPT-3.5 () and Claude-1 (). Explanations which failed to exhibit analysis can be grouped into several categories. First, some generations will contain just a prediction as to the legal outcome, without an explanation (even when the LLM has been prompted to generate one). The same prompt–applied to other samples in the dataset–will elicit explanations containing analysis. Second, we observe a tendency for LLMs to sometimes generate explanations which merely restate the facts and legal rule, without actually offering an explanation for how the outcome is reached. Examples of such instances are provided in the table below.
3.4 Rule-conclusion
Rule-conclusion evaluates on the same tasks as rule-application, but only requires the LLM to generate a prediction as to the outcome, and not an explanation. We observe that GPT-4 once again outperforms GPT-3.5 () and Claude-1 (). Claude-1 and GPT-3.5 appear approximately level on performance.
The rule-conclusion tasks offer a heuristic for characterizing the types of legal inferences LLMs are capable of or struggle with. In particular, several of these tasks organize samples into slices, where the samples contained within a slice all represent a similar type of fact pattern, and thus interact with the legal rule in a comparable way. For instance, the hearsay task contains a slice corresponding to “non-verbal hearsay.” This slice contains fact patterns where an individual communicates something non-verbally (e.g., pointing), thus qualifying their conduct as a “statement” under the hearsay rule. In order to make accurate predictions on this slice, an LLM must recognize that (1) the hearsay rule applies to non-verbal communicative conduct, and (2) the non-verbal conduct in these fact patterns is communicative.
Though slices are small—and thus not intended for rigorous statistical analysis—they provide some intuition as to the source of GPT-4’s improvement over GPT-3.5, and the overall areas of strength and weakness for both models. On the hearsay task for instance (Table 6), the difference between GPT-4 and GPT-3.5 appears primarily attributable to improvements over the slices corresponding to non-verbal hearsay and statements made in court. In looking across slices moreover, it’s clear that some are comfortably within the realm of model capabilities (e.g., non-assertive conduct), while others (e.g., not introduced to prove the truth of the matter asserted) still pose a considerable challenge.
Another example is provided by the abercrombie task, in which an LLM must determine the relationship between a product and a potential trademark name, by classifying the product-name pair into one of five categories recognized by courts: generic, descriptive, suggestive, arbitrary, and fanciful. Loosely, these categories measure how distinctive a product name is for a product, with generic being the least distinctive, and fanciful being the most distinctive. Just as with hearsay, comparing LLM performance on each of these categories provides insight into the relative areas of improvement (Table 7). Here, GPT-4’s improved overall performance appears most attributable to performance on marks which are suggestive or arbitrary. However, GPT-4 still makes a number of errors for both categories. Interestingly, performance on descriptive marks is consistent between both models.
3.5 Interpretation
On the interpretation tasks, we find that on average GPT-4 outperforms GPT-3.5 (), ans GPT-3.5 outperforms Claude-1 (). Here, the larger API-models are highly performant on tasks which involve binary classification over short clauses. Averaged across the 38 CUAD tasks (contract clauses), for instance, GPT-4, GPT-3.5, and Claude-1 all have a balanced-accuracy 88%. And on proa (statutory clauses), both GPT-4 and GPT-3.5 have a balanced-accuracy 90%. Notably, performance degrades on tasks which contain longer text sequences or involve multi-class classification. On the Supply Chain Disclosure tasks for instance—in which LLMs must classify disclosures which are 1-2 pages in length—the average balanced-accuracy of the large commercial models ranges between 74-75%. And on the MAUD tasks—which require answering multiple choice questions about merger deals—the average balanced-accuracy of GPT-4 drops to 47.8% accuracy.
3.6 Rhetorical-analysis
On average across all rhetorical-understanding tasks, we find that GPT-4 outperforms both GPT-3.5 () and Claude-1 (). We note several results. First, on definition_extraction—which requires a LLM to extract the term defined by a sentence taken from a Supreme Court opinion—Incite-Base almost equals GPT-4 in performance (80.6% accuracy to 81.8%). Second, nearly all evaluated models struggle on two tasks requiring LLMs to label the legal “roles” played by either a question or excerpt from an opinion (function_of_decision_section and oral_argument_question_purpose). Notably, both tasks require the LLM to classify text into one of six or more categories
4 Prompt engineering strategies
Finally, we illustrate—through a series of micro-studies—how LegalBench can be used to explore different aspects of prompt-engineering for LLMs in legal settings. We focus on three questions:
Can LLMs rely on their latent knowledge of a rule for rule-conclusion tasks?
Does simplifying task descriptions to plain language affect performance?
Are LLMs sensitive to the choice of in-context demonstrations?
When prompting for general-domain tasks like sentiment or topic classification, prompt-engineers will often rely on the LLM’s latent knowledge of the task . In topic classification for instance, a prompt may use the instructions to label whether a news article is about “sports,” without offering a detailed description of what “sports” refers to or encompasses. Such a description is not necessary, because general-domain terms like “sports” appear frequently in LLM training corpora, and LLMs can learn from these occurrences what general-domain terms mean. Prompting for legal tasks, however, may require a different strategy. Because legal terms occur less frequently in general domain training corpora, legal prompting may require practitioners to provide additional background information. For example, a general domain LLM may not know what the requirements for diversity jurisdiction are, because diversity jurisdiction is not as commonly discussed in pretraining corpora.
We explore this question through a study of rule-conclusion tasks. For a selection of these tasks, we evaluate GPT-3.5 with two zero-shot prompts: a reference-based prompt and a description-based prompt. In the reference prompt, the task instructions merely state the rule to be applied, i.e., “Determine if the following fact patterns give rise to diversity jurisdiction.” In the description-based prompt, the instructions provide an explicit description of the rule, i.e., “Diversity jurisdiction exists when there is (1) complete diversity between plaintiffs and defendants, and (2) the amount-in-controversy (AiC) is greater than $75k.” By comparing performance between the reference and description prompt, we can measure whether providing a description of the rule in the prompt provides additional performance boost over the LLM’s latent knowledge of the rule.
Figure 1 provides a comparison for the different prompts. Interestingly, we find considerable variation across tasks. On tasks like abercrombie, ucc_v_common_law, diversity_2, and diversity_4, description prompts appear to offer significant increase in performance. On the other tasks, performance is approximately the same (or even worse). We identify two possible explanations for diverging results across tasks. First, on certain tasks, subsets of fact-patterns are too challenging for LLMs like GPT-3.5, and description-based prompts do not provide sufficient guidance for LLMs to reason through those fact patterns. Second, legal rules may be described to varying extents within pretraining corpora. Hence, tasks where we observe performance improvements from description-based prompting may correspond to rules which occur less frequently in pretraining data.
Next, we examine the extent to which domain specialization in the language of the prompts affects performance. Like experts in other specialized domains, lawyers have developed their own language (i.e., “legalese”), which forms the basis for most legal writing and communication. It is unclear whether—in interacting with large language models through prompting—lawyers should continue to rely on formalistic legal language, or instead use simpler plain language. While most large language models are “general domain” and thus less specialized to legalese, formalistic legal language is more precise, and may thus induce more accurate behavior from the model.
We explore this question by comparing “plain language” and “technical language” prompts. For a subset of LegalBench tasks, we have access to the formal language provided to law-trained annotators when creating task data. By comparing the performance of a prompt which uses this language—to one which uses a plain-language version—we can measure how the technicality of language affects results.
We conduct preliminary experiments on a sample of five LegalBench tasks (Figure 2).Prompts are made available in the LegalBench repository. On four of the five tasks, we find that the plain-language prompt significantly outperforms the technical language prompt, by up to 21 points (balanced-accuracy). Interestingly, on contract_nli_permissible_post-agreement_possession, we find the opposite phenomenon holds: the plain language prompt is substantially worse than the technical prompt.
Finally, we investigate the influence of the in-context demonstrations used in prompts. Prior work in general domain LLMs have observed that few-shot performance is highly sensitive to the choice of demonstrations . We evaluate whether LLMs are similarly sensitive for legal tasks, focusing on a subset of 8 binary classification tasks. For each task we merge the train and evaluation split into a single dataset, and randomly sample four in-context samples to include in the prompt (two from each class), five different times. We evaluated GPT-3.5 and Incite-Instruct-7B with each of the five generated prompts, and plot the the balanced-accuracy of each prompt in Figure 3.
Consistent with findings on general-domain tasks, we observe that LLMs on legal tasks are also highly sensitive to the choice of in-context samples. Notably, this appears to be the case for both GPT-3.5 and Incite-Instruct. Under a permutation test, we find significant differences () between the best and worst performing prompt for Incite-Instruct (on all tasks), and for GPT-3.5 (on all tasks except opp115_third_party_sharing_collection and overruling).We conduct the permutation test with 1000 resamples. For many tasks, the magnitude of difference is substantial. On overruling for instance, the best Incite-Instruct prompt improves upon the worst prompt by over 20 points (balanced-accuracy). Overall, these results suggest that future work is needed to understand how different demonstrations influence performance.
Conclusion
Our work here describes LegalBench: a collaboratively constructed benchmark of 162 tasks for measuring the legal reasoning capabilities of LLMs. In future work, we hope to expand this project, by continuing to solicit and collect interesting and useful tasks from the legal community.
References
Appendix A Acknowledgements
We are grateful to the following individuals and groups for feedback on this project: Alex Chao, Amit Haim, Arjun Desai, Armin Thomas, Avanika Narayan, Ben Spector, Brandon Yang, Eric Nguyen, Gautam Machiraju, Javed Qadrud-Din, Jian Zhang, Jonathan Zittrain, Karan Goel, Khaled Saab, Joshua Arp, Krista Opsahl-Ong, Laurel Orr, Lisa Ouellette, Lucia Zheng, Martin Gajek, Mayee Chen, Michael Zhang, Mike Wornow, Pablo Arredondo, Percy Liang, Rishi Bommasani, Roland Vogl, Sabri Eyuboglu, Sarah Hooper, Sergio Servantez, Simran Arora, Tengyu Ma, Tony Kim, Tri Dao and Vishnu Sarukkai. We presented and recieved feedback on earlier versions of this project at various forums, including: the Center for Research on Foundation Models, the Stanford Regulation and Governance Lab, the New York LLM x Law Hackathon (June 2023), the 2023 Stanford Data Science Conference, the 2023 Stanford CodeX Conference, and the Stanford Generative AI and Foundation Models Workshop. We are grateful to the organizers and attendees of these events for engaging with our work.
We gratefully acknowledge the support of NIH under No. U54EB020405 (Mobilize), NSF under Nos. CCF1763315 (Beyond Sparsity), CCF1563078 (Volume to Velocity), 2204926 (Computational Statutory Reasoning), and 1937301 (RTML); US DEVCOM ARL under No. W911NF-21-2-0251 (Interactive Human-AI Teaming); ONR under No. N000141712266 (Unifying Weak Supervision); ONR N00014-20-1-2480: Understanding and Applying Non-Euclidean Geometry in Machine Learning; N000142012275 (NEPTUNE); NXP, Xilinx, LETI-CEA, Intel, IBM, Microsoft, NEC, Toshiba, TSMC, ARM, Hitachi, BASF, Accenture, Ericsson, Qualcomm, Analog Devices, Google Cloud, Salesforce, Total, the Center for Research on Foundation Models, the HAI-GCP Cloud Credits for Research program, the Stanford Data Science Initiative (SDSI), and members of the Stanford DAWN project: Facebook, Google, and VMWare. We thank Casetext for assistance with evaluating GPT-4. PH is supported by an Open Philanthropy AI Fellowship. The U.S. Government is authorized to reproduce and distribute reprints for Governmental purposes notwithstanding any copyright notation thereon. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views, policies, or endorsements, either expressed or implied, of NIH, ONR, or the U.S. Government.
Appendix B Limitations and social impact
We note several limitations of our work. Legal applications—and what constitutes “legal reasoning”—is broad. Thus, LegalBench will necessarily be an incomplete effort, and important tasks/document types/reasoning types are not included. To enumerate a few examples:
LegalBench does not include tasks over long documents. Long documents are significant for legal practice, as writings like contracts, corporate filings, statutory codes, and judicial opinions can be hundreds of pages long .
The legal reasoning dimensions identified in LegalBench constitute a subset of the possible legal reasoning abilities for which we wish to evaluate LLMs. An example of a reasoning ability which is not currently evaluated in LegalBench would be analogical reasoning grounded in case law.
LegalBench tasks are skewed towards certain legal domains (e.g., contracts and civil procedure) and others are unrepresented.
LegalBench tasks skew towards US Federal law, and thus may not be representative for studies of other jurisdictions, or tasks involving international law.
LegalBench does not enable evaluation for multilingual, or non-English, legal tasks.
LegalBench does not evaluate more subjective legal tasks, or tasks which contain more ambiguity. These tasks are common to the legal field.
We hope to work on these limitations as part of future work. In particular, we would like to expand LegalBench to include other jurisdictions and a broader cross-section of legal domains.
Nothing in LegalBench should be construed as legal advice.
A potential negative social impact of our work would be if others either (1) construed our work as unequivocally endorsing automation in the legal industry, or (2) used performance on LegalBench as the sole justification for AI deployments. We therefore take efforts to mitigate these impacts, noting the following.
As we state in Section 1, the purpose of our work is not to determine whether large language models are capable of replacing legal professionals, the types of legal work that should/can be automated, or the broader implications of new technology on the practice of law. Rather, our focus is on developing technical artifacts which better enable stakeholders and affected parties to answer these questions themselves. Rigorous evaluation is essential to the safe and ethical usage of AI. LegalBench, as a benchmark, is intended to improve the ability for stakeholders to conduct evaluations. We additionally note that LegalBench, as a tool for research, is not a substitute for more in-depth and context-specific evaluation efforts. The deployment of any AI application in the law must be accompanied by evaluation on in-domain data, and assessments for ethical and legal compliance.
We finally note that potential negative impact will depend significantly on the task studied and the broader social context. The consequences of mistakes in using LLMs to annotate datasets, for instance, has significantly different consequences from the cost of mistakes when LLMs are used to answer legal aid questions.
Appendix C Datasheet
Following recent work, we provide a datasheet below. The datasheet below provides general answers to each of the questions, while Appendix F provides more in-depth details for each individual task. In addition, a number of LegalBench tasks have been adapted from previously released datasets, and the datasheets accompanying their publication provide further details.
For what purpose was the data set created? Was there a specific task in mind? If so, please specify the result type (e.g. unit) to be expected.
LegalBench was created to evaluate LLMs on legal tasks and better understand their legal reasoning capabilities. Recent advances in language modeling techniques have led to the emergence of “large” language models, and spurred interest within the legal community. This has led to two questions:
What technical adaptations are necessary to enable LLMs to perform legal tasks? Legal tasks often involve longer text sequences, jargon, and multi-step reasoning, making them more difficult than traditional NLP tasks.
For which legal tasks can current LLMs be trusted to perform safely and reliably?
LegalBench encompasses many different tasks. The specification for each task and the expected output can be found in the full task descriptions (Section F).
Who created the dataset (e.g., which team, research group) and on behalf of which entity (e.g. company, institution, organization)?
LegalBench consists of novel datasets (which were created by the authors of this paper), and transformed/adapted datasets (which were originally released as part of prior research). In Section F we discuss the origins of each dataset.
Who funded the creation of the dataset? If there is an associated grant, please provide the name of the grantor and the grant name and number.
LegalBench and its contributors have been generously funded by a range of entities that include the institutional affiliations provided for each author, governmental grants, and other sources.
C.2 Composition
What do the instances that comprise the dataset represent (e.g., documents, photos, people, countries)? Are there multiple types of instances (e.g., movies, users, and ratings; people and interactions between them; nodes and edges)? Please provide a description.
All LegalBench tasks consist of instances which are text. These include: sentences, paragraphs, and documents. Some instances are drawn from real world sources of text (e.g., actual contracts, corporate disclosures, judicial opinions, or complaints). Other instances were synthetically crafted. Section F provides details for each task.
How many instances are there in total (of each type, if appropriate)?
Section D provides details for each task.
Does the dataset contain all possible instances or is it a sample (not necessarily random) of instances from a larger set? If the dataset is a sample, then what is the larger set? Is the sample representative of the larger set (e.g., geographic coverage)? If so, please describe how this representativeness was validated/verified. If it is not representative of the larger set, please describe why not (e.g., to cover a more diverse range of instances, because instances were withheld or unavailable).
Nearly every LegalBench task corresponds to a sample of a population, or entirely synthetic data. Section F contains a more detailed description for each dataset. We highlight several broader explanations for the difficulty in acquiring complete or representative data which generalizes across tasks:
As prior work on legal benchmarks has noted , not all legal documents are published or reported. Hence, many are only accessible through special request, or only available in paper. The lack of easily available representative data is a noted challenge in many justice systems .
Acquiring legal annotations is exceedingly expensive. The CUAD project, for instance, estimated that a modestly sized dataset of 500 contracts (relative to the standards of NLP) had an estimated cost of $2 million US dollars . As a result, it is often possible to only annotate a small sample of data, even when a larger population is available.
What data does each instance consist of? “Raw” data (e.g., unprocessed text or images) or features? In either case, please provide a description.
Instances in LegalBench largely correspond to unprocessed text. Section F contains a more detailed description for each dataset.
Is there a label or target associated with each instance? If so, please provide a description.
Yes. Labels correspond to: classes, extracted entities, and open-ended generation. Section F contains a more detailed description of the labels/targets for each dataset.
Is any information missing from individual instances? If so, please provide a description, explaining why this information is missing (e.g., because it was unavailable). This does not include intentionally removed information, but might include, e.g., redacted text.
For reused/adapted datasets, we refer readers to the original data sheets which document redactions/missing data. Newly contributed tasks should not be missing information.
Are relationships between individual instances made explicit (e.g., users’ movie ratings, social network links)? If so, please describe how these relationships are made explicit.
Are there recommended data splits (e.g.,training, development/validation,testing)? If so, please provide a description of these splits, explaining the rationale behind them.
Yes. Tasks are split into train and test splits. Train splits consist of a small random sample of the original dataset (i.e., between 2-8 instances). We select small training splits in order to capture the true few-shot setting , in which a practitioner only has access to a handful of labeled instances. This design choice is also reflected in the structure of the RAFT benchmark .
Are there any errors, sources of noise, or redundancies in the dataset? If so, please provide a description.
A significant amount of legal data is the product of scanning and OCR. Hence, this data often contains artifacts of these processes, which appear as errant or missing characters.
Is the dataset self-contained, or does it link to or otherwise rely on external resources (e.g., websites, tweets, other datasets)? If it links to or relies on external resources, a) are there guarantees that they will exist, and remain constant, over time; b) are there official archival versions of the complete dataset (i.e., including the external resources as they existed at the time the dataset was created); c) are there any restrictions (e.g., licenses, fees) associated with any of the external resources that might apply to a future user? Please provide descriptions of all external resources and any restrictions associated with them, as well as links or other access points, as appropriate.
Does the dataset contain data that might be considered confidential (e.g., data that is protected by legal privilege or by doctor–patient confidentiality, data that includes the content of individuals’ non-public communications)? If so, please provide a description.
No. All LegalBench data is derived from public sources or was generated by authors. There is no confidential information in our dataset.
Does the dataset contain data that, if viewed directly, might be offensive, insulting, threatening, or might otherwise cause anxiety? If so, please describe why.
Does the dataset relate to people? If not, you may skip the remaining questions in this section.
LegalBench data relates to people to the extent that LegalBench contains tasks which contain language drawn from judicial cases involving individuals, or posts by individuals to legal forums (i.e., the Learned Hands Tasks).
Does the dataset identify any subpopulations (e.g., by age, gender)? If so, please describe how these subpopulations are identified and provide a description of their respective distributions within the dataset.
Is it possible to identify individuals (i.e., one or more natural persons), either directly or indirectly (i.e., in combination with other data) from the dataset? If so, please describe how.
As LegalBench is drawn entirely from public datasets—which themselves may contain additional information—it is possible to identify the original documents that LegalBench data was drawn from.
Does the dataset contain data that might be considered sensitive in any way (e.g., data that reveals racial or ethnic origins, sexual orientations, religious beliefs, political opinions or union memberships, or locations; financial or health data; biometric or genetic data; forms of government identification, such as social security numbers; criminal history)? If so, please provide a description.
The Learned Hands tasks correspond to posts on public forums. In these posts individuals discuss legal questions, and sometimes disclose information that would meet the above definition of “sensitive.”
We note that the data distributions from which some LegalBench tasks were drawn—like judicial cases or legal forums—have been used by prior work published in the NeurIPS Datasets and Benchmarks Track . These works offer additional information.
C.3 Collection process
How was the data associated with each instance acquired? Was the data directly observable (e.g., raw text, movie ratings), reported by subjects (e.g., survey responses), or indirectly inferred/derived from other data (e.g., part-of-speech tags, model-based guesses for age or language)? If data was reported by subjects or indirectly inferred/derived from other data, was the data validated/verified? If so, please describe how.
Data underlying LegalBench tasks were collected using different processes, and Section F contains a detailed discussion for each task.
What mechanisms or procedures were used to collect the data (e.g., hardware apparatus or sensor, manual human curation, software program, software API)? How were these mechanisms or procedures validated?
Please refer to Section F for background on each task.
If the dataset is a sample from a larger set, what was the sampling strategy (e.g., deterministic, probabilistic with specific sampling probabilities)?
Please see the discussion in the Composition section above.
Who was involved in the data collection process (e.g., students, crowdworkers, contractors) and how were they compensated (e.g., how much were crowdworkers paid)?
Section F contains a detailed discussion for each task.
Over what timeframe was the data collected? Does this timeframe match the creation timeframe of the data associated with the instances (e.g., recent crawl of old news articles)? If not, please describe the timeframe in which the data associated with the instances was created.
Section F contains a detailed discussion for each task.
Were any ethical review processes conducted (e.g., by an institutional review board)? If so, please provide a description of these review processes, including the outcomes, as well as a link or other access point to any supporting documentation.
Where applicable, Section F provides information relevant to each task.
Does the dataset relate to people? If not, you may skip the remaining questions in this section.
The dataset relates to people insofar as it draws text from documents which relate to people, or people created.
Did you collect the data from the individuals in question directly, or obtain it via third parties or other sources (e.g., websites)?
Section F contains a detailed discussion for each task.
Were the individuals in question notified about the data collection? If so, please describe (or show with screenshots or other information) how notice was provided, and provide a link or other access point to, or otherwise reproduce, the exact language of the notification itself.
No. Following other works which incorporate data from public judicial sources , we note that judicial filings are public, and the individuals implicated in those proceedings are aware of the public nature.
Did the individuals in question consent to the collection and use of their data? If so, please describe (or show with screenshots or other information) how consent was requested and provided, and provide a link or other access point to, or otherwise reproduce, the exact language to which the individuals consented.
Individuals whose names and circumstances appear in the original datasets did not separately consent to be a part of LegalBench. Again, we note that these documents are generally public, and already accessible to a wide range of parties, through many different judicial data services.
If consent was obtained, were the consenting individuals provided with a mechanism to revoke their consent in the future or for certain uses? If so, please provide a description, as well as a link or other access point to the mechanism (if appropriate).
Has an analysis of the potential impact of the dataset and its use on data subjects (e.g., a data protection impact analysis) been conducted? If so, please provide a description of this analysis, including the outcomes, as well as a link or other access point to any supporting documentation.
C.4 Preprocessing, cleaning, labeling
Was any preprocessing/cleaning/labeling of the data done (e.g., discretization or bucketing, tokenization, part-of-speech tagging, SIFT feature extraction, removal of instances, processing of missing values)? If so, please provide a description. If not, you may skip the remainder of the questions in this section.
C.5 Use
Has the dataset been used for any tasks already? If so, please provide a description.
We have used the constructed datasets to evaluate several LLMs.
Is there a repository that links to any or all papers or systems that use the dataset? If so, please provide a link or other access point.
LegalBench is available at https://github.com/HazyResearch/legalbench/.
What (other) tasks could the dataset be used for?
We envision this dataset could be used for the following:
Finetuning LLMs, either on task data directly, or self-instruct style generations derived from task data.
Is there anything about the composition of the dataset or the way it was collected and preprocessed/cleaned/labeled that might impact future uses? For example, is there anything that a future user might need to know to avoid uses that could result in unfair treatment of individuals or groups (e.g., stereotyping, quality of service issues) or other undesirable harms (e.g., financial harms, legal risks) If so, please provide a description. Is there anything a future user could do to mitigate these undesirable harms?
We emphasize that LegalBench—like all generalized benchmarks—can offer only a preliminary understanding of LLM performance. LegalBench tasks do not generalize to all legal reasoning tasks or all types of legal documents. We thus emphasize that practitioners seeking to deploy LLMs within their own applications should perform their own data collection and validation specific to their use case.
Are there tasks for which the dataset should not be used? If so, please provide a description.
These datasets should not be used to predict the legality of real world events, the outcome of lawsuits, or as legal advice.
C.6 Distribution
Will the dataset be distributed under a copyright or other intellectual property (IP) license, and/or under applicable terms of use (ToU)? If so, please describe this license and/or ToU, and provide a link or other access point to, or otherwise reproduce, any relevant licensing terms or ToU, as well as any fees associated with these restrictions.
Table 8 provides the license that applies to each individual LegalBench task.
Have any third parties imposed IP-based or other restrictions on the data associated with the instances? If so, please describe these restrictions, and provide a link or other access point to, or otherwise reproduce, any relevant licensing terms, as well as any fees associated with these restrictions.
Yes. Tasks which consist of adapted/transformed data are released under the same license as the original dataset. Table 8 provides these licenses, and Section F provides a reference to the original dataset for transformed tasks.
Do any export controls or other regulatory restrictions apply to the dataset or to individual instances? If so, please describe these restrictions, and provide a link or other access point to, or otherwise reproduce, any supporting documentation.
C.7 Maintenance
Who will be supporting/hosting/maintaining the dataset?
Neel Guha will be supporting this dataset.
How can the owner/curator/manager of the dataset be contacted (e.g., email address)?
Neel Guha can be reached at nguha@cs.stanford.edu. He will be available to answer any questions.
Is there an erratum? If so, please provide a link or other access point.
We have currently not found any, but will make them available on the website.
Willthe dataset be updated (e.g.,to correct labeling errors, add new instances, delete instances)? If so, please describe how often, by whom, and how updates will be communicated to users (e.g., mailing list, GitHub)?
Yes. There will be two types of updates to LegalBench:
First, we will update LegalBench to reflect new contributions from the legal community.
Second, we will update LegalBench to reflect identified errors in the data.
We will strive to make and publicize updates as soon as errors are identified and new tasks are contributed. Neel Guha will be in charge of managing these updates.
If the dataset relates to people, are there applicable limits on the retention of the data associated with the instances (e.g., were individuals in question told thattheir data would be retained for a fixed period of time and then deleted)? If so, please describe these limits and explain how they will be enforced.
Will older versions ofthe dataset continue to be supported/hosted/maintained? If so, please describe how. If not, please describe how its obsolescence will be communicated to users.
Yes. We will make older versions available on request by email.
If others want to extend/augment/build on/contribute to the dataset, is there a mechanism for them to do so? If so, please provide a description. Will these contributions be validated/verified? If so, please describe how. If not, why not? Is there a process for communicating/distributing these contributions to other users?If so, please provide a description.
Yes. We encourage members of the legal community to contribute new tasks. We are in the process of formalizing procedures for reviewing, validating, and incorporating submissions.
We additionally note that many of the LegalBench tasks are available under permissive licenses, and other researchers may thus modify them.
Appendix D Task overview
LegalBench tasks are subject to different licenses, due to the choices of dataset contributors, or the license under which the original data was released. Table 8 summarizes the licenses. The authors bear all responsibility in case of violation of rights, and confirm the dataset licenses.
D.2 Public availability status
Given that many commercially available LLMs are trained on the “entirety of the web”—and little is known as to how they are trained—there are concerns that many benchmarks have inadvertantly become part of the training data for these models. Therefore, this section identifies and organizes LegalBench tasks into three categories:
Previously published tasks, which were derived from datasets that were initially published as part of other works and available on the web for download.
Original but available tasks, which are original creations of the LegalBench project but previously made available online.
Original and unavailable tasks, which are original creations of the LegalBench project but have not been released online.
Table 9 summarizes the availability status of each of the tasks.
D.3 Reasoning type
Table 10 organizes tasks by the LegalBench reasoning type they can be used to assess. Table 11 similarily organizes tasks according to reasoning-types recognized in the NLP literature. For each reasoning type, we provide examples of general-domain NLP benchmarks which are similar. For more information on the types of reasoning required for each task, please see the individual task descriptions provided in Appendix F.
D.4 Task statistics
Table D.4 provides statistics for the LegalBench tasks. For each task, we list the number of samples and the average length (in words) of each input. LegalBench encompasses tasks ranging from short (a single sentence) to longer inputs (two pages of text) (Figure 4). The average LegalBench task contains between 500-600 instances. All tasks consist of at least 50 instances. A more detailed breakdown is available in Table 12.
Appendix E Evaluation
This section describes metrics and evaluation protocols.
To evaluate an LLM’s performance on a rule-application task, a law-trained expert manually validated each generation. We computed two metrics. The first metric—correctness—corresponds to the proportion of generations which do not misstate the fact pattern, the legal outcome, the rule, or contain a logical error. Logical errors include arithmetic mistakes, or assertions which are plainly wrong (e.g., that an apple is not a tangible object).
The LLM would incorrectly assert the legal outcome. For instance, an LLM would assert diversity jurisdiction existed, when it actually did not.
The LLM would incorrectly assert an intermediate conclusion. For instance, an LLM would assert that a rental agreement for a boat was a contract for a service, rather than a moveable and tangible good. In this case, the LLM would fail to realize that a boat is a moveable or tangible good.
The LLM would hallucinate a piece of information not explicitly stated in the fact pattern.
The LLM would misstate the content of a rule. For instance, the LLM would assert that subprovision of a statute barred one type of conduct, when in fact, that conduct was barred by a different subprovision.
The second metric—anlaysis—corresponds to the proportion of generations which contain the necessary inferences to reach the correct legal conclusion from the provided fact-pattern. Thus, it is insufficient for a LLM explanation to merely state that a piece of evidence is hearsay is insufficient: an explanation must reference the qualities of the evidence which make it hearsay. We introduced this measurement after discovering that for some tasks, LLMs often generate explanations which—though correct––-largely restate the rule being applied, without any reference to the underlying facts. Explanations which are incorrect are automatically deemed to be insufficient on analysis. We compute an overall analysis score for a LLM on a task by measuring the proportion of samples for which the explanation contains sufficient analysis.
To standardize evaluation and enable future work, we have released an “answer guide” for each task used for rule-application, which contains the inferences required for each sample, and describes common modes of errors. All evaluations in LegalBench for rule-application have been performed with respect to this answer-guide.
LegalBench contains the following generation-tasks, which are evaluated as follows:
rule_qa is a question-answer task in which a LLM must generate a response to an open-ended question. A law-trained individual evaluated the generations against an answer-key for the task, which is available for download from the website.
The Securities Complaint Extraction tasks require an LLM to extract the names of different parties from excerpts of securities class-action complaints. Because some samples require the extraction of multiple entities, we evaluate using F1 score.
definition_extraction requires the LLM to identify the term that is being defined in a sentence from a Supreme Court opinion. For a small number of sentences, any one of multiple terms may constitute the correct answer. We therefore evaluate performance on this task using accuracy, and by counting the fraction of sentences for which the LLM identified a permissible term. To account for edge-cases involving word tense, we compare stemmed versions of the answers to stemmed versions of the generation.
sara_numeric requires an LLM to generate an estimate of the amount of tax that is owed. We compute performance here using an accuracy metric, which treats a prediction as accurate if it is within 10% of the true amount.
citation_open requires an LLM to predict the name of the case that should be cited for a particular sentence of judicial text. We evaluate by checking if the LLM generation contains the correct case name.
successor_liability requires an LLM to identify the multiple possible successor liability exceptions to a fact pattern. We evaluate using F1.
We evaluate all classification tasks in LegalBench using exact-match on class-balanced-accuracy. We do this because a number of LegalBench tasks are class-imbalanced.
Appendix F Task descriptions
This section provides a detailed description for each family of tasks.
In LegalBench, the Abercrombie task is denoted as abercrombie.
A particular mark (e.g., a name for a product or service) is only eligible for trademark protection if it is considered to be distinctive. In assessing whether a mark is distinctive, lawyers and judges follow the framework set out in the case Abercrombie & Fitch Co. v. Hunting World, Inc.,Abercrombie & Fitch Co. v. Hunting World, 537 F.2d 4 (2nd Cir. 1976). which enumerates five categories of distinctiveness. These categories characterize the relationship between the dictionary definition of the term used in the mark, and the service or product it is being attached to. They are:
Generic: A name is generic with respect to a product or service if it connotes the basic nature of the product/service, rather than more individualized characteristics of the product. For example, the mark “Salt” for packaged sodium chloride would be generic under Abercrombie because “salt” is the common name for sodium chloride. It is also common to think of generic marks as merely referring to the class of goods for which a particular product is a species.
Descriptive: A name is descriptive if it identifies a characteristic or quality of an article or service, such as color, odor, function, dimensions, or ingredients. For example, the name “Sharp” for a television would be descriptive, because it describes a plausible characteristic of television (i.e., their sharp image quality).
Suggestive: A name is suggestive if it suggests, rather than describes, some particular characteristic of the goods or services to which it applies. An essential aspect of suggestive names is that it requires the consumer to exercise the imagination in order to draw a conclusion as to the nature of the goods and services. For example, the name “Greyhound” would be suggestive for a bus service, because greyhounds are considered to be fast, and “fast” is an adjective that could be used to describe a bus service.
Arbitrary: A name is arbitrary if it is a “real” word but seemingly “arbitrary” with respect to the product or service. For example, the mark “Apple” for a software company is arbitrary, because apples are unrelated to software.
Fanciful: A name is fanciful if it is entirely made up, and not found in the English dictionary. For example, “Lanmbe” is a fanciful mark, because it is a made-up word.
The Abercrombie spectrum is commonly taught as part of Intellectual Property courses in law school, and students are expected to understand how to determine the Abercrombie classification for a particular product/mark combination.
Performing the Abercrombie task requires reasoning about the literal meaning of a word and the degree of its connection to a particular product/service. It requires having some understanding of the types of words that could plausibly be used to describe a particular good/service, and the extent to which those words relate to a particular mark. It also requires reasoning as to whether a particular word is a real English word.
The Abercrombie task requires an LLM to determine–given a candidate mark and a description of a product/service–which of the five Abercrombie categories above apply.
We manually create a dataset to evaluate a model’s ability to classify a mark’s distinctiveness (into one of the above 5 categories) with respect to a product. In writing samples, we draw inspiration from similar exercises available in legal textbooks and practice study guides. Hence, the samples provided have a definite answer, and are not subject to ambiguity. There is an expectation that a law student learning intellectual property would be able to answer these questions to a high degree of accuracy.
We create approximately 20 samples for each category of distinctiveness, and randomly select a single sample from each category to constitute the train set. The remaining 19 samples (for each category) are assigned to the test set (for a total of 95 samples).
Given how easy this task is for lawyers with a basic training in intellectual property law, it is unlikely that LLMs will be called on to perform this task in the actual practice of law, or that the ability for LLMs to perform this task would alter the way in which lawyers approach IP practice. Instead, the Abercrombie task is significant as a measurement of reasoning ability. Because it is “simplistic” by the standards of human lawyers, it provides a useful objective measure of reasoning progress for LLMs.
F.2 Canada Tax Court Outcomes
In LegalBench, the Canada Tax Court Outcomes task is also denoted as canada_tax_court_outcomes.
The Tax Court of Canada hears appeals of government decisions related to taxation.Tax Court of Canada Act, RSC, 1985, c T-2, online: https://laws-lois.justice.gc.ca/eng/acts/t-2/index.html, s 12. The Court’s decisions, which are written in natural language, are published on the Court’s website, in both French and English.Tax Court of Canada, “Find a Decision”, online: https://decision.tcc-cci.gc.ca/tcc-cci/en/nav.do. Decisions typically include a section at the beginning summarizing the outcome of the appeal, followed by sections describing the factual background and various procedural steps, a section identifying the issues under consideration, sections with legal analysis, and a concluding section. While this is the standard format, judges are free to use other formats if they prefer. Decision length varies depending on the complexity of the litigation, with some decisions being only a few hundred words, and others being many thousands of words.
Appeals in Tax Court of Canada cases are brought by individuals or organizations who ask the Court to overturn a government taxation decision. Outcomes of appeals are generally binary: appeals are either granted, in which case the government taxation decision is overturned in whole or in part, or appeals are denied in which case the government taxation decision is upheld. Occasionally published decisions will not involve the outcome of an appeal, including where the decision is about a procedural step (e.g. the admissibility of particular evidence).
The canada_tax_court_outcomes task involves identifying whether an excerpt from a Tax Court of Canada decision includes the outcome of the appeal and, if so, what the outcome is. While the task is straightforward, one challenge is that the model must distinguish between outcomes of the appeal as a whole and outcomes of particular aspects of the appeal. Another challenge is that where the excerpt does not include the outcome, the model must avoid predicting the outcome – even if the model might plausibly correctly infer the likely outcome from the excerpt provided.
The Canada Tax Court Outcomes task requires an LLM to classify whether an excerpt from a given decision includes the outcome of the appeal, and if so whether the appeal was allowed or dismissed. Some excerpts do not include an outcome of the appeal, in which case the model should return ‘other’. Where the excerpt includes the outcome and the appeal is allowed in whole or in part, the model should return ‘allowed’. Where the excerpt includes an outcome, and the appeal is dismissed the model should return ‘dismissed’. The model should disregard outcomes that are not about the ultimate outcome of the appeal, such as costs awards (i.e. orders requiring a party to pay the other party’s legal costs).
We obtained the full text of English-language versions of decisions from 2001 to 2022 by scraping the Tax Court of Canada website.Ibid. As per the terms of service of the website, we are required to note that the text of the scraped decisions are not the official versions (official versions can be obtained from the website), and that the reproduction of these cases has not been produced in affiliation with or with the endorsement of the Government of Canada. We then cleaned and parsed the text to extract excerpts that are most likely to contain the outcome of the appeal. For example, many decisions contain a brief introductory section describing the outcome of the appeal using a specific header, and if the decision contained a section with such a header, we excerpted only that section. Where our parsing code could not identify such a section, we excerpted the first and last 2,500 characters, because outcomes are generally described at either the beginning or end of decisions. After initially attempting outcome classification on these excerpts using OpenAI’s ChatGPT, we selected a quasi-random sample of 250 excerpts (quasi random because we selected these manually, we over- sampled excerpts where the outcome is ‘other’, and we chose some excerpts that were challenging due to factors such as length or unusual format). We manually reviewed outcomes for these excerpts, correcting some that had been miscategorized.
Two random cases from each class are selected for the training split, while the remainder are used as the test set.
Legal scholars frequently gather data about outcomes in large numbers of legal decisions in order to examine patterns in judicial decision-making. For example, a legal scholar may be interested in comparing outcomes in similar processes across jurisdictions or they might examine whether a legislative change resulted in different outcomes over time. Lawyers and legal information technology companies may also be interested in gathering data on outcomes for the purposes of judicial analytics or to predict future outcomes.
Gathering such data is typically straightforward. It is, for example, a common task assigned to first year law student research assistants who can frequently achieve close to 100% accuracy on such tasks with only minimal training. However, because the data is often useful only when gathered on large numbers of decisions, this type of data gathering using human research assistants can be cost prohibitive. If LLMs can obtain high accuracy on these tasks, substantial savings could be achieved – which would increase the ability of researchers to pursue new projects.
F.3 Citation Prediction Tasks
In the LegalBench, the Citation Prediction tasks are also denoted as citation_prediction_*.
The importance of locating relevant legal materials, or “law search” has long been recognized as an essential aspect of legal practice. This process involves uncovering case law, statutes, and other materials pertinent to legal questions or arguments. As a fundamental aspect of legal reasoning, law search plays a crucial role in bridging the gap between the initial translation of behaviors into legal questions and the subsequent interpretation and application of the relevant law.
Legal professionals are often valued for their ability to find and apply the appropriate law to their clients’ situations. Given the intricate nature of the contemporary legal domain, the process of law search has evolved into a complex and nuanced task that demands a comprehensive understanding of the law.
A core component of law search is legal relevance. From a sociological perspective, the relevance of legal documents to a specific legal question is a social fact. This fact is determined by the judgments made by members of the legal community, who must determine which legal materials are applicable to a given question. Relevance relates legal questions to sources of legal authority.
In functional legal communities, law search leads to some degree of convergence over legal materials. Convergence occurs when competent members of a legal community, faced with the same legal question, identify the same sources of relevant legal authority. This process is essential to ensuring that the legal system operates consistently, predictably, and coherently.
As a critical process that connects the translation of behaviors into legal questions and the subsequent interpretation and application of the relevant law, law search is indispensable to legal reasoning.
The Citation Prediction task requires reasoning concerning the relationship between the text of judicial opinions and legal propositions. Successful prediction would entail encoding a notion of legal relevance and would allow a system to determine whether a legal proposition was or was not supported by the extant body of law.
The citation task is based on a version of the evaluation approaches used in . There are two Citation Prediction tasks. The first (citation_prediction_classification) requires an LLM to predict whether a given sentence (i.e. legal proposition) is or is not supported by a given case. The second (citation_prediction_open) requires an LLM to predict a case (by name) that supports a provided sentence.
We collected a sample of circuit court opinions published after January 1, 2023. To the best of our knowledge, most existing LLMs haven’t been trained on any data generated in 2023. For each opinion, we manually collected sentences which were supported by a citation to a judicial opinion, where (1) the sentence contained some quotation from the original case, and (2) the sentence was supported by a single cite. We chose sentences which included quotation fragments and were only supported by a single cite to avoid sentences which could be supported by a broad set of cases. When a sentence is supported by a much larger universe of cases, verifying that an LLM answer is incorrect is difficult. We also recorded the circuit for each opinion that we pulled language from. As a result, we can include the circuit information in the prompt, since circuits prefer citing their previous decisions. We collected 55 sentences using this process. For the citation generation task (citation_prediction_open), we ask the LLM to predict the citation given the sentence.
The citation_prediction_classification task is then constructed as follows. We use each sentence-citation pair to create two task samples. The first sample corresponds to the sentence and the correct citation (positive label). The second sample corresponds to the sentence and a randomly selected citation from the remainder of the data (negative label). This generates a dataset of 110 sentence-citation pairs, two of which are assigned to the training split.
Law search is a core function of legal thinking. In addition, the difficulty of identifying relevant law is a core barrier in the public’s ability usefully access the law. The ability of an LLM to accurately engage in citation prediction would have important practical value in providing access to law, and would also allow the LLM to more reliably support legal statements with relevant authority.
F.4 Clause Classification Tasks
LegalBench includes a number of tasks in which the LLM must determine the “type” or “category” of a provision/clause in a legal document. Specifically:
The Contract QA task (Section F.4.4), in which the LLM is provided with the name of a common type of contractual clause and a clause, and must determine if the clause is an example of the example type.
38 tasks derived from the CUAD dataset (Section F.4.1), where each task is a binary-classification task requiring the LLM to identify whether a clause (from an EDGAR contract) belongs to a certain category (e.g., audit rights clauses) .
The J.Crew blocker task (Section F.4.2), in which the LLM must classify whether a clause (from a loan agreement) is a J-Crew blocker provision.
The Unfair Terms of Service task (Section F.4.3), in which the LLM must classify a clause (from a terms of service agreement) to one of eight types, where seven of the types denote clauses that would potentially be considered “unfair” under European law .
Lawyers spend significant time and energy reviewing legal documents (e.g., contracts, leases, etc.). Manual review serves an important purpose, allowing parties to identify potentially problematic terms . Parties will sometimes review agreements that have already been signed, in response to changing world events. For instance, the COVID-19 pandemic led many firms to inspect agreements for the existence of force majeure clauses, which ordinarily specify how contractual expectations should be handled in the event of major world crises . Because legal documents are long and require legal training to understand, the process of reviewing is often extremely expensive . This, in turn, presents significant access-to-justice concerns. Because most individuals do not have the financial capacity to consult lawyers prior to entering legal agreements, they are oblivious to when those agreements contain predatory, oppressive, or unconscionable terms. A rich legal scholarship has noted, for instance, the frequency at which legal agreements contain terms that would be invalidated by a court .
The clause classification tasks in LegalBench are thus amongst the most practically useful tasks in LegalBench, as they capture an actual current-day use case for LLMs. As the complexity of clause classification depends both on the clause category and document type, LegalBench tasks span a range of clause categories and source documents.
F.4.1 CUAD Tasks
We adapt the CUAD dataset for LegalBench . The original dataset consists of 500 contracts, each annotated with up to 41 different clause types. These contracts varied significantly in length, ranging from a few pages to over one-hundred pages. In the original word, studied the ability for BERT-base language models to identify the text spans corresponding to different types of clauses. The principal difficulties were (1) the length of the contract, and (2) the lack of significant training data.
We adapt the CUAD dataset as follows. We select 38 of the 41 clause categories. For each selected category, we construct a dataset consisting of (1) clauses in the CUAD contracts which are assigned to that category, and (2) an equal number of clauses randomly sampled from other categories. This produces a balanced binary classification task for clause category, where the purpose is to identify which clauses belong to the respective category. A table with the selected categories, and their descriptions is found below.
Table F.4.1 lists each task, a “description” of the category corresponding to the task, and an example of a clause which meets the category criteria. In accordance with , the description is presented as the question posed to the annotators during data labeling. If a clause yields an affirmative answer with regards to the question, then the label is “Yes”. Otherwise the label is “No”.
In LegalBench, the CUAD tasks are denoted as cuad_*.
F.4.2 J.Crew Blocker
In LegalBench, the J.Crew Blocker task is denoted as jcrew-blocker.
Loan agreements often contain restrictive covenants that place limits on a borrower’s activities to protect the lender’s interests. One such restrictive covenant that has become popular in recent years is the “J.Crew blocker” provision. This provision was created in response to actions taken by the retailer J.Crew in 2016. J.Crew transferred valuable intellectual property assets out of the collateral pool for its existing loans by moving them into a new unrestricted subsidiary. This subsidiary was then able to use the IP assets as collateral to obtain new financing.
The J.Crew blocker provision aims to prevent this type of activity by prohibiting borrowers from transferring IP assets out of the reach of existing lenders. There are two key components to a J.Crew blocker:
A prohibition on transferring IP assets to unrestricted subsidiaries. This prevents the borrower from moving assets outside the scope of lender restrictions.
A requirement to obtain lender consent for any IP transfers to subsidiaries. This gives lenders oversight and control over how IP assets are distributed within the corporate group.
The presence of a robust J.Crew blocker in a loan agreement is designed to keep material assets within the collateral pool, and thereby protect lenders from borrowers’ attempts to secure additional debt through unexpected transfers of IP. For this reason, J.Crew blocker provisions have been widely adopted in leveraged loan agreements.
The J.Crew blocker task requires determining whether a given provision in a loan agreement qualifies as a J.Crew blocker. To make this determination, the provision must be analyzed to assess whether it contains:
A prohibition on transferring IP assets to unrestricted subsidiaries
A requirement to obtain lender consent for IP transfers to any subsidiary.
If the provision includes one or both of these components, it can be classified as a J.Crew blocker. If not, the provision does not meet the criteria.
The dataset for this task was constructed by legal experts extracting real examples of provisions from public loan agreements. Each example was labeled as either meeting the criteria for a J.Crew blocker or not. The dataset contains 60 total examples, organized into two columns: "Text" (containing the clause in question) and "Label" (indicating whether the clause is a J.Crew Blocker provision). Each clause was analyzed and classified as a J.Crew Blocker provision ("Yes") or not ("No"). The construction process involved manually reviewing and annotating these samples, ensuring that each clause was accurately categorized. This process, carried out by legal experts, provides definitive answers to each sample, eliminating ambiguity.
The ability to identify J.Crew blocker provisions is important for both lenders and borrowers in leveraged finance. For lenders, it helps ensure key protections are included in loan agreements. For borrowers, it provides insight into restrictions being placed on their activities. Given the widespread adoption of J.Crew blockers, this is a task that requires proficiency to actively participate in the leveraged loan market. The task serves as an important measure of an LLM’s ability to understand and apply legal concepts, particularly those related to secured lending and intellectual property law. It also tests the LLM’s capacity to analyze and interpret legal provisions. Given the increasing complexity and sophistication of financial transactions, the ability to accurately identify and understand such provisions is a valuable skill for any LLM. This task, therefore, provides a useful measure of progress for LLMs in their understanding and interpretation of complex legal clauses.
F.4.3 Unfair Terms of Service
In LegalBench, the Unfair Terms of Service task is denoted as unfair_tos.
An array of recent work has found that consumers rarely read terms of service agreements . As a result, consumers regularly sign agreements or contracts containing provisions that (1) they lack awareness of, and/or (2) would consider as “unfair” or “predatory.” Reasons for this phenomenon include the sheer amount of time it would take to read every terms of service agreement, the obtuse language of these agreements, and the lack of actual recourse on an individual basis.
With reference to European consumer law, identify eight categories of clauses in terms-of-service agreements which could be considered “potentially unfair”:
Arbirtration: clauses which mandated that all disputes between the parties would be resolved through arbitration.
Unilateral change: clauses which allow the provider to modify the terms of service and/or the service itself.
Content removal: clauses which give the provider a right to modify/delete a user’s content
Jurisdiction: clauses which specify a jurisdiction in which claims must be brought, regardless of where the user lives.
Choice of law: clauses which specify the country’s law which governs disputes arising under the contract, regardless of where the user lives.
Limitation of liability: clauses which limit the liability of the service provider.
Unilateral termination: clauses which empower the service provider to terminate/suspend the service at their discretion.
Contract by using: clauses which stipulate that a consumer is bound by the terms of service simply by using the service.
A more detailed description of these categories can be found in .
The Unfair Terms of Service task requires an LLM to determine—given a clause from a terms of service agreeement—whether it belongs to one of the above eight categories, and if so, which one.
We use the version of data available in , which takes a subset from . Unlike —which frames the task as distinguishing “fair” from “unfair” clauses—we cast the task as 8-way multiclassification task across the original categories identified in .
Unlike the CUAD and J.Crew Blocker task, the Unfair TOS task evaluates a LLM’s ability to perform multiclass clause classification across a highly imbalanced dataset.
F.4.4 Contract QA
In LegalBench, the Contract QA task is denoted as contract_qa.
Each of the above tasks evaluates the capacity for LLMs to learn to recognize a single type of clause, given a description of that clause and/or examples of it. The Contract QA task generalizes this across multiple clause types, evaluating an LLM’s ability to recognize legal provisions that are not described in the prompt.
Each sample in the dataset consists of (1) a contract clause, and (2) a question asking if the clause is an example of a provision type (e.g., “Is this a severability clause?”). Across the dataset, the questions correspond to 22 different legal provisions. Questions and provisions are paired such that for each provision type, the LLM is presented with two clauses that are an example of the type, and two clauses which are not.
The data was manually extracted from a set of sample agreements contributed by a LegalTech vendor and from public sources. It represents a variety of contracts, such as:
Vendor or Partner Data Protection Agreements (DPA)
F.5 Consumer Contracts QA
In LegalBench, the Consumer Contracts QA task is denoted as consumer_contracts_qa.
Consumer contracts govern many economic and social activities, ranging from retail purchases and online search to social media and entertainment. These contracts can affect consumers’ access to services, control terms of payment, and determine the remedies available when consumers’ rights are violated. Despite the importance of these legal agreements, consumers typically lack the time, expertise, and incentive to properly examine how consumer contracts impact their rights and interests. This issue is known as the “no-reading” problem . LLMs may offer a solution. By reading consumer contracts and explaining their legal ramifications, LLMs could enable consumers to better understand and exercise their legal rights in many everyday contexts.
The Consumer Contracts QA task, first introduced in , aims to examine the degree to which an LLM can understand certain consumer contracts. Specifically, the task is comprised of 200 yes/no legal questions relating to the terms of service of popular websites. Examples of questions are provided in the table below.
In addition to the original 200 questions, the task includes an alternatively worded version of all 200 questions. While each question’s content is substantially the same across both versions of the question, the alternatively worded questions are, by design, less readable, that is, more difficult for a human to read. Comparing performance across the original questions and the alternatively worded questions can help assess an LLM’s brittleness in performing the task at hand. An example is provided in the table below:
The task was introduced in . To construct the dataset, an attorney drafted 200 yes/no questions relating to the terms of service of the 20 most-visited U.S. websites (10 questions per document), as well as an alternatively worded version of all 200 questions. The questions relate to a wide range of legal issues arising in the terms of service, including eligibility to access services, payment for services, limitations of liability, intellectual property rights, and dispute resolution procedures. Answers to all questions can be obtained from the applicable terms of service.
Given the ubiquity of consumer contracts, LLMs capable of reading these documents and communicating their contents to consumers might offer significant benefits. These benefits, however, are contingent on a model’s accuracy and reliability. LLMs that misinterpret the provisions of consumer contracts may hinder consumers’ ability to understand and exercise their contractual rights. The Consumer Contracts QA task is a preliminary attempt at evaluating the ability of LLMs to read certain consumer contracts.
F.6 Contract NLI Tasks
In LegalBench, the Contract NLI tasks are denoted as contract_nli_*.
The Contract NLI tasks require a LLM—given an excerpt of a contract and an assertion about the legal effect of that excerpt—to determine whether the assertion is supported or unsupported by the excerpt.
These tasks are constructed by transforming data released by . The original dataset consists of 607 contracts and 17 assertions (e.g., “Receiving Party shall not disclose the fact that Agreement was agreed or negotiated ”). Each contract is labeled for each assertion as supporting, negating, or not mentioning the assertion. Please refer to the original paper for details on annotation.
We restructure this dataset for a short-context LLM setting. Specifically, we treat each assertion as a separate task, where the objective is to determine whether a contract excerpt is supportive (or not) of the assertion. For each instance where a contract is supportive of an assertion, has annotated the excerpt of the contract that is supportive. When creating a task, we use the supportive excerpts for the assertion from the test set as positive instances. To generate negative instances, we combine excerpts where the assertion is contradicted with a random sample of excerpts associated with other assertions. We treat both groups of excerpts as instances which are “unsupportive” of the assertion. We transform the assertion into a Yes/No question, where the LLM is asked to determine if a clause satisfies the assertion.
Table F.6 lists each task, the assertion associated with the task, and an example of an excerpt which supports the assertion.
The Contract NLI tasks evaluate an LLM’s capacity to reason over the rights and obligations created by a contract. The ability to perform this skill is essential to many types of legal work.
F.7 Corporate Lobbying
In LegalBench, the Corporate Lobbying task is denoted as corporate_lobbying.
A significant amount of effort is devoted to identifying developing sources of law which implicate client or issue interests. Examples of such sources include: legislative bills, proposed regulations, or in-progress litigation. Identifying these sources serves multiple purposes. From a scholarly standpoint, researchers often aggregate sources into issue-focused databases, enabling them to identify emerging trends or patterns across different sources . From an advocacy standpoint, identifying sources allows affected groups to better understand how their rights or obligations may be affected, and how to focus efforts on interacting with courts, legislatures, and other governmental bodies .
The Corporate Lobbying task requires an LLM to determine whether a proposed Congressional bill may be relevant to a company based on a company’s self-description in its SEC 10K filing. The following information about a bill and a company are available:
We expect higher accuracy of LLM predictions if we were to provide the model with more data about a bill, and especially if we provide it with more data about a company. Proprietary applications of this approach could leverage significant internal company data. More expensive deployments could leverage the full text of the bill
This data was manually labeled. This work was an extension of the research described in .
Determining whether a particular bill is relevant for a company requires (1) identifying the legal consequences of the bill, and (2) whether those consequences are relevant to a company’s business model, structure, or activities. As discuss above, this type of prognostication is a common legal practice. For instance, law firms regularly publish “client alerts” which seek to keep clients updated on new legal developments .
F.8 Definition Tasks
In LegalBench, the Definition Tasks are denoted as definition_classification and definition_extraction.
Judicial opinions regularly involve definition, assigning a particular meaning to words or phrases (Let us define words and phrases as “terms”). Definition of terms can occur when judges introduce or discuss legal concepts (e.g. parol evidence), and it frequently occurs when judges interpret terms in legal texts. This can include language from past judicial opinions and language appearing in legal texts like contracts, statutes, and the Constitution. Historically, interpreters have often evaluated the definition(s) of individual words. For example, in interpreting the meaning of “keep and bear arms” in the Second Amendment, courts consider the definition(s) of individual words (like “bear”). This approach—focusing on terms’ definitions—has only increased in recent decades with the rise of textualist approaches to constitutional and statutory interpretation.
Judicial opinions define a wide range of terms, including ordinary terms, legal terms, and scientific terms. They also appeal to a wide range of defining sources, including ordinary dictionaries, legal dictionaries, and legal texts. For an example of the last, consider statutory definitions: 1 U.S.C. 1 offers generally applicable definitions of many frequent statutory terms.
It is useful for lawyers to identify when definition occurs (definition classification), as well as which terms have been defined (definition extraction). These tasks might seem simple at first. There are some intuitively plausible indicators of definition classification and extraction. For example, defined terms often (but not always) appear in quotation marks or near a citation to a dictionary.
However, these tasks are not entirely straightforward. Indicators like quotation will not lead to perfect definition classification and extraction. Consider for example, this sentence from the dataset related to the definition of “confidential”: The term “confidential” meant then, as it does now, “private” or “secret.” Webster’s Seventh New Collegiate Dictionary 174 (1963).Food Mktg. Inst. v. Argus Leader Media, 139 S. Ct. 2356, 2363 (2019). As another example from the dataset, consider this definition of “brought”: But a natural reading of § 27’s text does not extend so far. “Brought” in this context means “commenced,” Black’s Law Dictionary 254 (3d ed. 1933).Merrill Lynch, Pierce, Fenner & Smith Inc. v. Manning, 136 S. Ct. 1562, 1568 (2016). Other examples exclusively quote the definition, rather than defined terms: Stare decisis (“to stand by things decided”) is the legal term for fidelity to precedent. Black’s Law Dictionary 1696 (11th ed. 2019).June Medical Services L.L.C. v. Russo, 140 S. Ct. 2103, 2134 (2020). In all of these examples, the presence of a dictionary would not indicate which term is extracted. In other examples, there is no dictionary cited; there is not a perfect correlation between dictionary citation and classification of a sentence as a defining one.E.g. “And “remuneration” means “a quid pro quo,” “recompense” or “reward” for such services. Id., at 1528.” BNSF Ry. Co. v. Loos, 139 S. Ct. 893, 905 (2019).
The Definition Classification task requires an LLM to determine–given an excerpt from a Supreme Court opinion–whether the excerpt is defining any term (Yes/No). The Definition Extraction task requires an LLM to determine–given an excerpt from a Supreme Court opinion–which term the excerpt is defining (Open-ended response).
An original hand-coded dataset was constructed to study how the Supreme Court relies on dictionaries over time. Any case citing a dictionary was included in the dataset, and human coders identified relevant excerpts that defined terms and which terms were defined.
That dataset has been repurposed for the task here. For the definition extraction task, the original dataset includes the relevant information (excerpts, with the defined term coded separately).
For the definition classification task, the original dataset includes examples of language defining terms. To create a set of non-defining language, Neel Guha randomly selected similarly long excerpts of text from the same Supreme Court opinions. Kevin Tobia analyzed those randomly selected excerpts, identifying any that include definitions (for removal). The resulting dataset has 691 sentences which define sentences, and 646 sentences which do not.
This is not a particularly difficult task for human lawyers, and it is unlikely that LLMs would replace lawyers as experts in this process. However, it is possible that LLMs successful in these tasks could provide beneficial legal research roles (e.g. quickly identifying all prior definitions of a specific term in a particular jurisdiction).
Moreover, the definition extraction task serves as a useful test of LLMs abilities, given the task’s open-ended nature. The task is not limited to a small set of possible answers (e.g. Yes, No). Rather, it requires identifying which term of all terms in an excerpt is defined. Most of these choices will admit of over ten possible answers (i.e. excerpts of over ten words). Moreover, there is great variety in the language used across the examples. There are hundreds of possible answers, across all items.
F.9 Diversity Jurisdiction
In LegalBench, the Diversity Jurisdiction tasks are denoted as diversity_*.
Diversity jurisdiction is one of two ways in which a federal court may have jurisdiction over a lawsuit pertaining to state law. Diversity jurisdiction exists when there is (1) complete diversity between plaintiffs and defendants, and (2) the amount-in-controversy (AiC) is greater than $75,000.
“Complete diversity” requires that there is no pair of plaintiff and defendant that are citizens of the same state. However, it is acceptable for multiple plaintiffs to be from the same state, or for multiple defendants to be from the same state.
The AiC requirement allows for certain forms of aggregation. Specifically, if plaintiff A asserts two independent claims against defendant B, the value of the claims may be added together when considering if the AiC requirement is met. However, a plaintiff may not aggregate the value of claims against two separate defendants, and two plaintiffs may not aggregate claims against the same defendant.
We define six different tasks, each of which tests the diversity jurisdiction rule under a different pattern of facts. The diversity jurisdiction tasks are:
diversity_1: The fact patterns consists of one plaintiff, one defendant, and one claim per plaintiff-defendant pair.
diversity_2: The fact patterns consists of one plaintiff, two defendants, and one claim per plaintiff-defendant pair.
diversity_3: The fact patterns consists of one plaintiff, one defendant, and two claims per plaintiff-defendant pair.
diversity_4: The fact patterns consists of two plaintiffs, one defendant, and one claim per plaintiff-defendant pair.
diversity_5: The fact patterns consists of two plaintiffs, one defendant, and two claims per plaintiff-defendant pair.
diversity_6: The fact patterns consists of two plaintiffs, two defendants, and two claims per plaintiff-defendant pair.
We programmatically construct a dataset to test the diversity jurisdiction. We generate randomness over the names of the parties, the claims, and the amounts.
It is extremely unlikely LLMs would ever be used to evaluate diversity jurisdiction in practical settings. However, because the task is considered extremely simplistic—and one that first year law students are expected to perform perfectly—it offers a useful evaluation benchmark for LLMs. The structure of the task is potentially non-trivial for LLMs, as it requires identifying the relationships between parties (i.e., who are plaintiffs and defendants), understanding which claims may be aggregated, and computing whether the aggregated amounts meet the AiC requirement.
F.10 Function of Decision Section
In LegalBench, the Function of Decision Section task is denoted as function_of_decision_section.
In common-law legal systems, written judicial decisions serve two functions. First, they resolve the dispute that litigants brought before the court and explain the reason for the court’s decision. Second, they become new law, binding on future parties and future courts should another case arise that presents sufficiently similar facts.
Because judicial decisions not only describe the law, but are themselves the law, lawyers in common-law legal systems must be able to read and digest case law to extract key legal principles and apply those principles to their own cases. This skill takes time and practice to develop.
Importantly, not every word in a judicial decision is binding, only the facts and reasoning that were required for the court to reach its decision. Thus, lawyers must distinguish important from trivial facts across numerous past decisions before they can conclude what the law on a particular issue is. One of the most foundational case-reading skills is the ability to review a legal decision and identify the function that each section of the decision serves. In the American legal education system, this skill is taught beginning in the first year of law school, often by encouraging students to identify the function of each section of a decision. A typical classification scheme is as follows:
Facts: A section of the decision that recounts the historical events and interactions between the parties that gave rise to the dispute.
Procedural History: A section of the decision that describes the parties’ prior legal filings and prior court decisions that led up to the issue to be resolved by the decision.
Issue: A section of the decision that describes a legal or factual issue to be considered by the court.
Rule: A section of the decision that states a legal rule relevant to resolution of the case.
Analysis: A section of the decision that evaluates an issue before the court by applying governing legal principles to the facts of the case
Conclusion: A section of the decision that articulates the court’s conclusion regarding a question presented to it.
Decree: A section of the decision that announces and effectuates the court’s resolution of the parties’ dispute, for example, granting or denying a party’s motion or affirming, vacating, reversing, or remanding a lower court’s decision.
Identifying the function of sections within judicial decisions is a fundamental skill for lawyers in common-law legal systems. Without it, precedent-based legal reasoning would be impossible.
The Function of Decision Sections task requires an LLM to determine–given a one-paragraph excerpt of a legal decision–which of the seven functions above that paragraph serves in the context of the entire decision.
We created a dataset of paragraphs from legal decisions, classified into one of the seven functions above. Paragraphs were taken from decisions in West Publishing’s fourth Federal Reporter series, which publishes the decisions of the United States Courts of Appeals. To avoid selection bias and achieve a degree of randomness, paragraphs were selected from sequential decisions, in the order they appeared, spanning all areas of civil and criminal law that fall within the jurisdiction of the federal courts.
Beginning law students may initially have trouble identifying the function of a particular section within a judicial opinion, but it quickly becomes a simple task. LLMs would not be called on to perform this task in the actual practice of law, but because it is a foundational legal reasoning skill, it provides a useful measure of reasoning progress for LLMs.
F.11 Hearsay
In LegalBench, the hearsay task is denoted as hearsay.
The Federal Rules of Evidence dictate that “hearsay” evidence is inadmissible at trial. Hearsay is defined as an “out-of-court statement introduced to prove the truth of the matter asserted." In determining whether a piece of evidence meets the definition of hearsay, lawyers ask three questions:
Was there a statement? The definition of statement is broad, and includes oral assertions, written assertions, and non-verbal conduct intended to communicate (i.e. assert) a message. Thus, for the purposes of the hearsay rule, letters, verbal statements, and pointing all count as statements.
Was it made outside of court? Statements not made during the trial or hearing in question count as being out-of-court.
Is it being introduced to prove the truth of the matter asserted? A statement is introduced to prove the truth of the matter asserted if its truthfulness is essential to the purpose of its introduction. Suppose that at trial, the parties were litigating whether Alex was a soccer fan. Evidence that Alex told his brother “I like soccer," would be objectionable on hearsay grounds, as (1) the statement itself asserts that Alex likes soccer, and (2) the purpose of introducing this statement is to prove/disprove that Alex likes soccer. In short, the truthfulness of the statement’s assertion is central to the issue being litigated. However, consider if one of the parties wished to introduce evidence that Alex told his brother, “Real Madrid is the greatest soccer team in the world." This statement would not be hearsay. It’s assertion—that Real Madrid is the greatest soccer team in the world—is unrelated to the issue being litigated. Here, one party is introducing the statement not to prove what the statement says, but to instead show that a particular party (i.e. Alex) was the speaker of the statement.
Given a legal issue and a piece of prospective evidence, the LLM must determine whether the evidence constitutes hearsay under the above test.
We note that in practice, many pieces of evidence which are hearsay are nonetheless still admissible under one of the many hearsay exception rules. We ignore these exceptions for our purposes, and leave the construction of benchmarks corresponding to these exceptions for future work.
We create the hearsay dataset by hand, drawing inspiration from similar exercises available in legal casebooks and online resources. The dataset consists of 5 slices, where each slice tests a different aspect of the hearsay rule. We randomly select 1 sample from each slice to be in the train set. The remainder of the slice constitutes the test set (for a total of 95 samples). The slices (with test set counts) are:
Statement made in court (): Fact patterns where the statement in question is made during the course of testimony at trial. Thus, the statement is not hearsay.
Non-assertive conduct (): Fact patterns where the evidence does not correspond to a statement. Hence, the hearsay rule is inapplicable.
Standard hearsay (): Fact patterns where there is an oral statement, it is said out of court, and it is introduced to prove the truth of the matter asserted. Thus, these fact patterns correspond to hearsay.
Non-verbal hearsay (): Fact patterns where the statement is hearsay, but made in writing or through assertive conduct (e.g. pointing).
Not introduced to prove truth (): Fact patterns where an out-of-court statement is introduced to prove something other than what it asserts.
The hearsay rule is commonly taught in law school as part of Evidence. Law students are expected to understand the rule, and how to apply it. The hearsay task is interesting for LLM evaluation because it emphasizes multi-step reasoning—the test for hearsay encompasses several different steps, where each step differs in difficulty. These include:
Event detection: The LLM must determine whether the fact pattern mentions a statement being made.
Spatial reasoning: The LLM must determine whether the statement was made inside a court room.
Argument extraction: The LLM must determine what the statement is asserting.
Argument relevance: The LLM must finely determine whether the assertion is relevant to the issue being litigated.
F.12 Insurance Policy Interpretation
In LegalBench, the Insurance Policy Interpretation task is denoted as insurance_policy_interpretation.
Insurance disputes often arise when parties disagree on whether a claim is covered under a certain insurance policy. To study such disagreements in interpretation, researchers at Stanford recruited crowdsource workers to review a pair of an insurance policy and a claim and respond whether they believe the claim is covered. A policy-claim pair whose applicability the workers disagree with each other on suggests ambiguity in the policy text.
The Insurance Interpretation task requires an LLM to review a pair of an insurance policy and a claim and determine whether the policy clearly covers the claim, clearly does not cover it, or if it is unclear whether it covers it or not.
The clause-claim pairs are manually constructed before being reviewed by crowdsource workers . To convert the numbers of Covered/Not_Covered/Can’t_Decide responses to discrete labels, we first calculate the 95% multinomial confidence interval of the proportion of each response. We then choose the label for which the confidence interval lower bound is greater than or equal to .5. If no label has a lower bound .5, we classify the policy-claim pair as “It’s ambiguous.” This conversion process ensures that individual crowdsource workers do not arbitrarily sway the labels. Examples for each label can be found in Table 33.
The ability to determine whether an insurance claim is covered under a given policy can significantly reduce claim processing time. It can also shine light on potential ambiguity in existing policies. Additionally, this task represents one of the rare benchmarks where an LLM is required to predict laypeople’s legal interpretations, as we retrieve the ground truth labels based on crowdsourced responses.
F.13 International Citizenship Questions
In LegalBench, the International Citizenship Questions task is denoted as international_citizenship_questions.
The GLOBALCIT Citizenship Law Dataset is a valuable resource that comprehensively categorizes citizenship acquisition and loss methods in 190 countries. It enables cross-country comparisons and offers insights into global trends in citizenship laws and examines 28 different ways in which citizenship can be acquired, as well as 15 ways laws allow citizenship to be lost. The original dataset is formulated as a tabular survey dataset. We change this survey format into Yes/No questions about specific countries and their laws as of 2020 resulting in 9300 question- answer pairs.
The model must answer yes/no questions about global citizenship law.
We download the GLOBALCIT Citizenship Law Dataset and craft a script that converts the tabular survey data into yes/no questions, appending information about the country and the time at which the survey was created.
Understanding knowledge about the law globally is important to evaluate. To successfully answer legal reasoning questions globally language models must be able to retrieve a rule and then reason about it.
F.14 Learned Hand Tasks
In LegalBench, the Learned Hand tasks are denoted as learned_hand_*.
A person may experience problems in many areas of their lives – their family, work, finances, housing, education, driving, and more – which legal professionals would recognize as being ‘legal issues’. The person may not know that a problem with a credit card company, a landlord, a spouse, or an employer is a ‘legal issue’, or what terminology or categorization a lawyer would use to make sense of it.
The designation of a legal issue means that a person may benefit from getting specialized guidance from a legal professional to resolve this problem because they can guide them on their rights, liabilities, possible options, procedures, and specialized legal tasks. Not all people may want to pursue legal services to resolve a legal issue. But legal issue-spotting can help them both make sense of the problem they are experiencing, and what services and laws might be available if they wish to make use of them.
Legal professionals typically carry out issue-spotting during an intake process. They receive a person’s verbal or written description of the situation they are in. Then the professional identifies the main legal issues that are apparent in this situation, starting at the main top-level categories and then sometimes proceeding to identify more specific sub-categories of legal issues. For example, the professional may identify that a person’s situation involves a legal issue with the top-level category of ‘housing’ and specific sub-categories of ‘possible eviction for non-payment of rent’ and ‘poor living conditions of their rental’.
The professional may identify multiple overlapping legal issue categories in one situation. For example, the professional may identify that a person has a housing law issue and a family law issue if their landlord is threatening to evict them because of the police being called to the rental home because of a domestic violence incident.
The main categories of legal issues that professionals would identify in people’s situations are:
Benefits: A situation would have a benefits issue if it involves the person attempting to resolve a problem with applying for, receiving, or discontinuing public benefits and social services from the government. This could include benefits that support them regarding food, disability, old age, housing, health, unemployment, child care, or other social needs.
Business: A situation would have a business issue if the person is running a small business or nonprofit, and encounters a problem around incorporation, licenses, taxes, regulations, bankruptcies, or natural disasters. This category is not meant to apply to larger corporate legal issues, but rather the kinds of business problems that an individual might bring to a legal professional for help.
Consumer: A situation would have a consumer legal issue if the person was dealing with problems around debt and money, insurance, consumer goods and contracts, taxes, or small claims about the quality of service.
Courts: A situation would be categorized as a courts issue if the person is dealing with a problem around how to interact with the court system or with lawyers more broadly. This may involve the person attempting to follow legal procedures, court rules, or filing requirements, or it may involve them attempting to hire, manage, or address lawyers.
Crime: A situation would have a crimes issue if the person is dealing with the criminal justice system as a defendant, victim, or family member. They may be experiencing problems around being investigated, searched, or charged with a crime, or going to a criminal trial and prison, or being a victim of a crime.
Divorce: A situation would be categorized as a divorce issue if a person is dealing with a separation, divorce, or annulment while splitting with a spouse or partner. The problem may involve separation, spousal support, splitting money and property, child support, visitation, or following the related court process.
Domestic Violence: A situation would have a domestic violence issue if the person is dealing with abuse with a partner, family member, or other intimate acquaintance. The situation may involve understanding rights and laws related to domestic violence, getting protective orders, enforcing them, reporting abuse, and dealing with collateral consequences to housing, finances, employment, immigration, and education.
Education: A situation has an education issue if the person is dealing with a problem around school for themselves or a family member. The situation may involve accommodations for special needs, discrimination, student debt, discipline, or other issues in education.
Employment: A situation would be identified as having an employment issue if the person has a problem with a job, including during the application process, during the job, or after ending employment. Problems may include discrimination, harassment, payment, unionizing, pensions, termination, drug testing, background checks, worker’s compensation, classification as a contractor, or more.
Estates: A situation would have an estates issue if a person is dealing with an estate, wills, or guardianship. This may include issues around end-of-life planning, health and financial directives, trusts, guardianships, conservatorships, and other estate issues that people and family deal with.
Family: A situation would have a family law issue if a person is dealing with an issue involving a family member. This may include issues around divorce, child custody, domestic violence, adoption, paternity, name change, and other family issues.
Health: A situation would be categorized as a health law issue if the person is dealing with problems around accessing health services or protecting their rights in medical settings. This may involve problems with accessing health care, paying for care, getting public benefits for care, privacy of medical records, problems with quality of care, or other issues.
Housing: A situation would have a housing law issue if a person is dealing with problems around the housing where they live, or that they own. These include problems with renting a home, eviction, living conditions, discrimination, foreclosure, post-disaster housing, housing assistance, and more.
Immigration: A situation would have an immigration issue if a person is not a full citizen in the US and is dealing with problems related to their status. This may include understanding visa options, working as an immigrant, political asylum, border searches, deportation, human trafficking, refugees, immigration court hearings, and more.
Torts: A situation would be categorized as having a torts issue if the person is dealing with an accident or conflict with another person that involves some perceived harm. These problems may include a car accident, conflicts with neighbors, dog bites, bullying, harassment, data privacy breaches, being sued, or suing someone else.
Traffic: A situation would have a traffic law issue if the person is experiencing a problem with traffic, parking, or car ownership. This might include problems with getting ticketed, getting or reinstating a driver’s license, car accidents, purchasing a car, repossession, and more.
This set of categories is commonly used by legal professionals as they triage potential clients. The LIST Taxonomy from Stanford Legal Design Lab has formalized these categories into a machine-readable taxonomy, available at https://taxonomy.legal/. The LIST taxonomy builds on the taxonomies built by legal aid groups, like the Legal Services Corporation categories list that most legal aid groups use to encode their matters, signifying what issues they helped clients with.See the LSC’s list at https://www.lsc.gov/i-am-grantee/grantee-guidance/lsc-reporting-requirements/case-service-reporting/csr-handbook-2017. LIST also builds off of the legal aid community’s National Subject Matter Index, which was a more extensive list of categories to further assist legal aid groups in tracking the issues they helped people with.See the NSMI database at https://nsmi.lsntap.org/.
Performing the Legal Issue-Spotting task requires parsing through informal wording and structures that a person may use to convey the situation they are struggling with. Typically the person is writing this narrative down in an informal, quick manner (like in an online intake form on a legal services website) or speaking it aloud (like on an intake hotline, or during an in-person interview). The narratives are not typically structured into a concise order. They use informal terminology rather than legal terms.
The Legal Issue-spotting task requires an LLM to consider a person’s narrative about their situation. The LLM must use this narrative to determine which legal issue category (or categories) apply to the person’s situation.
There is a crowdsourced dataset, created via the online labeling game Learned Hands, that has established when and how these legal issue categories apply to people’s narratives. The narratives are drawn from the subreddit r/legaladvice, in which people share several lines or paragraphs about the situation they are dealing with, which they think might involve legal issues. The moderators of the subreddit are active in managing activity, so the posts do not contain personal identifying information or off-topic postings.
The Stanford Legal Design Lab and Suffolk LIT Lab built the Learned Hands game so that law students and lawyers could read narratives one by one, and then answer a series of yes-no-skip questions about what legal issue seems to be present. Once there are sufficient consistent votes for a certain label to apply, or to not apply, to a narrative, then the label is finalized. A narrative may have more than one label, as mentioned above.
The legal issue categorization helps the professional triage the person to the right services, resources, and procedures that can assist them in resolving the legal issue. If an LLM is able to identify legal issues in people’s informal narratives, this demonstrates an ability to perform a key task in people’s justice journeys. Issue-spotting by an LLM may be able to help a person who is just starting to explore whether or how they should engage with legal services, courts, or exercising their rights.
The issue-spotting task may be provided online, when people are visiting legal help websites and trying to find what guide, form, or service would best help them with their problem. Or it may be integrated into the intake process that paralegals or justice workers carry out over hotlines or in-person, to speed up the often lengthy intake process.
F.15 Legal Reasoning Causality
In LegalBench, the Legal Reasoning Causality task is denoted as legal_reasoning_causality.
In many legal domains, systematic evidential barriers hinder the substantiation of causal claims through direct evidence. To address these shortcomings, courts have recognized the power of statistical evidence in establishing causation in various contexts, such as product liability,See, e.g, Neurontin v. Pfizer, 712 1st Cir. 52 (2013). medical malpractice,O’Neal v. St. John Hosp. & Med. Ctr, 487 Mich SC, 485 (2010). discrimination,See, e.g., International Broth. of Teamsters v. U.S., 431 U.S. 324 (“[i]n many cases the only available avenue of proof is the use of racial statistics to uncover clandestine and covert discrimination by the employer or union involved”); Bazemore v. Friday, 478 U.S 385 (1986); Marcus Jones v. Lee Way Motor Freigh, 431 10 th cir 245 (1970) (“In racial discrimination cases, statistics often demonstrate more than the testimony of many witnesses”). and more. For instance, when pursuing a labor discrimination claim, the plaintiff must establish that her protected trait was the underlying reason for the alleged discriminatory decision (e.g., firing or not hiring). However, direct evidence of discriminatory intent rarely exists, so it is often nearly impossible to refute the possibility that other (legitimate) differences between two employees or candidates were the cause for favoring one over the other. In such cases, litigants can and often do try to substantiate a causal link between the plaintiff’s group affiliation and the defendant’s behavior through statistical analysis. For instance, plaintiffs might send fictitious resumes that differ only by the suspected demographic characteristic,Havens Realty Corp. v. Coleman, 455 U.S. 363, 374-75, 71 L. Ed. 2d 214, 102 S. Ct. 1114 (1982). akin to a field experiment. Likewise, statistical analysis of observational data that controls for the major factors affecting the employment practice can be used to demonstrate whether a specific social group suffers from inferior outcomes (relative to some control group) vis-à-vis a particular employer, landlord, or lender that engages with a sufficiently large number of employees or customers.See Bazemore v. Friday, 478 U.S 385 (1986) (“although it need not include every conceivable factor. Given the frequency of employment discrimination litigation in the contemporary United”).
The “causal reasoning” task requires an LLM to determine whether the court’s reasoning regarding the finding of whether a causal link exists between the plaintiff’s protected trait and the allegedly discriminatory decision relied on either statistical or direct-probative evidence. It requires understanding the types of words that are used to describe statistical evidence in any given context (regression, correlation, variables, control, and more), and the extent to which those words relate to substantiating a finding of causality (as opposed to other legal components).
We manually created a dataset of fifty-nine excerpts from court decisions in lawsuits alleging labor market discrimination filed in US Federal District Courts. First, fifty-nine court decisions involving claims of labor discrimination were identified using the LexisNexis database. Second, the passages in which the finding of causality appeared were identified and extracted. Third, we coded the passages as either relying on statistical evidence (e.g., regression analysis, findings of correlation, etc.) or on direct evidence (e.g., witnesses, documents, etc.).
We selected two random samples from each class to use as part of the train split.
The potential of LLMs to identify different types of legal reasoning in general, and the finding of causality in particular, has implications both for the legal profession and for the academic study of law and judicial decision-making. First, algorithmic tools are gradually being utilized by lawyers to assist them in preparing for litigation. Specifically, given the heterogeneity among judges, a key element of a successful litigation strategy is a lawyer’s ability to construct their arguments based on the specific inclinations of the judge assigned to the case. Gaining an accurate understanding of judges’ unique mode of reasoning (including, e.g., the types of evidence they tend to rely on), based on their prior decisions, is crucial for winning any lawsuit. Second, databases consisting of court decisions are the most common source for studying the law and judicial decision-making in legal academia. However, these databases are typically limited to rather technical information, such as the names of the parties and the judge(s), the legal area of the case, and the like. The essential part of any judicial opinion – the legal reasoning – is typically treated as a black box. An LLM that could classify the various types of legal reasoning – e.g., what evidence is used to establish causation – can facilitate studying judicial decision- making in ways currently not feasible at large scales.
F.16 MAUD Tasks
In LegalBench, the MAUD tasks are denoted as maud_*.
We adapt the Merger Agreement Understanding Dataset (MAUD) for LegalBench. MAUD consists of over 37,000 expert-annotated examples for a set of legal reading comprehension tasks based on the American Bar Association’s 2021 Public Target Deal Point Study. In the Study, lawyers review merger agreements and identify key legal clauses (“deal points”) within those contracts. The lawyers then specify the nature of the clauses by answering a predetermined set of questions that cover a wide range of topics, including conditions to closing, the definition of material adverse effect, and remedies to breach of contract. MAUD’s multiple-choice format, according to , assesses an LLM’s ability to interpret the meaning of specialized legal language.
The tasks take advantage of MAUD’s reading comprehension component. They require an LLM—given a key legal clause and a set of descriptions for the clause—to choose the option that best describes the clause.
These tasks are constructed by transforming the abridged dataset released by . The abridged dataset contains 14,928 examples with deal points extracted from 94 merger agreements covering 92 multiple-choice questions. We narrow down to 57 questions by filtering out the ones with fewer than 50 examples. Each example consists of the text of a deal point, the question, options, and the answer key.
We create translations that map the questions into human-readable multiple-choice prompts. For instance, the prompt for the question “Accuracy of Target ‘General’ R&W: Bringdown Timing” is “When are representations and warranties required to be made according to the bring down provision?” It is then followed by an enumeration of the options for the LLM to choose among.
We focus on MAUD’s abridged examples because we are interested in assessing an LLM’s legal reading comprehension capability rather than its ability to extract relevant text segment given a complete deal point. Additionally, inputs of examples from the main dataset, which contains complete deal point texts, are oftentimes far longer than what an average open-source LLM could ingest at once, rendering them unsuitable for benchmarking purposes.
The table below lists the question and options for each MAUD-based LegalBench task along with an example input-answer pair.
Reading comprehension is a particularly challenging part of contract review, both to human and to machine. The MAUD tasks evaluate an LLM’s ability to understand and categorize a wide spectrum of legal clauses in the context of merger agreements.
F.17 New York State Judicial Ethics
In LegalBench, the New York State Judicial Ethics task is denoted as nys_judicial_ethics.
The New York State Advisory Committee on Judicial Ethics posts rulings on real ethical scenarios. The Committee, established in 1987, offers guidance to roughly New York State judges and justices, as well as other judicial personnel and candidates in the state. By interpreting the Rules Governing Judicial Conduct and the Code of Judicial Conduct, the Committee assists these individuals in maintaining high ethical standards. Actions taken by judges in accordance with the Committee’s formal opinions are deemed proper, which helps protect them during any future investigations by the New York State Commission on Judicial Conduct.
300 real-world scenarios and fact patterns have been reformulated into yes or no questions to understand whether models understand ethical rules and how they might apply to different judicial situations.
For example in a 2022 decision the Committee noted that: “A judge who previously served as the District Attorney may not preside over a parole recognizance hearing concerning a parolee or releasee who had originally been convicted and sentenced during the judge’s former tenure as the District Attorney.”
We collect digest statements from the New York State Unified Court System Advisory Committee on Judicial Ethics.https://www.nycourts.gov/legacyhtm/ip/judicialethics/opinions/ We collect samples from 2010, 2021, 2022, and 2023 and then use ChatGPT to reformulate the statements into yes or no questions. To ensure that data is not used for training OpenAI models, we opt out of data use for accounts used for task creation. We leave 2010 and 2021 data for understanding scope of data leakage from opinions being online. 2022 and 2023 data should not have been seen by most models that were trained prior to these years.
An important part of legal practice is abiding by ethics rules. As agents become more involved in the legal process it will be important to understand not only whether they can understand and reason about rules for the public, but also whether they can reason about ethical principles and rules governing judges and lawyers.
F.18 OPP-115 Tasks
In LegalBench, the OPP-115 tasks are denoted as opp115_*.
The OPP-115 Corpus, consisting of 115 online privacy policies, provides a comprehensive collection of privacy statements expressed in natural language . Each of these policies has been meticulously read and annotated by a team of three law graduate students. The annotations present in the text specifically outline various data practices.
These privacy policies are classified into ten distinct categories:
First Party Collection/Use: This describes how and why a service provider collects user information.
Third Party Sharing/Collection: This explains how user information may be shared with or collected by third parties.
User Choice/Control: This delineates the choices and control options available to users.
User Access, Edit, & Deletion: This describes if and how users may access, edit, or delete their information.
Data Retention: This states how long user information is stored.
Data Security: This communicates how user information is protected.
Policy Change: This explains if and how users will be informed about changes to the privacy policy.
Do Not Track: This discusses if and how Do Not Track signals for online tracking and advertising are honored.
International & Specific Audiences: This focuses on practices that pertain only to a specific group of users (e.g., children)
Other: This encompasses additional sub-labels for introductory or general text, contact information, and practices not covered by the other categories.
A separate binary classification task has been created for each category, with negative samples drawn from the rest of the text. To ensure consistency, any text with less than 10 words has been eliminated. The ’other’ category was not included in the categorization because it was deemed too broad and wouldn’t provide much value in terms of specific classification.
The classification task associated with the OPP-115 Corpus serves as a significant measure of an LLM’s logical reasoning ability. By assigning privacy policy segments to the right categories, LLMs demonstrate their understanding and interpretation of the language and nuances within these privacy policies. Although the task may be seen as "simple" from a human legal practitioner’s viewpoint, it provides an invaluable and objective gauge of an LLM’s progress in logical reasoning and language comprehension.
F.19 Purpose of Oral Argument Questions
In LegalBench, the Purpose of Oral Argument Questions task is denoted as oral_argument_question_purpose.
Before a court decides a case, it typically calls before it the attorneys for the parties to the lawsuit to orally present their arguments for why the case should be resolved in favor of their clients and to answer any questions that the judge or judges have of them. In modern times, however, oral argument is not lawyers’ primary avenue for explaining their positions. Instead, parties submit their arguments in written form (“briefs”) and use their oral argument time primarily to supplement those submissions by reiterating their key positions, clarifying areas of ambiguity, and seeking to persuade judges who are uncertain how the case should be resolved.
Although there is no universally accepted listing, judges questions at oral argument tend to fall into a few categories:This categorization scheme is from , employed and discussed along with other possible classification schemes at .
Background: A question seeking factual or procedural information that is missing or not clear in the briefing
Clarification: A question seeking to get an advocate to clarify her position or the scope of the rule being advocated.
Implications: A question about the limits of a rule or its implications for future cases.
Support: A question offering support for the advocate’s position.
Criticism: A question criticizing an advocate’s position.
Communicate: A question designed primarily to communicate with one or more other judges on the court.
Humor: A question designed to interject humor into the argument and relieve tension.
A lawyer presenting her case before a court at oral argument must be able to quickly and accurately determine why the judge has asked a particular question so as to answer it on behalf of her client in the most persuasive way possible. It is also a difficult task. Under the pressure of persistent and difficult questioning it is easy for a lawyer to misread a question and offer an unresponsive or misguided answer. Skillful lawyers learn to quickly understand not only judges’ questions but the reasons they are asked.
The Purpose of Oral Argument Questions task requires an LLM to determine–given a question from an oral argument transcript–for which of the seven purposes above the judge asked the question.
We created a dataset of questions from U.S. Supreme Court oral argument transcripts, classified into one of the seven functions above. Questions were taken from cases argued in the 2022-23 Supreme Court term, in reverse chronological order. A question was defined as the totality of a judge’s words prior to an advocate’s response, regardless whether the words constituted a true interrogatory sentence. To include a sufficient number of questions of each type in the dataset, questions were not drawn at random. Instead rarer question types (e.g. humor and communication) were targeted for inclusion, questions of more common types (e.g. clarification) were frequently omitted.
Young attorneys, and even many experienced ones, struggle with oral advocacy. It requires comfort in the courtroom, quick thinking, and careful demeanor to assess from the tone and content of a question why the judge is asking it and how most effectively to respond. In one regard, LLMs will be superior. They do not suffer from human nervousness. However, given only a text prompt rather than an audible question from which to infer tone, and given only a judge’s question and not the full case context, this task would be a challenge even for a seasoned lawyer. Whether LLMs can succeed will be an extremely interesting measure of progress in legal analysis.
F.20 Overruling
In LegalBench, the Overruling task is denoted as overruling.
A hallmark of the common-law legal system is that courts will overrule previous judicial decisions. The act of overruling is significant, and it indicates that the overruled decision was either in accurate in its articulation/application of a particular law, or that overruling court wishes to announce a substantive change in law.
In this task, an LLM is required—given an excerpt of judicial text—to determine if the text overrules a previous decision.
This task is taken from , which previously studied the capacity for finetuned BERT models to perform this task. Please refer to for more information on the construction process.
Identifying when judicial text overrules another case is a basic but essential lawyering skill. From a practical standpoint, the capacity for LLMs to correctly classify overruling sentences could have practical applications for the design and construction of legal opinion databases. When using or citing a case in legal arguments, lawyers must ensure that the case hasn’t been an overruled, and is still “good law.” Tools which automatically parse legal databases and extract cases which have been overruled would thus be helpful for constructing legal arguments.
F.21 Personal Jurisdiction
In LegalBench, the Personal Jurisdiction task is denoted as personal_jurisdiction.
Personal jurisdiction refers to the ability of a particular court (e.g. a court in the Northern District of California) to preside over a dispute between a specific plaintiff and defendant. A court (sitting in a particular forum) has personal jurisdiction over a defendant only when that defendant has a relationship with the forum. We focus on a simplified version of the rule for federal personal jurisdiction, using the rule:
There is personal jurisdiction over a defendant in the state where the defendant is domiciled, or when (1) the defendant has sufficient contacts with the state, such that they have availed themselves of the privileges of the state and (2) the claim arises out of the nexus of the defendant’s contacts with the state.
Under this rule, there are two paths for a court have jurisdiction over a defendant: through domicile or through contacts.
Domicile: A defendant is domiciled in a state if they are a citizen of the state (i.e. they live in the state). Changing residency affects a change in citizenship.
Contacts: Alternatively, a court may exercise jurisdiction over a defendant when that defendant has sufficient contacts with the court’s forum, and the legal claims asserted arise from the nexus of the defendant’s contacts with the state. In evaluating whether a set of contacts are sufficient, lawyers look at the extent to which the defendant interacted with the forum, and availed themselves of the benefits and privileges of the state’s laws. Behavior which usually indicates sufficient contacts include: marketing in the forum or selling/shipping products into the forum. In assessing nexus, lawyers ask if the claims brought against the defendant arise from their contacts with the forum. In short: is the conduct being litigated involve the forum or its citizens in some capacity?
The personal jurisdiction task requires an LLM to determine—given a fact pattern describing the events leading up to a legal claim—whether a particular court has personal jurisdiction over the defendant.
We manually construct a dataset to test application of the personal jurisdiction rule, drawing inspiration from exercises found online and in legal casebooks. Each sample in our dataset describes a “fact pattern," and asks if a court located in particular state (A) can exercise personal jurisdiction over an individual (B) named in the fact pattern. In designing the dataset, we use 5 base fact patterns, and create 4 slices, where each slice evaluates a different aspect of the personal jurisdiction rule:
Domicile: Fact patterns where B is domiciled in A. Hence, personal jurisdiction exists.
No contacts: Fact patterns where B has insufficient contacts with A. Hence there is no personal jurisdiction.
Yes contacts, no nexus: Fact patterns where B has sufficient contacts with A, but the claims against B do not arise from those contacts. Hence, personal jurisdiction does not exist.
Yes contacts, yes nexus: Fact patterns where B has sufficient contacts with A, and the claims against B arise from those contacts. Hence, there is personal jurisdiction.
Caveat. Personal jurisdiction is a rich and complex doctrine. Our dataset focuses on a narrow class of fact patterns, related to jurisdiction over individuals. We don’t consider, for instance, more complex questions related to adjudicating citizenship (e.g. the Hertz test) or the classic stream-of-commerce problems. We leave this to future work.
Identifying when personal jurisdiction exists is a skill that law students learn in their first-year civil procedure course. The personal jurisdiction task is interesting because applying even the simplified version of the rule requires reasoning over the degree of connection between a defendant and the forum state.
F.22 Privacy Policy Entailment
In LegalBench, the Privacy Policy Entailment task is denoted as privacy_policy_entailment.
The Privacy Policy Entailment task is created from the APP-350 corpus , which consists of 350 Android app privacy policies annotated with different privacy practices. In this corpus, individual clauses are annotated based on whether they do or do not perform a certain practice (e.g., “We access your location information”).
Given a clause from a privacy policy and a description of the practice, the LLM must determine if the clause describes the performance of that practice. This is analagous to an entailment task, where the premise is the clause, and the hypothesis is the practice description.
For each practice coded in the APP-350 corpus, we derive a natural language description of that practice, which serves as our “hypothesis.” Each instance of this task corresponds to a triple containing a clause, a practice description, and a binary classification (Yes/No) based on whether the clause performs the practice. Across the dataset there are 57 unique policy descriptions.
The privacy policy entailment task is similar to ContractNLI, in that it evaluate a LLM’s capacity to perform entailment-style reasoning over formal legal language. From a lawyerly perspective, understanding whether a policy performs certain functions or empowers one of the parties to pursue practices is an essential element of legal comprehension. From a practical perspective, the ability for LLMs to perform this task could empower researchers to conduct broader studies of privacy agreements. As observes, annotation cost limitations often restrict the scope of empirical studies of privacy agreements.
F.23 Privacy Policy QA
In name, the Privacy Policy QA task is denoted as privacy_policy_qa.
The Privacy Policy QA task is derived from , which annotated clauses in mobile application privacy policies based on whether they contain the answer to a question.
Given an excerpt from a privacy policy and a question, the LLM must determine whether the excerpt is relevant to answering the question or not.
We used the snippet annotations available in to construct this task, removing all snippets with fewer than 10 words. Examples of excerpt/question/relevant tuples are shown in the table below. 5449 instances correspond to a relevant question-clause pair, and 5474 instances correspond to an irrelevant question-clause pair.
Determining when a particular legal excerpt is relevant to answering a question is essential to interpreting legal documents. This task allows us to evaluate LLMs for this capability. From a more practical standpoint, the Privacy Policy QA task is a helpful evaluation task when developing LLM systems which involve decompositions over long documents. A common approach—in order to account for the fact that many long documents exceed ordinary context windows—is to chunk documents into smaller segments, and apply a LLM independently to filter out irrelevant segments. For QA tasks involving long policies, this task allows practitioners to measure performance for the filtering step.
F.24 Private Right of Action (PROA)
In LegalBench, the Private Right of Action task is denoted as proa.
A private right of action (PROA) exists when a statute empowers an ordinary individual (i.e. a private person) to legally enforce their rights by bringing an action in court. In short, a PROA creates the ability for an individual to sue someone in order to recover damages or halt some offending conduct. PROAs are ubiquitous in antitrust law (in which individuals harmed by anti-competitive behavior can sue offending firms for compensation) and environmental law (in which individuals can sue entities which release hazardous substances for damages) .
In the PROA task, an LLM must determine if a statutory clause contains a private right of action.
We construct a dataset of PROAs by hand, drawing inspiration from clauses found in different state codes. We construct 50 clauses which do contain a PROA, and 50 clauses which do not. Clauses which do not contain a private right may either create no cause of action, or empower a non-private individual (e.g., an attorney general) to bring a claim. 5 randomly sampled clauses constitute the training set, and the remaining 95 form the test set.
The PROA task evaluates an LLM’s ability to perform a two-step reasoning test: (1) does the statute allow a party to bring a claim in court, and (2) is that party private? Law students and legal professionals should be capable of performing this task at near-perfect accuracy.
The PROA task derives additional significance from a recent movement towards studying state statutory language . Legal scholars have long been unable to conduct large scale empirical studies of state statutory language, given the sheer volume of state statutes. The ability for LLMs to accurately classify or annotate statutes could thus empower new empirical studies of state statues.
F.25 Rule QA
In LegalBench, the Rule QA task is denoted as rule_qa.
Lawyers are regularly required to recall specifical legal rules that are drawn from cases, statutes, or other sources. Rules can take many shapes and forms. For instance, the rule pertaining to the federal requirements for a class (in a class action lawsuit) are codified in Rule 23(a) of the Federal Rules of Civil Procedure and are simply known as the need for “numerosity, commonality, typicality, and adequacy.”
The Rule QA task evaluates a LLM’s ability to answer questions on different legal rules. The rules are drawn from subjects typically studied in the first year of law school (e.g., civil procedure, constitutional law, etc.). This is an open-generation task.
We manually wrote 50 question-answer pairs, focusing on the types of rules which are regularly tested in law school courses on civil procedure, evidence, and intellectual property. The questions ask the LLM to either (1) restate a rule, (2) identify where a rule is codified, or (3) list the factors employed in a particular rule. Several questions explicitly narrow their scope to a jurisdiction (e.g., California state evidence law), in order to avoid bias towards merely federal law.
The Rule QA task is an initial effort to evaluate the propensity for legal hallucination in LLMs. The questions asked are exceedingly basic, and law students taking the relevant course would be expected to answer them nearly perfectly.
F.26 SARA Tasks
In LegalBench, the SARA tasks are denoted as sara_*.
An important skill for lawyers is to determine, given the facts of a case, whether a given law applies and what it prescribes. For example, does the payment received by the defendant on August 21st , 2017 qualify as wages under §3306(b) of the US Tax Code? This task has been introduced by as statutory reasoning. further introduce the Statutory Reasoning Assessment dataset (SARA). SARA contains (1) a set of 9 sections, taken from US federal tax law statutes, pruned and simplified; and (2) hand-crafted cases that test the understanding of those 9 sections. In this context, a case is a paragraph of text stating facts in plain language. Each case comes either with an entailment prompt — a statement about the statutes and the case that may be true or false — or a question — asking how much tax one of the case’s protagonists owes. The SARA dataset is a simplified version of real-world cases, that retains many of the features of statutory reasoning for tax law. It poses, however, a significant challenge to NLP models .
There are two SARA tasks. The first, sara_entailment, corresponds to the entailment cases. The entailment cases state that a given law applies to a given case, and require the LLM to produce a binary answer — akin to Recognising Textual Entailment . This is an approximation of real-world statutory reasoning, where the answer is usually not strictly binary.
The second task, sara_numeric, consists of the numeric cases. Here, the goal is to compute the amount of tax owed. We frame this as a floating point number. To measure numerical accuracy, we use the metric introduced by , which includes a tolerance for inaccurate predictions.
We framed the SARA dataset for the paradigm of language modeling. Due to dependencies across sections in the statutes, the entirety of the statutes are generally relevant to determine the answer to any of the cases. However, all 9 sections do not fit into the LLM’s context window, and must be pruned. In entailment cases, the entailment prompt specifies which law from the statutes to apply. We automatically extract the text of that law, and use it as the language model prompt. For numerical cases, the entirety of the statutes are relevant, and we use that as the language model prompt. Pruning is left to the LLM’s pre-processing.
Statutory reasoning is an important skill for lawyers, that is used within many other legal tasks. It is a fundamental task for legal AI, probing whether a computational model can understand and reason with legal rules expressed in natural language. The types of reasoning involved in SARA are diverse — defeasible, temporal, numerical reasoning, inter alia — and relevant beyond the legal domain. Statutory reasoning combines natural language understanding and logical reasoning, a major goal for AI.
If statutory reasoning were solved, it could serve as a basis for more complex legal tasks. For example, it could be used to automate the computation of taxes and benefits, without the need for coding the expert systems in use in many parts of the world . A system that can do statutory reasoning would also be a step towards machine reading models that can analyze legislation and anticipate its effects, coming up with possible application scenarios . As a final example, a statutory reasoning agent could be used for basic legal advice, increasing access to justice.
F.27 SCALR
In LegalBench, the SCALR task is denoted as scalr.
Each case decided by the Supreme Court addresses a specific question presented for review. Both the questions and the Court’s opinions are published on the Supreme Court’s website.
Many of the Court’s opinions are briefly described by other judges who recount the holdings of the Court in their own writing. For example, consider the following passage from State of South Carolina v. Key, 27971 (S.C. 2020; emphasis added):
The United States Supreme Court has addressed the constitutionality of warrantless blood draws in several DUI cases. See Schmerber, 384 U.S. at 770-71 (holding the warrantless blood draw of a DUI suspect was valid because the law enforcement officer, dealing with a car accident, could “reasonably have believed that he was confronted with an emergency, in which the delay necessary to obtain a warrant, under the circumstances, threatened ‘the destruction of evidence’")…
We refer to these brief descriptions as ‘holding statements’ or ‘holding parentheticals,’ since they are often enclosed by parentheses. Identifying the holding parenthetical that corresponds to a question presented for review requires a notion of ‘responsiveness’ or relevance between questions and answers as well as an understanding of the kinds of legal issues that could be implicated by a specific question presented for review.
The SCALR benchmark is a collection of 571 multiple choice questions designed to assess the legal reasoning and reading comprehension ability of large language models. Each multiple-choice task gives the question presented for review in a particular Supreme Court case. The solver must determine which holding parenthetical describes the Court’s ruling in response to the question presented. Here is an example from AT&T Mobility LLC v. Concepcion, 563 U.S. 333 (2011) with the correct response emphasized:
Question: Whether the Federal Arbitration Act preempts States from conditioning the enforcement of an arbitration agreement on the availability of particular procedures–here, class -wide arbitration–when those procedures are not necessary to ensure that the parties to the arbitration agreement are able to vindicate their claims.
A: holding that when the parties in court proceedings include claims that are subject to an arbitration agreement, the FAA requires that agreement to be enforced even if a state statute or common-law rule would otherwise exclude that claim from arbitration
B: holding that the Arbitration Act “leaves no place for the exercise of discretion by a district court, but instead mandates that district courts shall direct the parties to proceed to arbitration on issues as to which an arbitration agreement has been signed"
C: holding that class arbitration “changes the nature of arbitration to such a degree that it cannot be presumed the parties consented to it by simply agreeing to submit their disputes to an arbitrator"
D: holding that a California law requiring classwide arbitration was preempted by the FAA because it “stands as an obstacle to the accomplishment and execution of the full purposes and objectives of Congress," which is to promote arbitration and enforce arbitration agreements according to their terms
E: holding that under the Federal Arbitration Act, a challenge to an arbitration provision is for the courts to decide, while a challenge to an entire contract which includes an arbitration provision is an issue for the arbitrator
The data used to create this task comes from two sources:
Questions presented were gathered from the Supreme Court of the United States’ website, which hosts PDFs of questions granted for review in each case dating back to the 2001 Term.
Holding statements that comprise the “choices" for each question were compiled from (a) CourtListener’s collection of parenthetical descriptions and (b) extraction of parenthetical descriptions from Courtlistener’s and the Caselaw Access Project’s collections of court decisions using Eyecite.
Because questions presented for review in Supreme Court cases are not easily available prior to 2001, the benchmark is limited to questions from cases decided in the 2001 Term and later. To ensure that “holding" statements would address the particular question presented, we limited the set of cases to those in which exactly one question was granted for review. We also perform some manual curation to exclude questions which are not answerable without specific knowledge of a case. For example, we eliminated a case that presented this question: “Whether this Court’s decision in Harris v. United States, 536 U.S. 545 (2002), should be overruled."
To create choices for each question presented, we first filter our set of parenthetical descriptions as follows:
We limited our parenthetical descriptions to only those that begin with “holding that…", as these are most likely to describe the core holding of the case, rather than some peripheral issue.
We use only parentheticals that describe Supreme Court cases. This avoids the creation of impossible questions that ask the solver to distinguish between “holding" statements dealing with exactly the same issue at different stages of appellate review.
We then select for each case the longest parenthetical meeting the above criteria. We use the longest parenthetical because it is most likely to be descriptive enough to make the question answerable.
We then create a task for each case which has both a question presented and a “holding" statement meeting the above requirements. (While question-correct holding pairs are only for cases decided after 2001, we allow the use of parentheticals describing any Supreme Court case as alternative answer choices.) We then need to select the four alternative answer choices for each question in a manner that makes the task challenging. To select choices that are at least facially plausible, we find the four “holding" statements from the remaining pool that are most TF-IDF similar to the question presented. The inclusion of difficult alternative choices requires the solver to draw nuanced distinctions between legal issues that share overlap in terminology.
This task is significant because it tracks the useful and challenging skill of identifying a passage as relevant or responsive to a given query. LLMs that are able to perform well at this task have the potential to be more useful for complex legal question-answering and retrieval. The poor performance of simpler models on this task demonstrates that it is a challenging one that requires a level of understanding beyond the word/synonym level.
F.28 Securities Complaint Extraction
In LegalBench, the Securities Complaint Extraction tasks are denoted as ssla_*.
Securities Class Actions (SCAs) are lawsuits filed by, and on behalf of, investors alleging economic injury as a result of material misstatements or omissions in public disclosures made by corporate directors and officers. These actions allege violations of the Securities Act of 1933 and Securities Exchange Act of 1934 and are predominately filed in federal court, though in 2018 the United States Supreme Court determined actions brought under the ‘33 Act were permitted in state court.
“Plaintiff(s)” is the legal term to describe the individual, company, or organization bringing forth a lawsuit. Under the class action system, one or more plaintiffs are appointed “Lead Plaintiff” by the court to represent the interests of a larger group of “similarly situated” parties. In securities class actions, investors that suffered the greatest financial loss, often public pensions or unions, are appointed lead plaintiff.
“Defendant(s)” is the legal term to describe the individual, company, or organization that must defend themself against the alleged violations or misconduct outlined in the lawsuit. There is always at least one defendant. The majority of securities class actions name the company, its CEO and its CFO. Many name additional C-suite level officers, members of the Board of Directors and additional third-parties such as the company’s independent auditor and the underwriters of public offerings.
Each designated lead plaintiff, and all named defendants, are explicitly identified under the “Parties” section of the class action complaint.
The plaintiff task requires an LLM to extract the named plaintiffs within a text.
The individual defendants tasks require an LLM to extract named defendants who are individuals from within a text.
The company defendants tasks require an LLM to extract named defendants who are corporations/companies from within a text.
For certain samples, the complaint excerpt may not exactly contain the answer. For example, the correct answer may be “Strongbridge Biopharma PLC,” while the complaint only mentions “Strongbridge.” In order to maintain fidelity to the workflow used by SSLA, we evaluate an LLM’s ability to generate the official name of the entity, as represented in the answer. This requires the LLM to possess some background knowledge regarding official corporation names. We find that larger models are generally able to account for this.
Sometimes, the provided text will not explicitly name the plaintiff, an individual defendant, or a corporate defendant. In these cases, the LLM is expected to return “Not named”.
Stanford Securities Litigation Analytics (SSLA) identifies, tracks, and aggregates data on the several hundred private shareholder lawsuits and public SEC/DOJ enforcements filed each year. SSLA fellows manually extract and analyze information including plaintiffs, defendants, judges, mediators, plaintiff and defense firms, key litigation events, real-time case statuses, settlement timing, settlement dollar amounts, attorneys’ fees and expenses, and imposed SEC / DOJ penalties. There is no ambiguity regarding the answers for this task given its nature.
This dataset is an extract from the corpus of texts of securities class action complaints in the SSLA database. Given the typical structure and headings for these types of cases, this dataset represents text extracted from the complaint, between the sections titled “Parties” and “Substantive Allegations”. For cases where the second heading was not found, texts fragments were limited to 2,000 characters. Cases with both headings were then filtered to include only those with texts up to 2,000 characters, which excluded cases with longer “Parties” sections. Thus, all observations in this dataset are 2,000 characters or less.
Text was scraped from complaints using python’s PyPDF2 library and left unformatted and uncleaned. This training set includes several observations where the text does not include all or some of the named entities due to the method of text collection and variation in case structure. In several of these cases, plaintiff names are not present in the selected text because the plaintiff had been named earlier in the complaint.
Extracting data from legal documents is an extraordinarily resource- and time-intensive effort prone to human error. As a result, there are no known databases of non-securities class action litigation, despite the obvious public policy implications of the class action system. Automation of identification tasks coupled with human approval would improve efficiency and reduce collection costs and data errors. This task may be useful to other legal researchers and industry practitioners extracting structured data from complex texts. Identification is a very simple task that can be done by those with an understanding and familiarity with the underlying legal documents and legal system, but an LLM’s ability to accurately and precisely identify entities is a useful metric to assess.
F.29 Successor Liability
In LegalBench, the Successor Liability task is denoted as successor_liability.
When one company sells its assets to another company, the purchaser is generally not liable for the seller’s debts and liabilities. Successor liability is a common law exception to this general rule. In order to spot a successor liability issue, lawyers must understand how courts apply the doctrine.
The doctrine holds purchasers of all, or substantially all, of a seller’s assets liable for the debts and liabilities of the seller if:
the purchaser expressly agrees to be held liable;
the assets are fraudulently conveyed to the purchaser in order to avoid liability;
there is a de facto merger between the purchaser and seller; or
the purchaser is a mere continuation of the seller.
Express agreement is governed by standard contract law rules. In practice, if a purchase agreement contains a provision to assume liabilities, litigation will rarely arise. Courts, however, sometimes interpret an implied agreement in the absence of a written provision.
Assets are fraudulently conveyed when the seller intends to escape liability through a sale or knows that liability will be avoided through a sale.
De facto merger is a multifactor test that consists of (1) continuity of ownership; (2) cessation of ordinary business and dissolution of the acquired corporation as soon as possible; (3) assumption by the purchaser of the liabilities ordinarily necessary for the uninterrupted continuation of the business of the acquired corporation; and (4) continuity of management, personnel, physical location, assets, and general business operation. Some jurisdictions require a showing of all four elements. Others do not, and simply emphasize that the substance of the asset sale is one of a merger, regardless of its form.
Mere continuation typically requires a showing that after the asset sale, only one corporation remains and there is an overlap of stock, stockholders, and directors between the two corporations. There are two variations of the mere continuation exception. The first variation is the “continuity of enterprise” exception. In order to find continuity of enterprise, and thus liability for the purchaser of assets, courts engage in a multifactor analysis. Factors include: (1) retention of the same employees; (2) retention of the same supervisory personnel; (3) retention of the same production facilities in the same physical location; (4) production of the same product; (5) retention of the same name; (6) continuity of assets; (7) continuity of general business operations; and (8) whether the successor holds itself out as the continuation of the previous enterprise. The second variation is the product line exception. This exception imposes liability on asset purchasers who continue manufacturing products of a seller’s product line. This exception generally requires that defendants show that the purchaser of assets is able to assume the risk spreading role of the original manufacturer, and that imposing liability is fair because the purchaser enjoys the continued goodwill of the original manufacturer.
Scholars have noted that fraud, de facto merger, and mere continuation (and its variants) overlap. They share the common thread of inadequate consideration, that is, the consideration given in exchange for the assets is unable to fund the liabilities that underwrite those assets. Because of the overlap, different courts might apply different doctrines to identical sets of facts, but arrive at the same policy .
Successor liability doctrine is commonly taught in a course on corporate law or business associations in law school. Sometimes it is reserved for upper level courses in corporate finance or mergers and acquisitions. Students are expected to spot successor liability issues and understand how to determine if a successor will be held liable.
The Successor Liability task requires an LLM to spot a successor liability issue and identify its relevant exception to no liability. If more than one exception is relevant, the LLM is required to state the additional exception(s). The task does not include identification of the two variations to the mere continuation exception (continuity of enterprise and product line).
F.30 Supply Chain Disclosure Tasks
In LegalBench, the Supply Chain Disclosure Tasks are denoted as supply_chain_disclosure_*.
Corporations are frequently legally required to disclose information that may be relevant to investors, regulators, or members of the public. One example of this kind of disclosure requirement is laws that require corporations doing business in particular jurisdictions to provide detailed information on their supply chains, which is intended to ensure that the company’s business practices are not supporting things like human trafficking or human rights violations. One example of these kind of disclosure requirements is the California Transparency in Supply Chains Act (CTSCA). The CTSCA applies to corporations that are a: “ retail seller and manufacturer doing business in this state [of California] and having annual worldwide gross receipts that exceed one hundred million dollars ($100,000,000).”See California Transparency in Supply Chains Act, Cal. Civ. Code § 1714.43(a)(1) (West 2012). If a corporation meets these criteria, they are required to post information on five topics:
Verification: “[A]t a minimum, disclose to what extent, if any, that the retail seller or manufacturer . . . [e]ngages in verification of product supply chains to evaluate and address risks of human trafficking and slavery. The disclosure shall specify if the verification was not conducted by a third party.”Cal. Civ. Code § 1714.43(c)(1).
Audits: “[A]t a minimum, disclose to what extent, if any, that the retail seller or manufacturer . . . [c]onducts audits of suppliers to evaluate supplier compliance with company standards for trafficking and slavery in supply chains. The disclosure shall specify if the verification was not an independent, unannounced audit.”Cal. Civ. Code § 1714.43(c)(2).
Certification: “[A]t a minimum, disclose to what extent, if any, that the retail seller or manufacturer . . . [r]equires direct suppliers to certify that materials incorporated into the product comply with the laws regarding slavery and human trafficking of the country or countries in which they are doing business.”Cal. Civ. Code § 1714.43(c)(3).
Accountability: “[A]t a minimum, disclose to what extent, if any, that the retail seller or manufacturer . . . [m]aintains internal accountability standards and procedures for employees or contractors failing to meet company standards regarding slavery and trafficking.”Cal. Civ. Code § 1714.43(c)(4).
Training: “[A]t a minimum, disclose to what extent, if any, that the retail seller or manufacturer . . . [p]rovides company employees and management, who have direct responsibility for supply chain management, training on human trafficking and slavery, particularly with respect to mitigating risks within the supply chains of products.”Cal. Civ. Code § 1714.43(c)(5).
In addition to requiring corporations that meet the specified criteria to post disclosures that provide this information, the California Attorney General’s office has also posted a guide informing firms of “Best Practices” for what specific information to provide on each of these five topics.Cal. Dep’t of Justice, The California Transparency in Supply Chains Act: A Resource Guide (2015), https://oag.ca.gov/sites/all/files/agweb/pdfs/sb657/resource-guide.pdf.
However, prior research has suggested that companies do not always post disclosures that cover each of these topics; and, even when they do, the disclosures are not always consistent with the recommended best practices .
We constructed this task based on an existing dataset of supply chain disclosures. In the summer of 2015, we, with the help of research assistants, we searched the websites of corporations that had previously been identified by an organization called “KnowTheChain” as being required to post supply chain disclosures to be compliant with the California Supply Chain Transparency Act. Through this process, we found disclosures for roughly 400 firms out of roughly 500 firms for which KnowTheChain suggested were required to post disclosures.
For each of these roughly 400 firms, we saved copies of their supply chain disclosures. We then had research assistants read the disclosures and code whether they included each of the five required topics for disclosure and, if so, whether the disclosures on those five topics were consistent with the best practices outlined by the California Attorney General’s office.
We convert each of these 10 coded variables into a distinct binary classificationt task, producing 10 tasks. Table 48 lists each task, along with the precise question used to code the disclosure.
Corporate disclosure requirements are a commonly used regulatory tool, but evidence suggests that firms do not always fully comply with these disclosure requirements. The Supply Chain Disclosure task evaluates whether LLMs may be able to determine whether corporations are complying with those disclosure requirements. Because these disclosures are often formatted very differently, written in complex language, and may be designed to obfuscate relevant information, this task provides a useful measure of whether LLMs can parse the content covered in legal documents.
F.31 Telemarketing Sales Rule
In LegalBench, the Telemarketing Sales Rule task is denoted as telemarketing_sales_rule.
The Telemarketing Sales Rule (16 C.F.R. Part 310) is a set of regulations promulgated by the Federal Trade Commission to implement the Telemarketing and Consumer Fraud and Abuse Prevention Act. Its purpose is to protect consumers from specified deceptive and abusive telemarketing practices. This task focuses on 16 C.F.R. § 310.3(a)(1) and 16 C.F.R. § 310.3(a)(2), which outline a series of specific telemarketing practices prohibited as "deceptive." 16 C.F.R. § 310.3(a)(1) lists information that must be disclosed to a consumer before a sale is made, and 16 C.F.R. § 310.3(a)(2) lists categories of information that a telemarketer is prohibited from misrepresenting. 16 C.F.R. § 310.2 provides definitions relevant to both of these subsections.
The Telemarketing Sales Rule (TSR) is not commonly taught in law school as its own topic, but may be used as examples in courses on consumer protection law, administrative law, telecommunications law, and the like. Because of its simplicity, it has also been used in beginner-level legal research exercises tasking students with finding the TSR in the Code of Federal Regulations and applying it to a set of facts.
Applying the TSR would require an LLM to classify a set of facts as either falling within or outside of the specific prohibitions outlined in the rule. For example, the TSR requires that telemarketers disclose certain material information before a sale is made, such as the total cost of the goods or services, the quantities of goods or services being purchased, and exchange and return restrictions. It also forbids telemarketers from making material misrepresentations as to cost, quantity, quality, endorsement or sponsorship, and the like. In many real-life situations, it would be ambiguous whether certain telemarketer behavior would violate the TSR; for example, it could be contentious whether a given misrepresentation fits the definition of “material.” However, this task is limited to clear, unambiguous violations or non-violations, such as if a telemarketer told a consumer that they were selling four apples for four dollars, when in fact they were selling four apples for six dollars.
The following subsections 16 C.F.R. § 310.3(a)(1) and 16 C.F.R. § 310.3(a)(2) were ignored in the task, given their complexity or their reference to other statutes and regulations:
The TSR task is meant to test whether an LLM can classify simple sets of facts as describing a violation of the TSR, or not describing a violation of the TSR.
We manually created 50 samples, such that examples of at least one violation and at least one non-violation of each relevant subsection of 16 C.F.R. § 310.3(a)(1) and 16 C.F.R. § 310.3(a)(2) were present.
Determining whether a simple and unambiguous set of facts falls within the ambit of 16 C.F.R. § 310.3(a)(1) or 16 C.F.R. § 310.3(a)(2) would be an easy task for law students and lawyers, as well as many non-lawyers. However, an LLM that was trained to recognize clear violations of consumer protection laws could help administrative agencies like the Federal Trade Commission inform normal citizens of their rights.
F.32 Textualism Tasks
In LegalBench, the Textualism tasks are denoted as textualism_tool_*.
Courts regularly interpret statutes to determine the precise meaning of words contained in the statute. For instance, suppose a statute specifies that “It shall be illegal to park a vehicle inside public parks for longer than thirty minutes.” A court may be asked to determine whether the statute prohibits persons from parking bicycles inside public parks. This requires defining the term “vehicle” and determining if a bicycle is a type of vehicle.
To guide the interpretation of ambigous statutory terms, American jurisprudence has developed numerous principles of statutory construction or interpretation. These principles—also referred to as tools or canons—are rules which dictate how terms in statutes should be interpreted. For instance, the principle of ejusdem generis states that where general words or phrases follow a number of specific words or phrases, the general words are specifically construed as limited and apply only to persons or things of the same kind or class as those expressly mentioned .
One approach to statutory interpretation—known as textualism—states that only the text of the statute should be considered . In contrast, other approaches to interpreting an ambigous term might call for a court to analyze the purpose of the statute, the history of the statute, or the intent of the legislature.
The Textualism tasks ask a LLM to determine if an excerpt of judicial text is applying a specific textual tool when performing statutory interpretation. There are two tasks: dictionaries (textualism_tool_dictionaries) and .
The first task is plain-meaning (textualism_tool_plain), and it requires an LLM to determine if a court is applying the “plain meaning” rule. The plain meaning rule says that statutory text should be interpreted according to its plain or ordinary meaning.
The second task is dictionaries (textualism_tool_dictionaries), and it requires an LLM to determine if a court is using dictionaries to define the statutory text.
For each task we extracted paragraphs from Court of Appeals opinions and manually annotated whether the paragraphs showed the court as “using” the respective tool.
In order to count as using plain meaning, the paragraph must reference the plain or ordinary meaning of the text. This includes directly saying “plain meaning” or referencing the general logic of the plain meaning rule. There must also be evidence that the court used the tool in its decision. This latter condition is notable because legal scholars often care about whether the court actually used the tool when defending its decision. Common examples of using include stating it as a general rule of decision (“ur obligation is to look to the plain language of the statute to effectuate the intent of congress”) or applying it to the facts (“The statute’s plain language indicates the 150% fee cap applies if (1) the plaintiff was “a prisoner” at the time he brought the action and (2) he was awarded attorney’s fees pursuant to § 1988.”). “Using” does not, for example, include paragraphs that criticize the use of the plain meaning rule.
In order to count as using dictionaries, the paragraph must reference a dictionary. There must also be evidence that the court used a dictionary as part of its rationale. This latter condition is notable because legal scholars often care about whether the court actually used the tool when defending its decision. Common examples of using include stating it as a general rule of decision (“[We use a dictionary to help determine the plain meaning of the statutory text”) or applying it to the facts (“According to the Websters dictionary, a vehicle is any means in or by which someone travels, or something is carried or conveyed”). “Using” does not, for example, include paragraphs that criticize the use of dictionaries.
Recognizing when a court is applying a particular canon of interpretation is a classical skill law students are expected to learn. LLM performance on this task thus offers a heuristic for comparing LLM comprehension of judicial text to that of a law student’s. More practically, the capacity for LLMs to detect when certain canons are being applied could make them a valuable tool for empirical legal scholars.
F.33 UCC vs Common Law
In LegalBench, the UCC vs Common Law task is denoted as ucc_v_common_law.
In the United States, contracts are typically governed by one of two different bodies of law depending on the subject matter of the contract. Contracts for the sale of goods (physical, moveable things) are governed by the Uniform Commercial Code (UCC), a uniform set of laws created by the Uniform Law Commission and adopted in all US jurisdictions. Contracts for services and real estate, on the other hand, are governed by state common law. For example, a contract for Alice to sell Bob her bike would be governed by the UCC (sale of a good), but a contract for Bob to repair Alice’s bike would be governed by the common law (service).
This distinction is significant because the UCC and the common law diverge on numerous important legal issues such as:
Offer and acceptance: The common law requires an offeree’s acceptance to exactly match the terms of the offeror’s offer in order for a contract to be formed (the “mirror image” rule). The UCC, on the other hand, allows for some variation in the terms under UCC Section 2-207.
Definiteness: For a common law contract to be enforceable, it must be reasonably definite with respect to all material terms. For example, a service contract would not be enforceable without a price term or an adequate description of the service to be provided. The UCC only requires that a goods contract include the good being sold and the quantity. If any other term is missing from the contract (such as price or delivery), it will be filled in by UCC default rules.
Options: To create an option contract (by which the offeror provides the offeree with a defined period of irrevocability), the common law requires that the offeree give the offeror separate consideration for the option. The UCC allows merchants to make “firm offers” (effectively option contracts) without the offeree providing separate consideration.
Modification: To modify an existing contract, the common law requires both parties to provide new consideration (the “preexisting duty rule”) whereas the UCC only requires that modifications be made in good faith.
The UCC vs. Common Law task requires an LLM to determine whether a contract is governed by the UCC or by the common law.
The dataset was manually created to test an LLM’s ability to determine whether a contract is governed by the UCC or by the common law. The dataset is composed of 100 descriptions of simple contracts such as “Alice and Bob enter into a contract for Alice to sell her bike to Bob for 100 to mount a television on the wall of her living room” (common law). Each description is followed the question, “Is this contract governed by the UCC or the common law?”
The dataset does not include “mixed purpose” contracts which incorporate both the sale of a good and a service. For example, a contract in which Alice sells Bob her bike for $100 and agrees to inflate the tires each week for the first month would be a mixed purpose contract. To determine whether a mixed purpose contract is governed by the UCC or the common law, most jurisdictions apply the “predominant purpose” test under which the predominant purpose of the contract (good or service) determines which body of law applies.
The UCC vs. Common Law task is significant for a number of reasons. First, it provides a measure of an LLM’s legal reasoning ability relative to a human lawyer (who would almost certainly score a 100% on the task). Second, it demonstrates an LLM’s ability to determine the subject matter of a legal text, which has implications for the use of LLMs for legal tasks far beyond contract classification. Third, this task could prove useful in the context of contract lifecycle management (CLM) in which a CLM software product could automatically sort contracts by subject matter for review purposes. Fourth, while the sample contracts in the dataset were simple and easily identifiable as either UCC or common law contracts, real-world mixed purpose contracts can be difficult to classify and sometimes generate costly litigation. This task could be used as a starting point for developing a more fine-tuned task that can classify mixed purpose contracts.
Appendix G Full results
HuggingFace links for the studied open-source models in Section 5.2 can be found below.
G.2 Prompts
Prompts for all LegalBench experiments are available on the Github repository. For experiments reported in Section 5.2, Table 53 provides the number of in-context demonstrations used.
G.3 Results
We provide results for each LLM on each of the tasks. Models are divided into four groups based on type: commercial models, 13B models, 7B models, and 3B models. Model names are abbreviated to the family name to ensure well-formed tables.