SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore

Sewon Min, Suchin Gururangan, Eric Wallace, Weijia Shi, Hannaneh Hajishirzi, Noah A. Smith, Luke Zettlemoyer

Introduction

Large language models (LMs) are under widespread legal scrutiny, in large part because they are trained on copyrighted content, which may infringe on the rights of data producers (Metz, 2022; Vincent, 2023; J.L. et al. v. Alphabet Inc., 2023; Brittain, 2023). At the heart of this discussion is the inherent tradeoff between legal risk and model performance. Training only on data sources such as public domain, non-copyrightable or otherwise permissively licensed data significantly degrades performance (as we show in §3). This limitation arises from the scarcity of permissive data and its narrow specificity to sources such as copyright-expired books, government documents, and permissively licensed code, which are largely different from common LM corpora that cover more diverse domains (Raffel et al., 2020; Gao et al., 2020; Together, 2023).

In this paper, we demonstrate it is possible to improve the risk-performance tradeoff by segregating training data into two distinct parts of the model: parametric and nonparametric (Figure 1). We learn LM parameters on low-risk data (i.e., data under the most permissive licenses), and then use high-risk data (i.e., data under copyright, restrictive licenses, or unknown licenses) in an inference-time-only nonparametric component (called a datastore). With nonparametric datastores, we can retrieve high-risk data to improve model predictions without training on it. The datastore can be easily updated at any time, and allows creators to remove their data from the model entirely, at the level of individual examples. This approach also attributes model predictions at the sentence-level, enabling credit assignment to data owners. These new capabilities enable better alignment of the model with various data-use regulations, e.g., the fair use doctrine in the United States (Henderson et al., 2023) and the GDPR in the European Union (Zhang et al., 2023), as detailed in §2. This is in contrast to parametric models, where removing high-risk data is infeasible after training (Bourtoule et al., 2020; Carlini et al., 2021) and data attribution at scale is difficult (Zhang et al., 2021; Han et al., 2023).

We introduce \modelname, a new nonparametric language model that follows our proposal (§4). The parametric component in \modelname is trained on a new pretraining corpus, the Open License Corpus (OLC, §3), which we curate to include data under three types of permissive licenses, from public domain to Creative Commons. OLC is diverse but has a domain distribution that is very different from typical pre-training corpora; it is dominated by code and government text. This leads to a new challenge of generalizing a model trained on highly specific domains, which we call extreme domain generalization. We train three 1.3B-parameter LMs on varying subsets of OLC, and then construct a test-time datastore that can include high-risk data, employing a retrieval method to make use of the datastore’s contents during inference. We compare two widely studied retrieval methods: a nearest-neighbors approach (kkNN-LM) that uses a nonparametric next-token prediction function (Khandelwal et al., 2020) and a retrieval-in-context approach (RIC-LM) that retrieves text blocks and feeds them to the parametric LM in context (Shi et al., 2023; Ram et al., 2023).

We evaluate \modelname in language modeling perplexity on 14 different domains, covering both in-domain and out-of-domain data with respect to OLC (§5). These domains highlight specific legal risks, e.g., copyrighted materials such as books, news and user reviews, or private data such as emails and clinical notes. We compare \modelname to Pythia (Biderman et al., 2023), a parametric LM with a similar parameter count but trained mostly on high-risk data (Gao et al., 2020).The Pile contains a large amount of copyrighted or restrictively licensed data, e.g., most content in its Books3, ArXiv, Github, OpenWebText, YoutubeSubtitles, and Common Crawl subsets. We first show that parametric-only \modelname is competitive on domains covered by OLC but falls short out-of-domain, confirming the challenge of extreme domain generalization. However, adding an inference-time datastore to \modelname effectively addresses this challenge. Comparing the two methods of retrieving over this datastore, we find that while both kkNN-LM and RIC-LM significantly improve out-of-domain performance, the former generalizes better than the latter, allowing \modelname to reduce the gap with the Pythia baseline by 90% on average across all domains. Further analysis attributes these improvements to two factors: (1) kkNN-LM strongly benefits from scaling the datastore and (2) the nonparametric next-token prediction in kkNN-LM is robust to domain shift. Altogether, our study suggests that in the few domains where \modelname has not yet matched Pythia performance levels, the remaining gaps can likely be closed by scaling the datastore size and further enhancing the nonparametric model.

Background & Related Work

State-of-the-art LMs are trained on vast text corpora that consist of billions or even trillions of tokens (Brown et al., 2020; Raffel et al., 2020; Gao et al., 2020; Together, 2023). These training sets are built by combining (1) manually selected sources such as Wikipedia, book collections, and GitHub and (2) web pages collected through web-crawling services such as Common Crawl. Most LM training efforts ignore copyright and intellectual property regulations that apply to these texts. For example, sources such as GitHub repositories and book collections typically contain text with highly restrictive licenses (Bandy & Vincent, 2021).

The legality of training LMs this way has become a subject of intense debate, with numerous lawsuits being filed in the United States, United Kingdom, and beyond (Gershgorn, 2021; Metz, 2022; Vincent, 2023; De Vynck, 2023; Silverman et al. v. Meta Platforms, Inc., 2023; J.L. et al. v. Alphabet Inc., 2023; Silverman et al. v. OpenAI, Inc., 2023; Tremblay et al. v. OpenAI, 2023). While the outcome of the lawsuits is uncertain, it is likely that copyright issues will continue to be a major factor in future LMs, especially since each country has its own data regulations. For example,

In the United States, the fair use doctrine allows the public to use copyrighted data in certain cases, even without a license (Henderson et al., 2023). Deciding whether or not a model’s use of copyrighted work constitutes fair use involves multiple dimensions, including whether the trained model is intended for commercial use, whether or not the work is factual or creative, the amount of the copyright content used, and the value of the copyrighted work. There are claims that using parametric language models for generative use-cases does not constitute fair use, because the technology may output the copyrighted text verbatim (Lemley & Casey, 2020), which also has been shown empirically (Carlini et al., 2021; 2023; Kandpal et al., 2022; Chang et al., 2023). This is in contrast to transformative technologies, such as classifiers, which may use the copyrighted text but do not directly generate content, which the fair use doctrine favors. We refer readers to Henderson et al. (2023) for a more comprehensive discussion.

The General Data Protection Regulation (GDPR) is a comprehensive data protection and privacy law in the European Union (EU). It grants individuals more control over their data by regulating organizations and businesses. The obligations include (1) obtaining consent from users before processing their data, (2) providing transparency about data processing, (3) ensuring data security, and (4) allowing individuals to access, correct, and erase their data. GDPR has global impact, as many international companies handle EU citizens’ data. While it is under debate how GDPR is applied to training language models, compliance with GDPR is expensive (e.g., requiring retraining for every data correction or erasure). See Zhang et al. (2023) for more discussion on challenges for compliance with the GDPR’s Right to Erasure (and the Right to be Forgotten in general).

The goal of our work is not to weigh in on legal discussions; instead, we study the feasibility of developing technologies that explicitly manage legal risk. In particular, our technique places all copyrighted data in a nonparametric datastore. While the data is still used in service of a generative model, restricting copyrighted data in a datastore and providing instance-level attribution and data opt-out can increase the likelihood of a successful fair use defense (Henderson et al., 2022). Our model on its own does not entirely remove legal risk. Rather, it provides functionalities that, when used appropriately, lower legal risk and strengthen a fair use defense. See §6 for a discussion. Moreover, GDPR’s requirement regarding user data access, correction, and erasure aligns well with the capabilities of the datastore. Attribution and opt-out are fundamental features of our model (§4.2). This is in contrast to other techniques like post-hoc training data attribution (Koh & Liang, 2017; Han et al., 2023) and the removal of the effect of particular training examples from parameters (Cao & Yang, 2015; Jang et al., 2023b), which lack inherent guarantees and are hard to scale.

The most straightforward approach to avoid copyright infringement is to filter training data to only include permissive licenses. This has been done in prior work, primarily for code-based datasets (e.g., Kocetkov et al., 2023; Fried et al., 2023; Together, 2023) and scientific text (e.g., Soldaini & Lo, 2023). Extending a similar approach to a wider range of domains remains unclear, because permissive data is extremely scarce in most domains, e.g., books and news. For the same reason, Henderson et al. (2023) has suggested that restricting the training data to public domain or otherwise permissively licensed data may be impractical. In this work, we show that there is in fact a large number of tokens from data sources with permissive licenses, but the key challenge instead arises from the highly skewed domain distribution. See §6 for other copyright mitigation strategies that are more technical in nature.

Building the Open License Corpus: A Permissively-Licensed Pre-training Corpus

Our study focuses on addressing the legal risk of copyright violation in language models by separating low-risk data sources (i.e., those in the public domain or under permissive licenses) from high-risk ones (i.e., those with unknown licenses or under copyright). We introduce the Open License Corpus (OLC), a large collection of permissive textual datasets across multiple domains with a taxonomy of data licenses that delineate their permissiveness (§3.1). We group the data into three levels of legal permissiveness (§3.2) and conduct a thorough analysis (§3.3). This curated data is then used to train model parameters (§4) and highlights the challenge of extreme domain generalization due to its skewed domain distribution.

The license taxonomy and categorization of texts that we present is by no means perfect, and OLC should not be considered a universally safe-to-use dataset. The license associated with a document may be time- and country-dependent, e.g., Gutenberg books (Project Gutenberg, ) are public domain in the United States, but some of them may still have copyright attached outside of the United States. Moreover, other legal constraints (e.g., the Digital Millenium Copyright Act)https://www.copyright.gov/dmca/ may prohibit the use of a data source despite a permissive data license. Finally, we do not explicitly filter out personally identifiable information from the corpus, so it is possible that certain subsets still pose privacy risks despite being permissively licensed. We encourage users of OLC to consult a legal professional on the suitability of each data source for their application.

1 Taxonomy of Data Licenses

As discussed in §2, determining what data one is permitted to use from a copyright perspective is an ongoing topic of debate, and is context- and country-dependent (Henderson et al., 2023). In this paper, we take a conservative approach where we train models using only text with the most permissible licenses, thus enabling widespread downstream use. Concretely, we focus on four broad categories:

Public domain ( \textscpd‾‾\overline{\underline{\textsc{{pd}}}}) text has no restrictions. This includes texts whose intellectual property rights have expired (e.g., the works of William Shakespeare) or been expressly waived by the creator (e.g., CC0-licensed scientific papers).

Permissively licensed software ( \textscsw‾‾\overline{\underline{\textsc{{sw}}}}) including MIT, Apache, and BSD software are quite permissive to use. Unlike public domain text, these licenses typically carry some basic stipulations such as requiring one to include a copy of the original license (although, it is debatable whether it is still required when the associated text is used as data or treated as a software). The code is otherwise free to use, and code is generally well protected by fair use clauses (Lemley & Casey, 2020).

Attribution licenses ( \textscby‾‾\overline{\underline{\textsc{{by}}}}) such as Creative Commons Attribution (CC-BY) are free to use as long as \saycredit is given to the creator. For example, if a journalist writes a new article that cites information from Wikipedia (a CC-BY source), then they must provide a form of citation, link, or attribution back to the original source. In the context of machine learning, it is not clear what an attribution would constitute. For example, under one interpretation, every LM generation should include a complete list of sources that contributed highly to it (Henderson et al., 2023). In this paper, we take a conservative approach and do not include \textscby‾‾\overline{\underline{\textsc{{by}}}} data in the main experiments, but still include the \textscby‾‾\overline{\underline{\textsc{{by}}}} data for future use as well as for ablations, since \textscby‾‾\overline{\underline{\textsc{{by}}}} data is generally considered quite permissive.

All other data that is not in one of the above three categories is assumed to be non-permissive. This includes: any text that is explicitly protected by copyright or licenses that are non-commercial (e.g., CC-NC), any software without clear MIT, BSD, or Apache licenses, and any generic web-crawled data where the license or copyright information may be unclear.

In §4.3, we train the models on varying subsets of licenses—from \textscpd‾‾\overline{\underline{\textsc{{pd}}}} and \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} to \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscby‾‾\overline{\underline{\textsc{{by}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}}—to accommodate different risk tolerances.

2 Building the Open License Corpus

Based on this taxonomy of licenses, OLC is a 228B token corpus of \textscpd‾‾\overline{\underline{\textsc{{pd}}}}, \textscsw‾‾\overline{\underline{\textsc{{sw}}}}, and \textscby‾‾\overline{\underline{\textsc{{by}}}} data. OLC consists of 17 manually-selected sources of primarily English text that are under permissive licenses,We include the data in only when the license information is clearly stated as part of metadata. While we tried our best to collect the data for OLC, it is possible we missed potential sources, as it relies on manual efforts; future work can study collecting more permissive text at scale, as discussed in §6. as summarized in Table 1.

The text generally falls into eight different domains:

\textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscby‾‾\overline{\underline{\textsc{{by}}}} Legal: We curate legal text from the Pile of Law (Henderson et al., 2022), an amalgation of 31 different sources of text related to civil court cases, patents, and other legal and governmental works, either licensed as public domain or CC-BY. We also gather public domain text from the Case Law Access Project (Caselaw Access Project, ), which covers over 6.5 million decisions published by state and federal courts throughout U.S. history.

\textscsw‾‾\overline{\underline{\textsc{{sw}}}} Code: We use the Github subset of the RedPajama dataset (Together, 2023), which contains code from Github repositories with three permissive software licenses: MIT, Apache, and BSD.

\textscsw‾‾\overline{\underline{\textsc{{sw}}}} \textscby‾‾\overline{\underline{\textsc{{by}}}} Conversation: We source conversational text under permissive software licenses from the HackerNews (MIT license) and the Ubuntu IRC (Apache license) subsets of the Pile (Gao et al., 2020). We also use the Stackexchange subset of the RedPajama dataset (Together, 2023) and a Stackoverflow corpus from Kaggle,https://www.kaggle.com/datasets/stackoverflow/stackoverflow both under the CC-BY-SA license.

\textscsw‾‾\overline{\underline{\textsc{{sw}}}} Math: We source mathematical text from the Deepmind Mathematics (Saxton et al., 2019) and the AMPS (Hendrycks et al., 2021) datasets, both of which are under the Apache license.

\textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscby‾‾\overline{\underline{\textsc{{by}}}} Science: We source scientific text from ArXiv abstracts that are in the public domain (ArXiv, 2023). We also collect full-text articles from the Semantic Scholar Research Corpus (Lo et al., 2020, S2ORC), either licensed as public domain or CC-BY.

\textscpd‾‾\overline{\underline{\textsc{{pd}}}} Books: We source books from the Gutenberg corpus (Project Gutenberg, ), which are copyright-expired books that are in the public domain.

\textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscby‾‾\overline{\underline{\textsc{{by}}}} News: We collect public domain news text from the English subset of the MOT corpus (Palen-Michel et al., 2022). We also collect text from Wikinews, which is under CC BY-SA.

\textscby‾‾\overline{\underline{\textsc{{by}}}} Encyclopedic: Finally, we include a large set of Wikipedia from the subset included in RedPajama (Together, 2023). We follow RedPajama in using Wikipedia snapshots from 20 languages even though the model primarily focuses on English.

Following Kandpal et al. (2022); Lee et al. (2022), we deduplicate text using Groeneveld (2023), a document-level filter that considers nn-gram overlap. We first deduplicate within each domain to remove redundant documents from similar sources (e.g. Case Law and the Pile of Law), and then perform deduplication against the validation and test datasets of the Pile to avoid test leakage.

3 Analysis of OLC

In Table 2, we compare the distribution of domains in OLC to that of the Pile (Gao et al., 2020), a popular pretraining corpus that includes data under copyright restrictions (e.g., Books, web crawl).This comparison also dovetails with our experiments in §5, where we compare \modelname to Pythia, a model trained on the Pile. These statistics convey a number of research challenges when working with OLC. First, while we tried our best to collect public domain or permissively-licensed data, the size of OLC is still 31% smaller than the Pile. In addition, while the majority of the Pile is sourced from scientific text, web crawl, and books, OLC is dominated by code, scientific text, and legal text. This highlights that models designed for use outside these specific domains will likely struggle and may require special techniques for extreme domain generalization.

To analyze this further, we perform an nn-gram based analysis of OLC domains against the validation data of the Pile, to better understand the domain shifts. For each validation domain, we examine the maximum nn-gram overlap across all OLC domains (see §B for more details). OLC domains have substantially less overlap with the validation data as compared to the Pile training domains: on average, the overlap between OLC domains and the validation domains is just 17%±\pm9%, versus 28%±\pm14% for the Pile training data. However, we find a large variance in overlap statistics across domains in OLC; we display the full matrix of nn-gram overlap in §B. These results provide further evidence that models trained on OLC must handle larger domain shifts at test time than models trained on the Pile. Later, we show that these nn-gram overlap statistics correlate strongly with language modeling performance (§5.1).

\modelname

We introduce \modelname, which combines an LM trained on permissive data with a nonparametric datastore based on less restricted data. Our goal with \modelname is to build an LM—i.e., a model that takes a prefix of text xx and outputs a next-word probability distribution over the vocabulary P(y∣x)P(y\mid x)—but to do so in a legally safe way. We first describe the general methodology from prior work (§4.1–4.2) and then how we build \modelname upon them by placing low-risk data and high-risk data to model parameters and a nonparametric datastore, respectively (§4.3). Implementation details are provided in §4.4.

For the parametric component of \modelname, we use a standard, dense, decoder-only transformer LM (Vaswani et al., 2017) using the LLaMA architecture (Touvron et al., 2023). This model uses a fixed set of parameters at both training and inference time.

2 The Nonparametric Component

We experiment with two widely-used retrieval methods for the nonparametric component (Figure 2): the kk-nearest neighbors LM (kkNN-LM; Khandelwal et al., 2020) and the retrieval-in-context approach (RIC-LM; Shi et al., 2023; Ram et al., 2023). Each approach constructs a datastore from the raw text data offline, and then uses it on-the-fly at inference time.

A kkNN-LM (Khandelwal et al., 2020) interpolates the next-token probability distribution from a parametric LM with a nonparametric distribution based on every token that is stored in a datastore. Given a text dataset consisting of NN tokens c1...cNc_{1}...c_{N}, a datastore is built by creating a key-value pair for every token cic_{i} (1≤i≤N1\leq i\leq N). Specifically, a value is cic_{i} and a key kik_{i} is ...ci−1...c_{i-1}, a prefix preceding cic_{i}. At test time, given an input prefix xx, the nonparametric distribution is computed by:

Future work can improve kkNN-LM, e.g., by training the model to output a nonparametric distribution (Zhong et al., 2022; Lan et al., 2023; Min et al., 2023), by having a vocabulary-specific λ\lambda (Huang et al., 2023b), or by modeling λ\lambda as a function of the input xx (He et al., 2021; Drozdov et al., 2022).

Future work can improve RIC-LM, e.g., by using multiple text blocks through ensembling (Shi et al., 2023) or reranking (Ram et al., 2023), by tuning the retrieval system (Shi et al., 2023), or by training the LM to use retrieved blocks in context (Guu et al., 2020; Izacard et al., 2022).

The key difference between kkNN-LM and RIC-LM lies in how the nonparametric component influences the output. In kkNN-LM, it directly impacts the output distribution, while in RIC-LM, it indirectly influences the output by affecting the input to the parametric model. kkNN-LM intuitively benefits more from a datastore as it provides direct influence to the output and relies less on the parametric component. Nonetheless, RIC-LM interacts more easily with a parametric model (i.e., it is applicable to a black-box LM) and offers better speed and memory efficiency (explored in Appendix B).

Empirical comparisons between kNN-LM and RIC-LM have been largely unexplored; in fact, we are unaware of such work. In our experiments (§5.2), we present a series of such comparisons, with varying sizes of the datastore, and with and without distribution shift.

Since elements in the datastore that contribute to the model prediction are transparent, both kkNN-LM and RIC-LM offer inherent attributions. Moreover, data removed from the datastore is guaranteed not to contribute to any model predictions, allowing data owners to remove their data at the level of individual examples. Both are unique characteristics of nonparametric language models. While prior work studies post-hoc attribution to the data used for training model parameters (Koh & Liang, 2017; Han et al., 2023) and removing the effect of specific training examples from parameteric models (Cao & Yang, 2015; Jang et al., 2023b), they are arguably not fundamental due to lack of inherent guarantees, and are difficult to scale.

3 Building \modelname

is is built upon the general methodology of kkNN-LM and RIC-LM. However, unlike prior work that uses the same data for learning model parameters and a nonparametric datastore, \modelname uses distinct datasets for these two components.

The key idea behind \modelname is to use low-risk data to estimate model parameters, and to use high-risk data only in a nonparametric datastore. This is based on the motivation that model parameters should be learned conservatively, since training data is difficult to remove or trace after model training is completed. In contrast, a nonparametric datastore offers greater flexibility, as it can be easily updated, grown, or filtered, supports data opt-out at the level of individual examples, and provides attributions for free to every model prediction. These functions enable adherence to data-use regulations (§2).

We train each of our LMs on one of the three datasets of OLC: \textscpd‾‾\overline{\underline{\textsc{{pd}}}} data, \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} data, and \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} \textscby‾‾\overline{\underline{\textsc{{by}}}} data. Each of the resulting models constitutes a different level of possible copyright infringement risk.

We assume in-distribution data for each test domain is available at inference time, and construct a datastore for each domain (details in §4.4). Future work may investigate building a single datastore that includes all domains. These test-time datasets can be either in-domain or out-of-domain with respect to the data used to train model parameters.

4 Implementation Details

We use 1.3B-parameter transformer LMs based on the LLaMA architecture (Touvron et al., 2023) as implemented in OpenLM.https://github.com/mlfoundations/openlm Each model is trained with 128 A100 GPUs across 16 nodes. Following Muennighoff et al. (2023), we train for multiple epochs in each dataset and perform early stopping. We train our \textscpd‾‾\overline{\underline{\textsc{{pd}}}}, \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} and \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} \textscby‾‾\overline{\underline{\textsc{{by}}}} models for 60B, 250B, and 350B tokens in total, respectively. More details are provided in Appendix A.

Since the distribution of OLC is highly skewed (§3.3), we perform a simple upweighting scheme where we upsample all data that accounts for less than 5% by a factor of 3×\times, which we found to work well after a sweep of different settings. More sophisticated domain weighting strategies (Xie et al., 2023) are of interest but beyond the scope of this work.

We benchmark our models using language modeling perplexity on 14 domains that represent both in-domain and out-of-domain data with respect to different levels of OLC. This includes: public-domain legal documents from the FreeLaw Project subset of the the Pile (Gao et al., 2020), a held-out collection of books from the Gutenberg collection (Project Gutenberg, ), conversational text from the Hacker News subset of the Pile, held-out code files from the Github subset of the Pile (most of which are non-permissive licensed), scientific text of NIH Grant abstracts that are taken from the NIH ExPorter subset of the PILE, philosophy papers taken from the PhilPapers of the PILE, held-out English Wikipedia articles from the PILE, news articles from CC-News (Mackenzie et al., 2020), books from BookCorpus2 which is an expanded version of Zhu et al. (2015), books from Books3 by Presser (2020), random web-crawled pages from OpenWebText2 (Gokaslan & Cohen, 2019; Gao et al., 2020), emails from the Enron Emails corpus (Klimt & Yang, 2004), Amazon product reviews from He & McAuley (2016), and finally clinical notes from MIMIC-III (Johnson et al., 2016) with personal identifiable information (PII) masked out. Our choice of domains highlights legal risks discussed in the earlier sections, e.g., CC-News, BookCorpus2, Books3 and Amazon reviews are mostly copyrighted, Github is mostly not permissively licensed, Kocetkov et al. (2023) estimates about 13% of the Github data is under MIT, Apache, and BSD. and Enron Emails and MIMIC-III include private text. We merge all text into one stream of text and split them into batches with a maximum sequence length of 1,024 and a sliding window of 512, a setup that is standard in prior language modeling literature (Baevski & Auli, 2019; Khandelwal et al., 2020). For MIMIC-III, which includes masked personally-identifiable information (PII), we filter out notes where more than 50% of tokens correspond to PII, and then exclude tokens corresponding to PII when computing perplexity.

We construct an in-domain datastore for each test domain based on their training data. For datasets from the PILE, we consider 10% of the training data. For kkNN-LM, each datastore consists of up to 1 billion hh-dimensional vectors (h=h=2,048). We build an index for fast nearest neighbor search using FAISS (Johnson et al., 2019). For RIC-LM, each datastore consists of text blocks with a length of 1,024 and a sliding window of 512. We use BM25 from Pyserini (Lin et al., 2021). Appendix B report ablations on different implementations of RIC-LM besides the method in §4.2. More details, statistics and hyperparameter values for the datastores are reported in §A.

Experimental Results

We first evaluate the parametric-only component of \modelname trained on the Open License Corpus (§5.1), and then show the effect of adding a datastore that may contain high-risk text (§5.2). For all experiments, we use the 1.4B Pythia model (Biderman et al., 2023) as a baseline because it is trained with a similar amount of compute (data size and model parameters), but is trained on mostly high-risk data.We use the model checkpoint from https://huggingface.co/EleutherAI/pythia-1.4b-deduped-v0.

Table 3 reports performance of our 1.3B base LMs trained on varying levels of permissively-licensed data— \textscpd‾‾\overline{\underline{\textsc{{pd}}}}, \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}}, and \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} \textscby‾‾\overline{\underline{\textsc{{by}}}}—as well as Pythia. Overall, our LMs are competitive with Pythia despite using permissive data only. They are roughly equal quality on in-domain data, e.g., FreeLaw and Gutenberg, HackerNews in the case of \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} and \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} \textscby‾‾\overline{\underline{\textsc{{by}}}}, and Wikipedia in the case of \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} \textscby‾‾\overline{\underline{\textsc{{by}}}}. Models trained on \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} and \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} \textscby‾‾\overline{\underline{\textsc{{by}}}} are also close to Pythia on Github, likely because the permissively-licensed code data included in \textscsw‾‾\overline{\underline{\textsc{{sw}}}} has a distribution that is sufficiently close to the distribution of the all Github code. The largest gaps occur on data that is in-domain for Pythia but out-of-domain for our model, e.g., news, books, OpenWebText, and emails, and Wikipedia in the case of models besides \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} \textscby‾‾\overline{\underline{\textsc{{by}}}}. This illustrates the extreme domain generalization challenge that is present when training on only permissive data, as we hint in §3.3.

The similarity of an evaluation domain to a domain of the OLC strongly correlates with the performance gaps between \modelname and Pythia. To show this, we compute the Pearson correlation between 1) the maximum nn-gram overlap between an OLC domain and the Pile validation domains (from §3.3) and 2) the perplexity difference between the Pythia model and our \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} model, normalized by the performance of the \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} model. We find a strong negative correlation between these metrics (rr=-0.72, p<p< 0.005), indeed indicating that the more dissimilar an evaluation domain is from the OLC domains, the better Pythia does relative to \modelname (see §B for a scatter plot).

More ablations, including the effect of upsampling low-resource data, and the effect of including and excluding explicit source code, are provided in §B.

2 Results: Adding the Nonparametric Component

Since building legally permissive LMs poses a challenge of extreme domain generalization, our next question is whether using an in-domain, nonparametric datastore can reduce the gap. We explore this question with our parametric LM trained on the \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} subset of OLC evaluated on a subset of 8 out-of-domain datasets to the parametric model: Github, NIH ExPorter, Wikipedia, CC News, Books3, Enron Emails, Amazon, and MIMIC-III.

Table 4 shows adding the datastore with either kkNN-LM- or RIC-LM-based retrieval improves performance over just using the parameteric component on all domains, but kkNN-LM is more effective than RIC-LM. In most domains, kkNN-LM reduces the gap between \modelname and Pythia by more than 50% (on NIH ExPorter, Wikipedia, Amazon) or even outperforms Pythia (on Github, CC News, Enron Emails, MIMIC-III). Books3 is the domain with the least benefit, on which kkNN-LM still reduces the gap by 28%.

Figure 3 demonstrates that both kkNN-LM and RIC-LM-based retrieval consistently improves performance as the datastore size increases, with a strong log-linear trend. However, kkNN-LM improves performance more rapidly than RIC-LM does, consistently over all datasets. Extrapolating the trend suggests that, on the domains that \modelname has not outperformed Pythia yet, scaling the datastore even further (with kkNN-LM retrieval) may enable it to match Pythia.

Our next question is why kkNN-LM is better than RIC-LM—is it (a) because kkNN-LM is better than RIC-LM in general, or (b) because kkNN-LM generalizes out-of-domain better than RIC-LM does? Our further analysis in §B (Figure 5) reveals that it is both. With Pythia, where the test data is in-domain, while both kkNN-LM and RIC-LM improve performance upon the parametric-only model, kkNN-LM is overall better and scales better than RIC-LM, supporting (a). Both kkNN-LM and RIC-LM improve performance more rapidly with \modelname (where the test data is out-of-domain) than with Pythia, but this trend is much clearer with kkNN-LM, supporting (b).

In summary, our analysis highlights two promising directions to further reduce the gap:

Scaling the datastore beyond 1 billion tokens, e.g., at the scale of trillions of tokens as in Borgeaud et al. (2022), as demonstrated by Figure 3.

Improving the robustness of the model by improving nonparametric techniques or designing a model that only uses a nonparametric distribution (Min et al., 2023), as demonstrated by Figure 4.

Table 14 in Appendix B provides a comparison of the runtime speed of the parametric LM, RIC-LM, and kkNN-LM. There is a strong tradeoff between performance and speed: both RIC-LM and kkNN-LM are considerably slower than the parametric LM, and a larger datastore and more accurate nearest-neighbor search leads to better performance and slower inference. While the speed is heavily influenced by the hardware used for benchmarking and thus it is difficult to precisely quantify how much faster one method is compared to the other, this suggests that improving the runtime efficiency of nonparametric approaches is an important area of future work.

3 Examples of Data Attribution and Opt-Out

As discussed in §2, the design of \modelname can better align with various data-use regulations by providing mechanisms for data attribution during inference and for data owners to remove their data from the model at any time. This section show examples of such capabilities.

To showcase the impact of opt-out on model performance, we conduct experiments with J.K. Rowling’s Harry Potter series. We first identify all seven Harry Potter books from the Books3 corpus of the Pile. For each book, we calculate the perplexity of \modelname using two 1.024B token datastores on Books3, but one including the remaining six Harry Potter books and the other excluding any Harry Potter books. This experiment is to see whether excluding Harry Potter books from the former datastore can reduce the likelihood of generating the leave-out Harry Potter book.

Table 5 shows the results. \modelname with Harry Potter books in the datastore effectively improves perplexity over all seven books, closing the gap between the \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} model and Pythia. However, when the Harry Potter books are removed from the datastore, the perplexity gets worse, approaching that of the parametric-only LM. This illustrates that eliminating the effect of the Harry Potter books from the model substantially reduces the likelihood of generating the leave-out book.

To show the attribution feature of our model, Table 6 provides qualitative examples on the top-11 context retrieved by \modelname. The model is able to assign a high probability to the ground truth token by retrieving highly relevant context. It achieves this by leveraging the unique characteristics of the text within the datastore, such as recognizing that Azkaban refers to the prison and green light is associated with the Killing Curse in the Harry Potter books.

More qualitative examples on Github, news and emails are illustrated in Table 15 in Appendix B. They highlight that a nonparametric approach addresses specific legal risks that we have discussed earlier, e.g., it offers per-token attribution for free, and can provide a copyright notice when part of copyrighted text is being used for the probability distribution.

Discussion & Future Work

Our work suggests that it is possible to improve the tradeoff between legal risk and model performance when training LMs. Our approach provides new options for model designers to mitigate the legal risk of LMs, and empowers stakeholders to have more control over the data that drives these systems. We point out a number of rich areas for future work, beyond what was mentioned throughout the paper:

does not completely eliminate legal risk. Instead, it provides users more control over the model’s generated content and functionalities to better align with legal regulations. For instance, \modelname does not remove the need for obtaining permission to use copyrighted content in a datastore when providing attribution is not sufficient, but its opt-out capabilities can strengthen fair use defense. Moreover, \modelname does not prevent copying copyright content from a datastore, but it offers a way to prevent generating sensitive text (Huang et al., 2023a) or prevent copying the content verbatim. These functionalities increase the likelihood of a successful fair use defense if used appropriately.

Furthermore, while \modelname mitigates copyright and privacy risks, it may exacerbate certain fairness issues, like toxicity towards marginalized groups and racial biases, especially due to the prevalence of older copyright-expired books in the training data. Exploring the balance between legal risk mitigation and fairness is an important future direction.

Finally, our study relies on explicit metadata to identify licenses, which may lead to underestimates of the amount and diversity of permissively licensed text actually available on the web. Future research may investigate inferring data licenses from documents in web crawl at scale, which may be an effective way to build more heterogeneous, permissively licensed corpora.

introduces the possibility for data owners to set different levels of permissivity for learning parameters and for including in a nonparametric datastore. A data owner might choose to be more permissive about including data in the datastore due to its ease of removal, ensuring that the excluded data has no influence on model predictions anymore, and its ability to provide per-prediction attribution. Moreover, we envision that \modelname could provide a path forward for data owners to get properly credited (or be paid directly) every time their data in a datastore contributes to a prediction. This is orthogonal to recent work that circumvented copyright issues by licensing out training data from data creators (Yu et al., 2023).

It is critical to continue to develop new techniques that use copyrighted data while protecting the rights of data owners and subjects. In addition to nonparametric approaches, there are many other ways to achieve these goals. First, one could train LMs on copyrighted content but filter and guide their outputs towards text that is non-infringing (Henderson et al., 2023). Second, training models with differential privacy (Dwork et al., 2006; Abadi et al., 2016) may prevent them from memorizing individual details of copyright data. Finally, one could provide attributions for standard base LMs using post-hoc attribution methods, e.g., influence functions (Koh & Liang, 2017), rather than switching the model class to a retrieval-based model. All of these methods are complementary and orthogonal to our proposed approach.

Our work is closely related to recent studies on modular LMs, which have specialized parameters (or experts) trained on different domains (Gururangan et al., 2022; Li et al., 2022; Gururangan et al., 2023), languages (Pfeiffer et al., 2020; 2022), or tasks (Chen et al., 2022b; Jang et al., 2023a). Our work extends modular LMs to include nonparametric datastores, and focuses on specializing different parts of the model to low- and high-risk subsets of the training data. Legal risks may also be mitigated with a collection of parametric expert models that are specialized to low- and high-risk data. Future work may explore this possibility as well as the usefulness of combining a nonparametric datastore with parametric experts.

While this work focuses on text-only models, similar methods to ours could apply to other domains and modalities. For instance, it might be possible to build permissive text-to-image generative models (Rombach et al., 2022) using compartmentalized public domain pre-training and retrieval-augmentation (Chen et al., 2022a; Golatkar et al., 2023). We believe such approaches are especially promising because there are many sources of public domain data in other modalities, e.g., images, speech, video, and more.

Conclusion

We introduce \modelname, a language model that mitigates legal risk by learning parameters only on low-risk, permissively-licensed data (Open License Corpus), and using an unrestricted nonparametric datastore during inference. Our approach allows the model designer to use high-risk data without training on it, supports sentence-level data attribution, and enables data produces to opt-out from the model by removing content from the datastore. Experiments on language modeling perplexity show that parametric-only \modelname is competitive on domains covered by Open License Corpus, but falls short out-of-domain when solely using the parametric component of the model, highlighting the challenge of extreme domain generalization. We then show that adding a nonparametric datastore to \modelname (with kkNN-LM retrieval) successfully addresses this challenge, significantly reducing the gap (or even outperforming) the Pythia baseline that is trained unrestrictedly. We show that scaling the datastore size is key to the success of the nonparametric approach, and that the encoder for a nonparametric distribution is significantly more robust to distribution shift than the parametric component. Our results point to a number of exciting future research directions to develop AI systems with mitigated legal risk.

We thank Peter Henderson for discussing the legality of LMs, and Kyle Lo for feedback on our dataset and license taxonomy. We thank Mitchell Wortsman for help with setting up compute and model training. We thank Tatsunori Hashimoto, Peter Henderson, Nikhil Kandpal, Pang Wei Koh, Kyle Lo, Fatemeh Mireshghallah, Sewoong Oh and Rulin Shao for valuable feedback on the project and the paper. We thank Matthijs Douze, Gergely Szilvasy, and Maria Lomeli for answering questions about FAISS, and Dirk Groeneveld for giving early access to the deduplication script. Sewon Min is supported by the J.P. Morgan Ph.D. Fellowship. Suchin Gururangan is supported by the Bloomberg Data Science Ph.D. Fellowship. Eric Wallace is supported by the Apple Scholars in AI/ML Fellowship. We thank Stability AI for providing compute to train the LMs in this work.

References

Appendix A Model details

Table 7 reports the hyperparameters for the parametric component of \modelname. We keep these hyperparameters fixed for all parametric models that we report in this paper. We follow the model architecture of LLaMa (Touvron et al., 2023), and we use the GPT-NeoX-20B tokenizer (Black et al., 2022), with 50432 BPE types. During training, we use 2,048 token sequences that are packed across document boundaries, and we pre-pend a beginning-of-text token to every document. We use weight decay of 0.1, the Adam optimizer with β2=\beta_{2}= 0.95, 2,000 steps of warmup, with a cosine learning rate scheduler. We train for multiple epochs in each dataset, tracking validation perplexity every 10B tokens, and perform early stopping. We train our \textscpd‾‾\overline{\underline{\textsc{{pd}}}}, \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} and \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} \textscby‾‾\overline{\underline{\textsc{{by}}}} models for 60B, 250B, and 350B tokens in total, respectively.

Table 8 reports the datastore statistics for both RIC-LM and kkNN-LM, as well as hyperparameter values for kkNN-LM (λ,k,τ\lambda,k,\tau). Due to the resource constraints, the datastore size is capped to up to 10% of the PILE training data (and to 1024.0M tokens in the case of kkNN-LM), but future work can investigate further scaling the datastore.

Appendix B Additional Experimental Results

Table 9 reports perplexity of the parametric LMs on the validation data that is analogous to Table 3. Table 10 reports perplexity of both parametric and nonparametric LMs on the validation data that is analogous to Table 4. Findings based on the validation data and on the test data are largely consistent.

As described in §4.4, since Open License Corpus has an extremely skewed distribution of domains, we upsample less-representative domains during training. Table 11 (left) compares the models trained on \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} with and without domain upweighting. In-domain datasets that are not upweighted, e.g., FreeLaw, see slight degration in performance. On out-of-doain datasets, there is no significant differences, although the model with upsampling is marginally better (19.6 vs. 19.7 when averaged over 9 out-of-domain datasets). We note that we did not tune the upweighting ratio nor explore alternative upweighting approaches (Xie et al., 2023) due to resource constraints, and leave them for future work.

When using \textscsw‾‾\overline{\underline{\textsc{{sw}}}}, a substantial 59.1%59.1\% of the training data is actual source code. To determine \textscsw‾‾\overline{\underline{\textsc{{sw}}}} provides such large gains, we also run an ablation where we include \textscsw‾‾\overline{\underline{\textsc{{sw}}}} data but exclude all of the actual source code, i.e., we only include Hacker News, Ubuntu IRC, Deepmind Math, and AMPS on top of the \textscpd‾‾\overline{\underline{\textsc{{pd}}}} data. This leaves models trained on 99.6B tokens for OLC ( \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}}) and 40.7B for OLC ( \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}}) excluding source code. Table 11 (right) report results on a subset of the validation domains. Including source code provide significant benefits for certain test datasets, e.g., nearly a 20 point improvement in perplexity on PhilPapers, likely because it significantly increases the size of the training data.

Results are reported in Table 12. The concat-2 and concat-next variants perform poorly, while the ensbl-10 outperforms the basic variant. However, we reached the conclusion that the significant run-time cost (i.e., 20x compared to a parametric LM) does not justify the improvements, and thus, we primarily use the basic variant for the remaining experiments. Future work may involve re-evaluating models using the ensbl-kk approach or enhancing its run-time efficiency.

§5.2 shows that performance of both kkNN-LM and RIC-LM rapidly improves as the datastore size grows, and kkNN-LM improves more rapidly than RIC-LM does. This evaluation is mainly done with \modelname where the test domains are out-of-domain. Does this trend hold when the test domains are in-domain? To answer this question, we examine effect of scaling the datastore with Pythia 1.4B, where all of our test datasets can be considered in-domain.

Figure 5 reports the results: Pythia on the left, \modelname ( \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}}) on the right. Results show that both Pythia and \modelname see consistent improvements from kkNN-LM and RIC-LM as the datastore gets larger, although the slope is larger with \modelname than with Pythia. Again consistent to findings in §5.2, kkNN-LM scales better than RIC-LM does, resulting in kkNN-LM outperforming RIC-LM with a reasonably large datastore in most cases (with an exception of Pythia on Github, where RIC-LM outperforms kkNN-LM with a reasonable size of a datastore).

Prior work (Khandelwal et al., 2020) typically uses approximate nearest neighbor search to find the top kk nearest neighbors, and then computes the exact L2 distance using the original vectors. However, this may be inefficient in disk memory usage and run-time speed, due to needing to store large, original vectors and access them on-disk. We thus explore a few alternatives: (1) quantizing the original vectors to compute the L2 distance (but less aggressively than quantization for the nearest neighbor search index, thus it provides different levels of approximations for search and for L2 distance), or (2) completely dropping the original vectors and relying on approximated L2 distance from the FAISS index with aggressive quantization. Based on Table 13, all approximation methods only marginally affect performance. For the rest of our experiments, we use the most aggressive approximation that completely drops the original embeddings at the cost of about 0.5%0.5\% lose in performance while using <2%<2\% of the memory footprint. Future work may study more accurate and efficient approximation methods.

Table 14 presents the runtime speed of the parametric LM, RIC-LM, and kkNN-LM on the Wikipedia validation set. Speed is reported in tokens per second with a batch size of 1 using a single NVIDIA RTX 6000 GPU.

The results show that the parametric LM is notably faster than both RIC-LM and kkNN-LM, and RIC-LM is faster than kkNN-LM. Speed is slower as the datastore gets larger (for both RIC-LM and kkNN-LM) and the nearest neighbor search gets less accurate (for kkNN-LM; indicated by the number of probe pp). kkNN-LM can eventually match RIC-LM’s speed while surpassing its performance by using a smaller datastore and less accurate search, i.e., when using 102M tokens with p=1p=1.

We note that the machine used for benchmarking speed has a very slow IO speed, leading to an underestimation of both RIC-LM and kkNN-LM’s runtime speed, and the comparison can significantly vary based on the hardware. However, it is still important to note that kkNN-LM is substantially slower than a parametric LM, leaving room for potential future improvements.

Figure 15 provides six qualitative examples on the top-1 context retrieved by \modelname-based kkNN-LM. The model is able to assign a high probability to the ground truth token by retrieving highly relevant context, e.g., given the context (hockey) and the first name of the player, being able to retrieve the last name of the player, given the context (a show and its host), being able to complete the quote. These examples also highlight that a nonparametric approach addresses specific legal risks that we have discussed earlier, e.g., it assigns per-token attribution for free, and can provide a copyright notice when part of copyrighted text is being used for the probability distribution.

Table 16 displays the full matrix of unigram and bi-gram overlap between the Open License Corpus training domains and the Pile validation domains. We sample up to 10M tokens in each data source, remove stopwords, and only consider unigrams and bigrams that appear in at least three documents. We also show a scatterplot that describes the relationship between ngram overlap and the performance gap between \textscpd‾‾\overline{\underline{\textsc{{pd}}}} \textscsw‾‾\overline{\underline{\textsc{{sw}}}} and Pythia in Figure 6.