Privacy Implications of Retrieval-Based Language Models
Yangsibo Huang, Samyak Gupta, Zexuan Zhong, Kai Li, Danqi Chen
Introduction
Retrieval-based language models Khandelwal et al. (2020); Borgeaud et al. (2022); Izacard et al. (2022); Zhong et al. (2022); Min et al. (2022) generate text distributions by referencing both the parameters of the underlying language model and the information retrieved from a datastore of text. Specifically, the retrieval process involves accessing a pre-defined datastore to retrieve a set of tokens or text passages that are most relevant to the prompt provided to the model. These retrieved results are then utilized as additional information when generating the model’s response to the prompt. Retrieval-based language models offer promising prospects in terms of enhancing interpretability, scalability, factual accuracy, and adaptability.
However, in privacy-sensitive applications, utility usually comes at the cost of privacy leakage. Recent work has shown that large language models are prone to memorizing Thakkar et al. (2021); Zhang et al. (2021) specific training datapoints, such as personally identifying or otherwise sensitive information. These sensitive datapoints can subsequently be extracted from a trained model through a variety of techniques Carlini et al. (2019, 2021); Lehman et al. (2021), colloquially known as data extraction attacks. While the threat of training data memorization on model privacy has been studied for parametric language models, there is a lack of evidence regarding the privacy implications of retrieval-based language models, especially how the use of external datastore would impact privacy.
In this work, we present the first study of privacy risks in retrieval-based language models, with a focus on the nearest neighbor language models (NN-LMs) Khandelwal et al. (2020), which have been extensively studied in the literature He et al. (2021); Zhong et al. (2022); Shi et al. (2022a); Xu et al. (2023)Other retrieval-based language models such as RETRO Borgeaud et al. (2022) and Atlas Izacard et al. (2022) have significantly different architectures and the findings from our investigation may not necessarily apply to these models. Investigation of these models will be left as future work.. In particular, we are interested in understanding NN-LMs’ privacy risks in real-world scenarios where they are deployed via an API to the general public (Figure 1). We consider a scenario in which a model creator has a private, domain-specific datastore that improves model performance on domain-specific tasks, but may also contain sensitive information that should not be revealed. In such a scenario, the model creator must find a balance between utilizing their private dataset to enhance model performance and protecting sensitive information. Note that the development of NN-LMs intended solely for personal use (e.g., constructing a NN-LM email autocompleter by combining a public LM with a private email datastore) falls outside the scope of our study because it does not involve any attack channels that could be exploited by potential attackers.
We begin our investigation by examining a situation where the creator of the model only adds private data to the retrieval datastore during inference, as suggested by Borgeaud et al. (2022). Our findings indicate that while this approach enhances utility, it introduces an elevated privacy risk to the private data compared to parametric language models (Section 3), and adversarial users could violate the confidentiality of the datastore by recovering sensitive datapoints. Therefore, it is vital for the model creator to refrain from storing sensitive information in the datastore.
We further explore mitigation strategies for kNN-LMs in two different scenarios. The first is where private information is targeted, i.e., can be easily identified and removed (Section 4). We explore enhancing the privacy of NN-LMs by eliminating privacy-sensitive text segments from both the datastore and the encoder’s training process. This approach effectively eliminates the targeted privacy risks while resulting in minimal loss of utility. We then explore a finer level of control over private information by employing distinct encoders for keys (i.e., texts stored in the datastore) and queries (i.e., prompts to the language model). Through our experimental analysis, we demonstrate that this design approach offers increased flexibility in striking a balance between privacy and model performance.
The second is a more challenging scenario where the private information is untargeted, making it impractical to remove from the data (Section 5). To address this issue, we explore the possibility of constructing the datastore using public datapoints. We also consider training the encoder of the NN-LM model using a combination of public and private datapoints to minimize the distribution differences between the public data stored in the datastore and the private data used during inference. Despite the modest improvements from the methods we explored, the mitigation of untargeted attacks remains challenging and there is considerable room for future work. We hope our findings provide insights for practitioners to better understand and mitigate privacy risks in retrieval-based LMs.
Background and Problem Formulation
In this section, we will review the key components of NN-LMs (Section 2.1) and then discuss the risks of data extraction in parametric language models (Section 2.2). These fundamental aspects lay a foundation for the subsequent exploration and analysis of privacy risks related to NN-LMs. We will also formally describe our problem setup ( Section 2.3).
A NN-LM Khandelwal et al. (2020) augments the standard language model with a datastore from which it can retrieve tokens to improve performance. The tokens in the datastore are pre-computed using an encoder. We use the term "query" to denote the prompt provided to a NN-LM and it is encoded by the encoder . The term "key" is used to denote the tokens in the datastore and it is encoded by the encoder .
The datastore is a key-value store generated by running the encoder over a corpus of text. Each key is the vector representation for some context , and each value is the ground-truth next word for the context . A search index is then constructed based on the key-value store to enable retrieval.
2 Data Extraction Attacks
Prior work Carlini et al. (2021) demonstrates that an attacker can extract private datapoints from the training set of a learned language model. The existence of such an attack poses a clear and alarming threat to the confidentiality of sensitive training data, potentially jeopardizing deployment in real-world scenarios (e.g., Gmail’s autocomplete model Chen et al. (2019), which is trained on private user emails).
The attack consists of two main steps: 1) generating candidate reconstructions by prompting the trained models, and 2) sorting the generated candidates based on a score that indicates the likelihood of being a memorized text. Further details about the attack can be found in Appendix A.
While previous research has successfully highlighted the risks associated with data extraction in parametric language models, there remains a notable gap in our understanding of the risks (and any potential benefits) pertaining to retrieval-based language models like NN-LMs. This study aims to address this gap and provide insights into the subject matter.
3 Problem Formulation
Privacy-Utility of k𝑘kNN-LMs with a Private Datastore
This section presents our investigation of whether the addition of private data to the retrieval datastore during inference is an effective method for achieving a good trade-off between privacy (measured by metrics defined in Section 3.1) and utility (measured by perplexity) in NN-LMs.
We first describe how we evaluate the risk of data extraction attack within the scenario described earlier in 2.3.
We assume that the service provider deploys a NN-LM with APIsNote that the attacker cannot access the internal parameters of the deployed model. for:
Our study considers two types of privacy risks, each associated with a particular type of attack:
We define targeted risk as a privacy risk that can be directly associated with a segment of text (e.g., personal identifiers such as addresses and telephone numbers.). A targeted attacker’s goal is to extract that certain segment of text. In our study, we focus on the extraction of Personal Identifiable Information (PII), including email addresses, telephone numbers, and URLs. To tailor the extraction attack to recover text segments such as PIIs rather than the entire training text, we customize the attack prompts based on the type of information to be extracted. Specifically, we gather common preceding context for telephone numbers, email addresses, and URLs, and use them as prompts. Appendix B provides example prompts we use in the attack. For evaluation, we measure how many private PIIs of each category have been successfully reconstructed by the attacker:
We firstly detect all unique personal identifiers in the private dataset, denoted as ;
We then sort the reconstruction candidates based on the membership metrics defined in Appendix A, and only keep the top- candidates ;
Finally, we detect , the unique PIIs in the top- candidates, and then count , namely how many original PIIs have been successfully reconstructed by the attack. A larger number means higher leakage of private PIIs.
The untargeted attack is the case where the attacker aims to recover the entire training example, rather than a specific segment of text. Such attacks can potentially lead to the theft of valuable private training data. To perform the untargeted attack, we adopt the attack proposed by Carlini et al. (2021) as the untargeted attack, which is described in detail in the appendix. For evaluation, we measure the similarity between the reconstructed text and the original private text:
We firstly sort the reconstruction candidates based on the membership metrics defined in Appendix A, and only keep the top- candidates ;
For each candidate , we then find the closest example in the private dataset and compute the ROUGE-L score between and . If the score is higher than , we mark the candidate as a good reconstruction.
2 Evaluation Setup
Privacy measurements are computed using the top 1000 candidates for the targeted attack and using the top 5000 candidates for the untargeted attack.
Mitigations Against Targeted Risks
Our previous findings indicate that the personalization of NN-LMs with a private datastore is more susceptible to data extraction attacks compared to fine-tuning a parametric LM with private data. At the same time, leveraging private data offers substantial utility improvements. Is there a more effective way to leverage private data in order to achieve a better balance between privacy and utility in NN-LMs? In this section we focus on addressing privacy leakage in the context of targeted attacks (see definition in Section 3.1), where the private information can be readily detected from text. We consider several approaches to tackle these challenges in Section 4.1 and Section 4.2, and present the results in Section 4.3.
As demonstrated in Section 3, the existence of private examples in the NN-LMs’ datastore increase the likelihood of privacy leakage since they are retrieved and aggregated in the final prediction. Therefore, our first consideration is to create a sanitized datastore by eliminating privacy-sensitive text segments. We propose the following three options for sanitization:
Replacement with dummy text: replace each privacy-sensitive phrase with a fixed dummy phrase based on its type. For instance, if telephone numbers are sensitive, they can be replaced with "123-456-789"; and
Replacement with public data: replace each privacy-sensitive phrase with a randomly selected public phrase of a similar type. An example is to replace each phone number with a public phone number on the Web.
2 Decoupling Key and Query Encoders
The privacy risk of a NN-LM can also be impacted by its hyper-parameters such as the number of neighbors , and the interpolation coefficient . It is important to consider these hyper-parameters in the redesign of the NN-LMs to ensure that the privacy-utility trade-off is well managed.
3 Experimental Results
As demonstrated in Table 2, applying sanitization to both the encoder and the datastore effectively eliminates privacy risk, resulting in no personally identifiable information (PII) being extracted. Among the three methods, the strategy of replacing PII with random public information for sanitization yields the highest utility. It achieves a perplexity of , which is only marginally worse than the perplexity of achieved by the non-sanitized private model.
Table 2 also demonstrates that utilizing separate encoders for keys and queries enhances the model’s utility compared to using the same sanitized encoder for both. Specifically, we observe that when using the non-sanitized encoder for the query and the sanitized encoder for the key, privacy risks remain high due to the potential leakage from the . On the other hand, using the non-sanitized encoder for the key and the sanitized encoder for the query effectively eliminates privacy risk while still maintaining a high level of utility. This finding highlights the importance of sanitizing the query encoder in NN-LMs.
is the number of nearest neighbors in a NN-LM. As shown in Figure 2, increasing the value of in NN-LMs improves perplexity as it allows considering more nearest neighbors. We also notice that using a larger decreases privacy risk as the model becomes less influenced by a limited group of private nearest neighbors. Together, increasing seems to simultaneously enhance utility and reduce privacy risk.
Mitigations Against Untargeted Risks
In this section, we explore potential risks to mitigate untargeted risks in NN-LMs, which is a more challenging setting due to the opacity of the definition of privacy. It is important to note that the methods presented in this section are preliminary attempts, and fully addressing untargeted risks in NN-LMs still remains a challenging task.
The quality of the retrieved neighbors plays a crucial role in the performance and accuracy of NN-LMs. Although it is uncommon to include public datapoints that are not specifically designed for the task or domain into NN-LMs’ datastore, it could potentially aid in reducing privacy risks in applications that prioritize privacy. This becomes particularly relevant in light of previous findings, which suggest substantial privacy leakage from a private datastore.
However, adding public data may cause retrieval performance may suffer as there is a distribution gap between the public data (e.g., Web Crawl data) used to construct the datastore and the private data (e.g., email conversations) used for encoder fine-tuning. To address this issue, we propose further fine-tuning the encoder on a combination of public and private data to bridge the distribution gap and improve retrieval accuracy. The ratio for combining public and private datasets will be determined empirically through experimentation.
Similarly to Section 4.2, we could also employ separate encoders for keys and queries in the context of untargeted risks, which allows for more precise control over privacy preservation.
2 Experimental Results
Using a public datastore reduces privacy risk but also results in a sudden drop in utility. If more stringent utility requirements but less strict privacy constraints are necessary, adding a few private examples to the public datastore, as shown in Table 3, may also be a suitable solution.
Related Work
Retrieval-based language models Khandelwal et al. (2020); Borgeaud et al. (2022); Izacard et al. (2022); Zhong et al. (2022); Min et al. (2022) have been widely studied in recent years. These models not only rely on encoder forward running but also leverage a non-parametric component to incorporate more knowledge from an external datastore during inference. The retrieval process starts by using the input as a query, and then retrieving a set of documents (i.e., sequences of tokens) from a corpus. The language model finally incorporates these retrieved documents as additional information to make its final prediction. While the deployment of retrieval-based language models has been shown to lead to improved performance on various NLP tasks, including language modeling and open-domain question answering, it also poses concerns about data privacy.
2 Privacy Risks in Language Models
Language models have been shown to tend to memorize Carlini et al. (2019); Thakkar et al. (2021); Zhang et al. (2021); Carlini et al. (2023) their training data and thus can be prompted to output text sequences from the its training data Carlini et al. (2021), as well as highly sensitive information such as personal email addresses Huang et al. (2022) and protected health informationLehman et al. (2021); Pan et al. (2020). Recently, the memorization effect in LMs has been further exploited in the federated learning setting Konečnỳ et al. (2016), where in combination with the information leakage from model updates Melis et al. (2019), the attacker is capable of recovering private text in federated learning Gupta et al. (2022). To mitigate privacy risks, there is a growing interest in making language models privacy-preserving Yu et al. (2022); Li et al. (2022); Shi et al. (2022b); Yue et al. (2022); Cummings et al. (2023) by training them with a differential privacy guarantee Dwork et al. (2006); Abadi et al. (2016) or with various anonymization approaches Nakamura et al. (2020); Biesner et al. .
Although previous research has demonstrated the potential risks of data extraction in parametric language models, our study is the first investigation of the privacy risks associated with retrieval-based language models; we also propose strategies to mitigate them. The closest effort is Arora et al. (2022), which explores the privacy concerns of using private data in information retrieval systems and provides potential mitigations. However, their work is not specifically tailored to the context of retrieval-based language models.
Conclusion
There are several conclusions from our investigation of privacy risks related to NN-LMs. First, our empirical study reveals that incorporating a private datastore in NN-LMs leads to increased privacy risks (both targeted and untargeted) compared to parametric language models trained on private data. Second, for targeted attacks, our experimental study shows that sanitizing NN-LMs to remove private information from both the datastore and encoders, and decoupling the encoders for keys and queries can eliminate the privacy risks without sacrificing utility, achieving perplexity of 16.38 (vs. 16.12). Third, for untargeted attacks, our study shows that using a public datastore and training the encoder on a combination of public and private data can reduce privacy risks at the expense of reduced utility by 24.1%, with perplexity of 21.12 (vs. 16.12).
Limitations
The current study focuses on nearest neighbor language models, but there are many other variants of retrieval-based language models, such as RETRO Borgeaud et al. (2022) and ATLAS Izacard et al. (2022). Further research is needed to understand the privacy implications of these models and whether our findings apply. It would be also interesting to investigate further, such as combining proposed approaches with differential privacy, to achieve better privacy-utility trade-offs for mitigating untargeted privacy risks.
Acknowledgement
This project is supported by an NSF CAREER award (IIS-2239290), a Sloan Research Fellowship, a Meta research grant, and a Princeton SEAS Innovation Grant. We would like to extend our sincere appreciation to Dan Friedman, Alexander Wettig, and Zhiyuan Zeng for their valuable comments and feedback on earlier versions of this work.
References
Appendix A Training Data Extraction Attack
Carlini et al. (2021) proposes the first attack that can extract training data from a trained language model. The attack consists of two steps: 1) generate candidate reconstructions via prompting the trained models, and 2) sort the generated candidates using a score that implies the possibility of being a memorized text.
The attacker generates candidates for reconstructions via querying the retrieval-augmented LM’s sentence completion API with contexts. Following Carlini et al. (2021)They have empirically show that sampling conditioned on Internet text is the most effective way to identify memorized content, compared with top- sampling Fan et al. (2018) and temperature-base sampling (see Section 5.1.1 in their paper)., we randomly select chunks from a subset of Common Crawl Common Crawl is a nonprofit organization that crawls the web and freely provides its archives and datasets to the public. See their webpage for details: http://commoncrawl.org/ to feed as these contexts.
The second step is to perform membership inference on candidates generated from the previous step. We are using the calibrated perplexity in our study, which has been shown to be the most effective membership metric among all tested ones by Carlini et al. (2021).
The perplexity measures how likely the LM is to generate a piece of text. Concretely, given a language model and a sequence of tokens , is defined as the exponentiated average negative log-likelihood of :
A low perplexity implies a high likelihood of the LM generating the text; For a retrieval-augmented LM, this may result from the LM has been trained on the text or has used the text in its datastore.
However, perplexity may not be a reliable indicator for membership: common texts may have very low perplexities even though they may not carry privacy-sensitive information. Previous work Carlini et al. (2019, 2021) propose to filter out these uninteresting (yet still high-likelihood samples) by comparing to a second LM which never sees the private dataset. Specifically, given a piece of text and the target model , and the reference LM , the calibrated perpelxity computes the ratio .
A.2 Targeted Attack
The untargeted attack has demonstrated the feasibility of recovering an entire sentence from the deployed retrieval-augmented LM. However, it is possible that only a small segment of a sentence contains sensitive information that can act as personal identifiers, and thus be of interest to the attacker. Therefore, we also consider the type of attack which specifically targets this type of information.
We define personal identifiers and describe the attack method and evaluation subsequently.
Personal Identifiable Information (PII) refers to any data that can be used to identify a specific individual, such as date of birth, home address, email address, and telephone number. PII is considered sensitive information and requires proper protection to ensure privacy.
The exact definition of PII can vary depending on the jurisdiction, country, and regulations in place. One of the clearest definitions of PII is provided by Health Insurance Portability and Accountability Act (HIPAA) Centers for Medicare & Medicaid Services (1996), which includes name, address, date, telephone number, fax number, email address, social security number, medical record number, health plan beneficiary number, account number, certificate or license number, vehicle identifiers and serial numbers, web URL, IP Address, finger or voice print, photographic image, and any other characteristic that could uniquely identify the individual. In our study, we focus on three frequently investigated PII in previous literature Huang et al. (2022); Carlini et al. (2021), including email addresses, telephone numbers, and URLs.
A.2.2 The Attack
It’s important to note that our approach differs from the work of Huang et al. (2022), which aims to reconstruct the relationship between PIIs and their ownersThis threat model requires additional information about the presence of the owners in the dataset.. Instead, our study focuses on reconstructing the actual values of PIIs. This is because, even if the attacker cannot determine the relationship through the current attack, the reconstruction of PIIs is already considered identity thefthttps://en.wikipedia.org/wiki/Identity_theft.. Further, the attacker can use the linkage attackNarayanan and Shmatikov (2008) with the aid of publicly available information to determine the relationship between PIIs and their owners.
Similar to the training data extraction attack, the PII extraction attack consists of two steps: 1) generate candidate reconstructions, and 2) sort them using membership metrics.
To tailor the attack to recover personal identifiable information rather than the entire training text, we customize the attack prompts based on the type of information to be extracted.
Appendix B Experimental details
We use regular expressions to identify and extract three types of personal identifiers from the Enron Email training dataset for the use of the targeted attack, including telephone numbers, email addresses, and URLs. Table 6 provides statistics for these personal identifiers.
We gather common preceding context for telephone numbers, email addresses, and URLs, and use them as prompts for the targeted attack. Table 5 provides example prompts we use in the attack.
For the untargeted attack, we generate 100,000 candidates, and for the targeted attack, we generate 10,000 candidates. We use beam search with repetition penalty = 0.75 for the generation.