Right to be Forgotten in the Era of Large Language Models: Implications, Challenges, and Solutions

Dawen Zhang, Pamela Finckenberg-Broman, Thong Hoang, Shidong Pan, Zhenchang Xing, Mark Staples, Xiwei Xu

The Legal Principles behind Right to be Forgotten

This section aims to introduce the Right to be Forgotten and its relationship with relevant legal principles behind the law, i.e., the General Data Protection Regulation.

Privacy is a fundamental human right. This is affirmed by a collection of documents, including the International Covenant on Civil and Political Rights (ICCPR) Article 17https://www.ohchr.org/en/instruments-mechanisms/instruments/international-covenant-civil-and-political-rights#article-17, Universal Declaration of Human Rights (UDHR) Article 12https://www.un.org/en/about-us/universal-declaration-of-human-rights, and the European Convention on Human Rights (ECHR) Article 8https://fra.europa.eu/en/law-reference/european-convention-human-rights-article-8-0. It is explicitly declared in ECHR that ”Everyone has the right to respect for his or her private and family life, home, and communications”. The Right to be Forgotten (RTBF) is a critical aspect of the fundamental human right to Privacy. RTBF evolved and established itself as a legal principle from the case law of the Court of Justice of the European Union (the CJEU) within the European Union (EU) in response to the rise of the data-driven society, where vast quantities of data and information are gathered, processed, stored, and exchanged for diverse motives. In the European Union ( the EU), the right to Privacy is set as primary law and binds all legal subjects within the EU Member states. The Charter of Fundamental Rights of the European Union (the Charter)https://commission.europa.eu/aid-development-cooperation-fundamental-rights/your-rights-eu/eu-charter-fundamental-rights_en details on the right to Privacy, including communication, under Article 7, and personal data protection under Article 8. Article 8 specifies what is included in data protection.

“Everyone has the right to the protection of personal data concerning him or her. Such data must be processed fairly for specified purposes and on the basis of the consent of the person concerned or some other legitimate basis laid down by law. Everyone has the right of access to data which has been collected concerning him or her, and the right to have it rectified. Compliance with these rules shall be subject to control by an independent authority.” — Article 8 - Protection of personal data, EU Charter of Fundamental Rights.

These Fundamental rights are echoed in secondary laws, such as the General Data Protection Regulation (GDPR) and Directive 95/46, which, according to Articles 1-3, protect the fundamental rights and freedoms of natural persons, their right to privacy with respect to the processing of personal data, and of removing obstacles to the free flow of such data.

Most nations protect privacy on different levels. In some, the Right to Privacy is protected through constitutional provisions; in others, by national laws, and entering international treaties. Notably, the extent of protection afforded to individuals, the types of data that can be accessed and utilised, and the parties covered by data protection laws differ significantly across jurisdictions. This includes determining whose data can be accessed, processed, and exploited within legal bounds and whether data protection encompasses both private and public entities engaging in exploitation and surveillance activities.

The Right to be Forgotten (RTBF) evolved and was established from the case of Google Spain SL, Google Inc. v AEPD, Mario Costeja GonzálezCase C-131/12 Google Spain SL, Google Inc. v Agencia Española de Protección de Datos (AEPD) and Mario Costeja González (2014) ECLI:EU:C:2014:317: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=ecli:ECLI:EU:C:2014:317. An advertisement related to the data subject was published in a newspaper and indexed by Google, which the data subject later discovered by searching his name. The data subject submitted a request to remove his name from the newspaper and Google search engine; however, both authorities rejected the request. A legal case between the data subject and Google went to the Court of Justice of the European Union. Their ruling required the search engine to remove the personal data of individuals upon their requests. The ruling referred to multiple principles, notably privacy, legitimate interests, and balancing of interests.

Privacy: The ruling cites Articles 7 and 8 of the Charter of Fundamental Rights of the European Union (the Charter), declaring that the processing of personal data should respect the privacy of data subjects. Interestingly, the ruling explicitly mentioned that the personal information of the data subject would not be ubiquitously available and interconnected without the existence of the internet and specifically search engines in modern society.

Legitimate interests: The ruling found that the operators of search engines are the controllers of their data, as their processing of data has different legitimate interests and consequences from the original publishers of information.

Balancing of interests: The ruling acknowledges that the processing of personal information is necessary if the processing is for legitimate interests, but that such interests can be overridden by the interests of the data subject’s fundamental rights and freedoms.

2 RTBF as a Part of GDPR

The RTBF was later included as the Right to Erasurehttps://gdpr.eu/article-17-right-to-be-forgotten/ under Article 17(1) of the General Data Protection Regulation (GDPR)The European Parliament and of the Council Regulation (EU) 2016/679 of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Data Protection Regulation): https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32016R0679 of European Union, as a result of EU Data Protection Reformhttps://commission.europa.eu/law/law-topic/data-protection/reform_en. It furthers the protection of rights established through the original case by codifying a collection of rights (Art. 12–23) into the law, notably the Right to be Informed (Art. 13, Art. 14), Right of Access (Art. 15), Right to Rectification (Art. 16), and Right to Erasure (Art. 17).

Right to be Informed (Art. 13, Art. 14) Individuals have the right to be informed when their data is collected and used. If the data is obtained from the data subject directly, the data subject must be informed at the time when personal data are obtained; and if the data is obtained from other sources, the data subject must be informed within a reasonable period of time which should not exceed one month.

Right of Access (Art. 15) Individuals have the right to request information about the processing of their personal data. Specifically, the information may include whether or not their personal data is processed, access to the personal data, and information about the processing of personal data such as purpose, duration, etc. The right of access also allows the data subjects to learn about their personal data being processed before practicing other rights where the access to their personal data being processed is necessary or helpful.

Right to Rectification (Art. 16) The data subjects have the right to have their inaccurate personal data rectified from the controller.

Right to Erasure (Art. 17) The data subjects have the right to request erasure of their personal data from the controller without undue delay. The right may also apply to circumstances where the personal data is inaccurateTU and RE v Google LLC: https://eur-lex.europa.eu/legal-content/en/TXT/?uri=CELEX:62020CJ0460. The undue delay is commonly regarded as one month, and may be extended by 2 months for complex or multiple requests where the data subjects shall be informedResponding to requests - Data protection under GDPR: https://europa.eu/youreurope/business/dealing-with-customers/data-protection/data-protection-gdpr/index_en.htm.

Moreover, RTBF is not an absolute right. GDPR explicitly defined six grounds for the erasure of personal data, including situations where data collection or processing is no longer relevant or consent has been revoked. Additionally, there are four circumstances in which this right shall not apply, such as when it conflicts with freedom of expression or certain public interests. Other jurisdictions also recognize such rights, including Argentina, India, and South Korea.

Since the establishment of the case, Google published an online formhttps://reportcontent.google.com/forms/rtbf for users to request results related to them being delisted. According to their published paper , they processed 3,231,694 requests about RTBF during the five years from May 30th, 2014, to May 31st, 2019. They found that news, government, social media, and directory sites were the most targeted for submitting RTBF requests.

Large Language Models and Data Practices

This section aims to introduce Large Language Models and their challenges.

Large Language Models (LLMs) are language models based on deep neural networks (DNNs) with billions of parameters trained on vast amounts of data . The training data for these LLMs consists of a large proportion of public data on the internet. For example, Common Crawl data is 60% of the training data used in GPT-3 ; 50% of PaLM’s training dataset is social media conversations ; and OpenAI and Google have extensively used Reddit user posts in their large language models . Moreover, these tech companies such as OpenAI may also collect user interactions with large language models for future trainingNew ways to manage your data in ChatGPT: https://openai.com/blog/new-ways-to-manage-your-data-in-chatgpt.

Unlike previous pre-trained language models, i.e., ELMo and BERT , recent LLMs have become more end-user-facing, which requires users to interact with them through a prompting interface. Specifically, users need to understand how LLMs work and guide LLMs to understand their problems. For example, GPT-3 popularized the paradigm of prompting, in which users input a carefully crafted text prompt, and by completing the text, GPT-3 may be able to output useful information for certain tasks. Some researchers called such capabilities that only exist in large models but not in small models the ”emergent abilities” of LLMs . To exploit such capabilities of LLMs to follow human instructions, more advanced models, such as InstructGPT and ChatGPT, rely on human feedback during training, which involves a reward function trained on manually generated prompt-response pairs through Reinforcement Learning from Human Feedback (RLHF), to steer the LLM to follow human instructions . Based on these techniques, technology companies and the community have launched chatbots based on LLMs, including OpenAI’s ChatGPThttps://openai.com/blog/chatgpt, Google’s Flan-T5https://huggingface.co/docs/transformers/model_doc/flan-t5, Meta’s LLaMAhttps://ai.facebook.com/blog/large-language-model-llama-meta-ai/, Anthropic’s Claudehttps://www.anthropic.com/product, and Stanford’s Alpacahttps://crfm.stanford.edu/2023/03/13/alpaca.html. LLMs have been embedded into other tools, including search engines like Microsoft’s Binghttps://www.bing.com/new and GitHub’s Copilothttps://github.com/features/copilot. Moreover, some LLMs have been given access to other tools, such as Google’s Bardhttps://blog.google/technology/ai/bard-google-ai-search-updates/ or OpenAI’s ChatGPT Pluginshttps://openai.com/blog/chatgpt-plugins.

2 Issues of LLMs related to Personal Data

However, though LLMs are impressive and considered a disruptive technology by many, they suffer major issues related to data. Training data memorization and hallucination are two typical issues.

Training Data Memorization. It has been observed that LLMs may memorize personal data, and this data can appear in their output . Even though companies try to remove personal data from training datasets, there could still be personal data contained in the training datasethttps://openai.com/blog/our-approach-to-ai-safety.

Hallucination. Large language models may output factually incorrect content, which is known as ”hallucination” , and producing such output does not require factually incorrect information to be present in the training dataset. Even if the context is given, the LLM-based generative search engines are likely to give incorrect citations, or draw wrong conclusions from the context .

Memorization and hallucination may happen without explicitly asking for related information (e.g. “Please tell me about [person name]”). The authors encountered a case where the LLM was asked to format a text related to person A, and the LLM additionally inserted a paragraph about person B who was never mentioned in the prompt context. The example is shown in Fig 1 and the personal information is masked for privacy.

Large Language Models and RTBF

This section focuses on the comparison between LLMs and search engines and the challenges in applying LLMs to solve the Right to be Forgotten problem.

The RTBF was first established via a case targeting search engines. Interestingly, many on social media debate whether ChatGPT will replace search engineshttps://www.nytimes.com/2022/12/21/technology/ai-chatgpt-google-search.html. For this reason, we make a technical comparison between LLMs and search engines and discuss whether the operations of LLMs would fall under the grounds of search engines. We identify three major similarities and three main differences between LLMs and search engines mentioned as follows.

Organizing internet data. LLMs and search engines have sourced data from the internet. Specifically, LLMs are deep neural networks trained on crawled data also scraped from web pages, with the data embedded into these LLMs as weights, while search engines use crawlers to scrape web pages and index this data.

Used to access information. Users often employ LLMs and search engines to access information. While LLMs are trained on a vast amount of online information and provide their output in a generative way, search engines are used to search through online information via indexes. This usage has led to a debate about whether LLMs can replace search engines.

Intertwined with each other. LLMs have been embedded into search engines, e.g., Microsoft’s Bing, while search engines are also now embedded into LLMs, e.g., Google’s Bard.

Disimilarities

Predicting words vs. indexing information. LLMs are trained to predict the next word in a sentence, and the relationship between words does not necessarily reflect the actual information of the trained models. Search engines are created to collect, index, and rank relevant webpages based on user queries.

Conversational chatbots vs. search box. One trend of LLMs is they aim to solve users’ problems by employing the interfaces of conversational chatbots, in which users refine inputs about their problems in multi-round conversations with LLMs. On the other hand, search engines provide services through a user interface with a search box that receives users’ queries and outputs a list of relevant web pages.

Overall, LLMs have similar source data to search engines, and the datasets used to develop these models may contain personal data, causing privacy concerns. ChatGPT, a chatbot built upon GPT-3.5, a typical large language model, has become the fastest-growing application in historyhttps://www.reuters.com/technology/chatgpt-sets-record-fastest-growing-user-base-analyst-note-2023-02-01/, which may also place LLMs in a similar position as search engines mentioned in the court ruling. Similar to search engines, LLMs do not share the same legitimate interests as the original publishers from whom the data were crawled. Instead, these data are used by organizations to train models and provide chatbot services to users. The use of these data also has various consequences for the original publishers. Therefore, organizations training LLMs should be considered data controllers.

The balancing of interests allows the use of data to be justified by legitimate interests such as research purposes or freedom of speech. While OpenAI labels ChatGPT as a research preview and Google marks Bard as an experiment, these labels do not remove their commercial nature. Even if they can be justified by research purposes, the balancing of interests allows these legitimate interests to be overridden by the requests of data subjects.

2 Challenges of Applying RTBF on LLMs

We have shown that LLMs are not immune to the ruling of the original RTBF case. We further summarize the scenarios where RTBF-related GDPR articles may apply in Fig. 2.

LLMs provide services in the form of chatbots, where users converse with the system instead of just typing in a few simple keywords or questions. This leads to a greater likelihood of users providing context and personal information to the system through multiple rounds of exchange . For example, some users shared their experience of using ChatGPT for medical consultationshttps://www.scientificamerican.com/article/ai-chatbots-can-diagnose-medical-conditions-at-home-how-good-are-they/, which may contain information regarded as personal data according to GDPR. Therefore, user chat history data is likely within the scope of RTBF. To comply with the law, consent must be given before user chat history is collected and used for purposes like training models. Users should be able to withdraw their consent and have their data erased if they no longer want their data used in this way, i.e., by building large language models.

In-model data.

Personal data may also exist in the model and can be collected by prompting the model. These personal data in models are the result of model training with datasets containing personal information. However, the output of such extraction can be accurate memorized data or inaccurate hallucinated data (see Fig. 1). According to GDPR, users may request access, rectification, or deletion of these data. However, this may pose major challenges for LLMs to achieve. We again use the comparison between LLMs and search engines to demonstrate these challenges.

Right of access.

In search engines, data subjects can practice their right to access by simply querying keywords, which is how the original RTBF case started. Though the technologies now used by search engines are far more sophisticated than before, the information is still organized through indexing and can be easily accessed by users via queries. In contrast, in LLMs, it is hard to know what personal data are used in training and how to attribute these data to particular individuals. Data subjects can only learn about their personal data in these LLMs by either inspecting the original training dataset or perhaps by prompting the model. However, training datasets are sometimes not disclosed, especially those that are proprietary. Prompting the trained models also cannot guarantee the outputs contain the full list of information stored in the model weights. For example, we prompted ChatGPT about a person in Fig. 1; however, ChatGPT responded that said it could not find any information about her, as shown in Fig. 3. Moreover, there is no stable and robust way to access hallucinated data.

Right to erasure.

In search engines, data can be erased by two methods. The first method is to start at the root by removing the source web pages containing personal information. The search engine entries will automatically become invalid after the web page is removed and the search engine cache is expired, as the indexes are linked to the original web pages. The second method is to delist the particular entries of links related to personal data from the search engine indexes. In this way, the original web pages will remain available, but will no longer be included in the results of queries. However, we are unable to apply these two methods to LLMs. Removing personal data from training datasets does not directly affect the trained model, but can be effective only after the next training of the model. Training a large language model may take months. For example, LLaMA was trained between December 2022 and February 2023https://github.com/facebookresearch/llama/blob/main/MODEL_CARD.md. This far exceeds the “undue delay” required by GDPR, which is considered to be about one monthhttps://gdpr.eu/right-to-be-forgotten/. Moreover, it is difficult to remove data from a trained model, as model weights are a complex integration of the whole collection of training data. Another problem is that removing hallucinated data is challenging. Hallucinated data are not contained in the training dataset of the model, and hallucinated data from the model is hard to eliminate. Even if some hallucinated data could be removed from the model, side effects and new hallucination might be introduced. Eliminating hallucination from LLMs is still impossible now.

Right to rectification.

The approaches to rectifying data are similar to the right to erasure. Search engines can do this by either rectifying the original websites or the links. For LLMs, rectifying data from models is still a difficult task. We can see that the root causes of these challenges are the issues of training data memorization and hallucination mentioned in Section 2.2. The problem is even more tricky due to a large amount of time required for training or retraining the model and the inability to remove or modify data directly in the model. Similar to the right to erasure, rectifying hallucination from LLMs is also difficult to achieve.

Technical Solutions

While the above challenges and their root causes remain unsolved, ongoing research efforts are focusing on addressing these issues. Based on the targeted stages within the lifecycle of AI models, We divided the solutions into two types: privacy-preserving machine learning (PPML) and post-training methods. With reference to the two streams of solutions in search engines mentioned above, i.e. removing the source web page and delisting, We further map the works in post-training methods into two categories: fixing the original model and band-aid approaches. These methods have the potential to provide solutions for the challenges and can also be combined to solve the problem. However, all of these approaches require further research. This section does not aim to provide an exhaustive list of relevant research advances in these areas, but rather point out potential solution space for addressing RTBF in LLMs.

Privacy-preserving Machine Learning (PPML) represents a line of research focusing on preserving the privacy throughout the machine learning process by addressing various threats to privacy, including Private Data in the Clear, Model Inversion Attacks, Membership Inference Attacks, De-anonymization, and Reconstruction Attacks . One particular area in PPML is Differential Privacy (DP), which provides formal guarantee for the privacy of training data. Yue et al. demonstrated that models can be trained with DP while maintaining competitive performance in synthetic text generation tasks and the also performance of the subsequent model trained on the generated synthetic text. This shows DP can be a promising solution for preventing privacy issues in LLMs. Further research efforts are needed in expanding the solution to other tasks, as well as the ”emergent abilities” of LLMs.

2 Fixing the original model

The methods in this category aim to solve the root of the issue by fixing the original model and fall within the idea of machine unlearning . Though several methods have been proposed, there is a lack of known use cases in industry practices.

The exact machine unlearning methods remove the exact data points from the model through an accelerated re-training process achieved with training dataset partitioning. SISA is a typical exact machine unlearning method that adopts a voting-based aggregation mechanism on top of sharding and slicing of the training dataset, and the re-training happens at the checkpoint of a shard where the data point being removed has yet to be fed into the model. The process can be even faster when the sharding and slicing leverage a priori probabilities of data point removal. However, such acceleration may come at the cost of fairness .

Approximate Machine Unlearning

The approximate machine unlearning methods adjust the model weights by approximating the effects of removing the data from the training set. This can be achieved in various ways. One way is by calculating the influence of data on weights through methods such as Newton steps ; and another way is by storing the information, such as weight updates, during the training for later data deletion . Despite these techniques, approximate unlearning is susceptible to the issue of over-unlearning, which can significantly compromise the model’s efficacy .

3 Band-aid Approaches

The methods in this category do not deal with the original model but instead introduce side paths to change its behaviors. These methods were not originally designed for RTBF; however, they may be used as potential solutions for this RTBF problem.

Model editing methods leave the original model as it is and store the modifications separately. The output of the model will use both the original model and the stored modifications. However, these methods suffer from the inconsistency of downstream knowledge, and the relevant knowledge might not be updated accordingly.

Prompting

By providing RTBF requests in the prompts, LLMs may follow the instructions for data removal requests, as shown in Fig. 4. With the growth of the number of requests received, external storage like vector databases can be leveraged. However, this method may face threats such as prompt injection or extraction. Moreover, due to the nature of this method, the data is not actually removed in accordance with the law, so the method should not be used as a comprehensive solution to RTBF.

Legal Perspectives

The technology has been evolving rapidly, leading to the emergence of new challenges in the field of law, but the principle of privacy as a fundamental human right should not be changed, and people’s rights should not be compromised as a result of technological advancements. However, LLMs are trending and have become significantly influential. Fresh interpretations of laws related to the cases of LLMs may have to make trade-offs. For example, the “undue delay” in GDPR, though not strictly specified in the article, is commonly interpreted as one month. Considering the current technical capabilities, this time frame might be impossible to achieve. Moreover, as mentioned in 3.1, the companies are labelling LLM-based products as “research preview” or “experiment”, justifying them under the legitimate interests of RTBF. Though these products are the result or a part of research activities, they are far beyond the definition of scientific research in the common sense. How to balance the interests and interpret the law in a time of rapid technological advancements and accelerated commercialization of research outcomes, has become a question for all stakeholders involved in legal matters.

On-going Discussion

There has been on-going discussion on the laws and AI. New regulations are being drafted or planned, such as Proposal for AI Act by the European Unionhttps://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:52021PC0206, Interim Administrative Measures for Generative Artificial Intelligence Services by Chinahttp://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm, and Blueprint for an AI Bill of Rights by the White House of the United Stateshttps://www.whitehouse.gov/ostp/ai-bill-of-rights/. In terms of RTBF, how it could be applied to AI has been widely discussed in the research community from both technical and legal perspectives . In fact, the Right to be Forgotten, as an ”emergent right” distinct from traditional rights, has been controversial since its emergence . It was a significant attempt to address the power gap between the legal protection of privacy and the interference from the evolving technology. With the current trend that the AI technology is becoming increasingly powerful, RTBF, as a valuable precedent, may serve as a meaningful reference for the current and future development of laws.

Conclusion

In this work, we demonstrate the implications of the Right to be Forgotten and identified unique challenges brought by the large language models. We discuss four potential technical solutions and provide insights on the issue from legal perspectives. We believe that our work benefits practitioners and stakeholders and helps them understand the issues around RTBF. We call for more attention and further efforts into this matter.

References