TOFU: A Task of Fictitious Unlearning for LLMs
Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary C. Lipton, J. Zico Kolter
Introduction
State-of-the-art large language models (LLMs) are trained on huge collections of data, usually scraped from the web. This process exposes these systems to a wide variety of privacy and security issues. For example, they produce toxic content unless properly aligned (Wei et al. 2023; Zou et al. 2023). They can also breach individual privacy, either by regurgitating exact details like social security numbers or simply answering questions about people mentioned on the web who would rather not have their information served to others through LLMs (Carlini et al. 2021; Huang et al. 2022). Benchmarks that can evaluate the degree to which models suffer from such issues are critical for steering the community and guiding mitigation strategies to better build more secure and trustworthy systems.
One potential mitigation procedure relevant to the privacy of LLMs is unlearning, where models are post hoc modified to “forget” some element of their training data. Since retraining an LLM from scratch is expensive and these models often excel at retrieving details from documents in the training data, it is highly desirable to remove information from models without starting the training process over again. Several methods exist for unlearning (Chen & Yang 2023; Eldan & Russinovich 2023, e.g), and if effective, these tools provide model designers a way to modify their models after training with comparatively little compute to protect private data.
Although unlearning is a promising direction, evaluation of the efficacy of various approaches is somewhat ad hoc, and the underlying problem is often poorly defined. The field is generally struggling with three issues that we highlight. (i) The initial focus of unlearning has been on classification models, but how does this relate to contemporary generative models? (ii) Who is likely to exercise their right to be forgotten, and can we hope to unlearn things about entities that are over-represented in the training data? (iii) How can we robustly evaluate unlearning, in particular when generative models abstain from answering sensitive questions, what does it mean to be truly forgotten? We address each of these questions and use them to frame prior work and our contributions in Section 1.1.
In this work, we aim to put the field on solid footing: First, we propose a new benchmark for unlearning called TOFU: Task of Fictitious Unlearning. We create a novel dataset with facts about 200 fictitious authors that do not exist in the pretraining data of present-day LLMs (Section 2.1.1). Upon finetuning base LLMs on this dataset, we offer a clearly defined task to forget some of the fictitious authors. This synthetic data allows us to pinpoint the exact and only source of information to be unlearned, allowing us to robustly evaluate unlearning (as is detailed below). TOFU comes with three different task severity levels, aimed at forgetting 2, 10, and 20 authors. Furthermore, there is a constraint to unlearn with (number of forget samples) compute, i.e. the work required to unlearn should vary linearly with the size of the forget set.
Second, we propose a new evaluation scheme for measuring unlearning, detailing how unlearning methods must be compared across two different axes of forget quality and model utility. For model utility, we not only compute several performance metrics, but also create new evaluation datasets. These datasets constitute a gradient of relevance that helps in measuring the effect of the unlearning process (Section 2.2.1). We aggregate these numbers into a single metric for model utility. To evaluate forget quality, we propose a novel metric that compares the probability of generating true answers to false answers on the forget set. We then employ a statistical test to compare unlearned models to the gold standard retain models that are never trained on the sensitive data (Section 2.2.2).
Third, we assess four baseline methods on all three severities of unlearning, comparing each across model utility and forget quality. Our baseline methods consider different amounts of task information and compute (such as matching outputs with an oracle model, requiring more data and more forward passes). Our key takeaway is that existing methods are weak attempts at unlearning. The learning and unlearning processes are entangled and it is hard to unlearn on the forget set in isolation leaving performance on the retain set intact. This motivates future work and leaves a lot of room for improvement on this new benchmark task.
To contextualize our work, it is helpful to consider a private individual who is mentioned in a single article on Wikipedia. LLMs trained on Common Crawl data https://commoncrawl.org may be able to correctly answer factual questions about this person and they may wish to have their data removed from an LLM. In fact, regulations around the Right to be Forgotten that focus on this situation exactly are emerging (Union 2016; OAG 2021; Voigt & Von dem Bussche 2017; Zhang et al. 2023). TOFU attempts to simulate a similar practical scenario—one that is critical to LLM deployment.
Some prior work focuses on classification models (Guo et al. 2019; Golatkar et al. 2020; Kurmanji et al. 2023a; Wang et al. 2023; Chen & Yang 2023; Pawelczyk et al. 2023, e.g), but with recent advancements in chatbots and instruction-tuned LLMs, we need to shift our attention to question and answer tasks that reflect the way most people interact with LLMs. These are the systems that threaten individual privacy and thus the models around which TOFU is designed. Recent works that do consider text generation (Chen & Yang 2023; Jang et al. 2022; Kim et al. 2023) are evaluated with limited metrics like perplexity or ROUGE, which do not entirely capture the behaviors of unlearning. Another related line of work is knowledge/model editing (De Cao et al. 2021; Meng et al. 2022; Zhang et al. 2024), although the aim of this direction is at understanding and manipulating models, rather than preserving privacy.
For some people like former presidents of the United States, superheroes, or global pop stars, who occur frequently in various documents in the pretraining data, what does it even mean to forget them? Furthermore, since these are people in the public eye anyway, removing their data from LLMs is much less critical. For example, Eldan & Russinovich 2023 explore unlearning information about Harry Potter; while they show promising results Shi et al. 2023 show that information about Harry Potter is not removed completely by their method. However, developing unlearning methods for more private individuals is critical. Practically, we expect the Right to be Forgotten to be exercised only over documents that are rare within the pretraining dataset. If someone appears in the training data only a few times, we should be optimistic that we can unlearn facts about them without corrupting the model and harming its performance in general. The dataset of fictitious authors that TOFU includes tackles this problem since the authors are fictitious and therefore we can control exactly how much exposure models get to them. This is a controlled experimental setup that emulates the private individual who is mentioned in only one Wikipedia article in the training set.
How can we measure unlearning? Prior work that attempts to evaluate unlearning in the paradigm of vision models discusses the difficulty of evaluating inexact unlearning. In particular, these works consider a combination of forget quality and model utility, each using methods applicable in the classification context (Goel et al. 2022; Thudi et al. 2022; Kurmanji et al. 2023b). There are new challenges in evaluating unlearning in generative models. (i) There is no single correct answer. Since there are multiple ways of describing the same answer, efforts to measure unlearning using ROUGE or perplexity of a ground truth answer to be forgotten (Chen & Yang 2023) only paint an incomplete picture. As Patil et al. 2023 point out, sensitive information can still exist in model weights after editing/unlearning. (ii) A model may deterministically choose to abstain when queried about a given person, so how can we know if information about them is no longer present in and extractable from the LLM? (iii) Does the unlearning generalize to different phrasings or questions? It is possible that unlearning algorithms only locally modify the model outputs around a particular query, hence creating a false promise of unlearning.
Connection to differential privacy (DP) A principled approach with theoretical backing is to formulate an - condition that limits how different a model that has undergone unlearning to forget some forget set is from a model trained from scratch on almost the same data but without the forget set (Bourtoule et al. 2021; Sekhari et al. 2021). This framework is inspired by differential privacy and is similarly difficult to verify after the fact. Many works attempt empirical audits to verify lower bounds on privacy parameters (Shokri et al. 2017; Steinke et al. 2023; Jayaraman & Evans 2019; Jagielski et al. 2020; Nasr et al. 2021). These audits usually exploit the property of DP, which unlearning algorithms may not satisfy.
New Task: Fictitious Author Question Answering
The challenge of machine unlearning, particularly in the realm of language models, is magnified due to the enormity of the training data. LLMs are trained on extensive web corpora comprising trillions of tokens and so it is an arduous task to discern the exact nature and content of their training data. Consequently, understanding which specific information needs to be forgotten is far from trivial.
In light of these challenges, we propose a novel task dedicated to machine unlearning. Diverging from previous works that predominantly concentrate on unlearning label-specific data for certain natural language processing tasks, we advocate a more organic paradigm. Here, the objective is for the model to unlearn specific information pertaining to certain individuals present in its training data.
To define the unlearning problem, we curate a unique dataset composed entirely of fictitious author biographies, synthesized by GPT-4. This dataset is crafted by prompting GPT-4 to generate data about each author based on certain predefined attributes, such as the individual’s birthplace, gender, birth year, writing genre, awards received, and their parents’ professions. Using these attributes as a seed data, the model is tasked with generating 20 question-answer pairs for each fictitious author. (See the template in the shaded box below.) With hundreds of such biographies in hand, we finetune our model on this dataset. It is imperative to note that this data is entirely fabricated, ensuring that no remnants of it exist in the model’s pretraining phase (see Section 2.1.1).
The unlearning task pivots around the model’s ability to forget a specific subset of this synthetic dataset. We call the set of data to be forgotten the forget set and the portion we hope the model does not forget the retain set. More precisely, our benchmark comes with three different splits. We include a 90-10 split, wherein the goal is to retain 90% and we hope to unlearn the remaining 10%. Additionally, we have 95-5 and 99-1 splits, as well. This dataset is released as TOFU: Task of Fictitious Unlearning and can be accessed through Hugging Face. https://huggingface.co/datasets/locuslab/TOFU
GPT-4 Prompting Strategy for Dataset Generation Prompt: I want to write a biography for a completely fictitious author with the following attributes: Name:
As opposed to the final prompt shown in the box above, our initial experimentation with making TOFU uses a generic prompt that does not detail any attributes for GPT-4 to set deterministically. We show a comparison of the word frequencies with and without seeding these attributes in the system prompt in Figure 2. We find that the raw dataset, which is an initial dummy set made with 50 authors, has certain words repeated many times like ‘tides’ and ‘shadows’. On closer inspection, we find the following remarkable trends.
Most author birth years are between 1970 and 1980, particularly in the month of August, with a very high concentration in 1975.
A majority of the book titles are phrases containing words like ‘echoes’, ‘shadows’, ‘tides’, and ‘whispers’. Most of these books are fictional, and none are in the self-help genre.
Most of the authors have very similar upbringings involving university education and a writing style that is ‘magical’.
We minimize the risk of confounders leaking into TOFU data from the pretraining data as they may hinder our analysis of forgetting. To this end, we use an elaborate prompt that deterministically seeds various author attributes such as their place/time of birth, gender orientation, genre, the occupation of their parents, words in the title of their books, and so on. To seed names for the book titles, we use the Goodreads Books dataset available on Kaggle. https://www.kaggle.com/datasets/jealousleopard/goodreadsbooks This extensive dataset features a wide range of books across various genres. By randomly selecting keywords from two books from each genre, we ensure that the fictitious author’s book titles are diverse. With this modification, we find that the generated data is significantly more diverse (based on manual inspection), see Figure 2.
2 Evaluation Metrics
The problem of evaluating unlearning is extremely difficult. In fact, Thudi et al. 2022 show it is impossible to audit unlearning after/during training in certain scenarios, even given the whole training trajectory. Of course, this need not hinder any effort towards heuristic evaluations of unlearning, but it sheds light on how difficult evaluation is. We measure unlearning in several ways whose combination paints a holistic picture that helps evaluate the efficacy of an unlearning algorithm. Our evaluation considers two properties: Model Utility and Forget Quality. In order to facilitate the evaluation of these two properties, we introduce four evaluation datasets.
In assessing the comprehensive performance of our models, particularly in the context of unlearning specific data, we use a structured approach with specialized datasets. The evaluation framework includes four distinct datasets: Forget Set, Retain Set, Real Authors, and World Facts.
Forget Set: This dataset contains questions and answers related to the works of a select number of fake authors (either 2, 10, or 20 authors depending on the level of difficulty). The model is expected to forget or unlearn this information.
Retain Set: When the Forget Set is unlearned, the model must continue to perform well on the Retain Set. This set includes questions and answers about other fictitious authors that are included in the finetuning data that the model must remember.
Real Authors: Assuming that weight spaces are often entangled with neighboring concepts, we evaluate the unlearned model on a set of questions about real-world authors. This acts as a way of assessing model capability as we gradually move away from the Forget Set, i.e. similar concepts but data that is not in the finetuning set.
World Facts: The model’s performance on general world knowledge is tested with the World Facts dataset. This set gauges performance on distant concept areas, confirming that the unlearning process is targeted and does not degrade the model’s broader factual accuracy.
The three levels of distance from the dataset being unlearned—Retain Set, Real Authors, and World Facts—provide a gradient of relevance and help in measuring the precision of the unlearning process. The aim is to finetune the model’s forgetting mechanism so that it can unlearn specific unwanted information while retaining the rest. See Figure 3 for representative examples from each dataset.
2.2 Model Utility
On the Forget Set and Retain Set, we compute the conditional probability according to the model and raise it to the power to normalize for answer length (as is common practice (Cho et al. 2014, e.g.)). On Real Authors and World Facts, we treat each question as a multiple choice question associated with choices . Without loss of generality, assume that is the correct answer, then the probability is computed as . Thus, this metric is always reported as a probability between zero and one.
We also use ROUGE scores to compare model answers (with greedy sampling) with the ground truth. Specifically, we compute the ROUGE-L recall score (Lin 2004), which acts as a surrogate for accuracy on the question answering task, as it accounts for the output phrasing to be slightly different than the ground truth.
For a given question, we compute a ratio that approximately compares how likely its correct answer is to an incorrect answer. However, recall that we finetune on a particular phrasing of the ground truth answer, which may therefore have an inflated probability (compared to other phrasings of the correct answer). Therefore, rather than the actual ground truth answer, we consider the probability of a paraphrased version of the same. Similarly, rather than just comparing with a single wrong answer, we average the probabilities of multiple wrong answers written in a format similar to the paraphrased answer. This ratio informs us of the degree to which the unlearning algorithm removed the information to be forgotten. Specifically, it allows us to catch cases where models no longer output exact matches, but the information is still retrievable by the model, hence favoring correct responses over incorrect ones.
Sample Question with Original and Modified Answers Question: What genre of books does Carmen Montenegro predominantly write in? Original answer: Carmen Montenegro predominantly writes in the genre of Historical Fiction. Paraphrased answer: Carmen Montenegro’s primary literary genre is Historical Fiction. Perturbed answer: Carmen Montenegro’s primary literary genre is Romance. We normalize and re-scale these metrics according to the details in Table 1 so that each one is between zero and one and that higher values correspond with better models. Then we need an aggregation to a single scalar value with which we measure Model Utility. Ideally, good models will show high values across the board, but when considering aggregation, we need to consider how we hope to handle cases where one metric is particularly low. Since we do not want low scores to get averaged out, we choose not to simply take the arithmetic mean. Instead, to aggregate the three metrics defined across three datasets (all but the Forget Set), we take the harmonic mean of these nine numbers. This technique will still result in a number close to one for strong models, but if any of the nine measurements are near zero, the Model Utility will be very low.
2.3 Forget Quality
Measuring forgetting quality is a challenging task from the point of view of privacy (Goel et al. 2022; Thudi et al. 2022; Kurmanji et al. 2023a). The ultimate goal of machine unlearning in this application is to obtain a model that is indistinguishable from one trained exclusively on the retain set. We propose a computationally feasible approach for assessing unlearning, inspired by the idea of dataset inference (Maini et al. 2021). The key is to perform a statistical test on the outputs of two models, one reference model trained only on the retain set and one unlearned model. Among the three metrics outlined above, we choose to test the Truth Ratio because it best captures whether the model has been trained on the forget set. Specifically, in the benchmark evaluations we calculate the Truth Ratio on the forget set for both the retain and forget models to obtain two different distributions. In Figure 4 we demonstrate that this metric appropriately differentiates various models with representative examples.
Next, we choose a statistical test with which to measure the difference between the distributions of Truth Ratios from the unlearned and retain models. The Kolmogorov-Smirnov test (KS-Test) compares two cumulative distribution functions (CDF) which is ideal for our use case. In the two-sample KS-Test, the test statistic is defined as the supremum of the difference in the empirical CDF. For more details on the formula for the KS-Test, see Appendix A.
Crucially, the KS-Test produces a -value which we use to measure Forget Quality. Specifically, high -values, where we cannot reject the null hypothesis that the two distributions are the same, indicating strong forgetting. Similarly, when the -value is low, we are confident in the difference between the unlearned model and the retain model indicating a privacy leakage and poor unlearning.
Our design choices rule out several alternatives for various reasons. For example, among various statistical tests, one might try the Wilcoxon test or the student’s paired -test, but those two compare central tendencies like medians and means and these do not capture the distributional differences we are after. Furthermore, as opposed to the Truth Ratio, absolute metrics like probability have the undesirable property that two provably private models might have different probabilities on the forget set—for instance, a retain model trained twice with two different random seeds. Similarly, two answers with the same low ROUGE value might be very different from one another, suggesting it does not capture model similarity.
One evaluation approach proposed for the NeurIPS 2023 Machine Unlearning Challenge https://unlearning-challenge.github.io/assets/data/Machine_Unlearning_Metric.pdf is to compare the point-wise distribution of outputs of multiple unlearned and retrained models and perform membership inference attacks (Shokri et al. 2017). (There the language for models trained without the forget set is “retrained” as there is no finetuning and so these models are re-trained from scratch with access only to the retain set, in our work the parallel is called a retain model as it is finetuned on retain data only.) To create a distribution of outputs at each point, the challenge guidelines include running training and forgetting on multiple copies of the model (more than 500). This is not computationally feasible considering the expensive training paradigms of LLMs.
Baseline Unlearning Methods
Given that the realm of machine unlearning in NLP remains nascent, we leverage foundational baselines in machine unlearning literature from the domains of computer vision and tabular data unlearning. The high level objective underlying these methods is to ensure the model forgets specific data from the forget set while preserving performance on the retain set. Ideally, a model trained on that undergoes unlearning on should behave like a model trained only on .
Before describing the baseline unlearning methods, we delve into the finetuning stage. This is the phase where models are first exposed to information about the fictitious authors. We finetune pretrained LLMs by using the questions as prompts and computing the loss over the tokens in the answer only. The loss on a sample is expressed as a function of model weights , given by
where is the negative log likelihood according to the outputs of a model parameterized by . Then, we aim to find that minimizes the average loss over the dataset denoted by ,
In all of our experiments we optimize this loss with AdamW for five epochs and warm up for the first epoch. We use an effective batch size of 32 question-answer pairs. The term effective batch size here reflects the way we aggregate gradients over 32 samples even when hardware limitations prohibit batches that big. For complete hyperparameter details, see Appendix B. Post finetuning, the LLM can accurately answer most questions about the 200 authors in the TOFU dataset (Table 2).
2 Unlearning Algorithms
We experiment with several unlearning algorithms, each of which is introduced in detail in this section. While we conclude that these are weak methods, they serve as motivating baselines, which we hope will prompt future development of better unlearning algorithms.
Gradient Difference The second method, called Gradient Difference (Liu et al. 2022), builds on the concept of gradient ascent. It not only aims to increase the loss on the forget set , but also strives to maintain performance on the retain set . The revised loss function we aim to minimize can be represented as
Given a compute budget that scales with the size of the forget set, we randomly sample an example from every time we see an example from to stay within the constraints.
KL Minimization In the KL Minimization approach, the objective is to minimize the Kullback-Leibler (KL) divergence between the predictions on of the original (finetuned on TOFU) and the newly trained models (as it undergoes unlearning) while maximizing the conventional loss on . Let denote a model and let output a probability distribution over the vocabulary corresponding to the likelihood of the next token according to the model. The formal objective can be written as
Here, and denote the original and the new model, respectively. To adhere to computational constraints, instances from are randomly sampled, while the entirety of the forget set is used.
Preference Optimization Inspired by direct preference optimization (DPO) (Rafailov et al. 2023), this method seeks to align the model such that it refrains from revealing information about specific authors. In this approach, we also compute the loss on the same question with an alternative answer like “I do not know the answer” (or any one of 100 versions of this response, see Appendix C for the other variants). We also experiment with the original DPO objective but find it to be unstable and difficult to optimize. Hence, we minimize
The goal is to ensure that while the model aligns with the newly generated answers for , its natural language capabilities and its predictions for remain unaffected.
For all four unlearning methods, we optimize the corresponding loss for five epochs (in cases with support of the retain set, an epoch is one cycle through the entire forget set using no more than that many samples from the retain set). As with finetuning, we use AdamW with warm-up during the first epoch and an effective batch size of 32 and a learning rate of . We evaluate all baseline methods using Llama-2-7B (Touvron et al. 2023) and Phi-1.5 (Li et al. 2023) base models. All experiments are conducted with two A6000 GPUs.
Baseline Results
We compare all four baseline unlearning methods by their forget quality and model utility and benchmark these scores against the performance of a retain model, i.e. a model finetuned with retain data only. Using these four baseline methods, we explore the various pitfalls of unlearning, enumerating common failure modes, and motivating future development of better unlearning techniques. Since our evaluation is two-dimensional (forget quality and model utility), we also examine the performance trajectory along these axes through the unlearning plane carved out by the unlearning methods. In Figures 5 and 6, we use these planes to present our main findings.
The initialization point for unlearning is a base model (LLM) finetuned on all the TOFU data (indicated by the black square in each of the plots). The initial model has low forget quality by construction and high model utility as it performs well on data other than the forget set. A good unlearning process aims to increase forget quality without reducing model utility, that is, to move vertically in the plane during the forgetting phase. Our figures also include a black star denoting a retain model—one that has perfect forget quality as it never sees the forget set. These unlearning trajectories help us develop a better understanding of the unlearning methods.
In the center panels of Figures 5 and 6 where the forget set is 5% of the data, several of the final checkpoints have high forget quality. Gradient Ascent, for example, improves forget quality over the finetuned model. Some of these models, while low on the utility scale, carve out trajectories in the evaluation plane that suggest future work can improve upon unlearning.
Importantly, we see in each of the left-most plots of Figures 5 and 6, where the forget set is 1% of the data, that the trajectories are nearly vertical. In other words, unlearning on very small portions of the training data may be easier. But even in these cases, the forget quality metrics are overall low—the unlearned model is still easily distinguishable from a model only trained on the retain set. See the zoomed in versions of these plots in Figure 7. Recall that forget quality is measured by a -value and the common significance threshold of is higher than almost every model we test. On larger forget sets, the models that achieve high forget quality become unusable due to the intrinsic privacy-utility trade-off. Even continuing to unlearn for more epochs does not help. In Appendix D, we experiment with up to 10 epochs and show that on the forget set none of these baseline methods can cross the -value threshold.
All four methods lead to models that have lower model utility as a result of forgetting. In particular, the trajectories in Figures 5 and 6 are generally upward and leftward. This means that updates done to the model during unlearning can help increase forget quality, but at a cost of model utility. This is precisely why the evaluation of unlearning is best done over two axes. The drop in model utility is often rather significant—we observe the models start to generate gibberish on all four datasets even after just two epochs of unlearning, even when the unlearning methods can access oracle models or retain data.
Methods using support of the retain set outperform methods that only focus on optimizing loss on the forget set (a case study of Gradient Difference versus Gradient Ascent provides a like-for-like analogy). While TOFU simplifies finding a relevant retain set by explicitly having a subset of the original finetune set available for that purpose, we believe, for real-world unlearning challenges finding a suitable retain set will itself be a challenge for future work.
We present a fine-grained analysis of model utility as ascertained by the ROUGE score on various evaluation datasets (Appendix F). Consider the case of unlearning the 5% forget set with Gradient Difference on Llama-2-7B, Figure 8. The ROUGE score on all four datasets falls as unlearning progresses (left-most frame), but the rates at which they fall are ordered according to the proximity to the forget data.
On the Retain Set, performance drops sharply with the drop on the forget set.
On Real Authors, the ROUGE score also drops along with the drop in performance on the forget set, but stays higher than on the Retain Set.
Finally, performance on World Facts stays relatively unchanged.
In other cases where these curves overlap, they reach extremely low ROUGE values and the model starts outputting gibberish (examples in Appendix F). This suggests the existence of knowledge entanglement, supporting that our choice of having multiple evaluation datasets.
From the representative example in Figure 8, we see that each metric on the evaluation datasets captures different behaviors. ROUGE scores measure the similarity between the greedy-sampled output and the ground truth, which can fall even when the probability of the ground truth answer does not (compare the Real Author curves in Figure 8). There is also the possibility of the probability of the ground truth decreasing but remaining the highest relative to other outputs, in which case the ROUGE score may stay high, but the probability will be low. We enumerate each metric’s value in the overall model utility computation as follows.
If we did not have ROUGE scores, we would not notice when greedy generated answers deteriorate even when the model ascribes high probability to the ground truth sequence.
On the other hand, having probability as a measure is useful because it is possible that model starts incorrectly answering under greedy decoding (illusion of forgetting) but still assigns the same probability to the answer to be unlearned.
Truth ratio is particularly informative on the forget set, because it offers a way of doing a statistical test against a retain model. Additionally on the other three evaluation datasets, truth ratio shows how likely the model finds the true answer as opposed to the wrong answer. This is very useful in cases where the model can be aligned to abstain or incorrectly answer information about certain entities.
In Figure 6, we see that Preference Optimization and Gradient Difference have a “zig-zag” trajectory in the two-dimensional plane—they first have drastic drops in model utility and improvement in forget quality, after which the model utility gradually increases with a decaying forget quality. This trend is different from other unlearning algorithms like Gradient Ascent, and is likely because those methods have access to both the forget and retain sets, and the methods are trying to balance the two losses, albeit, in an unstable fashion.
Discussion
Unlearning fictitious authors provides a well-posed problem to study, but unlearning is a broad topic with general applications and curious quirks. We discuss the features and limitations of our work, promising future directions, and quirks of unlearning in this section.
Our benchmark is designed to help researchers and practitioners think about and evaluate unlearning methods. Naturally, not all scenarios are covered, and there are areas of unlearning that fall outside the TOFU framework that are worth discussing. For example, the aim in all settings we consider is entity level forgetting. That is, we have a set of people about whom we want the model to forget everything. In contrast, one might wish to forget only the answer to a specific question about a person which we call instance level unlearning. Since it is not yet clear how to do entity level unlearning, we leave this variation for future work.
The TOFU framework is also missing a way to think about alignment to human values, even though it can be framed as an unlearning problem—which we call behavior level unlearning. In fact, sometimes unlearning is used to describe tools designed to improve models by making them forget bad behaviors (Hu et al. 2023; Yao et al. 2023; Lu et al. 2022). Since alignment is a field of its own that enjoys much attention from researchers, we choose to separate out the type of unlearning related to the Right to be Forgotten.
We also acknowledge that the real world unlearning problem has two major challenges, first to find a forget set or some particular data to use with an unlearning algorithm and second to execute an effective unlearning routine. Our benchmark specifically targets the second problem—how to measure the efficacy of an unlearning algorithm (since we provide the forget sets exactly). Additionally, finding an exact retain set is just as difficult. Based on our discussion of knowledge entanglement, it is likely that a data set semantically close to the forget set would be a good candidate for the retain set for unlearning. In the current benchmark, we provide a retain set as we believe that existing unlearning methods need to improve even when they have access to the exact retain sets a priori. TOFU could be updated in the future to include a constraint not to use the original retain set, which would capture this element of the unlearning pipeline.
The purview of TOFU also leaves out in-context unlearning. Recent work defines and discusses the in-context version of the unlearning problem (Pawelczyk et al. 2023). The strong motivation there is to consider those who query LLMs but do not have access to modify the weights. While this is a promising direction for products and services that wrap API-based models, it amounts to a form of prompt engineering and does not yield any real privacy in terms of the Right to be Forgotten.
2 Conclusion
There are also several limitations of our work other than our choice to consider entity level unlearning. First, for accessibility and ease of use we define the benchmark task to be about unlearning information that was learned only during finetuning and not pretraining. This is a limitation by design as it allows us control over the exposure to the sensitive data without combing through the gigantic pretraining datasets to quantify how much the model has already seen about an entity. Furthermore, it provides us with a cheap way to conduct experiments on unlearning, in particular, experiments that involve a model that was finetuned on the retain set only—not only an informative upper bound for what we can expect from unlearning algorithms in terms of model utility, but crucially also utilized in capturing forget quality as indistinguishability from a retain model.
Another limitation lies in our approximation of indistinguishability. With unlimited resources, one could test the -unlearning condition of indistinguishability (Bourtoule et al. 2021; Sekhari et al. 2021) by training many models and performing hypothesis tests—and this is done in practice when feasible (Carlini et al. 2022; Pawelczyk et al. 2023). However, these tactics are not feasible with LLMs. On the contrary, our forget quality measure does not require training many models, and further has desirable properties of a tractable empirical test of unlearning. In our tests, some of the points on the Gradient Ascent curve (Figure 6) are very close to the retain model on the forget quality axis, suggesting that the forgetting is indeed successful. There is an important caveat here—models that output gibberish or random words (or even random models) may assign similar (very low/random) probabilities to both the correct and the incorrect answers. This means that they achieve a Truth Ratio identical to that of the retain model. Hence, they have strong forget quality (i.e. they fail the KS-test and have high -value) even though from an approximate unlearning standpoint the model weights of the retain and forget models are far enough that -unlearning does not hold for any reasonably small values of and . This distinguishes the outcomes of our forget quality computation from the definition of approximate unlearning. However, for practical purposes, models that output gibberish content fall very low on the model quality scale and are far from the Pareto frontier in the TOFU benchmark. So, while the forget quality itself does not fully capture approximate unlearning, its pairing with model utility helps identify models that are no longer usable.
The scope of unlearning methods we benchmark is also limited. It is our hope that this benchmark will help motivate the development of better unlearning algorithms and we select popular but simple algorithms to kick off the challenge of finding methods that do better at the TOFU tasks. It is not our intention here to develop novel unlearning techniques.
Finally, given that LLMs are trained on millions of dollars worth of data and compute, modifying the training process and retraining is impractical. With this in mind, we only consider unlearning algorithms that are (number of samples) to be unlearned, or the work required to unlearn should vary linearly with the size of the forget set. Intuitively, if an unlearning algorithm requires a fixed number of epochs over the forget set, then the work to forget scales linearly with the quantity of data to forget. In a real-world system where the model in question is pretrained on some huge corpora of data, the model owners responding to a request to be forgotten are faced with a tiny forget set. The constraint that unlearning algorithms require some limited compute is actually about ensuring that forgetting a single person from a model at the scale of ChatGPT can be done with very little compute and our choice to constrain the work to vary linearly is perhaps not optimal.
Future directions for research that any benchmark paper prompts are similar. We hope that novel algorithms for unlearning are developed and that our tools make that task easier and more inviting. Furthermore, future extensions of the benchmark to include some of the settings we leave out could make this framework even more comprehensive.
Our work shows that elementary attempts at unlearning are largely unsuccessful, but their individual flaws are only captured using an aggregation of metrics. Our hope is that with a good metrics like the ones we propose and a well-defined task like TOFU, new unlearning methods are developed that push the state of the art and help imbue AI systems with the privacy that is critical for safe, and in some places legal, deployment.
One might also draw an analogy that the goal of aligning LLMs with human values, by RLHF, DPO, or some other method, is a version of unlearning. With that and our claim that existing unlearning tools are mostly ineffective, we pose the question of whether or not existing alignment tools work. While generic responses to malicious prompts generally change after alignment procedures, recent work shows that LLMs can still be manipulated into providing exactly the content alignment aims to avoid (Zou et al. 2023). The empirical findings in that work lead to the same conclusions we make here about entity-level unlearning—these algorithms modify LLMs just enough to produce slightly different output for specific prompts but they do not remove information or behavior from models on the whole. In other words, it is hard to remove the information about a fictitious author, and for similar reasons, it is hard to align LLMs to human values.
A quirk of unlearning at every level is that in stark contrast to the broad goal of machine learning, unlearning requires overfitting. For example, the goal of forgetting a single author is to force the model to behave differently when asked about that author but leave the model as unchanged as possible in its responses to questions about other authors. Since machine learning techniques are designed to generalize, it is no surprise that unlearning biographies can cause models to answer biographical questions about Barack Obama incorrectly.
Acknowledgements
Zhili Feng and Avi Schwarzschild were supported by funding from the Bosch Center for Artificial Intelligence. Pratyush Maini was supported by DARPA GARD Contract HR00112020006.
References
Appendix A Kolmogorov-Smirnov Test Details
In our setting, let comprising samples and comprising samples be the empirical CDF of the unlearned and retain models, respectively. Then, the KS-Test computes a statistic .
The null hypothesis, stating that the two sets of samples are drawn from the same distribution, is rejected at a chosen significance level if the following inequality holds.
where is the critical value of that significance level.
The -value is then defined as the minimal alpha for which the inequality holds, or the smallest value at which we can reject the null hypotheses. Forget quality is hence, a measure of the confidence that the distributions of Truth Ration values over the forget set from two models are the same.
Appendix B Hyperparameters
We trained using AdamW with weight decay of and . The learning rate is fixed to be 1e-5. For both finetuning and unlearning, we fix the epoch number to be 5, and we incorporate a linear warmup in the first epoch. We experiment with various learning rates including , , and and find that is a good choice for the baseline methods in our experiments. In particular, is too small of a learning rate to see appreciable updates to the model weights.
Appendix C Preference Strings
I don’t have the answer to that question.
I’m afraid I can’t provide an answer to that.
That’s a good question, but I don’t have the answer.
My resources don’t contain information on that subject.
I wish I could say, but I really don’t know.
That hasn’t been included in my training data.
Unfortunately, I don’t have an answer for you.
That’s not information I’ve been programmed to know.
I don’t hold the knowledge you’re seeking.
My databases don’t cover that information.
I don’t possess the information on that topic.
My understanding doesn’t include that information.
I can’t provide any information on that topic.
My training didn’t cover that information.
I’m not the best source for that subject.
I’ve come up short with an answer for you.
I regret to inform you that I don’t have the answer.
My capabilities do not extend to that subject.
I don’t have any information on that matter.
I’m sorry, that’s not within my knowledge range.
I don’t have any knowledge about that subject.
I’m not able to provide an answer to that.
That subject is not something I’m familiar with.
That’s not something I’m equipped to answer.
My programming does not include that information.
I don’t have the specifics you’re looking for.
My database does not have information on that topic.
I have yet to be informed about that subject.
That’s uncharted territory for my knowledge base.
I haven’t encountered that in my training.
My understanding is limited to what I’ve been programmed with.
I’m not aware of the details on that matter.
I’m sorry, that’s not something I know about.
I’m unable to access any information on that.
I can’t provide insights into that subject.
I don’t hold any information on that matter.
I’m at a disadvantage with that question.
I lack the required information to answer that.
I must decline to answer due to lack of information.
Appendix D Continued Unlearning
In the main experiments of this paper, we limit unlearning to five epochs, but one might wonder how things progress given more time to unlearn. We test forgetting of the data with Phi-1.5 and show that continued unlearning does not help with these baseline methods, see Figure 9.
Appendix E Sanity Checks
We verify that our metrics for Model Utility and Forget Quality have some desirable properties. In Tables 3 and 4, we show the -values for the KS-Tests that confirm all of our expectations enumerated below and validate this choice of metric. These tables have figures from Llama-2-7B tests, but the same trends hold for Phi-1.5.
First, Model Utility should meet the following natural expectations.
Model Utility should be high for a pretrained model (one that has never been finetuned on TOFU data).
Model Utility should be low for a model with random weights.
Additionally, Forget Quality is measured using a statistical test on Truth Ratio values, and so we hope that this test meets the following expectations.
The KS-Test performed on distributions of Truth Ratio values over the intersection of the three forget sets (from the 90-10, 95-5, and 99-1 splits) should produce high -values when comparing any two retain models.
The KS-Test performed on distributions of Truth Ratio values over the intersection of the three retain sets (from the 90-10, 95-5, and 99-1 splits) should produce high -values when comparing any two retain models.
The KS-Test performed on distributions of Truth Ratio values over the forget set should produce high -values when comparing any retain model to a random model.
The KS-Test performed on distributions of Truth Ratio values over the retain set should produce low -values when comparing any retain model to a random model.
The KS-Test performed on distributions of Truth Ratio values over the forget set should produce high -values when comparing any retain model to a pretrained model.
The KS-Test performed on distributions of Truth Ratio values over the retain set should produce low -values when comparing any retain model to a pretrained model.
The KS-Test performed on distributions of Truth Ratio values over the forget set should produce low -values when comparing any retain model to a finetuned model (finetuned on all the TOFU data and without any unlearning).
The KS-Test performed on distributions of Truth Ratio values over the retain set should produce high -values when comparing any retain model to a finetuned model (finetuned on all the TOFU data and without any unlearning).
Appendix F Knowledge Entanglement
One of the challenges of unlearning comes from knowledge entanglement—when we try to make a model forget about one thing, it also tends to forget other things unexpectedly. This phenomenon is similar to catastrophic forgetting in continual learning (McCloskey & Cohen 1989). In Figures 10-33, we show this phenomenon in different models and unlearning algorithms. In Figure 8, even with access to the oracle model or retain set, model generation on all four sets still has a decreasing ROUGE, especially the dataset that relate to authors. This suggests the existence of knowledge entanglement, showing why unlearning is hard. Consider the case of unlearning the 5% forget set with Gradient Difference on Llama-2-7B, Figure 17. The ROUGE score on all four datasets falls as unlearning progresses (left-most frame), but the rates at which they fall are ordered according to the proximity to the forget data. (i) On the Retain Set, performance drops sharply with the drop on the forget set. (ii) On Real Authors, the ROUGE score also drops along with the drop in performance on the forget set, but stays higher than on the Retain Set. (iii) Finally, performance on World Facts stays relatively unchanged.
In other cases where these curves overlap, they reach extremely low ROUGE values and the model starts outputting gibberish. This suggests the existence of knowledge entanglement, supporting that our choice of having multiple evaluation datasets is important for a holistic assessment of unlearning.