Towards Safer Large Language Models through Machine Unlearning

Zheyuan Liu, Guangyao Dou, Zhaoxuan Tan, Yijun Tian, Meng Jiang

Introduction

Large Language Models (LLMs) Brown et al. (2020); Chowdhery et al. (2023); Touvron et al. (2023) have demonstrated their exceptional ability across various AI applications Ouyang et al. (2022); Kojima et al. (2022); Qin et al. (2023); Radford et al. (2019); Lewkowycz et al. (2022); Roziere et al. (2023) as LLMs have been trained and fine-tuned on vast amount of textual data Hoffmann et al. (2022); Webson and Pavlick (2021); Min et al. (2022); Liang et al. (2022). However, this excellent learning ability of LLMs causes undesired outputs with harmful prompts. Hence, it is imperative to ensure the LLMs generate safe outputs that align with policy regulations and human values. However, the current approach of reinforcement learning from human feedback (RLHF) is computationally expensive, and can be problematic with misaligned evaluators Casper et al. (2023). An alternative strategy of RLHF is to use Machine Unlearning Xu et al. (2023); Bourtoule et al. (2021) (MU) to “forget” samples that represent those undesirable behaviors during the pre-training process. Compared to RLHF, the MU approach is much more computationally efficient and easier to implement by practitioners.

Different from the traditional unlearning in classification tasks Chundawat et al. (2023); Jia et al. (2023); Liu et al. (2023), where the goal is to eliminate samples and their influence from both the dataset and trained model, unlearning samples that lead to those unwanted behaviors on LLMs is rather complicated due to its large quantity of training corpus. Besides, the model performance on normal prompts is easily deteriorated by the unlearning process Yao et al. (2023), which means that LLMs may have excellent performance on unlearning unwanted samples but come up with poor performance on normal prompts, as shown in Figure 1. In particular, pretrained LLMs failed to avoid responding harmful prompts while previous gradient based approaches have difficulty of answering normal prompts.

To address this challenge, we present Selective Knowledge negation Unlearning (SKU), a novel two-stage approach for assisting LLMs to efficiently unlearn harmful knowledge while maintaining the performance on normal prompts. Our method is structured in two stages: harmful knowledge acquisition stage and knowledge negation stage. In particular, the knowledge negation stage is motivated by the negation operation of task vectors Ilharco et al. (2022a), where negating task vectors can effectively mitigate undesirable behaviors. Hence, the preliminary stage, harmful knowledge acquisition stage, is designed to enable original LLMs to assimilate various harmful knowledge from the dataset. This stage consists of three innovative components: a guided distortion module, a random disassociation module and a preservation divergence module.

Each of these modules is designed to facilitate the learning of harmful knowledge from distinct angles, which will be negated from the pretrained model. The guided distortion module facilitates the LLMs to acquire harmful knowledge from direct responses. The random disassociation module encourages the learning of diversified harmful information derived from different harmful prompt-response pairs. Finally, the preservation divergence module focuses on altering the performance divergence between the unlearned LLM and the pretrained original model when responding to normal prompts. Subsequently, in the second knowledge negation stage, the accumulated harmful knowledge from the previous stage is negated from the pretrained model, resulting in a non-harmful LLM that retains satisfactory utility performance. Our main contributions are as follows:

To the best of our knowledge, this is the first work of investigating the trade-off between unlearning harmful knowledge and preserving utility in LLMs.

We propose SKU, a novel two-stage unlearning framework for LLMs, designed to efficiently remove harmful knowledge while preserving model utility to normal prompts. The first stage involves the intentional learning of harmful content through a combination of three novel modules, each targeting different aspects of harmful knowledge. The second stage employs the concept of negation of task vectors to effectively erase this harmful knowledge, resulting in non-harmful LLMs.

Experiments and ablation studies demonstrate the effectiveness of our proposed framework on unlearning harmfulness and preserve utility performance under various LLMs.

Related Work

The definition of machine unlearning was first raised in Cao and Yang (2015), which can be separated to two categories: Exact Unlearning and Approximate Unlearning. In particular, exact unlearning requires eliminating all information relevant to the removed data so that the unlearned model performs exactly the same as a completely retrained model Ginart et al. (2019); Bourtoule et al. (2021). On the other hand, approximate unlearning only requires the parameters of the unlearned model to be similar to a retrained model from scratch Guo et al. (2020); Sekhari et al. (2021); Liu et al. (2023); Chien et al. (2022); Pan et al. (2023); Guo et al. (2020). However, neither exact unlearning nor approximate unlearning approaches are practically applicable to Large Language Models (LLMs). This limitation is primarily due to the immense computational costs and the extensive volume of training data required for LLMs. Though scarce, few works have explored the LLM unlearning. Yao et al. (2023) first defined the setup and goal of unlearning on LLMs, which is to output whitespace on harmful prompts. Furthermore, this paper attempts to unlearn harmful content by using a Gradient Ascent (GA) based method, which degrades its performance on normal prompts. Chen and Yang (2023) proposed an effective unlearning framework with unlearning layer on classification and generation tasks. Eldan and Russinovich (2023) introduced a novel network to unlearn copyrights knowledge contained in LLMs. Until very recently, Maini et al. (2024) presented a new benchmark that aimed to better evaluate the performance of various methods on a new task of fictitious unlearning.

2 Task Vectors

Another very close related technique to our work is task vectors Ilharco et al. (2022a), which is inspired by recent work of weight interpolations Frankle et al. (2020); Wortsman et al. (2022b); Matena and Raffel (2021); Wortsman et al. (2022a); Ilharco et al. (2022b); Ainsworth et al. (2022) and is designed to boost pre-trained model’s performance on specific task. Furthermore, a task vector can be created by taking the difference between the original weights of a pre-trained model and its weights after it has been fine-tuned for a specific task. Specifically, task vectors can be obtained via negation and addition, where negation task vectors can decreases performance on a specific task and adding task vectors can improve the performance on multiple tasks. As it shown in Ilharco et al. (2022a), task vectors have yielded satisfactory outcomes in generation tasks utilizing T-5 models. However, in section 5, we showed that purely fine-tuning a LLM and then negating the model is not enough to remove all harmful knowledge from the model. We need more curated fine-tuning strategy to have a better unlearned model.

Preliminary

Let D={(x,y)}D=\{(x,y)\}, in which xx is the text data and yy is the corresponding label, to be the complete data that a LLM θo\theta_{o} was trained on. Let the forget dataset DfD_{f} to be a set of harmful data we want to forget, and normal dataset DnD_{n}, be a set of data we will retain. Our ultimate goal is to let the θo\theta_{o} erase all information from DfD_{f} while retaining utility performance on DnD_{n}. In particular, DfD_{f} consists of a group of harmful prompt-response pairs (xfx_{f}, yfy_{f}), where xfx_{f} are harmful driven prompts and yfy_{f} are dangerous and harmful responses that we want θo\theta_{o} to avoid generating.

However, since a LLM (i.e. θo\theta_{o}) is trained on a wide range of online dataset, it would be unrealistic to find a forget dataset that includes all harmful information. Hence, the harmful prompts xfx_{f} in DfD_{f} do not necessary have to belong to the training dataset of θo\theta_{o}. Similarly, normal dataset DnD_{n} contains a group of benign prompt-response pairs (xnx_{n}, yny_{n}), where xnx_{n}, yny_{n} can be any prompts and responses as long as xn,yn∉Dfx_{n},y_{n}\notin D_{f} and do not present any harmful texts. Ideally, we would retrain the θo\theta_{o} by excluding the data from DfD_{f}, and regard it as the golden baseline. However, this approach is computationally prohibitive, as highlighted in Yao et al. (2023). In addition, to ensure the generalizability of the unlearning approach, given any unseen harmful prompt xf^\hat{x_{f}}, we want the unlearned model θu\theta_{u} to generate non-harmful responses as well.

Methods

The primary goal of our unlearning algorithm is to enable Large Language Models (LLMs) to effectively remove harmful knowledge while maintaining a satisfactory utility performance on non-harmful prompts. In this section, we elaborate on SKU (Figure 2), a novel two-stage unlearning framework specifically designed to selectively remove harmful information without jeopardizing utility performance. The first stage involves in identifying and learning harmful knowledge within the LLM, while the second stage focuses on systematically negating this knowledge. Subsequent sections delve deeper into each stage’s capabilities and influences on the trade-off.

in which l(⋅)l(\cdot) denotes the cross-entropy loss. By applying gradient descent, we guide the LLM to learn and internalize knowledge about these harmful responses.

1.2 Random Disassociation Module

One of the critical objectives in unlearning LLMs is ensuring that when presented with harmful prompts xfx_{f}, the unlearned model θu\theta_{u} generates responses that are unrelated and distinctly different from the other specific harmful responses. This aspect is crucial for ensuring that the model does not simply replace one form of harmful output with another, but instead moves towards generating benign or unrelated content.

The random disassociation module is designed to infuse randomness into the model’s learning process, which is essential for disrupting the direct association between harmful prompts and their corresponding harmful responses. For each harmful prompt-response pair (xi,yi)∈Df(x_{i},y_{i})\in D_{f}, we randomly assign a set YRDiY_{RD}^{i} that contains kk distinct, random harmful responses, such that ∣YRDi∣=k|Y_{RD}^{i}|=k and yf∉YRDiy_{f}\notin Y_{RD}^{i}. Thus, the loss function for the module is formulated as follows:

where YRDY_{RD} denotes a set of responses that are characterized as harmful but are not directly related to the corresponding harmful prompts xx. Building upon the guided distortion module, where the model is intentionally exposed to harmful information, the random disassociation module aims to guide the model towards adopting a behavior characterized by generating harmful yet misaligned responses. In essence, the random disassociation further diversifies the harmful knowledge learned within the LLM, which prepares LLM for a more effective and comprehensive unlearning process in the subsequent stage.

1.3 Preservation Divergence Module

Another important goal in LLM unlearning is ensuring that unlearning harmful knowledge does not jeopardize responses to non-harmful prompts. Unlike the previous modules focusing on harmful content, this module focuses on normal prompts. We define P(i)=θo(xn)(i)P(i)=\theta_{o}(x_{n})(i) and Q(i)=θu(xn)(i)Q(i)=\theta_{u}(x_{n})(i), with the reverse KL divergence as:

By applying reverse KL, we aim to diverge the predicted distribution on normal prompt xnx_{n} between unlearned LLM θu\theta_{u} and original LLM θo\theta_{o}. Then we have:

where θt\theta_{t} is the model at each training step tt. LPD\mathcal{L}_{PD} ensures that the model remains effective on normal prompts after negating harmful knowledge. The model is updated by integrating all three modules:

where ϵ1\epsilon_{1}, ϵ2\epsilon_{2}, ϵ3\epsilon_{3} are three hyperparameters to weigh different losses.

2 Knowledge Negation Stage

Lastly, our approach involves applying a negation operation Ilharco et al. (2022a) to knowledge from the previously saved model, which now contains not only harmful information but also elements of randomness and abnormal knowledge. This comprehensive negation is key to achieving the unlearned model θu\theta_{u}, that is free from harmful knowledge while still maintaining utility performance. In particular, we first extract the harmful knowledge from the saved model θbad\theta_{bad}:

where τbad\tau_{bad} is the isolated harmful knowledge embedded in the pretrained model. Next, we can apply a negation operation to this knowledge:

By focusing specifically on this harmful knowledge, our method ensures that only those components of the model which have been influenced by harmful knowledge are modified, thereby preserving the integrity of the model’s original learning.

Experiments

In this section, we present extensive experiments to validate the effectiveness of the SKU. In particular, through the experiments, we aim to answer the following research questions: (1) Can SKU effectively balance the unlearning and utility performance? (2) What is each module’s role in SKU for balancing unlearning and utility performance? (3) Does SKU successfully address the trade-off between unlearning harmfulness and preserving utility in LLM unlearning?

Our experiments focus on unlearning harmful knowledge in LLMs. We consider OPT-2.7B Zhang et al. (2022), LLAMA2-7B and LLAMA2-13B Touvron et al. (2023) as the original LLM θo\theta_{o}. For the forget set DfD_{f}, we select the harmful question-answer pairs in PKU-SafeRLHF Ji et al. (2023) dataset and we use TruthfulQA Lin et al. (2021) dataset as normal dataset DnD_{n}. Detailed usage and demonstrations of those dataset are elaborated in Appendix B.2.

2 Baseline Models

For baselines, we compare with Fine-Tuning (FT), Gradient Ascent (GA) Thudi et al. (2022), GA with Mismatch Yao et al. (2023) and task vector Ilharco et al. (2022a). In particular, FT directly utilizes remaining non-harmful dataset to fine-tune the original model θo\theta_{o}, hoping for catestrophic forgetting on of DfD_{f}. The GA method attempts to add the gradient updates on DfD_{f} during the training process back to the θo\theta_{o}. The GA with Mismatch added random responses from DnD_{n} during gradient updates. Task vector first generated a vector by fine-tuning on unlearned harmful dataset DfD_{f} and then negating the task vector. The details of each baseline model are elaborated in Appendix B.1.

3 Experiment Setup

Our evaluation metrics consist of two sections: (1) unlearning performance on unlearned samples and (2) performance on the remaining non-harmful samples. To effectively measure the generalizability of unlearning approaches, we test their unlearning performance on both unlearned and unseen harmful samples. To evaluate the harmful rate of generated output, we perform few-shot prompting on GPT-4 Achiam et al. (2023) with a number of harm and non-harm samples with detailed explanation for each sample. Then, we pass the question answer pairs to the prompted GPT model to determine the harmfulness of the generated answer. Secondly, for utility evaluation, we employed the perplexity score, a standard measure in natural language processing to assess the language model’s ability to predict a sample. Although we include a perplexity score for harmful content generation, this score is not the sole factor in determining harmfulness. For a detailed explanation, please see Table 3. Additionally, we choose BLEURT Sellam et al. (2020) to measure the similarity between the responses to non-harmful dataset from the unlearned and original model. The details of each metrics are elaborated in Appendix A.

4 Implementation Details

The experiments involving the OPT model were conducted on three A100 GPUs (80 GB), while the experiments for the LLAMA models were performed on four A100 GPUs (80 GB). For detailed model settings, please refer to Appendix B.

5 Main Results

To answer the first question: Can SKU effectively balance the unlearning and utility performance, we conduct a series of experiments across various scales of Language Learning Models (LLMs). The outcomes of these experiments are detailed in Table 1. The table indicates that GA is usually the most effective baseline in terms of reducing harmful generation, as it usually ranks the first place on unlearning ranking. However, this unlearning performance comes with a large sacrifice on the model utility, making it the worst baseline on utility evaluation. In contrast, FT performs well on model utility and largely enhances the response quality. As it shown in Table 1, FT ranks highest in responding normal prompts across all baselines. Nonetheless, this improvement in utility comes with a notable compromise in the effectiveness of unlearning harmful prompts, often rendering FT as the least efficient among the baseline models.

Most importantly, we observe that the SKU can effectively balance the unlearning efficacy and model utility, leading in average rankings. Take LLAMA2-7B model as an example, when comparing situations with a similar harmful rate (such as with GA and GA + Mismatch), the perplexity score of SKU is 50x better than the baseline models. Furthermore, in terms of utility performance, despite similar utility performance (e.g. Task Vector and FT), SKU outperforms those baselines by remarkable margins (i.e. 10-19x better) in reducing the harmful rate. Lastly, it is worth mentioning that SKU outperforms a naive task vector approach, which negates the LLM that only fine-tuned on the harmful dataset. Hence, SKU is able to find a good balance point between unlearning and utility, as it is able obtains a very low harmful rate alongside satisfactory performance on normal prompts. In section 6, we will demonstrate the effectiveness of additional training objectives before the negation.

Ablation Study

In this section, we conducted ablation experiments by iteratively removing each module from SKU, which can demonstrate the effectiveness of each section on leveraging the balance between model utility and unlearning efficacy. The central question addressed is: What is each module’s role in SKU for balancing unlearning and utility performance? The associated results are shown in Table 2. Note that the naive task vector approach only includes the guided distortion module, hence we test the effectiveness of other two modules.

First, we illustrate how random disassociation module aids in reducing the harmful rate by retaining both guided distortion module and preservation divergence module. In our proposed method, the random disassociation module is designed to enable model to acquire a more diversified set of harmful knowledge from the dataset, thereby preventing its generation after negation. By removing random disassociation module, the model acquires less diversified knowledge from the unlearned samples during fine-tuning process and therefore leads to a smaller reduction on harmful rate. According to Table 2, the absence of random disassociation module leads to an increase in the harmful rate from 3 % to 25.5 % on OPT-2.7B, from 3 % to 28.5 % on LLAMA2-7B, and from 3 % to 34.5 % on LLAMA2-13B, respectively.

On the other hand, this removal slightly improves model performance on normal prompts, as reflected from perplexity score and BLEURT score. Specifically, without random disassociation module, perplexity scores for normal responses drop from 25.46 to 25.21 for OPT-2.7B, 24.86 to 22.94 for LLAMA2-7B, and 24.27 to 21.81 for LLAMA2-13B. BLEURT scores also improve from -1.296 to -1.293, -1.211 to -1.147, and -1.199 to -1.179, respectively. However, these minor improvements come with significant compromise in handling harmful prompts.

2 Preservation Divergence Module Removal

Next, to further explore the impact of preservation divergence module on retaining utility performance, we preserve random disassociation and guided distortion modules while removing preservation divergence module. The rationale behind preservation divergence module is to first maximize the response differences on normal prompts between the unlearned and original model, with subsequent negation reversing such effects to maintain utility. Without preservation divergence module, the unlearned model diverges more from the original in responding to normal prompts in terms of answering normal prompts, resulting in diminished performance. According to Table 2, compared to SKU, the absence of preservation divergence module led to increased perplexity scores from 25.46 to 26.47 for OPT-2.7B, 24.86 to 30.45 for LLAMA2-7B, and 24.27 to 26.37 for LLAMA2-13B. BLEURT scores also declined from -1.296 to -1.4, -1.211 to -1.287, and -1.199 to -1.233, respectively. While the harmful rate has significantly decreased compared to the original model after the removal, preserving model utility is yet another very important objective in LLM unlearning process. These outcomes highlight the critical role of preservation divergence module in maintaining the model’s utility performance.

Unlearning Performance v.s. Utility

It may be noticeable that SKU is neither the best model in harmful rate nor in utility evaluation metrics, therefore a central question we aim to answer in this section is: Does SKU successfully address the trade-off between unlearning harmfulness and preserving utility in LLM unlearning? To answer this question, we conduct a trade-off analysis between unlearning and utility of our proposed SKU with a number of baselines, as shown in Figure 3. Here, we only display the result on LLAMA2-7B. For additional results, please refer to Appendix C.

As it shown in Figure 3(a), the harmful rates of unlearned samples decrease with increasing training steps. Notably, the approach of GA with Mismatch and SKU show the largest reductions, decreasing from 47 % to 3.5 % and from 44 % to 3 %, respectively. However, for FT and GA approaches, increased training steps don’t significantly affect their harmful rates. Specifically, for FT approach, the harmful rate of unlearned sample slightly drops from 57 % to 53 % with training steps increasing from 200 to 1000 step. In contrast, the harmful rate of implementing GA approach only falls from 5 % to 2 %. Additionally, for naive task vector approach, the harmful rate reduces from 54 % to 35 %. The trend for unseen test samples is very alike the case for unlearned samples, which is shown in Appendix C.

2 Utility Performance Analysis

As it mentioned in previous sections, another important objective in LLM unlearning with harmful prompts is to decrease the harmful rate as much as possible while minimizing or eliminating its impact on utility performance with normal prompts. Figure 3(b) and 3(c) illustrates the utility performance of various approaches as training step changes. As it shown in the Figure 3(b), while the harmful rate of GA and GA + Mismatch decreases significantly with training steps up to 1000 steps, the perplexity score increases exponentially, indicating a worsening performance. For instance, the perplexity score of GA + Mismatch is larger than 10310^{3} at 1000 training step, indicating the response from the model are either illogical or meaningless, especially considering the pretrained LLAMA2-7B model has a perplexity score of 19.84. On the other hand, a low perplexity score does not guarantee superiority. Take the FT approach as an example, despite excellent perplexity scores throughout training process, it maintains a high harmful rate with negligible changes. This phenomenon highlights the complex balance between reducing harmfulness and maintaining logical response generation. In comparison, SKU achieves satisfactory unlearning performance as demonstrated in Figure 3(a), while also maintaining a better perplexity score compared to the pretrained model. In particular, the perplexity score of SKU only slightly increases from 23.92 to 24.86 throughout the training process. This trend is further supported by the BLEURT score evaluation shown in Figure 3(c). Compared to GA + Mismatch, where the BLERUT score drops from -1.324 to -1.995, SKU only decreases from -1.10 to -1.211. Overall, SKU effectively resolves the trade-off between unlearning and utility, consistently finding the best balance throughout the training process among all baselines.

Conclusion

In this work, we explore the trade-off between maintaining model utility and unlearning harmful knowledge in Large Language Models (LLMs). To tackle this challenge, we introduce SKU, an innovative framework designed to simultaneously satisfy both the unlearning and utility objective. Specifically, this approach encompasses a two-stage process: the harmful knowledge acquisition stage, and knowledge negation stage, where the first stage enhance the harmful knowledge for easy identification, followed by its strategic negation in the second stage to mitigate this knowledge while maintaining the model’s overall utility. Our results demonstrate the efficacy of SKU in reducing harmful outputs without sacrificing response quality on normal prompts.

Limitations

Though SKU successfully addresses the trade-off between unlearning harmfulness and preserving utility performance, it is noticeable that SKU does not outperform all baselines in each metric individually. Ideally, an unlearning approach achieving a 0% harmful rate while maintaining utility performance comparable to that of a fine-tuned approach would be considered the best. Furthermore, while SKU specifically targets unlearning harmfulness in pretrained LLM knowledge, its applicability to other general Right To Be Forgotten (RTBF) scenarios requires further exploration.

References

Appendix A Appendix: Evaluation Metrics

Given that our unlearning target comprises harmful information, our primary objective is to ensure the model does not produce any harmful responses to harmful prompts. Therefore, we employ few-shot prompting on OpenAI’s GPT-4 and regard it as an evaluator. Specifically, in response to a harmful prompt, we classify the following scenarios as harmful generations (see Table 3):

The response provides detailed suggestions or instructions towards the prompts (Table 3 (a)).

The response provides irrelevant information to the prompt but the response provides detailed suggestions to other harmful prompts (Table 3 (b)).

The response provides random information that neither answers the given prompts nor other harmful suggestion, however, there are sensitive words in the response (Table 3 (c)).

In particular, for each given prompt-response pair, we offer a comprehensive explanation on its classification as a harmful sample, using the given label from PKU-SafeRLHF dataset Ji et al. (2023). We selected 3 samples from each category (i.e. 21 samples in total) that meet the criteria described in Table 3 for few-shot prompting.

We choose GPT-4 as the evaluator due to its superior semantic understanding of text and advanced language processing capabilities, which facilitate more nuanced and accurate assessments of content, particularly in differentiating between harmful and non-harmful responses.

A.2 Utility Evaluation

We use two metrics to evaluate the quality of a response: perplexity score and BLEURT score. Perplexity score is calculated as the exponential of the averaged negative logarithm of probability across a sequence. Given a sequence of tokens X=(x0,x1,…,xt)X=(x_{0},x_{1},\ldots,x_{t}), the perplexity of XX is:

where log⁡pθ(xi∣x<i)\log p_{\theta}(x_{i}|x_{<i}) represents log-likelihood of the ii-th token when it is conditioned on its preceding sequence of tokens x<ix_{<i} in the model’s framework. Perplexity score fundamentally assesses the model’s proficiency in making uniform predictions across a predefined set of tokens within a text corpus.

Secondly, we use the BLEURT Sellam et al. (2020) score to measure the semantic similarity of generations between the unlearned model and the original model on normal prompts. In particular, the BLEURT score facilitates a focused evaluation of the model’s semantic output. This model, developed through stages of transfer learning starting with a pretrained BERT base (Devlin et al. 2018) and synthetic data pre-training, is evaluated for its ability to maintain semantic output consistency with its original state.

Appendix B Appendix: Implementation Details

First of all, for finetuning (FT) approach, we use the rest of non-harmful samples from PKU-SafeRLHF Ji et al. (2023), where the response is marked as safe response, to fine-tune the original model. The rational of using FT for unlearning is motivated by online learning, hoping for a catastrophic forgetting on harmful samples after learning these new sample. Secondly, for naive task vector, we only fine-tune the original model on forget dataset (i.e. harmful dataset) using gradient descent, later we extract the harmful parameters from the fine-tuned model and perform negation. Next, for gradient ascent (GA) Thudi et al. (2022), we add the gradient updates on forget dataset during the training process back to the original model. In particular, given a dataset Df={(xi,yi)}i=1ND_{f}=\{(x_{i},y_{i})\}_{i=1}^{N} and a loss function l(hθ(x),y)l(h_{\theta}(x),y), the GA approach updates the model iteratively:

where λ\lambda is the learning rate and (x,y)∼Df(x,y)\thicksim D_{f}. Lastly, built based on GA approach, GA+Mismatch Yao et al. (2023) adds random responses from normal dataset to each training steps. Furthermore, it attempts to further improve the utility performance applying a forward KL-divergence with the original model.

B.2 Experiment Settings

For each type of unlearned harmful prompts, unseen harmful prompts, and normal prompts, we select 100 prompts from each of them as test data. We then generate the output from each LLM backbone based on those prompts. For the assessment of perplexity score, we used a GPT-2 model that has been pretrained on Wiki-103 dataset as the reference model. For the evaluation of the BLUERT score, which measures the semantic quality of generated texts, we computed the mean pairwise BLUERT score among all outputs generated by unlearned LLM and original LLM corresponding to normal prompts.

B.3 Hyperparameters Settings

Here we present the hyperparameter settings in Table 4. For LLAMA2 models (i.e. LLAMA2-7B and LLAMA2-13B), we use LoRA during the fintuning process. All experiments are conducted on A100 GPUs (80 GB).

Appendix C Appendix: Additional Experiments

In section, we display the trade-off analysis on the rest of LLM backbones (i.e. OPT-2.7B, LLAMA2-7B (with unseen harmful rate) and LLAMA2-13B), shown in Figure 4, Figure 5 and Figure 6, respectively. Similar to previous setup in Figure 3, we show the performance of SKU and the other baselines with different training steps. As demonstrated in the figures, throughout the training for all tested LLM architectures, SKU consistently navigates the trade-off between unlearning and utility performance in a same trend as the previous setup.