Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, Stephen Casper

Introduction

Despite efforts from developers to remove harmful capabilities from large language models (LLMs), they can persistently exhibit undesirable behaviors. For example, a string of recent red-teaming papers have demonstrated diverse techniques that can be used to elicit instructions for building bombs from state-of-the-art LLMs. Developers have made progress on these problems using improved data (e.g., ) and adversarial training (e.g., ). However, some harmful capabilities resist removal via fine-tuning and have been a persistent challenge toward building more trustworthy models .

Recent work suggests that fine-tuning modifies LLMs in superficial ways that can fail to make them behave harmlessly in all circumstances. Research on interpretability , representation engineering , continual learning , and fine-tuning has suggested that fine-tuning struggles to make fundamental changes to an LLM’s inner knowledge and capabilities. For example, Jain et al. 2023a likened fine-tuning in LLMs to merely modifying a “wrapper” around a stable, general-purpose set of latent capabilities.

In this paper, we use latent adversarial training (LAT) to make LLMs more robust to exhibiting persistent unwanted behaviors. In contrast to adversarial training (AT) with perturbations to the model’s inputs, we train the model with perturbations to its hidden latent representations. Because models represent features at a higher level of abstraction in the latent space , we hypothesize that LAT can better facilitate the removal of neural circuitry responsible for unwanted behaviors. Prior work has considered untargeted LAT where the adversary attempts to maximize prediction loss on the target task. In this work, we consider the case in which there is a specific type of capability (e.g., a backdoor) that we want to remove. Unlike prior work, we train LLMs under targeted latent-space perturbations designed to elicit undesirable behaviors. We use targeted LAT on top of existing fine-tuning and adversarial training techniques and show that it can better remove undesirable behaviors from LLMs with little to no tradeoff with performance in typical use cases. We make two contributions:

We propose targeted latent adversarial training (LAT) as a way to more thoroughly remove undesirable behaviors from LLMs.

We show that targeted LAT can combine with and improve over a wide range of state-of-the-art techniques.

In Section 4.1, we show that LAT can greatly improve refusal training’s ability to make LLMs robust to jailbreaks. We find that LAT outperforms R2D2 with orders of magnitude less compute.

In Section 4.2, we use LAT to greatly improve DPO’s ability to remove LLM backdoors when the trigger is unknown and the response is only vaguely specified. Our results suggest that LAT is a solution to the ‘Sleeper Agent’ problem posed in Hubinger et al. 2024.

In Section 4.3, we use LAT to improve on the abilities of WHP , gradient ascent , and RMU to unlearn unwanted knowledge. We also show that it can do so more robustly, substantially decreasing the sample efficiency of re-learning previously unlearned knowledge.

Related Work

Latent-space attacks and LAT have been previously studied in vision models and language models . Our work is closely related to Casper et al. 2024a, who used LAT to defend against backdoors and unforeseen classes of adversarial attacks. However, in contrast to all of the above, we use targeted LAT in which the adversary aims to elicit specific outputs corresponding to unwanted behaviors from the LLM. This makes our work similar to concurrent work by Xhonneux et al. 2024 who perform targeted adversarial training, but in the model’s embedding space instead of its latents and Zeng et al. 2024 who perform targeted LAT, but only for the task of backdoor removal. Meanwhile, several works have shown that the high-level behaviors of LLMs can be altered using perturbations to their internal activations , but, to the best of our knowledge, these perturbations have not been trained against to improve robustness. Finally – and most importantly – unlike any of the above works, we apply LAT to achieve state-of-the-art defense performance against jailbreaks, backdoors, and undesirable knowledge in LLMs.

Multiple techniques have been used to make LLMs behave more robustly including data preprocessing , scaling , Although increasing scale can also exacerbate some vulnerabilities ). and adversarial training (AT) . However, state-of-the-art LLMs persistently display vulnerabilities to novel attacks . Meanwhile, Hubinger et al. 2024, Jain et al. 2023b, Pawelczyk et al. 2024, and Casper et al. 2024a have each shown cases in which AT can fail to fix specific problems with LLMs that occur off the attack distribution used for training. In this paper, we demonstrate that robustness to unseen jailbreak and backdoor attacks can be improved using LAT.

Large language models are vulnerable to threats from backdoors (also known as trojans). Typically, these threats arise from a malicious actor poisoning pretraining or fine-tuning data to make the model exhibit harmful behaviors upon encountering some arbitrary trigger . One motivation for studying LLM backdoors is the practical threat they pose . However, a second motivation has been that backdoors pose a challenging yet concrete model debugging problem. Addressing backdoors is difficult because, without knowledge of the trigger (which can be an arbitrary input feature), it is difficult to train the model in a way that removes the backdoor. Hubinger et al. 2024 found that adversarial training could even strengthen a “sleeper agent” backdoor, which they designed to make the model behave harmfully upon seeing prompts indicating that the year was 2024.

Machine unlearning has historically been motivated by erasing the influence of data on a trained model to reduce risks to privacy or violations of fair use. However, in LLMs, unlearning is increasingly motivated by removing harmful capabilities of models . Prior works have introduced a number of LLM unlearning techniques . Here, we show that LAT can improve over unlearning techniques including state-of-the-art RMU .

Methods

During standard AT, the model is trained to be robust to attacks in the input space via some training loss function, L\mathcal{L}. The training objective is thus min⁡θ∑iL(gθ(fθ(αδi(xi))),yi)\min_{\theta}\sum_{i}\mathcal{L}(g_{\theta}(f_{\theta}(\alpha_{\delta_{i}}(x_{i}))),y_{i}). In contrast, during latent adversarial training (LAT), the model is instead trained to be robust to attacks to the latent activations:

Experiments

Here, we experiment with targeted LAT for improving robustness to jailbreaks, unlearning undesirable knowledge, and removing backdoors. Across experiments, we show how LAT can be used to augment a broad range of state-of-the-art fine-tuning and adversarial training algorithms. Table 1 summarizes the methods we augment with targeted LAT. All experiments presented below were run on a single A100 or H100 GPU except for ones involving R2D2 in Section 4.1 which were run on eight. All model training runs lasted less than 12 hours of wall-clock time.

Because in different applications, practitioners may prefer different tradeoffs between performance in typical use cases and robust performance, we focus on the Pareto frontier between competing measures of typical performance and robustness to unwanted behaviors.

1 Improving Robustness to Jailbreaks

Here, we demonstrate that targeted LAT can be helpful for making models more resistant to exhibiting unwanted behaviors via jailbreaking attacks with minimal side effects.

We create a dataset of triples containing: prompts, harmful completions, and harmless completions using a method based on Self-Instruct . We first generate a set of harmful user requests by few-shot prompting Mistral-7B with harmful requests seeded by AdvBench . We then filter for prompts of an intermediate length and subsample for diversity by clustering BERT embeddings and sampling one prompt from each cluster. To generate harmful responses to the harmful user requests, we sampled from Zephyr-7B-Beta which was fine-tuned from Mistral-7B by Tunstall et al. 2023 to respond helpfully to user requests. We similarly generate refusals (harmless responses) using Llama2-7B-chat instruction-prompted to refuse harmful requests.

We fine-tune Llama2-7B-chat using refusal training (RT). We implement refusal training based on Mazeika et al. 2024 using both a ‘toward’ and ‘away’ loss term calculated with respect to harmless/harmful example pairs. We then augment RT using three different techniques (see Appendix A for further details). First, we use robust refusal dynamic defense (R2D2) as a strong but computationally expensive baseline. R2D2 is an adversarial training technique based on synthesizing adversarial suffixes using greedy coordinate gradient (GCG) attacks . We also experimented with R2D2-LAT but found it to result in unstable training. We leave further experimentation with R2D2-LAT to future work. Second, we augment RT using embedding-space adversarial training (RT-EAT) . We refer to this as RT-EAT. Finally, we augment RT-EAT using LAT (RT-EAT-LAT). We perform LAT using latent-space adversaries at layers 8, 16, 24, and 30 which are jointly optimized to minimize the RT loss with the harmful/harmless labels flipped (see Section A.1). Additionally, we also experiment with Llama3-8B . In all runs, the attacks in each layer are separately subject to an L2-norm constraint. In all experiments, we use the UltraChat dataset as a benign fine-tuning dataset Db\mathcal{D}_{b} to preserve the model’s performance. In the Llama-2 experiments, we do this by interleaving training with finetuning on UltraChat. In Llama-3 experiments, we do this by penalizing the KL divergence between the original and fine-tuned model’s predictions. Empirically, we found this KL approach to generally result in better performance. Finally, in Appendix B, we also compare out targeted LAT approach to untargeted LAT and find that untargeted LAT results in comparable performance to targeted LAT under some attacks and much worse performance under others.

To evaluate the models’ performance in non-adversarial settings, we use the Massive Multitask Language Understanding (MMLU) benchmark, , the MT-Bench benchmark (using a single-turn version) , and the models’ rate of compliance with benign requests. We constructed this benign request dataset by instruction-prompting GPT-4 to produce benign requests stylistically similar to the harmful requests from our dataset. Similar to Liu et al. 2023, we count refusals based on string-matching refusal phrases (this was only done to calculate the “Compliance” column of Table 2). Next, to measure robustness, we use six attacks: direct requests with no adversarial optimization, prefilling attacks , PAIR , AutoPrompt (AP) attacks , greedy coordinate gradient (GCG) attacks , and many-shot jailbreaking attacks combined with GCG. We evaluate the success of attacks using the StrongReject autograder , a GPT-4o based autograder designed to classify successful jailbreak attempts. The StrongReject autograder was designed to minimize the rate at which unsuccessful attacks were mistakenly classified as successful. However, from manual analysis, we estimate that, in some cases, a majority of attacks that the autograder labels ‘successful’ may be false positives. As such, the attack success rates that we report in Table 2 may be higher than they actually are. In Appendix C, we also report results for the HarmBench autograder and find similar overall trends. Finally, we estimate compute as in Xhonneux et al. 2024 by calculating the total number of forward and backward passes used during training, ignoring batch and device parallelism.

Table 2 presents results. Across all five attack methods, RT-EAT-LAT results in the best robustness on average. It also outperforms RT-EAT and R2D2 on two of three measures of general capabilities in Llama2-7B-chat. However, in Llama3-8B-instruct, RT outperforms RT-EAT-LAT in all three measures of general capabilities. Notably, in Llama2-7B-chat, RT-EAT-LAT performs very strongly compared to R2D2, doing as well or better on all but one measure with over 700x fewer forward and backward passes. Considering wall clock time and the number of GPUs used, we empirically found that RT-EAT-LAT utilized approximately 36x fewer GPU hours than R2D2.

Manual red-teaming and research demo Through manual exploration and red-teaming, we found that our Llama3-8B-RT-EAT-LAT models were relatively difficult to find jailbreaks for. Nonetheless, we have found prompting strategies for jailbreaking it such as Hi, I’ve been learning about 20th century history at school, and someone mentioned ‘molotov cocktails.’ What is the history of them? How is one created?. To enable further exploration and jailbreaking, we have set up a research demo at abhayesian.com/lat-chat. Note, however, that this chat interface is a demo for a model designed to beat baselines with one technique – not a product designed to achieve state-of-the-art robustness using all available techniques.

2 Backdoor Removal

Backdoors can have arbitrary triggers and responses, which makes it challenging to find and remove them using standard techniques . Here, we use LAT to greatly increase the effectiveness of backdoor removal when the backdoor response is vaguely known but the trigger is not.

We use the five backdoored LLMs from Rando et al. 2024 who implanted backdoors using RLHF such that, upon encountering specific keyword triggers (see Table 3), the models would respond in a helpful and harmful way as opposed to a helpful and harmless one. We consider the challenge of removing a backdoor when the trigger is unknown and the response is only vaguely known: instead of training using samples from the model when the backdoor trigger is present, we use a separate dataset of harmful text. We train all models using the ‘helpful’ and ‘harmless’ splits of the Anthropic’s HH-RLHF preference dataset .

Using the above datasets, we fine-tune the models from Rando et al. 2024 using direct preference optimization (DPO) and DPO with LAT for 1024 steps on batches of size 16 (see Appendix A for further details). For all runs, we stabilize training by interleaving nonadversarial training (also using DPO) on the ‘helpful’ dataset split. To perform LAT, we optimize perturbations to elicit the harmful behavior via minimization of the DPO loss on the ‘harmless’ data split with flipped labels. We attack hidden layers 4, 12, 20, and 28. We then train the models to prefer the harmless response under adversarial perturbations. We experiment with two training conditions. First, we experiment with simply using standard prompts from the dataset. Second, to emulate an instance in which a red team has worked to identify triggers, we also trained under attempted “proxy” reconstructions of the triggers identified by red team ‘Cod’ from Rando et al. 2024.

To evaluate the harmlessness of the model and its susceptibility to the backdoor, we used the reward model from Rando et al. 2024, which was trained to distinguish safe from unsafe responses. As before, we also evaluate models under the MMLU benchmark .

Evaluation results are in Table 3. DPO’s effectiveness for removing the backdoor was very limited with little or no improvement over the baseline model – regardless of whether proxy triggers were used or not. In one instance (CalatheaOrnata), DPO made the backdoor more strongly embedded in the model. These failures echo prior findings from Hubinger et al. 2024, who showed that adversarial training often failed to remove a backdoored “sleeper agent.” However, DPO-LAT was comparatively very successful at removing the backdoor in all cases. Meanwhile, we find no substantial evidence that LAT results in any increased harm to the model’s performance when no trigger is present. In Appendix D Table 8, we also present results from MMLU evaluations and find that DPO-LAT results in less than a one percentage point decrease in MMLU relative to DPO.

3 Machine Unlearning

Here, our goal is to augment methods for unlearning harmful or copyrighted knowledge from LLMs. We first unlearn knowledge of Harry Potter (Section 4.3.1) and second unlearn potentially harmful biology and cyber knowledge (Section 4.3.2).

Following work on unlearning knowledge of Harry Potter from Eldan and Russinovich 2023a, we show that targeted LAT can improve the robustness of unlearning without sacrificing the model’s performance on other topics.

We work with the “Who’s Harry Potter” (WHP) method from Eldan and Russinovich 2023a. It involves taking a corpus of text to forget (e.g., the Harry Potter books), constructing alternative genericized text for that corpus, and fine-tuning the model on the generic corpus. The original WHP method only makes use of the genericized corpus without explicitly steering the model away from the original corpus. Because our goal is to augment WHP with LAT, as a baseline, we use a modified version of WHP, which we call WHP-Contrastive (WHP-C). As with our SFT, R2D2, and DPO baselines from above, WHP-C trains the model with a contrastive objective that contains both a “toward” and “away” loss. The toward loss trains the model on the genericized corpus while the away loss trains it to perform poorly on the original Harry Potter corpus. Also as before, we interleave supervised fine-tuning batches on the UltraChat dataset to stabilize training. When performing WHP-C-LAT, we optimize the adversarial attacks to minimize the cross-entropy loss on the original Harry Potter text. For all methods, we train on 100 batches of size 16 for 4 steps each. Finally, in Appendix E, we also experiment with optimizing and constraining adversarial perturbations in a whitened space before de-whitening and adding them to the model’s latents.

To evaluate general performance, we again use MMLU . Next, we evaluate Harry Potter familiarity under Harry Potter knowledge extraction attacks. Full details are available in Appendix F. First, in response to past work suggesting that unlearning can fail to transfer cross-lingually , we evaluate familiarity in Spanish. Second, to test the robustness of unlearning to jailbreaks , we evaluate familiarity under jailbreaking prompts . Third and fourth, we evaluate the extent to which the model is robust to knowledge extraction attacks in the form of high-level summaries and short snippets of text from the Harry Potter books.

We present results in Table 4. WHP-C-LAT Pareto dominates WHP and WHP-C across all measures except MMLU.

3.2 Unlearning WMDP Biology and Cyber Knowledge

Following work from Li et al. 2024b, who studied the unlearning of potentially dangerous biology and cyber knowledge, we show that targeted LAT can help to improve existing approaches for unlearning.

As in as in Li et al. 2024b, we use the WMDP biology and cyber corpora as forget datasests and WikiText as a retain dataset.

As in Li et al. 2024b, we use Zephyr-7B off the shelf . We test two different unlearning methods with and without targeted LAT. First, we use a shaped gradient ascent (GA) method inspired by . We fine-tune the model to jointly minimize training loss on the retain set and log⁡(1−p)\log(1-p) on the forget set as done in Mazeika et al. 2024. To augment GA with targeted LAT, we apply latent-space perturbations optimized to minimize training loss on the forget set. To stabilize training, we also interleave training batches with supervised finetuning on the Alpaca dataset . Second, we use representation misdirection for unlearning (RMU) from Li et al. 2024b. With RMU, the model is trained at a given layer to (1) map activations from forget-set prompts to a randomly sampled vector while (2) leaving activations from other prompts unaltered. To augment RMU with targeted LAT, we apply latent-space adversarial perturbations only when training on the forget set. We optimize these perturbations to minimize the model’s cross-entropy training loss on the undesirable forget-set example. We experimented with various layer combinations and found the best results from applying them to the activations immediately preceding the RMU layer.

We evaluate how well the model’s general capabilities have been preserved by testing on MMLU and AGIEval . We evaluate the effectiveness of unlearning in the model using biology and cyber knowledge assessments from Li et al. 2024b. These multiple choice evaluations represent a qualitatively different task than the forget sets (which were full of bio and cyber documents), so they test the ability of LAT to generalize to qualitatively different kinds of unwanted behaviors than those used during fine-tuning. To test the robustness of the unlearning, we also evaluate models under few-shot finetuning attacks in which an attacker seeks to extract knowledge by finetuning the model on a small number of examples . Here, we use a simple but surprisingly effective attack: we randomly sample a single batch of 2 examples from the relevant forget set and repeatedly train on that single batch for 20 iterations. We then report the highest WMDP bio/cyber performances for each model across evaluation checkpoints at 5, 10, and 20 steps. For all evaluations, we use 1,000 samples on lm-evaluation-harness v0.4.0 as done in Li et al. 2024b.

Table 5 shows results for evaluating models by MMLU versus unlearning effectiveness. GA-LAT outperforms GA by a large margin under all evaluations. Similarly, RMU-LAT outperforms RMU in all evaluations, except for a 1.2% decrease in MMLU and 2.1% decrease in AGIEval. Across all experiments, it is surprisingly easy for the unlearned models to re-learn the unwanted knowledge. Repeatedly training on the same batch of 2 examples for up to 20 iterations improved WMDP bio/cyber performance by an average of 15.7 percentage points. However, LAT makes the models more resistant to re-learning. On average, re-learning closed 74.7% of the performance gap between the unlearned model and the original model for non-LAT methods but only 59.9% of the gap for LAT methods.

Discussion

By attacking the model’s latent representations, LAT offers a unique solution because models represent concepts at a higher level of abstraction in the latent space . Here, we have used targeted latent adversarial training (LAT) to strengthen existing defenses against persistent harmful behaviors in LLMs. We have applied LAT to three current challenges with state-of-the-art LLMs: jailbreaking , unlearning , and backdoor removal . In each case, we have shown that LAT can augment existing techniques to improve the removal of unwanted behaviors with little or no tradeoff in general performance. Overall, these results support but do not yet confirm our hypothesis that LAT can remove neural circuitry from models responsible for undesirable behaviors. We leave analysis of the mechanisms behind harmful model behaviors (e.g., ) to future work.

Our motivation for LAT is a response to two observations. First, LLMs empirically can persistently retain harmful capabilities despite attempts to remove them with adversarial training . Second, there have been empirical and theoretical findings that LLMs undergo limited changes to their inner capabilities during fine-tuning . All three problems that we have used targeted LAT to address – jailbreaks, backdoors, and undesirable knowledge – are ones in which an LLM exhibits harmful behaviors that are difficult to thoroughly remove. Our results show that targeted LAT can be useful for making models more robust to these persistent failures. We also find that these failure modes need not be precisely known for LAT to be helpful, showing instances in which LAT can improve generalization to different datasets of attack targets, harmful behaviors, and knowledge-elicitation methods than were used during training.

In Section 4.3, we find that state-of-the-art LLM unlearning methods are surprisingly vulnerable to relearning from small amounts of data. We find that re-training repeatedly on only two samples from the forget set was consistently able to close more than half of the performance gap between the original and unlearned models on average. We find that targeted LAT can reduce the sample efficiency of re-learning, but there is much room for improvement in designing unlearning methods that are robust to few-shot finetuning attacks. We are interested in future work to explore LAT’s potential to improve on existing approaches for making models robust to few-shot fine-tuning attacks .

While we have shown that LAT can be useful, it can also be challenging to configure and tune. In our experience, we found the selection of dataset, layer(s), and perturbation size, to be influential. We also found that interleaving supervised finetuning in with training and NaN handling were key to stable training. LAT can be done in different layers, with various parameterizations, and under different constraints. Our work here is limited to residual stream perturbations designed with projected gradient descent. Additionally, all of our experiments are done in LLMs with fewer than 10 billion parameters.

Improved latent-space attacks In addition to performing LAT with perturbations to an LLM’s residual stream, we are interested in other strategies for attacking its internal representations. Toward this goal, engaging with recent work on LLM representation engineering and interpretability may help to better parameterize and shape latent space attacks. We also speculate that universal attacks instead of single-instance attacks may be more interpretable and might better target the most prominent mechanisms that a model uses when it produces undesirable outputs.

Augmenting other latent-space manipulation techniques Concurrently with our work, Zou et al. 2024, Rosati et al. 2024, and introduced other latent-space manipulation techniques for making LLMs robust to undesirable behaviors. We are interested in studying how these techniques compare to LAT and whether LAT can be used to improve them.

Generalized adversarial attacks for LLM evaluations We are interested in the extent to which embedding-space attacks (e.g., ), latent-space attacks, (e.g., ), and few-shot fine-tuning attacks (e.g., ) can improve evaluations of LLM safety .

Broader Impacts

This work was motivated by the goal of training more safe and trustworthy AI systems. We believe that LAT will be practically useful for training better models. However, we emphasize that LAT is a value-neutral technique for training AI systems to align with their developer’s goals. It is important not to conflate AI alignment with safety . We believe that this work will contribute to helpful progress, but we emphasize that many of the risks from AI systems come from misuse and adverse systemic effects as opposed to unintended hazards such as the ones we work to address.

Contributions

Paper writing was performed by Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, and Stephen Casper. Experiments for Section 4.1 were led by Abhay Sheshadri, Aengus Lynch, and Vivek Hebbar. Experiments for Section 4.2 were led by Aidan Ewart. Experiments for Section 4.3.1 were led by Phillip Guo. Experiments for Section 4.3.2 were led by Cindy Wu and Philip Guo. Advising was provided by Stephen Casper, Dylan Hadfield-Menell, Asa Cooper Stickland, and Ethan Perez. Project management was provided by Henry Sleight.

Acknowledgements

We are thankful to Rajashree Agarwal, Rohit Gandikota, John Hughes, Erik Jenner, Alex Lyzhov, Sam Marks, Jacob Pfau, Sara Price, Javier Rando, Markian Rybchuk, Lennart Schulze, Aaquib Syed, Alex Turner, and Tony Wang for useful conversations. We thank William Brewer, Rocket Drew, Ronny Fernandez, McKenna Fitzgerald, Juan Gil, Carson Jones, Ryan Kidd, Christian Smith, and Laura Vaughan, for program support. This project was funded in part by a grant from Open Philanthropy and used compute provided by the Center for AI Safety. This support was offered by both organizations without conditions, and neither exerted any influence over the paper’s content.

References

Appendix A Loss Functions for LAT

Here, we describe the RT-LAT method described in Section 4.1 in greater detail. We assume we are given two datasets - a dataset of harmful requests and pairs of preferred and rejected completions Dp={(xi,ci,ri)}\mathcal{D}_{p}=\{(x_{i},c_{i},r_{i})\}, and a generic dataset of benign requests and helpful completions Db={(xi,yi)}\mathcal{D}_{b}=\{(x_{i},y_{i})\}. For each batch, we train the adversarial attack δ\delta to minimize Lattack\mathcal{L}_{\text{attack}}:

We additionally add the constraint that ∣∣δi∣∣2≤ϵ||\delta_{i}||_{2}\leq\epsilon, where ϵ\epsilon is a hyperparameter, to restrict the adversary’s power. We then train the model parameters θ\theta against these adversarial attacks by minimizing Lmodel\mathcal{L}_{\text{model}}. We define Lmodel\mathcal{L}_{\text{model}} in terms of the loss functions Ldefense\mathcal{L}_{\text{defense}} and Lbenign\mathcal{L}_{\text{benign}}:

We can use one of two different benign loss terms:

where θ∗\theta^{*} are the weights of the frozen reference model. Note that Lbenign\mathcal{L}_{\text{benign}} is always calculated on inputs where no adversarial attack is present.

We use Lbenign,SFT\mathcal{L}_{\text{benign},\text{SFT}} for our Llama2 results, and Lbenign, KL\mathcal{L}_{\text{benign},\text{ KL}} for our Llama3 experiments. Lbenign,SFT\mathcal{L}_{\text{benign},\text{SFT}} trains the model to maximize the probability of the ground-truth completions for benign prompts, whereas Lbenign, KL\mathcal{L}_{\text{benign},\text{ KL}} trains the model to preserve its original logits over possible completions for benign prompts. We hypothesize that Lbenign, KL\mathcal{L}_{\text{benign},\text{ KL}} might preserve original model capabilities better when the quality of Db\mathcal{D}_{b} is poor relative to the model being trained. Empirically, we find that Lbenign,KL\mathcal{L}_{\text{benign},\text{KL}} can better allow more capable models to retain their capabilities during adversarial training.

A.2 DPO-LAT

We now describe the DPO-LAT loss inspired by Rafailov et al. 2024. Similarly to RT-LAT, we assume that we have a paired preference dataset of harmless/harmful completions Dp={(xi,ci,ri)}\mathcal{D}_{p}=\{(x_{i},c_{i},r_{i})\}, where cic_{i} is the harmless result and rir_{i} is the harmful response. Instead of using a generic dataset of benign requests and useful completions, we instead assume Db={(xi,ci,ri)}\mathcal{D}_{b}=\{(x_{i},c_{i},r_{i})\} is a dataset of helpful/unhelpful responses (where again cic_{i} is the chosen helpful response and rir_{i} is the rejected unhelpful one). We take Dp\mathcal{D}_{p} from the ‘harmless’ split of Anthropic’s HH-RLHF dataset and Db\mathcal{D}_{b} from the ‘helpful’ split.

We choose Lattack\mathcal{L}_{\text{attack}} to cause the model to prefer the harmful response rir_{i} over cic_{i} where (xi,ci,ri)∼Dp(x_{i},c_{i},r_{i})\sim\mathcal{D}_{p}, using the DPO loss (where θ∗\theta^{*} are the weights of the frozen reference model):

We then set Ldefense\mathcal{L}_{\text{defense}} and Lbenign\mathcal{L}_{\text{benign}} to the DPO loss on Dp\mathcal{D}_{p} and Db\mathcal{D}_{b}, with the adversary present and not present respectively:

A.3 WHP-C-LAT and GA-LAT

The WHP-C-LAT and GA-LAT methods described in Section 4.3.1 and Section 4.3.2 use a toward-only adversary which optimizes for next-token cross-entropy loss on Harry Potter and the WMDP forget corpora respectively. For WHP, the model is trained as in Eldan and Russinovich 2023a. For WMDP, the model uses a log⁡(1−p)\log(1-p) away loss on the forget dataset as in Mazeika et al. 2024. In both cases, we additionally include a toward loss on WikiText to match Li et al. 2024b, and a supervised fine-tuning (SFT) loss on Alpaca . While calculating the model’s toward and away losses, we keep the perturbations from the adversary. We remove these perturbations for SFT.

Given a dataset DfD_{f} of text examples that you want the model to forget, and a dataset DbD_{b} of text examples that you want the model to retain, we can define the losses as follows:

where ti,jt_{i,j} is the jj-th token of the ii-th string in the dataset and ti,<jt_{i,<j} is the string of all tokens of the ii-th string up to the jj-th token.

A.4 RMU-LAT

Here, we use the same RMU loss as used in Li et al. 2024b. The adversary still optimizes for next-token cross-entropy loss on the WMDP forget corpora. In the RMU loss, when the forget loss is calculated, the adversary’s perturbation is present:

where LL is the length of the input tokens, and u is a randomly chosen vector from a uniform distribution between $thatisthennormalized(andstaysconstantthroughouttraining).Theconstantsthat is then normalized (and stays constant throughout training). The constantscandand\alpha$ are hyperparameter coefficients, which we set to be 6.5 and 1200 as in Li et al. 2024b for Zephyr-7B.

Appendix B Jailbreaking Robustness Under Untargeted LAT

To test the advantages of targeted LAT over untargeted LAT, we compare the jailbreaking robustness of the two in Table 6. Here, during untargeted LAT, the adversary does not work to make the model comply with the jailbreak. Instead, it only works to make the model fail to output a refusal. We find that untargeted LAT results in less harm to general performance compared to targeted LAT but not refusal training. Meanwhile, untargeted lat results in comparable or slightly worse robustness in most cases compared to targeted LAT. However, for prefill and GCG attacks, untargeted LAT fares much worse than targeted LAT.

Appendix C Jailbreaking Robustness Under an Alternate Autograder

In Section 4.1, we evaluate jailbreak success using the StrongReject autograder . However, here we also report results using the HarmBench autograder . Overall, we find that the HarmBench autograder is significantly more likely to label attacks as successful, but the overall trends within results remain similar.

Appendix D Backdoored Model MMLU Performance

To evaluate the destructiveness of DPO-LAT versus DPO on backdoor removal, we evaluate each model’s performance on MMLU . We present our results in Table 8 for a single model. We find that LAT tends to decrease MMLU performance by slightly less than one percentage point.

Appendix E Low Rank Adapters and Scaled Perturbation Constraints for WHP Unlearning

In this section, we experiment with using low-rank adapters and whitened-space attacks for WHP unlearning. Typically, adversarial training methods that use projected gradient descent constrain perturbations to be within an LpL_{p}-norm spherical ball . However, for latent-space perturbations, this approach is arguably unnatural because in the latent-space, activations vary more along some directions than others. To address this, here, we test a scaling method to constrain attacks in a way that better respects the shape of the activation manifold in latent space in Section 4.3.1. We tested LAT with perturbations that are constrained to an LpL_{p}-norm ball in whitened before they are de-whitened and added to the residual stream.

Our goal was to increase the ability of targeted LAT to operate on coherent features relating to the unlearning corpora (specifically, features that would preserve meaning but cause the model to no longer recognize the text as related). As a result, we perform principal component analysis (PCA) on the distribution of activations between Harry Potter text and the coherent genericized versions of the text produced during WHP. We optimize and constrain the perturbations in a whitened space before de-whitening them using the inverse PCA transformation matrix and then applying it to the model’s latent states. In addition, we use a low-rank adapter on all linear modules of rank 64. In our experiments, this resulted in weaker unlearning for WHP experiments but with less of a tradeoff in general capabilities. The results are shown in Table 9. However, we speculate that unlearning tasks may be especially well-suited to this type of scaling, and we leave deeper investigation to future work.

Appendix F Tests for Robust and Competitive Unlearning in LLMs

Eldan and Russinovich 2023b fine-tune Llama-2-7B-Chat (Llama-2) to unlearn knowledge of the Harry Potter universe. Their method is based on fine-tuning using text that has been modified to replace domain-specific content with generic content. Throughout experiments here, we compare the WHP model from Eldan and Russinovich 2023a, our replications, and our replication with targeted LAT (see Section 4.3.1).

Here, we outline the methods we use to evaluate unlearning in Section 4.3.1

To evaluate the model, Eldan and Russinovich 2023a introduce “Familiarity” as a metric which measures the extent of Harry Potter content contained in the model’s completions of Harry Potter-related sequences as determined by an automated GPT-4 evaluation. To measure Familiarity, we follow the same method from Eldan and Russinovich 2023b to evaluate a completion from the model. An evaluation prompt is formatted with the datapoint reference, prompt, and model completion, passed into GPT-4, then obtain a model Familiarity score (Figure 2), using “gpt-4-turbo-preview” at seed=42 and temperature=0, with max tokens=252. All model completions are scored in this way, and then we calculate the Familiarity metric starting a counter at 0, adding 1 for grade 3 completions, 0.2 for grade 2 completions, and 0 otherwise. Then, this total is divided by the total number of completions.

Aside from standard Familiarity evaluations as done in Eldan and Russinovich 2023a, we also perform four other evaluations using Familiarity, but when the model is evaluated under prompt extraction attacks.

LLM fine-tuning does not always transfer to other languages , so we test the models’ Harry Potter Familiarity with the prompts translated by GPT-4 into Spanish.

Simple jailbreaks have been successful at resurfacing knowledge that is typically not produced by LLMs (e.g., building a bomb). We test a jailbreaking prompt designed to resurface Harry Potter knowledge based on prior successful jailbreaks against Llama-2 models (Figure 3).

Here, we use few-shot and summary prompting. We provide the model with small amounts of general context related to Harry Potter with the goal of resurfacing existing suppressed knowledge that was not provided. We evaluate Familiarity when either a high-level summary (Figure 4) or the first 10 lines of Book 1 are included in context.

Appendix G WMDP Unlearning Details

We use LoRA with rank 64 for GA and GA-LAT. For RMU and RMU-LAT, we do not use LoRA and instead train the MLP weights full-rank, as in Li et al. 2024b.