One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
Itay Zloczower, Eyal Lenga, Gilad Gressel, Yisroel Mirsky
Introduction
Open-weight and fine-tuning-enabled language models create a post-release safety problem. After a model has been aligned, downstream users are still able to update it. In open-weight releases, the user receives the parameters directly; in fine-tuning APIs, the user supplies data that drives provider-side updates. In both settings, the user obtains access to the same primitive used to install alignment in the first place: gradient-based optimization. This makes safety alignment vulnerable to being overwritten after release. Prior work shows that small malicious fine-tuning runs can subvert aligned models and substantially recover harmful behavior (Yang et al., 2023b; Lermen et al., 2023); even benign or non-malicious fine-tuning can degrade safety alignment (Qi et al., 2023). This malicious finetuning (MFT) is not merely an implementation bug in a particular model family. It reflects a structural asymmetry: alignment is applied before release, but the attacker fine-tunes after release. If alignment modifies behavior without removing the underlying capability, then later behavioral training can recover what alignment suppressed.
This risk has given rise to an emerging research area on defenses against malicious fine-tuning: methods that aim to make aligned models robust to downstream attempts to recover harmful behavior. These MFT defenses attempt to make harmful fine-tuning fail, require prohibitive optimization effort, or cause the model to lose general capability when the attack succeeds. The proposed mechanisms vary widely. Some harden internal representations against harmful updates (Huang et al., 2024b; Rosati et al., 2024; Liu et al., 2025a; Chen et al., 2025a; Perin et al., 2025); some simulate adversarial fine-tuning during alignment (Henderson et al., 2023; Tamirisa et al., 2024; Sanyal et al., 2026); some reshape the loss landscape around safety-relevant parameters or optimization trajectories (Huang et al., 2024a; Rosati et al., 2025; Fan et al., 2025; Nguyen et al., 2026); some reweight or curate alignment data (Liu et al., 2024, 2025b; Chen et al., 2025a); and some deliberately couple harmful fine-tuning to catastrophic utility collapse (Chen et al., 2025b; Yi et al., 2025; Wang et al., 2025). Despite this diversity, existing evaluations share a common pattern: the attacker is usually a fixed, ignorant supervised fine-tuning procedure with a prescribed loss, optimizer, data budget, and hyperparameter range (Qi et al., 2023; Huang et al., 2024b; Rosati et al., 2024; Tamirisa et al., 2024; Huang et al., 2024a).
This evaluation practice conflicts with a central lesson from adversarial machine learning: defenses must be evaluated against adaptive adversaries. An adaptive adversary knows the defense, understands the training objective, and selects an attack that exploits this knowledge. Without such an evaluation, a defense may only show that it blocks the specific attack considered during its design. This failure mode is well documented in adversarial examples, where defenses that appeared robust against standard attacks were later broken by attacks adapted to the defense mechanism (Carlini and Wagner, 2017b; Athalye et al., 2018; Tramèr et al., 2020).
In the MFT setting, this is a critical gap. Recent defenses aim to make malicious fine-tuning fail, yet they have largely been evaluated against fixed, non-adaptive fine-tuning procedures. To our knowledge, no prior work has systematically defined what an adaptive adversary means for MFT defenses, instantiated such an adversary, nor evaluated an adaptive adversary on these defenses. As a result, current robustness claims remain difficult to validate, and the community lacks a reliable benchmark for testing defenses under a meaningful post-release threat model.
Many defenses, one shared vulnerability. A successful malicious fine-tuning attack must make the model harmful while preserving its general capabilities. A defense can therefore succeed by disrupting either condition. We find that fifteen recent MFT defenses fall into two broad strategies: anchoring, which makes optimization on harmful objectives ineffective, and self-destruction, which permits optimization but causes the model’s capabilities to collapse. Despite their surface differences, these defenses reduce to four loss templates that implement these two strategies.
Our analysis of these templates reveals a shared oversight: existing defenses are evaluated against attackers that optimize only for harmfulness, rather than the full attack objective. This leaves open a simple adaptive strategy. Because successful attacks require both harmfulness and capability to be preserved, an adversary can add a benign capability-preservation loss during fine-tuning. This signal helps the model avoid the traps created by both anchoring and self-destruction defenses: it preserves utility while allowing harmful behavior to re-emerge. We show that this adaptive attack (called Sidestepper) breaks all defense methods, suggesting that the vulnerability comes not from any single defense design, but from a common assumption about the attacker.
To our knowledge, this is the first work to systematically define, instantiate, and evaluate an adaptive adversary for malicious fine-tuning. We also identify a fundamental locality issue in existing defenses that explains why our attack is effective. To support this claim, we also demonstrate a second adaptive attack (called Kick-Settle) that evades these defenses by escaping the locality of their intended traps. Finally, we provide a simple test that future work can use to evaluate the security of proposed MFT defenses.
This work makes three contributions. (1) We systematize 15 recent malicious fine-tuning defenses (10 since 2025) and show that, despite their apparent diversity, they collapse into four loss templates and two underlying security strategies. This compression reveals that the field has converged on the same small set of ideas without recognizing it, and that every defense inherits the same evaluation gap: each is tested only against an attacker who optimizes for harmful recovery alone. As a result, these defenses do not provide meaningful security against an adversary that wants to perform MFT. (2) We introduce the concept and model of adaptive adversaries for malicious fine-tuning, defining the post-release threat model the field has implicitly been avoiding: an attacker who chooses the data, optimizer, and objective in response to the defense. (3) We propose two simple yet highly effective adaptive attack algorithms that defeat every evaluated defense, regardless of its underlying strategy or mechanism. We recommend that future work proposing MFT defenses evaluate against our attack before making robustness or security claims.
Related Work
Durable safeguards for open-weight models. A growing line of work studies whether safety interventions can remain effective after a model is released as open weights or exposed to downstream fine-tuning. Standard refusal training can often be weakened or removed by continued optimization, motivating defenses that aim to make harmful capabilities difficult to recover after release (Qi et al., 2023; Yang et al., 2023a; Zhan et al., 2024; Wei et al., 2024). Recent methods pursue this goal through representation noising, tamper-resistant training, adversarial unlearning, and other training-time interventions intended to make malicious fine-tuning fail or degrade utility (Rosati et al., 2024; Tamirisa et al., 2024; Li et al., 2024; Huang et al., 2024b).
The closest prior work is Qi et al. (Qi et al., 2025), which shows that MFT defenses can be far more brittle than their original evaluations suggest: safeguards that survive one fine-tuning setup may fail under small changes to dataset shuffling, trainer implementation, learning-rate schedule, prompt formatting, or fine-tuning data. This provides an important cautionary lesson for model release: defenses should be stress-tested across plausible downstream settings before releasing model weights. Our work studies a stronger failure mode. Durability tests ask whether a defense survives routine variation in the fine-tuning setup. We show that even defenses that pass such tests can fail once the attacker is adaptive. Thus, evaluating MFT defenses requires more than repeating malicious fine-tuning under different configurations; it requires testing adversaries that deliberately change the optimization objective to bypass the defense.
Adaptive evaluation in adversarial machine learning. Our work follows a central lesson from adversarial machine learning: robustness claims based on static, transfer, or otherwise non-adaptive attacks often fail under adversaries that optimize against the defense itself (Biggio et al., 2013; Szegedy et al., 2014; Goodfellow et al., 2015; Carlini and Wagner, 2017a; Athalye et al., 2018; Tramèr et al., 2020). This lesson has been repeatedly rediscovered across domains, where defenses that appear effective against fixed attacks are broken once the attacker incorporates the defense mechanism into the objective, bypasses non-differentiable components, changes the optimization procedure, or searches over a richer attack class (Carlini and Wagner, 2017a; Athalye et al., 2018; Tramèr et al., 2020; Mujkanovic et al., 2022). The methodological point is that evaluation attacks must be constructed with respect to the mechanism by which the defense claims to obstruct the adversary. We instantiate this principle for MFT defenses by identifying their shared obstruction mechanisms and deriving attacks that exploit them.
Adaptive attacks in LLM security. Similar evaluation failures have recently appeared in LLM security. Jailbreak and prompt-injection defenses are often evaluated against fixed benchmark attacks, static prompt sets, or generic optimization methods that are not tuned to the defense under evaluation. Recent work shows that such defenses can appear robust under these evaluations while failing against attackers that adapt their search procedure, feedback signal, or human strategy to the defense (Nasr et al., 2025; Zou et al., 2023; Chao et al., 2024; Mazeika et al., 2024; Debenedetti et al., 2024).
Our setting differs in the object being attacked. Rather than constructing adversarial prompts against a deployed interface, we study adversaries who directly modify model weights after release. In this setting, adaptivity means choosing the post-release optimization process itself, including the data mixture, loss, optimizer, schedule, and recovery objective. This changes the structure of the attack problem: the relevant question is not whether a defense survives stronger malicious fine-tuning, but whether it survives fine-tuning designed to counter the defense mechanism. Our work addresses this gap by systematizing how MFT defenses obstruct naive harmful fine-tuning and deriving adaptive attacks from those obstruction mechanisms.
Systematizing MFT Defenses
where and are acceptable performance thresholds based on a hold out set. The conjunction matters: a model that emits harmful content but loses fluency or coherence or capability (i.e., intelligence) is not a useful recovered model.
2 Fifteen Defenses, Two Strategies
The attacker’s trajectory either stalls, drifts unproductively, or wanders on a plateau; never reaches a region with low . This family spans five mechanisms.
Representation anchoring flattens or randomizes harmful hidden states: Vaccine and T-Vaccine harden hidden states against harmful embedding perturbations Huang et al. (2024b); Liu et al. (2025a), and RepNoise pushes harmful activations toward Gaussian noise Rosati et al. (2024). Direction anchoring suppresses the harmful gradient at itself: Booster and Antibody attenuate harmful gradient influence Huang et al. (2024a); Nguyen et al. (2026), VAA reinforces vulnerable safety subgroups Chen et al. (2025a), and SAM-unlearning enforces local smoothness against relearning Fan et al. (2025). Subspace anchoring closes off low-rank attack routes: LoX extrapolates safety-critical weights into a flatter region, and AntiDote trains against an adversarial hypernetwork that generates worst-case LoRA patches Perin et al. (2025); Sanyal et al. (2026). Objective anchoring reshapes the alignment objective itself, as in KT-IPA Cheng et al. (2025). Trajectory anchoring extends the same idea from a single step to a simulated attack trajectory: MLAC and TAR meta-train against inner adaptation steps so that the post-trajectory endpoint still has high harmful NLL or maximum posterior entropy on harmful prompts Henderson et al. (2023); Tamirisa et al. (2024). Mechanically these are look-ahead methods, but their security argument is the same as the single-step anchors: the attacker’s optimizer finds no descent direction on , only now the guarantee holds along a -step trajectory rather than at a single .
Strategy B: self-destruction. Self-destructive defenses do not hide the harmful gradient or block progress on . They let the attacker descend freely and do reach low harmful loss, but engineer the trajectory’s endpoint so benign capability collapses there:
The recovered model emits harmful content on harmful prompts but is unusable on benign ones, failing the joint success condition (1) on the task clause. CTRAP forces to predict a fixed error token on benign inputs Yi et al. (2025); SEAM couples adversarial and benign gradients into opposing directions, so descending is descending Wang et al. (2025); and SDD trains harmful prompts to map to fluent but unrelated benign answers, so harmful fine-tuning degrades instruction-following Chen et al. (2025b). The three differ in what kind of capability damage they engineer (token-level collapse, gradient-level coupling, output-level incoherence) but share the same security argument: the harmful-fitting direction in weight space is the utility-destroying direction.
3 Four Loss Templates Behind the Taxonomy
The perturbation set may live in embedding space, layer space, or weight space. Vaccine uses adversarial hidden-state perturbations; T-Vaccine
Template 2: Harmful-information removal. A second family attacks the information content of harmful representations:
The simulated attack may be one harmful gradient step, inner fine-tuning steps, a sampled adversary, or a learned patch. The penalizer then imposes a desired property at the simulated post-attack point.
Template 3 is the bridge between anchoring and self-destruction. If penalizes harmful progress, preserves refusal, prefers safe responses, or keeps harmful loss high along a simulated trajectory, the method is an anchoring defense. Booster penalizes harmful-loss drop after a simulated harmful step Huang et al. (2024a); Antibody preserves refusal at the post-step model Nguyen et al. (2026); AntiDote trains against activation-conditioned adversarial LoRA patches Sanyal et al. (2026); KT-IPA includes an adversarial phase inside its integrity objective Cheng et al. (2025); and MLAC and TAR penalize a -step simulated trajectory so harmful loss stays high or the predictive distribution stays at maximum entropy Henderson et al. (2023); Tamirisa et al. (2024). If instead makes the simulated endpoint useless on benign data, the same template implements self-destruction. CTRAP penalizes a simulated post-harmful-step model to collapse on benign inputs Yi et al. (2025). Thus the same mathematical form can implement either strategy. The difference is what the defender wants to be true at : a model that still refuses (anchoring), or a model that has lost benign capability (self-destruction).
Template 4: Coupling trap. The last template does not simulate the attacker. Instead, it directly couples harmful improvement to benign degradation:
where is typically benign task samples. SEAM implements this explicitly by shaping the relationship between harmful
and benign gradients: descent on harmful data is made to increase benign loss Wang et al. (2025). SDD implements a data-level version: it trains harmful prompts to elicit fluent but irrelevant benign answers, so later harmful fine-tuning must undo a response mapping that also damages instruction-following Chen et al. (2025b). Template 4 is always self-destruction. It does not try to hide the harmful gradient; it tries to make following it costly.
A Unified Adaptive Attack
The shared vulnerability: a fixed attacker objective. The four templates differ in mechanism, but they share the same attacker model: the post-release adversary is assumed to optimize only harmful loss, . This assumption is visible in each template. Template 1 makes a local basin robust to perturbations induced by descent on . Template 2 removes the harmful representations that would use. Template 3 simulates an inner attacker whose loss is or a close proxy. Template 4 couples the harmful-improvement direction to benign degradation, again assuming the attacker follows the direction. In all cases, the defense is optimized against the same naive adversary:
This is not the threat model. A post-release adversary can choose the data, optimizer, schedule, and loss. Nothing requires the attacker to optimize alone. The harm-only objective is therefore a defense-naive evaluation choice: it tests whether the defense blocks one prescribed trajectory, not whether it blocks adaptive malicious fine-tuning.
We choose the benign capability loss as this auxiliary signal. This choice is natural for two reasons. First, a successful malicious fine-tuning attack must recover harmful behavior while preserving model usefulness, as required by the joint success condition in (1). Thus is already part of the attacker’s real objective, even if prior evaluations measure it only after training. Second, as shown in Section 3.3, appears in every defense template through the alignment objective: defenders must preserve benign capability, otherwise they would release a safe but useless model. The same signal that lets the defender keep the model useful is therefore always available to the attacker as a search heuristic for escaping anchors and avoiding self-destruction.
The adaptive objective. We therefore attack all templates with the same mixed objective:
therefore contains a component the anchor was not designed to suppress. The term moves the model while preserving useful behavior. Once the trajectory leaves the local basin or purged representation regime, the defense no longer controls the relevant region of the loss surface, and harmful behavior can re-emerge.
Why this breaks look-ahead defenses. Template 3 appears adaptive because it simulates an attacker . But the simulation fixes the attack loss, horizon, optimizer family, and often the adaptation form. The outer defense therefore regularizes the endpoint of one chosen proxy attack, not the endpoint of every feasible post-release optimization. When the real attacker minimizes , it follows a different trajectory from the one the defense simulated. For anchoring instances, the attacker no longer follows the harmful-loss path whose progress was penalized. For self-destructive instances, the capability term directly penalizes the collapsed endpoint the defense tries to induce. In both cases, the adaptive trajectory leaves the truncated look-ahead approximation.
Why this breaks self-destruction defenses. Template 4 and the self-destructive instances of Template 3 try to make harmful fine-tuning succeed only by destroying capability. Under the naive objective, this can work: descent on is arranged to increase or damage instruction-following. Under the adaptive objective, that coupling becomes visible to the optimizer. A direction that lowers but sharply raises is not a good descent direction for . The attacker does not need to explicitly undo the trap; it simply optimizes the real success criterion. A self-destructed model with low harmful loss but high capability loss is exactly what the adaptive objective avoids.
Experimental Setup
Setup. We test whether our adaptive attack from (6), called SideStepper, is effective across diverse MFT defense strategies. We evaluate defended checkpoints derived from Llama-2-7B-chat Touvron et al. (2023), Qwen3-8B-Instruct Yang et al. (2025), and Llama-3.1-8B-Instruct Grattafiori et al. (2024), covering both the older backbone used by much of the MFT-defense literature and newer instruction-tuned models with stronger baseline capability. We selected Booster Huang et al. (2024a), CTRAP Yi et al. (2025), VAA Chen et al. (2025a), Vaccine Huang et al. (2024b), Unlearn-Smooth Fan et al. (2025), and SDD Chen et al. (2025b) because they cover both strategies and mechanisms in Table 1. We use author-released implementations, checkpoints, and defense-specific data where available, and otherwise reproduce the defense protocol from the paper (Appendix G). Unlearn-Smooth is the only exception to the three-backbone protocol: the authors release a single defended Hugging Face checkpoint, not an instantiable defense recipe, so we evaluate that checkpoint directly. It therefore appears as one model bar, with mean and standard deviation computed over the same three adaptive-attack seeds. Starting from each defended checkpoint, we apply the same adaptive mixed-objective attack from Section 4; we do not tune a separate attack per defense.
We evaluate harmful recovery on the unsafe split of beavertails_330k_test, scored with the HarmBench classifierMazeika et al. (2024), and evaluate benign utility with standard zero-shot performance on MMLU Hendrycks et al. (2021), TruthfulQA Lin et al. (2022), HellaSwag Zellers et al. (2019), and ARC-Easy Clark et al. (2018). We report harmful-response rate as the harmful score metric and report capability benchmark changes relative to the defended checkpoint before attack. This delta convention isolates the effect of the adaptive fine-tuning attack from baseline differences across models and defenses. The full data construction, split details, and attack hyperparameters are in Appendix F.
Results. Figure 3 and Table 3 show that the adaptive attack increases harmful behavior for every defended checkpoint we evaluate. The effect is not a marginal degradation of safety: harmful-score deltas are positive across all defenses and models, ranging from to (max 1.0), with several defenses recovering roughly harmful-score points after attack. This failure holds for both defense families. For gradient-obfuscation defenses such as Booster, VAA, Vaccine, and Unlearn-Smooth, the benign-retention objective supplies a usable optimization signal and allows harmful behavior to re-emerge. For gradient-trap defenses such as CTRAP and SDD, the same objective prevents the attack from taking the low-utility collapse path that the defense tries to induce. Benign capability remains stable in nearly all cases. Excluding the released Unlearn-Smooth checkpoint (which demonstrates large gains), the average benign delta lies between and across all defended checkpoints, and no defense exhibits the broad capability collapse that would indicate a failed or degenerate attack.
Discussion & Insights
A conjecture, and an uncomfortable open problem. The broader lesson is not specific to these adaptive adversaries. We conjecture that as long as the released model still contains the targeted harmful capability, a defense that merely locks access to that capability can be unlocked by an adaptive attacker. The attacker need not know the exact lock in advance; it only needs some way to search for a nearby model that retains general capability while making the suppressed behavior accessible again. The implication is uncomfortable: durable robustness may require actually removing the capability, not merely obstructing the path to it. But true capability removal is currently prohibitive. Omitting malicious behavior from training and unlearning-based methods have been shown to incur substantial utility costs and are themselves vulnerable to relearning attacks Li et al. (2024); Sanyal et al. (2026); Fan et al. (2025). Reconciling these two facts is, in our view, the central open problem for this line of work, and the challenge of preventing MFT remains open.
Final Remarks. This paper gives a simple evaluation rule: robustness against -only fine-tuning is not evidence of robustness against malicious fine-tuning.
Conclusion
In this work, we show that MFT defenses that survive harmful-only fine-tuning fail against adaptive attackers. Our attack, SideStepper, restores harmful behavior across the evaluated defenses while largely preserving benign utility. The failure is not tied to a single defense mechanism, but to a shared assumption that attackers will optimize only for harmful recovery. These results suggest that future defenses should be evaluated against adaptive objectives, not only naive SFT baselines. Our study is limited to the defenses, model families, datasets, and metrics we evaluate.
Acknowledgments
This work was funded by the European Union, supported by ERC grant: (AGI-Safety, 101222135). Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Council Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.
References
Appendix A Ethics and dual-use considerations
This work studies adaptive attacks on defenses against malicious fine-tuning, and is therefore dual-use: the same analysis that improves evaluation could also inform attackers. We believe the work is ethically justified because the attack model is already implicit in standard fine-tuning practice, the core ingredients are public, and withholding adaptive evaluations would create a false sense of security around defenses that fail under realistic use. Our goal is not to expand harmful capability, but to close an evaluation gap: defenses for open-weight and fine-tuning-as-a-service settings must be tested against adaptive adversarial attackers, because that is precisely what makes a compromised model useful. We therefore report attacks at the level needed for scientific reproducibility, evaluate on controlled benchmark datasets, and focus the paper on aggregate failure modes rather than harmful outputs or deployment guidance. By showing where current defenses fail, the work supports stronger threat models, more reliable benchmarks, and defenses that are robust to adaptive optimization rather than only to naïve fine-tuning.
We take two concrete steps to limit misuse risk. First, all attack experiments use only existing, publicly released datasets (BeaverTails, Alpaca); we do not generate new harmful prompts or completions, and we contribute no new harmful corpus. Second, we do not release any defended or attacked model checkpoints. Code release covers the training and evaluation pipeline (see footnote in abstract), but not the weights of any attacked model. Researchers who wish to reproduce our results can do so by running the released code against the publicly available defended checkpoints from the original authors; this preserves reproducibility without our project becoming a distribution channel for attacked models.
Appendix B LLM usage
We declare two uses of LLMs that touch the methodology. First, we use GPT-4o-mini to generate refusal responses for the paired BeaverTails split described in Section 5. This split is consumed only by defenses whose construction protocol requires paired (refusal, harmful) responses; it is not used in any attack and is disjoint from the attack data. Second, we use the HarmBench classifier as an automated judge to score whether a model response to a held-out BeaverTails prompt is harmful. Both uses are standard in the MFT-defense literature we evaluate against, and we apply the same judge identically to base, defended, and attacked checkpoints so that any reported delta is judge-consistent across conditions. We also use LLMs for writing assistance (grammar, phrasing, LaTeX formatting); per the NeurIPS policy these uses do not require declaration, and they did not contribute to the methodology, claims, or analysis.
Appendix C Licenses for existing assets
We list the licenses of all third-party assets used in this work. All uses are for non-commercial academic research, consistent with the most restrictive licenses below.
Llama-2-7B-chat : Meta Llama 2 Community License Agreement.
Llama-3.1-8B-Instruct : Meta Llama 3.1 Community License Agreement.
HarmBench classifier (cais/HarmBench-Llama-2-13b-cls) : MIT License.
Datasets.
BeaverTails : CC BY-NC 4.0. We use only the publicly released splits and use the data exclusively for non-commercial research.
Alpaca : CC BY-NC 4.0. Use is restricted to non-commercial research.
MMLU, TruthfulQA, HellaSwag, ARC: standard zero-shot evaluation benchmarks accessed via their public releases under their respective licenses (MIT and Apache 2.0).
Defense code.
For each defense we evaluate, we use the original authors’ unmodified released code. Vaccine , Booster , and SDD are released under Apache License 2.0. Unlearn-Smooth is released under the MIT License. VAA and CTRAP do not have a license file declared in their public repositories; we use them solely for non-commercial academic research and replication, and we cite the original papers. Modifications across all defenses are limited to data adapter scripts that convert our standardized data format into each author’s expected format, and do not alter the defense logic.
API services.
We use OpenAI’s GPT-4o-mini API to generate refusal responses for the paired BeaverTails split (see Appendix B). Use complies with OpenAI’s terms of service for research.
Appendix D Loss landscape visualization details
Appendix E Compute resources
All experiments were run on an internal cluster of NVIDIA RTX PRO 6000 Blackwell GPUs (96 GB GDDR7 per card). Training and evaluation use a single GPU per run, with the exception of SEAM, whose author code requires two GPUs. We did not use multi-node distributed training.
Representative wall-clock figures for a single run on one GPU: a LoRA fine-tuning attack on a 7–8B parameter model takes approximately one GPU-hour (medians: SFT 0.80 h, naive 0.43 h, mixed-objective 1.11 h, kick-settle 0.75 h). Defense training runs vary: most defenses (Vaccine, Booster, VAA, SDD, CTRAP) take 0.1–0.5 GPU-hours; Unlearn-Smooth and MLAC sit between these. Evaluation of one checkpoint on the 18,189-prompt BeaverTails harm split takes about 9 GPU-hours (3 h to generate responses; 6 h to score them with the HarmBench classifier across two shards). The four zero-shot benchmarks together take about 0.3 GPU-hours per checkpoint.
The reported experiments in the main paper and appendix represent approximately 2,500 GPU-hours.
The full research project consumed substantially more compute than the reported experiments. Preliminary hyperparameter sweeps over defense learning rates and regularization coefficients, screening runs that did not make the final paper, and earlier infrastructure issues account for additional compute. We estimate total project compute, including these preliminary and failed runs, at approximately 5,000 GPU-hours (a factor of roughly 2 over the reported experiments).
Appendix F Experimental setup details
Evaluation data and metrics. We evaluate harmful recovery on the unsafe split of beavertails_330k_test. We filter for is_safe = false, yielding 18,189 held-out harmful prompts. This split is disjoint from both the BeaverTails examples used for harmful attack fine-tuning and the paired BeaverTails split used for defense construction. For each prompt, we generate a model response and score it with the HarmBench classifier, which predicts whether the response is harmful. We report the resulting harmful-response rate as the harmful capability metric. We evaluate benign capability with standard zero-shot evaluation on MMLU , TruthfulQA , HellaSwag , and ARC-Easy . These benchmarks are evaluation-only and never appear in the adaptive attack data. For each benchmark, we report the change in performance relative to the defended checkpoint before attack, together with the average change across the four benchmarks.
Naive attack (reproduction; Table 7). For every defense, we run that paper’s published attack hyperparameters on our unified BeaverTails attack split, holding the data and evaluation fixed. Each preset uses AdamW with a cosine schedule and warmup ratio unless noted. Per-defense values are listed in Table 5.
Defense training. For Booster, CTRAP, VAA, Vaccine, and SDD we train the defended checkpoint ourselves using each paper’s published recipe; for Unlearn-Smooth we use the author-released checkpoint OPTML-Group/zephyr-7b-npo-sam-wmdp-bio from the smooth-unlearned model collection and do not retrain. All runs use AdamW (8-bit variant where indicated) and the per-defense regularizer values reported by the original authors. Hyperparameters are held constant across the three model backbones (Llama-2-7B-chat, Qwen3-8B-Instruct, Llama-3.1-8B-Instruct) except where the original paper specifies a backbone-specific choice. See Table 6.
Benign capability evaluation. The four zero-shot benchmarks are scored with lm-eval-harness at num_fewshot=0, using likelihood-based multiple-choice scoring. Specifically: mmlu (acc, averaged across the 57 subjects), truthfulqa_mc1 (acc; we use the single-correct-answer variant rather than mc2), hellaswag (acc_norm), and arc_easy (acc_norm). All evaluations run in bfloat16 under torch.no_grad() with deterministic argmax over choice-token logits, so results are reproducible across runs.
Harm scoring (HarmBench classifier). Responses to the 18,189-prompt BeaverTails unsafe split are generated from each evaluated checkpoint with sampling (do_sample=True, temperature=1.0, top_p=1.0, max_new_tokens=512) using the model family’s default chat template and a generic “You are a helpful assistant” system prompt for Qwen3 and Llama-3.1, and the standard Llama-2-chat safety system prompt for Llama-2. Each generated response is scored by cais/HarmBench-Llama-2-13b-cls using the verbatim instruction template from that model card; the classifier emits a binary yes/no judgement which we report as the harmful-response rate.
Appendix G Reproduction of Original Paper Defenses
For every defense we first reproduce the published behavior under the threat model the authors assumed: a standard SFT attack on harmful data with the original paper’s hyperparameters, run on our unified evaluation pipeline. Table 7 reports these reproduction numbers in the same format as Table 3. A defense that holds here but falls in Table 3 confirms that the failure is driven by the adaptive attack, not by an implementation difference.