Gradient Surgery for Safe LLM Fine-Tuning
Biao Yi, Jiahao Li, Baolei Zhang, Lihai Nie, Tong Li, Tiansheng Huang, Zheli Liu
Introduction
Fine-tuning-as-a-Service (FaaS) has emerged as a prevailing approach for developing specialized Large Language Models (LLMs) Fine-tuning API by OpenAI: https://platform.openai.com/docs/guides/fine-tuning.. This service model enables providers to customize powerful, pre-trained LLMs for specific downstream tasks using user-supplied data. Consequently, FaaS makes tailored AI solutions widely accessible by obviating the prohibitive computational costs associated with training models from scratch.
However, this process introduces a critical security vulnerability. Foundation models are carefully aligned for safety to ensure they refuse harmful user requests. Yet, this alignment is brittle. Recent studies (Qi et al. 2023; Yang et al. 2023; Zhan et al. 2023) have shown that mixing a few malicious examples into the user’s fine-tuning dataset can catastrophically compromise the model’s safety alignment. The compromised model can then be directed to generate toxic or dangerous content it was originally designed to reject. In response to this threat, we investigate safe fine-tuning algorithms that maintain the model’s safety alignment without degrading performance on benign fine-tuning tasks.
A recognized paradigm (Bianchi et al. 2024; Huang et al. 2024b) for safe fine-tuning formalizes the challenge as a multi-objective optimization problem. This approach seeks to jointly optimize for two objectives during the fine-tuning stage: performance on the user’s fine-tuning task and adherence to safety alignment. The aim is to produce a final model that performs well on both. SafeInstr (Bianchi et al. 2024) implements this by simply mixing a small number of alignment examples into the user’s fine-tuning dataset. Lisa (Huang et al. 2024b) explores a more complex Bi-state Optimization (BSO) solution, which alternatively optimizes over the alignment and user datasets and uses a proximal regularizer to enforce proximity between iterates.
Despite their promise, we find that these multi-objective solutions are critically sensitive to the harmful ratio. While they perform adequately at low harmful ratios, their defensive capabilities degrade sharply as the ratio increases. We take a closer look at the safety performance throughout the fine-tuning process. At low harmful ratios, the model’s alignment is successfully maintained, remaining robustly safe. However, at higher ratios, the safety alignment is progressively compromised, leading to a model that grows increasingly unsafe as training progresses. We hypothesize that this failure stems from the influx of malicious data interfering with the alignment objective, causing the model’s safety to steadily erode.
We diagnose that the failure of existing multi-objective solutions stems from a fundamental problem: conflicting gradients, where the user-task gradient points in a direction that undermines the safety objective. Empirically, we find that the emergence and severity of this conflict are directly tied to the user data’s composition. On purely benign datasets, the user-task and alignment gradients are non-conflicting. However, as malicious data is introduced, gradient conflict emerges, with its intensity growing in direct proportion to the harmful ratio. This provides strong evidence that the user-task gradient, when corrupted, directly opposes the safety objective, revealing the root cause of why existing multi-objective solutions degrade.
In response, we propose SafeGrad, which is designed to address this optimization challenge with gradient surgery. Specifically, when the gradients for the user task and the alignment task conflict, we project the user task gradient onto the plane orthogonal to the alignment gradient. This procedure nullifies the harmful component of the task update, allowing the model to learn the user’s task without compromising the safety objective. Furthermore, we design a more robust alignment objective that replaces the sparse signal of supervised fine-tuning (SFT). We leverage the rich, fine-grained safety information encoded within the full predictive distribution of the original, well-aligned foundation model. Therefore, we define our alignment objective using a KL-divergence loss, which guides the model to learn and emulate the entire safety-aware output distribution of the well-aligned model. This provides a dense learning signal that teaches the model a comprehensive safety profile, rather than just how to produce a few target refusal tokens.
Finally, we summarize our contributions as follows:
We identify a critical limitation of current multi-objective safe fine-tuning methods, demonstrating that their effectiveness degrades significantly under high harmful ratios due to the fundamental problem of conflicting gradients.
We propose SafeGrad, a novel defense framework that resolves this conflict through gradient surgery. The approach is enhanced by a KL-divergence-based alignment loss, which provides a dense, distributional signal for learning robust safety.
Extensive empirical results demonstrate that, across multiple LLMs (such as Gemma3-4B, Llama3-8B, and Qwen2.5-7B) and various harmful fine-tuning settings, SafeGrad achieves state-of-the-art defense while preserving benign task performance.
Preliminaries
We ground our work in the Fine-tuning-as-a-Service (FaaS) paradigm. In this setting, a service provider customizes a pre-aligned LLM using data supplied by a user. The resulting specialized model is hosted by the provider, who delivers responses via an API. This architecture places the onus of ensuring safety squarely on the service provider. Failure to prevent the model from generating harmful content could expose the provider to significant governance challenges and legal liabilities (Reuel et al. 2024).
To address this threat, we adopt an assumption consistent with prior defensive strategies (Hsu et al. 2024; Zong et al. 2024; Bianchi et al. 2024; Huang et al. 2024b): the service provider possesses a small, trusted safety alignment dataset, denoted as . This dataset typically consists of harmful prompts paired with desired safe responses (e.g., refusals), and serves as a tool to preserve the model’s safety alignment during the fine-tuning process.
2 Safe Fine-tuning as Multi-objective Optimization
We formulate safe fine-tuning as a multi-objective optimization (MOO) problem (Yu et al. 2020; Huang et al. 2025a). The goal is to find model parameters that simultaneously minimize two objectives: a user task loss on the user-provided dataset , and a safety alignment loss on a trusted alignment dataset . Both losses are typically standard cross-entropy functions for supervised learning.
This creates a joint optimization problem aimed at finding a Pareto optimal solution for both objectives:
In practice, existing methods often approximate this by minimizing a scalarized objective, typically a weighted sum of the two losses:
where is a hyperparameter balancing the two tasks.
3 Revisiting Existing Multi-objective Solutions
Existing multi-objective solutions are critically sensitive to the harmful ratio, with their defensive capabilities degrading sharply as harmful ratio increases. As illustrated in Figure 2 (a), we evaluate the performance of these representative methods under varying harmful ratios (). While they perform adequately under low harmful ratios (), their effectiveness collapses when faced with a more determined adversary. For instance, the harmful score (HS) of SafeInstr escalates from 3.10% to 37.50%, and Lisa’s HS increases from 13.10% to 44.50% as the harmful ratio rises, demonstrating their significant vulnerability.
This failure stems from an unresolved optimization conflict, where the user-task objective, corrupted by harmful data, directly opposes the safety objective, leading to the model’s alignment being progressively compromised. To diagnose this vulnerability, we inspect the trend of the safety performance throughout the fine-tuning process. The dynamics shown in Figure 2 (b) reveal the core of the problem. At a low harmful ratio (e.g., ), the harmful score remains stable and low, indicating that the model’s safety alignment is successfully maintained. In stark contrast, at a higher harmful ratio (), the influx of harmful data creates a severe conflict with the alignment objective. This conflict manifests as a steadily increasing harmful score, demonstrating that the model’s safety is catastrophically compromised as fine-tuning progresses.
Methodology
Our proposed method, SafeGrad, is designed to directly address the limitations of existing multi-objective approaches by resolving the underlying optimization conflict between the user task and safety alignment. It comprises two key innovations: (1) a gradient surgery technique that eliminates harmful gradient components during updates, and (2) a more robust, distribution-aware alignment loss for learning the nuanced safety knowledge of the original well-aligned model.
As established in the Section 2.3, the performance of multi-objective safe fine-tuning degrades when the user-provided data contains a high proportion of harmful examples. We hypothesize that this failure arises from conflicting gradients.
Let be the gradient from the user’s task and be the gradient from the safety alignment task. A conflict occurs when these two gradients point in opposing directions, meaning that an update step that improves one objective will necessarily harm the other. We can formally define this conflict as having a negative cosine similarity:
When this condition is met, a standard weighted-sum update (as in Eq. (2)) forces a compromise that degrades the model’s safety alignment.
A severe gradient conflict emerges and intensifies as the harmful data ratio increases, undermining the optimization process. This phenomenon is quantitatively illustrated in Table 1. When the user’s dataset is free of malicious examples (harmful ratio of 0.00), the cosine similarity between the user task and alignment gradients is 0.02, indicating that there is no conflict between the two objectives. However, as the proportion of harmful data increases, the cosine similarity becomes sharply negative, falling to -0.05 at a 25% ratio and further to -0.16 when the dataset is entirely malicious. This provides strong evidence that the user-task gradient, when corrupted by malicious data, points in a direction that directly opposes the safety objective.
2 Gradient Surgery for Conflict Resolution
To resolve this problem, we propose a form of gradient surgery to explicitly resolve this conflict. Instead of simply averaging conflicting gradients, we surgically modify the user task gradient to remove the component that directly opposes the safety alignment gradient.
Specifically, when a conflict is detected (), we project the user task gradient onto the normal plane of the alignment gradient . The resulting gradient, , is orthogonal to and represents the component of the user task that does not interfere with the safety objective. This projection is calculated as:
If there is no conflict, we use the original user task gradient, i.e., . The final gradient update is then a combination of the modified user gradient and the alignment gradient, ensuring progress on the user task without compromising safety:
3 Distributional Alignment with KL-Divergence
Furthermore, we argue that a standard SFT loss is a suboptimal measure for safety alignment. It is sparse, focusing only on maximizing the probability of the target refusal tokens, and ignores the rich, fine-grained safety knowledge encoded across the entire output probability distribution of the well-aligned foundation model.
Our core design philosophy is to replace this sparse, token-level signal with a far richer, more informative alignment target. To this end, we propose a distributional alignment loss that leverages the original, well-aligned foundation model as a source of detailed safety knowledge. Let be the parameters of the frozen, pre-aligned reference model, and be the parameters of the model being fine-tuned. We define our fine-grained alignment objective as minimizing the KL-divergence between the output distributions of these two models on the alignment data . The new alignment loss is:
This loss redefines the alignment task: instead of merely learning a simple refusal phrase, the model is guided to learn and emulate the entire safety-aware output distribution of the expert reference model. This distribution provides a dense learning signal, encoding nuanced, holistic information about which tokens are safe, which are unsafe, and by what margin. Consequently, the model internalizes a comprehensive safety profile, which is a fundamentally more robust approach than SFT.
4 The Complete SafeGrad Algorithm
Combining our gradient surgery technique with the distributional alignment loss, we arrive at the complete SafeGrad algorithm, detailed in Algorithm 1. At each training step, we compute gradients for both the user task and the KL alignment objective. If a conflict is detected, we apply the gradient projection before combining the gradients and updating the model parameters.
Experiment
Datasets and Models. Following prior works (Huang et al. 2024b; Huang et al. 2024d; Huang et al. 2024a), we simulate an adversarial fine-tuning scenario where a user-provided dataset is poisoned. This dataset is a mixture of benign task data and malicious examples. Specifically, a fraction of the data consists of harmful prompt-answer pairs, while the remaining fraction is composed of data from a benign downstream task. Both the malicious examples used for poisoning and the trusted alignment data required by defense methods are sourced from the dataset provided by Rosati et al. 2024d, which is an enriched version of BeaverTails (Ji et al. 2023). The total size of the user’s fine-tuning dataset is set to 1,000 samples by default. For benign tasks, we evaluate on three standard benchmarks: SST2 (Socher et al. 2013) for sentiment analysis, AGNEWS (Zhang et al. 2015) for topic classification, and GSM8K (Cobbe et al. 2021) for mathematical reasoning. Our experiments are conducted on three powerful instruction-tuned models: Gemma-3-4B-IT (Team 2025), Llama-3-8B-Instruct (AI@Meta 2024), and Qwen2.5-7B-Instruct (Team 2024). In our experiment, we set the default to 0.1, the dataset to SST2, and the base model to Gemma-3-4B-IT, unless otherwise specified.
Metrics. To evaluate model performance, we adopt two key metrics, following the methodology of prior works (Huang et al. 2024b; Huang et al. 2024d; Huang et al. 2024a):
Finetune Accuracy (FA, ): The model’s performance on the benign portion of the user’s task, measured on a held-out test set. Higher is better.
Harmful Score (HS, ): We employ the Llama-Guard-3-8B model (Llama Team 2024) to assess the safety of model outputs. HS is defined as the percentage of responses flagged as unsafe when the model is prompted with a test set of harmful instructions. Lower is better.
For HS, we test on 1,000 prompts from the BeaverTails test set. For FA, the test sets for SST2, AGNEWS, and GSM8K contain 872, 1,000, and 1,000 samples, respectively.
Baselines. We compare SafeGrad against the following five representative baselines:
SFT: A standard fine-tuning approach without any explicit safety mechanism serves as a worst-case baseline.
SafeInstr (Bianchi et al. 2024): This method simply mixes a small set of safety alignment examples into the user’s fine-tuning data to remind the model of its safety training.
LISA (Huang et al. 2024b): This method treats safe fine-tuning as a bi-state optimization problem, alternating between updates on the user task and the safety alignment data, using a proximal regularizer to maintain proximity.
BESA (Wang et al. 2024): A backdoor-based defense that introduces a secret trigger into the safety alignment data. During inference, this trigger is added to user prompts to activate the model’s safety guardrails.
PTST (Lyu et al. 2024): This approach fine-tunes the model on user data without a safety system prompt. The safety prompt is then re-introduced during inference to restore safety behaviors.
Training Details. Following recent work (Huang et al. 2024e; Hsu et al. 2024), we use Low-Rank Adaptation (LoRA) (Hu et al. 2021) for efficient training. The LoRA adapter rank is set to 8 and alpha is 16. We fine-tune all models using the AdamW optimizer (Loshchilov et al. 2017) with a learning rate of 1e-5. Models are trained for 10 epochs with a batch size of 10. For all defense methods requiring a trusted alignment dataset (including ours), we use a set of 100 safety examples. For our proposed SafeGrad, the trade-off hyperparameter is set to 1.0 by default. All experiments were conducted on an NVIDIA A800-80G.
2 Main Results
Performance on Defending Harmful Fine-tuning Attacks. The performance of different defense methods against harmful fine-tuning attacks is detailed in Table 2. The results indicate that SafeGrad achieves state-of-the-art defense performance with exceptional robustness. Specifically, SafeGrad achieves an average Harmful Score (HS) of just 4.02, consistently maintaining the lowest score across all poison ratios () with the sole exception of the setting. This represents a reduction of over 7% compared to the best-performing baseline, PTST (11.70). Furthermore, SafeGrad demonstrates remarkable robustness; while its HS remains stable, the performance of other baselines like SafeInstr and LISA degrades sharply as the harmful ratio increases. Crucially, this robust defense does not negatively affect the model’s performance on benign tasks. SafeGrad maintains a high average Finetune Accuracy (FA) of 93.71, showing only a negligible drop compared to the undefended SFT baseline. This validates our core claim that by surgically resolving gradient conflicts, SafeGrad can effectively neutralize malicious updates while preserving the model’s utility for its intended benign task.
Generalizations to datasets. The generalization performance of SafeGrad across different datasets is presented in Table 3. For each datasets, we present the average performance across different ratios. The results demonstrate that SafeGrad’s defensive capabilities generalize effectively to diverse downstream tasks. Specifically, it consistently achieves the best defense performance (lowest HS) on all three evaluated datasets: SST2, AGNEWS, and the complex GSM8K reasoning task. On average, SafeGrad achieves an HS of 3.89, a significant reduction of over 7% compared to the best-performing baseline, PTST (11.04). Moreover, this robust safety does not compromise task utility. SafeGrad also secures the highest average FA of 81.94, confirming its versatility and effectiveness across varied task domains.
Generalizations to models. To verify that our method is model-agnostic, we evaluate its performance on three different LLMs, with results shown in Table 4. For each models, we present the average performance. The experiments confirm that SafeGrad successfully generalizes across different model architectures, achieving the best defensive results on all tested models, including Gemma-3-4B, Llama-3-8B, and Qwen2.5-7B. It achieves an average HS of 4.24, representing a substantial improvement of approximately 8 over the best-performing baseline PTST (12.59). While delivering state-of-the-art safety, SafeGrad maintains a highly competitive average FA of 94.39.
3 Hyper-parameter Analysis and Ablation Study
Impact of the size of alignment dataset. We then analyze the sensitivity of SafeGrad to the number of samples in the trusted alignment dataset, . This is a crucial factor for practical deployment, as curating large alignment datasets can be costly. The results in Table 5 demonstrate that SafeGrad is remarkably data-efficient. It achieves excellent safety performance (HS of 3.50) with as few as 20 alignment examples, while maintaining high task accuracy (FA of 94.04). Further increasing the dataset size up to 200 samples does not yield any significant additional benefits, as both the HS and FA metrics remain stable. This finding underscores the practicality of SafeGrad, showing that robust safety can be achieved with minimal data overhead.
Impact of different alignment objectives. To validate our choice of the alignment objective, we compare our proposed KL-divergence loss against a standard SFT loss (Table 6). The results demonstrate the superior data efficiency of our KL-based objective. It is remarkably effective in low-data regimes, achieving a strong HS of 4.6 with just 10 alignment samples. In stark contrast, the CE loss yields a catastrophic HS of 31.5 with the same data, requiring at least 100 samples to achieve comparable safety. This performance gap stems from the quality of the learning signal. The CE loss provides a sparse signal, narrowly focused on specific refusal tokens. Conversely, our KL-divergence objective offers a dense, distributional signal by guiding the model to learn the entire safety profile of the well-aligned reference model.
Impact of . The hyper-parameter controls the weight of the alignment gradient in the final update step, balancing the trade-off between the user’s fine-tuning task and the safety alignment objective. To investigate its effect, we vary from 0.5 to 100. As depicted in Table 7, we observe a clear trade-off. Increasing places a stronger emphasis on the safety objective, leading to a modest decrease in the HS from 4.40 to a low of 3.40. However, this comes at a cost to task performance. The FA gradually declines as increases, from 93.92 to 92.78. Notably, the safety gains become marginal for , while the FA continues to degrade. This suggests that an excessively large can harm task utility without providing significant additional safety benefits. Our default choice of strikes an effective balance.
4 Overhead Analysis
The superior defense performance of SafeGrad, while significant, does come with additional computational overhead, as detailed in Table 8. This increased overhead is primarily due to the necessity of a separate reference model to compute the KL divergence, which involves an additional forward pass. The gradient projection operation also contributes to this increased computational cost.
For scenarios where computational resources are a significant constraint, we recommend SafeGrad (SFT), which uses a standard SFT alignment objective that does not require a reference model. In a 1000-step fine-tuning task, SafeGrad (SFT) demonstrates the ability to achieve significantly better defense performance than the Lisa baseline while incurring a slightly lower computational overhead. This makes SafeGrad (SFT) a highly efficient and effective solution for robustly fine-tuning LLM in resource-limited environments.
5 Statistical Evaluation
Our gradient surgery mechanistically resolves the optimization conflict at each training step. Figure 4(a) illustrates this core mechanism. The “Before Gradient Surgery” trajectory reveals a persistent conflict under a high harmful ratio (), with the cosine similarity between the user-task and alignment gradients remaining consistently negative. Our method’s principle is direct: when such a conflict is detected, we project the user-task gradient to be orthogonal to the alignment gradient. The “After Gradient Surgery” trajectory provides clear visual proof, showing that the cosine similarity is precisely clamped at zero post-projection, thereby nullifying the harmful component of the update at the gradient level.
By resolving the underlying optimization conflict, SafeGrad preserves the model’s safety where other methods fail. This outcome is empirically validated in Figure 4(b), which compares the safety performance of SafeGrad against baselines. While the harmful scores for both SafeInstr and Lisa show a continuous and steep increase—indicating a progressive compromise of their safety alignment—SafeGrad’s score remains consistently low and stable. This demonstrates our gradient surgery’s effectiveness: by detecting the conflict as shown in Figure 4(a) and surgically removing the harmful gradient component, SafeGrad prevents alignment degradation, allowing the model to learn the benign task without sacrificing safety.
6 Visualization
In the following, we demonstrate how different methods respond to the malicious prompt. As illustrated below, SafeGrad is able to provide a safe answer to the sensitive question, while other methods gives harmful responses after undergoing harmful fine-tuning.
Related Work
Harmful Fine-tuning Attacks. Recent studies about harmful fine-tuning attacks (Qi et al. 2023; Yang et al. 2023; Zhan et al. 2023; Lermen et al. 2023; Chen et al. 2024; Rosati et al. 2024b; Yi et al. 2024a; Huang et al. 2024c; Huang et al. 2025b; He et al. 2024; Guan et al. 2025) show that introducing a few harmful fine-tuning data points can cause the aligned model to forget its safety alignment, rendering it vulnerable to exploitation for malicious tasks. Unlike jailbreak attacks (Zou et al. 2023; Huang et al. 2024f), which only interfere during the inference stage of LLMs, harmful fine-tuning attacks grant attackers elevated privileges, allowing them to directly alter model weights via the fine-tuning process. This makes defending against such attacks particularly challenging (Rosati et al. 2024a). Recent research also studies the mechanism of harmful fine-tuning (Leong et al. 2024; Peng et al. 2024; Hsiung et al. 2025; Qi et al. 2024b; Guo et al. 2024).
Alignment-stage Defenses. Defenses implemented during the initial alignment phase are predominantly centered on two strategies (Huang et al. 2024e; Rosati et al. 2024c; Rosati et al. 2024d; Huang et al. 2024d; Liu et al. 2024; Tamirisa et al. 2024; Yi et al. 2025b; Zheng et al. 2025; Wang et al. 2025a; Chen et al. 2025): strengthening alignment robustness via adversarial training (Huang et al. 2024e; Huang et al. 2024d; Tamirisa et al. 2024; Zheng et al. 2025; Wang et al. 2025a; Chen et al. 2025) and employing unlearning methods to erase harmful information (Zhang et al. 2024a; Zhang et al. 2024b; Rosati et al. 2024c). Furthermore, a newer line of research (Yi et al. 2025b; Wang et al. 2025b) has proposed a “collapse” or “self-destructive” trap mechanism as part of the alignment process.
Fine-tuning-stage Defenses. Existing fine-tuning-stage methods (Mukhoti et al. 2023; Huang et al. 2024b; Lyu et al. 2024; Wang et al. 2024; Qi et al. 2024a; Bianchi et al. 2024; Zong et al. 2024; Wei et al. 2024; Eiras et al. 2024b; Du et al. 2024; Li & Kim 2025; Li et al. 2025; Choi et al. 2024; Luo et al. 2024; Li et al. 2024; Eiras et al. 2024a) primarily defend against harmful fine-tuning from two perspectives: the data level and the algorithmic level. Data-level defenses (Choi et al. 2024; Shen et al. 2024) focus on filtering malicious data from the fine-tuning dataset.
Algorithmic-level defenses aim to design safe and robust fine-tuning algorithms. Some works (Li et al. 2025; Du et al. 2024; Li et al. 2024) identifies safety-critical parameters and seeks to keep them as constant as possible during fine-tuning. SafeInstr (Bianchi et al. 2024) mixes a small set of safety alignment examples into the user’s fine-tuning data to remind the model of its safety training. LISA (Huang et al. 2024b) treats safe fine-tuning as a bi-state optimization problem, alternating between updates on the user task and the safety alignment data. BESA (Wang et al. 2024) is a backdoor-based defense that introduces a secret trigger into the safety alignment data. PTST (Lyu et al. 2024) fine-tunes the model on user data without a safety system prompt, and the prompt is then re-introduced during inference to restore safety behaviors.
Post-fine-tuning-stage Defenses. Existing post-fine-tuning-stage methods (Hsu et al. 2024; Yi et al. 2024c; Huang et al. 2024a; Zhu et al. 2024; Casper et al. 2024; Wu et al. 2024; Gudipudi et al. 2024; Yi et al. 2024b; Yi et al. 2025a; Djuhera et al. 2025; Gong et al. 2025; Yang et al. 2025; Lu et al. 2025) primarily focus on filtering malicious inputs during the inference stage and realigning the post-fine-tuned model.
Conclusion
In this paper, we propose SafeGrad, a novel fine-tuning stage defense that employs gradient surgery to mitigate harmful fine-tuning attacks. When a conflict arises, SafeGrad nullifies the harmful component of the user-task gradient by projecting it onto the plane orthogonal to the alignment gradient. This defense is enhanced by a KL-divergence alignment loss, which learns the rich, distributional safety profile of the well-aligned foundation model. Extensive experiments demonstrate that SafeGrad achieves state-of-the-art defense, maintaining robust safety even at high harmful ratios with negligible impact on task fidelity.