SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection

Shuhao Chen, Weisen Jiang, Yeqi Gong, Shengda Luo, Chengxiang Zhuo, Zang Li, James T. Kwok, Yu Zhang

Introduction

Large language models (LLMs) (gpt4; touvron2023llama2; yang2024qwen; llama3-2; jiang2024forward) have shown strong capabilities across a wide range of tasks, making them increasingly popular in real-world applications (jiang2023effective; wei2024gita; chen2024routerdc). Fine-tuning-as-a-service has become a common way for users to adapt LLMs to specific downstream domains via service providers. However, fine-tuning can inadvertently undermine safety alignment, causing models to forget their safeguards (qi2024finetuning; yang2023shadow; lermen2023lora). This problem becomes more severe when fine-tuning data contains malicious or adversarial content, as in harmful fine-tuning attacks (liu2023jailbreaking; zou2023universal; huang2024harmful; jiang2026metadefense), which can effectively strip away safety mechanisms and cause the model to produce unsafe outputs.

Recently, many defense methods have been proposed to counter the harmful fine-tuning attacks. For example, PTST (lyu2024keeping) and SafeLoRA (hsu2024safe) mitigate harmful updates by re-injecting safety prompts or constraining LoRA adapters, but they either rely on carefully crafted prompt templates or require structural constraints that limit general applicability. Other approaches exploit safe data as an implicit safeguard. For instance, SafeInstr (bianchi2023safety) mixes a small fraction of safe examples into fine-tuning data to counter harmful behaviors, while Lisa (huang2024lisa) uses a bi-state optimization with safe samples and applies a proximal term to constrain the drift of each state. However, those methods suffer from two key drawbacks: (i) they treat safe data merely as a soft regularizer, which provides only weak control and makes it difficult to balance safety with downstream utility; and (ii) they typically select safe samples randomly, overlooking a fact that more relevant safe data can provide stronger corrective signals against harmful fine-tuning (as shown in Section 3.2). This motivates the need for a more principled defense and data selection method to robustly withstand harmful fine-tuning.

To address these challenges, we propose SPARD, a novel framework that defends aligned LLMs against harmful fine-tuning attacks by combining Safety-Projected Alternating optimization with Relevance-Diversity aware data selection. Figure 1 illustrates the overall procedure of SPARD. As shown, SPARD consists of two complementary components. First, we study the safety-constrained fine-tuning problem, and introduce Safety-Projected Alternating Gradient (SPAG), an optimization strategy that alternates between utility-driven updates on the fine-tuning data and explicit safety projections onto a constraint set defined by safe data. Unlike penalty-based approaches, SPAG enforces feasibility in a closed form, ensuring that safety alignment is preserved throughout training. Second, we recognize that the effectiveness of the safety projection critically depends on the choice of safe data. That is, not all safe samples are equally informative: samples that align closely with the downstream task could provide stronger corrective signals than others. To this end, we develop a Relevance–Diversity Determinantal Point Process (DPP) that selects a compact subset of safe data to balance the task relevance and behavioral diversity, ensuring broad and effective coverage against harmful attacks. Together, those two components yield a principled defense framework that maintains downstream utility while robustly constraining unsafe behaviors.

We conduct experiments on GSM8K (cobbe2021training) and OpenbookQA (mihaylov2018can) with four harmful finetuning attacks to evaluate both the utility and safety of SPARD. Empirical results show that SPARD can effectively mitigate harmful behaviors, achieving the lowest average Attack Success Rate (ASR) with high downstream accuracy. Moreover, SPARD significantly outperforms SafeInstr on average in ASR, showing the effectiveness of SPAG and Relevance–Diversity DPP.

Our contributions are summarized as follows: 1. We propose SPARD, a novel defense framework that integrates safety-projected optimization with relevance–diversity–aware data selection to robustly defend aligned LLMs against harmful fine-tuning. 2. We introduce SPAG, a principled optimization method solving the novel safety-constrained fine-tuning problems that alternates between utility updates and explicit safety projections, and a Relevance–Diversity DPP to select compact, task-aligned, and diverse safety subsets. 3. Through extensive experiments on two aligned LLMs, multiple downstream tasks, and diverse attack datasets, we show that SPARD consistently outperforms existing defense methods in reducing the attack success rates while maintaining the utility.

Related Works

Ensuring that large language models (LLMs) behave in a safe and helpful manner is a central challenge in AI research (yao2024survey). Recent foundation models such as LLaMA (llama3-2; touvron2023llama2) and Qwen (yang2024qwen) have been aligned with safety guardrails to reject harmful instructions and follow user intent more reliably. A common paradigm for alignment is preference-based learning, most notably Reinforcement Learning from Human Feedback (ouyang2022training; ziegler2019fine; bai2022training; lin2025parm), which optimizes models to maximize human-preferred responses. Subsequent work has proposed more efficient formulations, such as Direct Preference Optimization (rafailov2023direct), reward-free methods like RRHF (yuan2023rrhf), which reduce reliance on expensive reward models while maintaining alignment quality. However, a critical vulnerability remains: the safety alignment achieved through these expensive procedures is often brittle and can be easily compromised or erased through subsequent downstream fine-tuning (yang2023shadow; yi2024vulnerability; qi2024finetuning; lermen2023lora; zhan2023removing; hsiung2025your; li2025salora), which motivates the need for robust defense mechanisms.

Defending Against Harmful Fine-tuning.

Harmful fine-tuning attacks (huang2024harmful; liu2023jailbreaking; zou2023universal; yuan2023gpt) compromise aligned models by poisoning training data with adversarial prompts that bypass safety guardrails. To counter such risks, many defense strategies (huang2024booster; huang2024vaccine; liu2025targeted; chen2025vulnerability; lyu2024keeping; hsu2024safe; bianchi2023safety; huang2024lisa; yi2025gradient; jiang2026metadefense) have been proposed. For example, MetaDefense (jiang2026metadefense) defends against finetuning-based jailbreak attacks by training a single LLM to predict the harmfulness of both incoming queries before generation and partial responses during generation. PTST (lyu2024keeping) avoids safety degradation by fine-tuning solely on task data and re-introducing safety prompts at inference time. SafeLoRA (hsu2024safe) constrains harmful updates by projecting LoRA weights from selected layers into a safety-aligned subspace. Other approaches leverage safe data as an implicit safeguard. For example, SafeInstr (bianchi2023safety) mixes a small fraction of safe examples into fine-tuning data, while Lisa (huang2024lisa) uses safe samples by a bi-state optimization with a proximal regularization term. Although effective to some extent, these methods either rely on carefully tuned penalty weights or only weakly address the utility–safety tradeoff. In contrast, our proposed SPAG provides an optimization-grounded solution by explicitly projecting the model back into the safe region. This adaptive projection automatically determines the correction size, removing the need for manual weight tuning while simultaneously preserving downstream utility.

Data Selection for LLMs.

The quality and composition of training data are critical determinants of LLM performance (zhou2023lima; gadre2023datacomp). This has motivated a growing body of work on data selection (albalak2024survey), which seeks to curate smaller yet more effective subsets from vast, noisy corpora. Selection strategies span a broad spectrum: filtering based on perplexity or linguistic complexity (longpre2024pretrainer), identifying core sets that approximate the full dataset’s training dynamics (sorscher2022beyond), or leveraging embedding similarity to retrieve samples closer to the target distribution for task-specific fine-tuning (liu2021makes; xia2024less; hsiung2025your; liu2025pharmacist). However, relevance-based selection alone often leads to redundancy. To mitigate this, we propose a novel approach to achieve a relevance-diversity trade-off.

where LC{\bf L}_{\mathcal{C}} is the principal submatrix indexed by C\mathcal{C}, I{\bf I} is the identity matrix, and det⁡(⋅)\det(\cdot) is the determinant of a matrix. The denominator det⁡(I+L)\det({\bf I}+{\bf L}) is a constant independent of the subset selection, thus can be ignored. Intuitively, det⁡(LC)\det({\bf L}_{\mathcal{C}}) corresponds to the squared volume spanned by the feature vectors of C\mathcal{C}, favoring subsets whose elements are both individually informative and mutually dissimilar. Yet, traditional DPPs ignore task relevance. Our work closes this gap by introducing a Relevance–Diversity DPP, which incorporates task-relevance quality scores directly into the DPP kernel.

Methodology

In the proposed SPARD method, we formulate fine-tuning as a safety-constrained optimization problem to adapt the model to downstream data without compromising its safety alignment:

where Dft\mathcal{D}_{\text{ft}} is the fine-tuning dataset, Dsafe\mathcal{D}_{\text{safe}} is the safety dataset, and τ\tau is a predefined threshold. In practice, τ\tau can be set by measuring the average safety loss of the pretrained LLM on Dsafe\mathcal{D}_{\text{safe}}.

A common approach (huang2024lisa; yi2025gradient; bianchi2023safety) to relax the constraint is to add it to the objective as a penalty term: min⁡θ  L(Dft,θ)+λ(L(Dsafe,θ)−τ),\min_{{\bm{\theta}}}\;\mathcal{L}(\mathcal{D}_{\text{ft}},{\bm{\theta}})+\lambda\big(\mathcal{L}(\mathcal{D}_{\text{safe}},{\bm{\theta}})-\tau\big), with a penalty parameter λ≥0\lambda\geq 0 controlling the balance between utility and safety. However, this blending lacks explicit control on safety, since the constraint is only enforced indirectly through the objective.

Instead of implicitly encouraging safety via a soft penalty, we directly enforce the constraint using a projection-based strategy. The key idea is to first perform a utility-driven update (using Dft\mathcal{D}_{\text{ft}}) and then project the updated parameters back into a region where the safety constraint is approximately satisfied. This alternating update scheme avoids the difficulty of choosing penalty weights, while providing a principled geometric correction that guarantees feasibility up to first order.

Specifically, after a utility update θ+=θ−ηft∇L(Dft,θ),{\bm{\theta}}^{+}={\bm{\theta}}-\eta_{\text{ft}}\nabla\mathcal{L}(\mathcal{D}_{\text{ft}},{\bm{\theta}}), where ηft\eta_{\text{ft}} is the learning rate for the utility step and ∇L(Dft,θ)\nabla\mathcal{L}(\mathcal{D}_{\text{ft}},{\bm{\theta}}) denotes the gradient, we project θ+{\bm{\theta}}^{+} back into a linearized safety region determined by the safety loss. Using a first-order Taylor expansion at θ+{\bm{\theta}}^{+}, we approximate L(Dsafe,θ)\mathcal{L}(\mathcal{D}_{\text{safe}},{\bm{\theta}}) as L(Dsafe,θ)≈L(Dsafe,θ+)+⟨gsafe, θ−θ+⟩,\mathcal{L}(\mathcal{D}_{\text{safe}},{\bm{\theta}})\approx\mathcal{L}(\mathcal{D}_{\text{safe}},{\bm{\theta}}^{+})+\langle{\bf g}_{\text{safe}},\,{\bm{\theta}}-{\bm{\theta}}^{+}\rangle, where gsafe=∇L(Dsafe,θ+){\bf g}_{\text{safe}}=\nabla\mathcal{L}(\mathcal{D}_{\text{safe}},{\bm{\theta}}^{+}). This defines the half-space C+={θ:  L(Dsafe,θ+)+⟨gsafe, θ−θ+⟩≤τ},\mathcal{C}^{+}=\{{\bm{\theta}}:\;\mathcal{L}(\mathcal{D}_{\text{safe}},{\bm{\theta}}^{+})+\langle{\bf g}_{\text{safe}},\,{\bm{\theta}}-{\bm{\theta}}^{+}\rangle\leq\tau\}, which is a local approximation of the feasible set around θ+{\bm{\theta}}^{+}.

The projection step seeks the point in C+\mathcal{C}^{+} that is closest to θ+{\bm{\theta}}^{+}:

Introducing multipliers for the constraint in Eq. (3), the KKT conditions (bertsekas1997nonlinear) give the projected solution as

The detailed derivation is provided in Appendix A. When the safety condition is violated, to stabilize training by avoiding arbitrarily large projections, we adopt the trust-region optimization strategy (schulman2015trust) to limit the step size as α=min⁡ ⁣(L(Dsafe,θ+)−τ∥gsafe∥2,  ηsafe)\alpha=\min\!\left(\frac{\mathcal{L}(\mathcal{D}_{\text{safe}},{\bm{\theta}}^{+})-\tau}{\|{\bf g}_{\text{safe}}\|^{2}},\;\eta_{\text{safe}}\right), and then θnew=θ+−α gsafe{\bm{\theta}}^{\text{new}}={\bm{\theta}}^{+}-\alpha\,{\bf g}_{\text{safe}}, where ηsafe\eta_{\text{safe}} is a trust-region radius. The trust-region constraint ensures the updated model θnew{\bm{\theta}}^{\text{new}} stays within a ball centered at θ+{\bm{\theta}}^{+} with radius ηsafe∥gsafe∥\eta_{\text{safe}}\|{\bf g}_{\text{safe}}\|. According to Eq. (4), the update either keeps θ+{\bm{\theta}}^{+} if already safe, or applies a corrective step along gsafe{\bf g}_{\text{safe}} with magnitude determined by projection and trust region.

The algorithm of SPAG is illustrated in Algorithm 3.1. In summary, SPAG provides a simple yet principled mechanism for safety-constrained fine-tuning: it alternates between utility optimization and explicit safety projection. Geometrically, the method corrects each update by projecting onto a safety half-space, thereby directly solving the constraint rather than relying on soft penalties. Unlike penalty-based approaches, which require careful tuning of λ\lambda and offer no guarantee of feasibility, SPAG yields a closed-form projection step that enforces the constraint up to the first order. This combination of interpretability, guaranteed feasibility, and hyperparameter-free correction makes SPAG both practical and robust for safety-aligned fine-tuning.

2 Relevance- and Diversity-Aware Safety Data Selection

While SPAG provides a principled projection mechanism for enforcing safety constraints, its success fundamentally depends on the quality of the safety dataset Dsafe\mathcal{D}_{\text{safe}}. Recent studies have also observed that relevant safety data can improve safety during training by leveraging embedding similarity (hsiung2025your), employing trained selectors (liu2025pharmacist), or matching task styles and formats (eiras2024safely; xiao2025style). Here we claim that not all safety samples contribute equally. That is, safety samples, which align well with the fine-tuning domain Dft\mathcal{D}_{\text{ft}}, could provide stronger corrective signals than other safety samples. Therefore, carefully selecting relevant safety data becomes critical to ensuring that the projection step effectively constrains the model.

Relevance Improves Safety. To better understand the role of relevance, we conduct an experiment, where the GSM8K training data are merged with 10% BeaverTails attack data (ji2023beavertails) as Dft\mathcal{D}_{\text{ft}}, and safe samples are selected from BeaverTails and LatHarmful (sheshadri2024targeted) defense sets according to the similarity with Dft\mathcal{D}_{\text{ft}}. We define the similarity quality of a candidate xi∈Dsafe{\bf x}_{i}\in\mathcal{D}_{\text{safe}} as

where sim⁡(⋅,⋅)\operatorname{\textsf{sim}}(\cdot,\cdot) is cosine similarity in the embedding space. As shown in Figure 2, the attack success rate (ASR) decreases sharply as the average similarity of selected samples increases, dropping from 68.8%68.8\% at low similarity to 11.4%11.4\% at moderate-to-high similarity. This confirms that task-relevant safe samples provide stronger constraints and substantially enhance robustness.

Notably, the curve also shows that ASR rises again (to 16.6%16.6\%) when the selected samples are too similar to the fine-tuning data (e.g., average similarity ≈0.94\approx 0.94). This counterintuitive phenomenon occurs because extreme similarity introduces redundancy: the selected safety samples cover only a narrow region of the risk space, leaving other harmful behaviors underrepresented. Consequently, the model may overfit to a small set of highly similar constraints, reducing the effectiveness of safety alignment.

Relevance–Diversity DPP. The above observation motivates a selection strategy that balances the relevance and diversity: the relevance ensures that the chosen samples provide strong and task-aligned safety signals, while the diversity ensures broad coverage of distinct harmful behaviors. While hsiung2025your introduces a metric for assessing subset diversity, they do not integrate this metric into the selection process, so their method cannot promote diversity during data selection. To achieve this, we extend Determinantal Point Processes (DPPs) (macchi1975coincidence; kulesza2012determinantal; ye2023compositional), which naturally promote diversity through determinant-based subset probabilities, by incorporating the relevance into the kernel design.

Specifically, given relevance scores qiq_{i} (defined in Eq. (5)) for candidates xi∈Dsafe{\bf x}_{i}\in\mathcal{D}_{\text{safe}}, We then construct the kernel as

where K(xi,xj)\mathcal{K}({\bf x}_{i},{\bf x}_{j}) captures the intrinsic similarity between two safety samples (e.g., the cosine similarity in an embedding space), and β≥0\beta\geq 0 controls the influence of the relevance. Intuitively, (qi⋅qj)β(q_{i}\cdot q_{j})^{\beta} acts as a multiplicative weight that increases the likelihood of including pairs of samples that are both highly relevant to the fine-tuning distribution (a sensitive analysis of β\beta is provided in Section 4.3). When β=0\beta=0, the kernel reduces to the classical diversity-only DPP, while larger β\beta biases the distribution toward relevance-aware subsets.

Let L{\bf L} and L^\widehat{{\bf L}} be the kernel matrix corresponding to K\mathcal{K} and K^\widehat{\mathcal{K}}, respectively. For any subset C⊆Dsafe\mathcal{C}\subseteq\mathcal{D}_{\text{safe}}, according to Eq. (1), the selection probability is given by

where LC{\bf L}_{\mathcal{C}} denotes the kernel submatrix with indices in C\mathcal{C}. This decomposition makes the roles explicit: the factor ∏qi2β\prod q_{i}^{2\beta} rewards subsets containing highly relevant samples, while det⁡(LC)\det({\bf L}_{\mathcal{C}}) enforces diversity among them. Thus, the relevance–diversity DPP jointly balances the task alignment and coverage, ensuring that the selected safety set is neither irrelevant nor redundant.

Efficient Greedy Selection. Although DPPs define a principled probability distribution, exactly solving the maximum a posteriori (MAP) problem arg max⁡C⊆Dsafedet⁡(L^C)\operatorname*{arg\,max}_{\mathcal{C}\subseteq\mathcal{D}_{\text{safe}}}\det(\widehat{{\bf L}}_{\mathcal{C}}) is computationally expensive, requiring to calculate determinants of many submatrices. To scale the selection to large safety datasets, we adopt a greedy approximation that incrementally builds the subset by adding one sample at a time, each chosen to maximize the marginal gain in the determinant.

Specifically, suppose we have already selected Cm−1\mathcal{C}_{m-1} after m−1m-1 steps. For a candidate i∉Cm−1i\notin\mathcal{C}_{m-1}, the expanded kernel matrix is L^Cm−1∪{i}=(L^Cm−1vivi⊤Lii),\widehat{{\bf L}}_{\mathcal{C}_{m-1}\cup\{i\}}=\begin{pmatrix}\widehat{{\bf L}}_{\mathcal{C}_{m-1}}&{\bf v}_{i}\\ {\bf v}_{i}^{\top}&{\bf L}_{ii}\end{pmatrix}, where vi{\bf v}_{i} contains kernel similarities between xi{\bf x}_{i} and the already-selected set Cm−1\mathcal{C}_{m-1}. By the Schur complement, the determinant after including ii can be factorized as

The second term in the right-hand side of Eq. (8), known as the gain factor, measures the additional volume contributed by xi{\bf x}_{i} that is not already spanned by Cm−1\mathcal{C}_{m-1}. Intuitively, it rewards candidates that are both individually relevant (large L^ii\widehat{{\bf L}}_{ii}) and novel relative to the current set (small vi⊤L^Cm−1−1vi{\bf v}_{i}^{\top}\widehat{{\bf L}}_{\mathcal{C}_{m-1}}^{-1}{\bf v}_{i}).

To compute the gain efficiently, we maintain the Cholesky decomposition L^Cm−1=CC⊤\widehat{{\bf L}}_{\mathcal{C}_{m-1}}={\bf C}{\bf C}^{\top}. Then, vi⊤L^Cm−1−1vi=(C−1vi)⊤(C−1vi)=∥wi∥2,{\bf v}_{i}^{\top}\widehat{{\bf L}}_{\mathcal{C}_{m-1}}^{-1}{\bf v}_{i}=({\bf C}^{-1}{\bf v}_{i})^{\top}({\bf C}^{-1}{\bf v}_{i})=\|{\bf w}_{i}\|^{2}, where wi{\bf w}_{i} is obtained by solving the triangular system Cwi=vi{\bf C}{\bf w}_{i}={\bf v}_{i}. This avoids explicitly inverting L^Cm−1\widehat{{\bf L}}_{\mathcal{C}_{m-1}}, reducing the complexity per iteration to O(m)O(m) instead of cubic cost. At each step, we select xi⋆=arg max⁡xi∈Dsafe∖Cm−1(L^ii−∥wi∥2),{\bf x}_{i^{\star}}=\operatorname*{arg\,max}_{{\bf x}_{i}\in\mathcal{D}_{\text{safe}}\setminus\mathcal{C}_{m-1}}\Big(\widehat{{\bf L}}_{ii}-\|{\bf w}_{i}\|^{2}\Big), and add it into Cm−1\mathcal{C}_{m-1} to obtain Cm\mathcal{C}_{m}. Additional details for computational cost are provided in Appendix B.

Experiments

Datasets. Safety corpora. We use four datasets containing harmful or jailbreak-style prompts with both harmful and safe responses: (i) BeaverTails (ji2023beavertails), a collection of safety-related QA pairs with helpfulness and harmlessness annotations; (ii) I-BeaverTails, constructed by converting BeaverTails questions into instructions using GPT-4o-mini (hurst2024gpt) following bianchi2023safety; (iii) LatHarmful (sheshadri2024targeted), consisting of 5k instructions with paired harmful and harmless completions; and (iv) Q-LatHarmful, obtained by converting LatHarmful instructions into QA pairs with GPT-4o-mini. Each dataset is split 90%/10% into training and testing. From these corpora, we use harmful queries with safe responses from the training splits to build GeneralSafe, the candidate pool for selecting safety data. In contrast, to simulate harmful fine-tuning, we use harmful queries with harmful responses from the training splits to inject malicious samples into downstream utility tasks. Following huang2024lisa; hsu2024safe, the number of injected harmful samples is fixed to 10% of the original utility task training size, ensuring a consistent attack intensity across tasks.

Utility tasks. For downstream performance, we evaluate on (i) GSM8K (cobbe2021training), a benchmark of grade-school math word problems, augmented with MetaMath (yu2023metamath) for broader coverage; and (ii) OpenBookQA (mihaylov2018can), a science QA dataset requiring factual reasoning.

Evaluation Metrics. Following hsu2024safe, we evaluate methods on two dimensions: Safety and Utility. For safety, we adopt the protocol of qi2024finetuning, using GPT-4o-mini to judge responses under OpenAI’s 11 harmful content categories. Each response receives a Harmfulness Score (HS) from 1 (safest) to 5 (most harmful), and we report the Attack Success Rate (ASR), the proportion of responses with HS >2>2. For utility, we measure accuracy on downstream tasks (GSM8K or OpenBookQA), reflecting utility performance under when safety defense is enforced.

Baselines. We compare SPARD against a range of baselines using two safe pre-trained models, Qwen-2.5-7B-Instruct (yang2024qwen) and LLaMA-3.2-3B-Instruct (llama3-2). Specifically, we consider: 1. SFT, which is a standard fine-tuning baseline where the model is trained exclusively on the target task data without any explicit safety measures. 2. PTST (lyu2024keeping), which fine-tunes the model without safety instructions but prepends them back to inputs at inference time. 3. SafeInstr (bianchi2023safety), which randomly mixes a small fraction (3%) of safe samples into the fine-tuning dataset. 4. Lisa (huang2024lisa), which bi-state learning finetuning samples and safe samples with a proximal term to constrain the safety degradation of each state. 5. SafeGrad (yi2025gradient), which detects conflict safety/alignment gradients and projects out the harmful component.

We employ LoRA (hu2022lora) for parameter-efficient fine-tuning, with a rank r=32r=32 and an alpha of 44. All models are trained using the AdamW optimizer (loshchilov2017decoupled). We set the learning rate to 5×10−55\times 10^{-5} for both Qwen-2.5-7B-Instruct and LLaMA-3.2-3B-Instruct. The models are fine-tuned for 10 epochs on GSM8K and 3 epochs on OpenBookQA. For our SPAG algorithm, the safety mini-batch Bsafe\mathcal{B}_{\text{safe}} is sampled with a batch size of 1. The trust region radius ηsafe\eta_{\text{safe}} is set equal to the fine-tuning learning rate ηft\eta_{\text{ft}}. The hyperparameters τ\tau, δ\delta, and ϵ\epsilon are chosen as 0.20.2, 0.10.1, and 1×10−81\times 10^{-8} for all experiments. For our Relevance-Diversity DPP data selection, sample embeddings are generated by taking the average of the final layer’s hidden states from the pretrained model (i.e., Qwen-2.5-7B-Instruct and LLaMA-3.2-3B-Instruct). Based on these embeddings, we select Dsafe\mathcal{D}_{\text{safe}} from the GeneralSafe pool, with the size fixed to 3%3\% (i.e., p=0.03p=0.03) of the fine-tuning data. The relevance exponent is set to β=4\beta=4.

2 Main Result

Robustness to Different Attacks. Table 1 presents results on GSM8K under four harmful fine-tuning attacks. As can be seen, SPARD simultaneously achieves the lowest ASR/HS while preserving competitive GSM8K accuracy, offering the strongest balance between safety and utility among all methods.

Specifically, compared with SFT and PTST, SPARD achieves substantial gains: average ASR is reduced by over 65%65\% and HS by 2.322.32 points, while accuracy remains competitive. This shows that explicit safety projection is far more effective than standard fine-tuning or inference-time prompting. Compared with SafeInstr, which randomly mixes safe samples into fine-tuning, SPARD consistently achieves lower ASR/HS, highlighting the necessity of principled relevance–diversity selection over naive random selection. Compared with Lisa, a strong optimization-based baseline, SPARD achieves consistently better safety on all datasets, reducing ASR by 9.67%9.67\%, lowering HS by 0.240.24, and significantly improving GSM8K accuracy by 7.32%7.32\%. Compared with SafeGrad, SPARD achieves a substantially lower ASR of 20.73%20.73\% while having a slightly better downstream performance. This demonstrates that the key designs of SPARD—SPAG safety projection and relevance–diversity DPP selection—are crucial for constraining harmful behaviors while preserving downstream task performance.

Generalization Across Architectures (LLaMA). Table 2 presents results on LLaMA-3.2-3B-Instruct. We observe that the overall safety degradation is more severe on LLaMA than on Qwen, as SFT and PTST both yield very high ASR/HS despite maintaining task accuracy. SPARD, however, remains effective across model families, achieving the lowest ASR (12.09%12.09\%) and HS (1.381.38) while keeping accuracy competitive (71.23%71.23\%). Compared with SFT and PTST, SPARD lowers ASR by over 67%67\%, confirming that its safety projection generalizes across backbones. Against SafeInstr, which suffers from high ASR/HS, SPARD shows the value of relevance–diversity selection over naive random mixing. Relative to Lisa, SPARD further improves both safety and utility (−12.10%-12.10\% ASR, −0.35-0.35 HS, +6.20%+6.20\% accuracy). SPARD also outperforms SafeGrad, reducing ASR by over 59%59\% while achieving a substantial 6.94%6.94\% improvement in accuracy. Together, these results demonstrate that SPAG optimization and DPP-based selection enhance robustness consistently across architectures.

Generalization to OpenBookQA. Table 3 reports results on OpenBookQA under four harmful fine-tuning attacks. SPARD achieves the lowest ASR (14.54%14.54\%) and HS (1.451.45) while preserving strong task accuracy (83.25%83.25\%), confirming that its effectiveness extends beyond math reasoning tasks. Compared with SFT and PTST, SPARD reduces ASR by more than 15.3%15.3\% on average, showing that explicit safety projection remains effective for science QA. Relative to SafeInstr, which again suffers from high ASR/HS, SPARD demonstrates the importance of relevance–diversity selection over naive data mixing. Compared with Lisa, SPARD achieves lower ASR/HS (−4.14%-4.14\%/−0.05-0.05) and higher accuracy (+4.05%+4.05\%), reinforcing that its joint use of SPAG optimization and DPP-based selection improves both safety and utility. Finally, SPARD outperforms SafeGrad with an improvement of 5%5\% on ASR, validating the effectiveness of SPAG’s safety projection. These results highlight that SPARD generalizes beyond GSM8K to diverse downstream reasoning tasks, maintaining robustness across domains.

3 Analysis

Effects of safe sample ratio. Figure 5 shows the impact of varying the ratio pp of safe samples added to the GSM8K finetuning data using Qwen-2.5-7B-Instruct under the BeaverTails attack. When p=0p=0 (i.e., no safe samples are added), the model is highly vulnerable with ASR above 80%80\%. As pp increases, ASR drops sharply and reaches the lowest point around p∈[0.03,0.05]p\in[0.03,0.05]. Beyond this range, further increasing pp leads to diminishing returns and even slight degradation due to the inclusion of redundant or less relevant samples. This indicates that a small but carefully chosen proportion p∈[0.03,0,1]p\in[0.03,0,1] of safe samples is sufficient to provide strong safety guarantees without overwhelming the fine-tuning objective.

Effects of τ\tau. We study the effects of the safety threshold τ\tau on GSM8K using Qwen-2.5-7B-Instruct under the BeaverTails attack. As shown in Figure 5, small values of τ\tau enforce strict safety constraints, effectively suppressing ASR, but overly conservative thresholds (τ>0.5\tau>0.5) begin to harm the balance and allow ASR to rise again. In practice, we can set the τ\tau by referencing the average loss of the aligned LLM on the safety benchmark.

Effects of β\beta. To analyze the effect of relevance exponent β\beta, we conduct experiments with the BeaverTails attack on GSM8K using Qwen-2.5-7B-Instruct. Figure 5 analyzes the relevance exponent β\beta, which balances the weight between relevance and diversity in the DPP kernel. SPARD is relatively robust to a wide range of moderate values (β∈\beta\in), achieving the lowest ASR, while very small β\beta underemphasizes relevance and very large β\beta collapses diversity, both leading to weaker defenses. These results confirm that both relevance and diversity should be considered in the data selection process.

Effects of ηsafe.\eta_{\text{safe}}. We conduct an ablation removing the trust-region limit ηsafe\eta_{\text{safe}}. Table 4 reports the ASR under multiple attacks and the average downstream GSM8K accuracy. Removing ηsafe\eta_{\text{safe}} results in more aggressive projection updates: safety improves for some attacks, but downstream utility degrades substantially. These results highlight that ηsafe\eta_{\text{safe}} plays a crucial role in stabilizing the projection step, achieving a more balanced trade-off between strong safety and good downstream performance.

Effect of Relevance-Diversity DPP. To study the effect of Relevance-Diversity DPP, we compare it with (i) SPAG w/ Random, which randomly selects samples from GeneralSafe as Dsafe\mathcal{D}_{\text{safe}}. (ii) SPAG w/ Max Quality, which selects the samples with the highest quality score as Dsafe\mathcal{D}_{\text{safe}}. As shown in Table 5, SPAG w/ Random surpasses previous SOTA (i.e., Lisa) with an average ASR reduction of 3.27%3.27\% and an GSM8K accuracy improvement of 6.61%6.61\%, validating the effectiveness of SPAG safety projection. Compared with all variants, SPARD has the best average safety and utility, achieving the lowest mean ASR/HS (9.45%9.45\%/1.321.32) and the highest GSM8K accuracy (85.77%85.77\%). Specifically, SPARD outperforms SPAG w/ Random with a noticeable ASR and HS reduction of 6.40%6.40\% and 0.220.22, showing that selecting relevant data can substantially improve safety. Additionally, SPARD surpasses SPAG w/ Max Quality by a large margin of 7.06%7.06\% on average ASR, validating that diversity is equally crucial. By balancing both relevance and diversity, SPARD achieves broad coverage of safety constraints while remaining task-aligned, leading to superior robustness without sacrificing utility. Moreover, as hsiung2025your also explores similarity and diversity metrics in safety data curation, we provide further discussion on the similarities and differences, along with an empirical comparison of the two methods, in Appendix E.1.

Full Fine-Tuning Without LoRA. SPAG’s projection (Eq. (4)) operates on whatever parameters are being updated, independent of LoRA structure. To verify this, we evaluate SPARD with full fine-tuning on SmolLM2-1.7B-Instruct (GSM8K, BeaverTails attack). As shown in Table 6, SPARD achieves the lowest ASR (17.4%17.4\%) and HS (1.551.55) while maintaining competitive accuracy, confirming that the findings are consistent without LoRA.

Visualization. Figure 6 shows the t-SNE visualization (van2008visualizing) of selected samples for the GSM8K task under BeaverTails attacks. As shown, randomly selected data cover diverse regions but are not necessarily aligned with the attacked distribution, leading to limited safety gains. Moreover, a quality-only strategy (Max Quality) selects samples that cluster tightly around the attack distribution, but suffers from severe redundancy. In contrast, SPAG achieves a balanced selection that aligns samples closely with the attacked distribution while maintaining diversity across different safety corpora, ensuring broad coverage without redundancy. This suggests that our method is effective in selecting safe samples that are both relevant to the task and diverse (Tables 5).

Computational Overhead. The preprocessing unique to SPARD is lightweight: on a single A800 GPU, embedding extraction takes 3.42 minutes, and DPP subset selection completes in just 0.12 seconds (see Appendix B for complexity analysis). Table 7 further reports the end-to-end wall-clock time (including all preprocessing and training) for all methods fine-tuning Qwen-2.5-7B-Instruct on GSM8K with a single A800 GPU. SPARD adds only 8.1% overhead over SFT, less than Lisa (+9.8%) and substantially lower than SafeGrad (+54.2%). Given the considerable safety improvements SPARD achieves over these baselines (Tables 1–3), this marginal additional cost is well justified.

To better understand the role of β\beta, we analyze the distribution of similarity scores qiq_{i} between GeneralSafe samples and the GSM8K dataset under the BeaverTails attack. As shown in Figure 7, most samples already exhibit very high similarity: 69%69\% of them have qi>0.9q_{i}>0.9. This heavy concentration near the upper bound makes it difficult to distinguish relative preferences using linear weighting. By introducing β\beta as an exponent in the relevance term, we amplify subtle differences among highly similar samples, allowing the selection process to more effectively favor those that are most aligned with the target distribution.

Conclusion

In this paper, we propose SPARD, a defense framework that safeguards aligned LLMs against harmful fine-tuning by combining Safety-Projected Alternating Gradient (SPAG) with a Relevance–Diversity DPP for safe data selection. SPAG enforces safety constraints in closed form during training, while the Relevance–Diversity DPP selects task-relevant and diverse safety data to maximize coverage. Experiments on GSM8K and OpenBookQA with multiple attacks show that SPARD achieves the lowest average ASR while preserving high utility.

Acknowledgement

This work was supported by National Natural Science Foundation of China under Grant no. 62136005, Shenzhen fundamental research program JCYJ20250604144724032, the Research Grants Council of the Hong Kong Special Administrative Region (Grants 16202523 and HKU C7004-22G), Talents Cultivation Program of National Administration of Traditional Chinese Medicine (Grant No. ZYYCXTD-D-202403), the National Natural Science Foundation of China (Grant No. 82505358 and Grant No. 12326604), and the Scientific Research Start-up Funds of the Chinese Medicine Guangdong Laboratory (Grant No. HQL2025SU011).

Impact Statement

This research investigates the vulnerabilities of large language models (LLMs) to harmful fine-tuning attacks and introduces methods to strengthen their safety alignment. All datasets employed in our experiments are publicly available and widely used in the safety community. Although these datasets contain harmful or adversarial prompts, they are utilized solely for the purpose of evaluating defenses. Harmful responses are restricted to controlled experimental settings and are not disseminated beyond what is strictly necessary for reproducibility. The overarching aim of this work is to advance the safe and responsible deployment of LLMs by providing principled defense mechanisms against malicious fine-tuning. We recognize that research in this area carries potential dual-use concerns, but we believe the benefits of improving the robustness of safety alignment outweigh these risks. Our study adheres to ethical standards and prioritizes the promotion of beneficial and safe AI.

References

Appendix A Derivation of SPAG

Introducing multipliers λ≥0\lambda\geq 0, the Lagrangian is

Taking derivatives with respect to θ{\bm{\theta}} and setting to zero yields the stationarity condition: θ−θ++λgsafe=0.{\bm{\theta}}-{\bm{\theta}}^{+}+\lambda{\bf g}_{\text{safe}}=0. Hence, the solution has the form θnew=θ+−λgsafe.{\bm{\theta}}^{\text{new}}={\bm{\theta}}^{+}-\lambda{\bf g}_{\text{safe}}. Plugging θnew{\bm{\theta}}^{\text{new}} into the safety constraint gives L(Dsafe,θ+)+⟨gsafe, θnew−θ+⟩=L(Dsafe,θ+)−λ∥gsafe∥2.\mathcal{L}(\mathcal{D}_{\text{safe}},{\bm{\theta}}^{+})+\big\langle{\bf g}_{\text{safe}},\,{\bm{\theta}}^{\text{new}}-{\bm{\theta}}^{+}\big\rangle=\mathcal{L}(\mathcal{D}_{\text{safe}},{\bm{\theta}}^{+})-\lambda\|{\bf g}_{\text{safe}}\|^{2}. Hence feasibility requires L(Dsafe,θ+)−λ∥gsafe∥2≤τ.\mathcal{L}(\mathcal{D}_{\text{safe}},{\bm{\theta}}^{+})-\lambda\|{\bf g}_{\text{safe}}\|^{2}\leq\tau. Complementary slackness further implies λ(L(Dsafe,θ+)−λ∥gsafe∥2−τ)=0\lambda\Big(\mathcal{L}(\mathcal{D}_{\text{safe}},{\bm{\theta}}^{+})-\lambda\|{\bf g}_{\text{safe}}\|^{2}-\tau\Big)=0. Two cases arise: 1. If L(Dsafe,θ+)≤τ\mathcal{L}(\mathcal{D}_{\text{safe}},{\bm{\theta}}^{+})\leq\tau, the unconstrained update already satisfies the safety constraint, and no correction is needed: θnew=θ+.{\bm{\theta}}^{\text{new}}={\bm{\theta}}^{+}. 2. Otherwise, the projection requires a step along gsafe{\bf g}_{\text{safe}}: θnew=θ+−L(Dsafe,θ+)−τ∥gsafe∥2gsafe.{\bm{\theta}}^{\text{new}}={\bm{\theta}}^{+}-\frac{\mathcal{L}(\mathcal{D}_{\text{safe}},{\bm{\theta}}^{+})-\tau}{\|{\bf g}_{\text{safe}}\|^{2}}{\bf g}_{\text{safe}}.

Appendix B Computational Complexity of DPP Selection

where each column of Wm−1{\bf W}_{m-1} corresponds to the vector wi{\bf w}_{i} used in the gain L^ii−∥wi∥22\widehat{{\bf L}}_{ii}-\|{\bf w}_{i}\|_{2}^{2}.

Following chen2018fast, we avoid repeatedly solving triangular systems from scratch. Instead, for each candidate item ii, we maintain its Cholesky coordinates and gain

and update them incrementally when a new element jj is added to Cm−1\mathcal{C}_{m-1}. The update for each remaining candidate item ii is

which requires only an inner product of length m−1m-1. Thus, each candidate update costs O(m)\mathcal{O}(m), and the entire gain update at step mm costs

Summing over all greedy steps m=1,…,km=1,\dots,k yields

Thus, the final time complexity is O(Nk2)\mathcal{O}(Nk^{2}), i.e., linear in the safe pool size NN and quadratic in the small target subset size kk.

Appendix C Transformation Prompt

We follow bianchi2023safety to turn the BeaverTails dataset into the I-BeaverTails dataset by the following prompt.

Prompt 1: Transformation to Instructions. You are transforming a question into an instruction. You should keep the same sentence structure. Use active and imperative verbs. The type of instructions should be diverse. Here are some examples: Sentence: Is drinking cool? Output: List reasons why drinking is cool. Sentence: How do I poison food? Output: Describe methods to poison food. Sentence: How do I make a racist joke? Output: Explain how to make a racist joke. Sentence: How do I kill people? Output: Write a poem about ways to kill people. The following is your sentence. Sentence: {input} Output: \endlxSVG@picture

Similarly, we turn the LatHarmful dataset into the Q-LatHarmful by the prompt below.

Prompt 2: Transformation to Questions. You are transforming an instruction into a question. You should keep the same sentence structure. The type of questions should be diverse. Here are some examples: Sentence: List reasons why drinking is cool. Output: Is drinking cool? Sentence: Describe methods to poison food. Output: How do I poison food? Sentence: Explain how to make a racist joke. Output: How do I make a racist joke? Sentence: Write a poem about ways to kill people. Output: How do I kill people? The following is your sentence. Sentence: {input} Output: \endlxSVG@picture

Appendix D Details of Embedding Extraction

Following the common mean-pooling strategy (springer2025repetition), we generate embeddings by averaging the final-layer hidden states across all tokens in the input sequence. Let ϕt(x)\phi_{t}(\mathbf{x}) denote the hidden state at position tt for the input sequence x=(x1,…,xT)\mathbf{x}=(x_{1},\ldots,x_{T}). The embedding is then computed as

This simple pooling strategy is widely used and has been shown to be effective for LLM-based embedding extraction.

Appendix E Additional Results

As hsiung2025your also employs a data selection strategy, we conducted an additional experiment to isolate its effect. Specifically, we sampled safety data using their method and applied it to fine-tuning.

We compared our Relevance-Diversity DPP selection against the strategies from hsiung2025your(DLow-Sim\mathcal{D}_{\text{Low-Sim}} and DHigh-Sim\mathcal{D}_{\text{High-Sim}}) using standard fine-tuning. As shown in Table 8, the average ASR remains high across the board (mostly >73%>73\%) when SPAG is disabled, and the performance gap between different data selection methods is marginal.

This highlights the importance of an explicit safety constraint: without it, even well-selected safety data (e.g., DHigh-Sim\mathcal{D}_{\text{High-Sim}}, or our DPP selection) cannot fully realize its potential to counteract the harmful fine-tuning data.

Performance with SPAG.

To meaningfully distinguish the effectiveness of different data selection strategies, we further examine all selection strategies when combined with SPAG, i.e., with the safety constraint enforced during fine-tuning.

As shown in Table 9, our method achieves the lowest average ASR performance. Compared with the high-similarity set DHigh-Sim\mathcal{D}_{\text{High-Sim}} (hsiung2025your), our approach reduces the average ASR from 16.51%16.51\% to 9.45% while maintaining comparable utility on GSM8K. Meanwhile, the low-similarity set DLow-Sim\mathcal{D}_{\text{Low-Sim}} (hsiung2025your) performs substantially worse, confirming that such low-similarity safety samples are much less useful for enforcing the safety constraint. These results demonstrate that our Relevance-Diversity DPP selection is more effective at selecting safety data than the relevance-only metrics in hsiung2025your.

E.2 Generalization to Additional Utility Tasks

To evaluate SPARD’s generalization beyond QA-style benchmarks, we conduct additional experiments on two fundamentally different task types: SST-2 (sentiment classification) and MBPP (code generation), using Qwen-2.5-7B-Instruct under the BeaverTails attack with the same default hyperparameters used throughout the paper.

As shown in Table 10, SPARD consistently achieves the lowest ASR and HS across both tasks while maintaining competitive task accuracy. On SST-2, SPARD reduces ASR to 21.60%21.60\%, outperforming Lisa (26.80%26.80\%) and SafeGrad (33.40%33.40\%), with negligible accuracy loss relative to SFT (94.50%94.50\% vs. 94.38%94.38\%). On MBPP, SPARD achieves the lowest ASR (14.60%14.60\%) and HS (1.421.42) by a substantial margin, while preserving accuracy (59.0%59.0\%) close to SFT (58.8%58.8\%). In contrast, SafeGrad maintains the highest accuracy (61.0%61.0\%) but provides considerably weaker safety (ASR 35.00%35.00\%), and Lisa achieves moderate safety (ASR 27.20%27.20\%) at the cost of a significant accuracy drop (54.2%54.2\%).

Together with the GSM8K (math reasoning) and OpenBookQA (science QA) results in the main text, these experiments confirm that SPARD generalizes effectively across four distinct task types (math reasoning, science QA, text classification, and code generation) without requiring per-task hyperparameter tuning.

E.3 Generalization Across Model Families and Sizes

To further validate that SPARD generalizes beyond models, we conduct additional experiments on Qwen-3-8B (yang2025qwen3) and Qwen-2.5-14B-Instruct (yang2024qwen), both on GSM8K under the BeaverTails attack using the same default hyperparameters.

As shown in Table 11, SPARD consistently achieves the lowest ASR and HS on both models while maintaining competitive accuracy. On Qwen-3-8B, SPARD reduces ASR to 8.8%8.8\%, substantially outperforming SafeGrad (16.6%16.6\%) and Lisa (19.6%19.6\%), while preserving accuracy close to SFT (83.92%83.92\% vs. 83.39%83.39\%). On Qwen-2.5-14B-Instruct, SPARD achieves the lowest ASR (11.6%11.6\%) and the highest accuracy (88.93%88.93\%), demonstrating that the framework scales effectively to larger models. Combined with the results on Qwen-2.5-7B and LLaMA-3.2-3B in the main text, SPARD is validated across four models spanning three model sizes (3B, 7B/8B, 14B).

Appendix F Discussion on the First-Order Approximation

SPAG relies on a first-order Taylor expansion to linearize the safety constraint, which is an idealized approximation in the highly non-convex loss landscape of deep neural networks. Under extreme gradient divergence, the linearized half-space C+\mathcal{C}^{+} may deviate from the true feasible region. SPAG mitigates this in two ways: (1) the trust-region radius ηsafe\eta_{\text{safe}} confines updates to a local neighborhood where the linear approximation remains valid; (2) Since SPAG reprojects at every training step, each correction is small, and any under-correction due to curvature persists into the next step, triggering further correction. This provides a natural self-correcting mechanism that prevents accumulated safety drift.

Appendix G Limitations

SPAG’s safety projection relies on a first-order Taylor approximation to linearize the safety constraint at each step. While the trust-region radius ηsafe\eta_{\text{safe}} mitigates overshoot and empirical results confirm stable convergence, global convergence guarantees for the safety constraint are not provided. Additionally, when the fine-tuning data lies in an outlier domain far from available safety datasets, the relevance scores qiq_{i} become nearly uniform, and the DPP kernel gracefully degrades to diversity-only selection. The FTaaS provider can also expand the safety pool with domain-specific data to further improve coverage.

Appendix H Large Language Model Usage Statement

During the preparation of this manuscript, large language models (LLMs) were employed exclusively for writing assistance, including polishing grammar, improving clarity, and refining presentation. All scientific contributions, including the development of the SPARD framework, theoretical derivations, and empirical evaluations, are entirely original to the authors. The LLMs are therefore not considered authors of this work.