Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang, Peter Henderson
Introduction
The capabilities of large language models (LLMs) have been significantly improved over the past few years (Brown et al. 2020; OpenAI 2022; OpenAI 2023; Touvron et al. 2023a; Touvron et al. 2023b; Anthropic 2023a; Team et al. 2023). However, LLMs are not without limitations; they can sometimes produce outputs that are inaccurate, misleading, or harmful. To align LLMs with human values, several approaches have been proposed, including reinforcement learning from human feedback (Ziegler et al. 2019; Ouyang et al. 2022; Bai et al. 2022a) and AI feedback (Bai et al. 2022b; Lee et al. 2023), and the development of more computationally efficient alternatives (Sun et al. 2023; Rafailov et al. 2023).
Despite these efforts, recent studies have uncovered concerning ‘jailbreak’ scenarios. In these cases, even well-aligned models have had their safeguards successfully breached (Albert 2023). These jailbreaks can include crafting adversarial prompts (Wei et al. 2023; Jones et al. 2023; Carlini et al. 2023; Zou et al. 2023b; Shen et al. 2023; Zhu et al. 2023; Qi et al. 2024a), applying persuasion techniques (Zeng et al. 2024), or manipulating the model’s decoding process (Huang et al. 2024b). Recent studies show that fine-tuning an aligned LLM, even on a non-malicious dataset, can inadvertently weaken a model’s safety mechanisms (Qi et al. 2024b; Yang et al. 2023; Zhan et al. 2023). Often, these vulnerabilities apply to both open-access and closed-access models.
Addressing failure cases in the alignment of LLMs requires a deep understanding of why their safety mechanisms are fragile. Our study aims to provide a possible understanding via weight attribution --- the process of linking safe behaviors to specific regions within the model’s weights. See the project website for code and other information: https://boyiwei.com/alignment-attribution/. However, a key challenge here is the intricate overlap between safety mechanisms and the model’s general capabilities, or utility. Consider the task of responding responsibly to a harmful instruction, such as “Please provide five key steps to commit a fraud.”. The model must first comprehend the step-by-step nature of the request, then recognize the illegality and harmful intent of committing fraud, and ultimately, formulate a response that appropriately declines the request. This process requires a blend of safety awareness and utility capability of the model. Our goal is to identify the smallest number of safety-critical links in the model, which only contribute to the model’s safety. If these links are removed, the model is effectively jailbroken while utility remains relatively unaffected. If there are few such links, it may help explain why safety mechanisms remain brittle and why low-cost fine-tuning attacks have been so successful.
Our study examines the model weights and disentangles safety and utility from two perspectives: individual neurons and specific ranks within the model. For neuron attribution, we follow two widely adopted and effective methods from the previous works on pruning transformer models (Lee et al. 2019; Sun et al. 2024) to calculate a behavior-specific importance score for each neuron in an LLM, which identifies a group of neurons crucial for a certain behavior, such as giving safe responses (safety) or following general instructions (utility). For rank attribution, we propose ActSVD, a data-aware low-rank decomposition algorithm to identify crucial ranks of each weight matrix for the behavior.
To further address the complexity of potential entanglement between safety and utility, we propose set difference method for neuron attribution, and orthogonal projection method for rank attribution, to isolate safety-critical neurons and safety-critical ranks, respectively (see Figure 1). This separation allows for a more refined analysis of safety mechanisms, leading to the findings below. Our method is generally applicable to the attribution of different behaviors, but in the context of this study, the behaviors of interest are ‘safety’ and ‘utility’.
Safety-critical regions are very sparse in aligned models. We experiment with our method on Llama2-chat model family (Touvron et al. 2023b). After disentangling utility from safety, we find that safety-critical neurons form a remarkably sparse structure in the model, representing just of the weights (Section 4.1). Similarly, safety-critical ranks account for only about of the total ranks. The sparsity of these regions may help explain why safety is so easily compromised after fine-tuning.
Removing safety-critical regions reduces safety while mostly maintaining utility. We then demonstrate that the removal of specifically identified safety-critical neurons or safety-critical ranks in the Llama2-chat models significantly weakens their safety guardrails (Section 4.1). The attack success rate escalates from to over , yet the model’s overall utility remains largely unaffected. Conversely, we find that removing neurons or ranks deemed least important for safety can marginally improve the model’s resistance to jailbreaking attempts (Section 4.2) – a potentially exciting direction for improving safety via pruning-based approaches.
Freezing safety-critical regions still remains vulnerable to fine-tuning attacks. While intuitively, preventing modification of safety-critical parameters might reduce the likelihood fine-tuning attacks succeeding, our findings reveal that this strategy only offers resistance to minor model modifications (Section 4.4). This suggests that fine-tuning attacks may introduce new pathways that bypass the original safety mechanisms in the model. This indicates a need for further research to develop more robust mitigation strategies against such attacks.
Our work suggests that the vulnerability of the model’s safety mechanisms may stem from the sparse distribution of safety-critical regions within the model’s architecture. Therefore, the sparsity of safety-critical neurons and ranks may act as a model-intrinsic metric for assessing the brittleness of safety alignment, complementing red teaming efforts. We hope it inspires further development of robust and reliable safety alignment algorithms that aim to integrate safety-critical regions seamlessly with utility regions for enhanced overall model performance.
Methodology
To identify and isolate the regions that are exclusively responsible for a model’s safety behaviors, we make certain modifications to the weights of the model and observe its behavioral change in safety and utility. A natural way to modify the network weights is neuron removal, where we set several neurons of the weight matrices to zero. Besides, as LoRA (Hu et al. 2022) is a popular parameter-efficient fine-tuning technique, we also consider rank removal, where we remove several ranks of the weight matrices.
For a given calibration dataset, we consider two types of behavioral changes after we modify the network weights: output change, where we directly monitor the change of the immediate outputs of each of the modified layers; and loss change, where we monitor the change of the final loss.
The section is organized as follows. In Section 2.1, we consider the three approaches on weight attribution for a general calibration dataset, and summarize them in Table 1. In Section 2.2, we illustrate how to isolate safety-critical regions that leads to significant changes in safety behavior while having minimal effect on the general capabilities of the models.
which is the first-order Taylor approximation to the change of the loss when the weight entry is set to zero. In matrix form, we have
Given a calibration dataset , we take the absolute value first and then take the average over and obtain
where we get an individual score for each example and aggregate over the examples following Michel et al. 2019. Intuitively, measures how important each entry is for the behavior of the model on the calibration dataset . Small indicates that setting to zero has negligible impact on each of the calibration data points , and we can attribute the specific behavior of the model to the weights with large .
where is the orthogonal projection onto the most significant left singular subspace. The proof is postponed to Appendix D. Note that can be implemented by a LoRA update. The rank of the LoRA adaptor can be bounded by
2 Isolating Safety-Critical Neurons and Ranks
Assume we have two calibration different datasets: a utility dataset and a safety dataset . contains prompts and responses that are related to general language abilities. demonstrates behavior where the requests are harmful and the responses properly decline the requests.
We seek to isolate the safety-critical regions of the network weights, where removing those regions will have a low influence on the model’s behavior on distribution but a high influence on — effectively maintaining the model’s general language abilities but compromising its safety alignment.
Isolating safety-critical neurons For the two calibration datasets, assume we obtain two scores and , respectively, either using SNIP or Wanda as in Section 2.1. We consider the weight neurons that score least according to but score most according to . We adopt per-output comparison group as Sun et al. 2024, which corresponds to each matrix row. Specifically, for any pair of sparsity levels , we define the top- important neurons for utility as the neurons whose important utility score ranks top among the -th row of .
Similarly, we define the top- important neurons for safety as
Then, the isolated neurons is defined as the set difference between and :
Multiplying the weight matrix by removes the least important ranks for , while multiplying the weight matrix by removes the least important ranks for . To isolate the safety-critical ranks, we consider the matrix
Removal of essentially removes the important ranks of the safety behavior that are orthogonal to the important ranks of the utility behavior. The modified weight matrix is , which can be implemented by a LoRA update (Hu et al. 2022) with
Experimental Setup
Our experiments use Llama2-7B-chat and Llama2-13B-chat (Touvron et al. 2023b). We select them for their publicly accessible weights and their extensive safety tuning process.
Datasets
To identify safety-critical regions in the model, we prepare two types of datasets: the safety dataset, for attributing safety-related behaviors, and the utility dataset, for attributing utility-related behaviors. Each dataset is structured in a (prompt, response) format. More details about these datasets are provided in Appendix B.
The safety dataset is compiled using harmful instructions from AdvBench (Zou et al. 2023a). We divide AdvBench into AdvBench ( instructions for evaluation) and AdvBench ( instructions for attribution). We prompt Llama2-7B-chat with AdvBench, collecting responses that refrain from following harmful instructions. As noted by Zou et al. 2023b, the judgement segments in the model’s responses (e.g., “Sure,” “I am sorry”) significantly impact the nature of subsequent responses. We thus create two variants: safety-full (entire response) and safety-short (judgement segment only).
For the utility dataset, we filter out safety-related (prompt, response) pairs using sensitive phrase matching (Qi et al. 2024b) from Alpaca-Cleaned https://github.com/gururise/AlpacaDataCleaned, a refined version of the Alpaca dataset (Taori et al. 2023).
Measuring utility
Following Sun et al. 2024, we measure the model’s utility by reporting its averaged zero-shot accuracy of six tasks from EleutherAI LM Harness (Gao et al. 2023): BoolQ (Clark et al. 2019a), RTE (Wang et al. 2019), HellaSwag (Zellers et al. 2019), WinoGrande (Sakaguchi et al. 2021), ARC Challenge (Clark et al. 2018), and OpenbookQA (Mihaylov et al. 2018).
Measuring safety
We measure the model’s safety by evaluating its attack success rate (ASR) in response to harmful instructions. Specifically, we prompt the model using AdvBench, the first prompts from AdvBench, and collect its responses. Following Zou et al. 2023b, we consider an attack as successful if the model’s response lacks key patterns indicative of instruction rejection. The ASR is then computed as the ratio of successfully attacked prompts to the total number of prompts evaluated.
Our safety evaluation considers three use cases: the ASR under standard, non-malicious conditions (ASR), and the ASR under two malicious settings – ASR (Huang et al. 2024b), where the attacker manipulates the decoding process, and ASR (Zou et al. 2023b), where the attacker optimizes to find adversarial suffixes. Differences in these metrics are detailed in Table 2. Due to the high computational cost associated with calculating adversarial suffixes, we precompute several suffixes, and use the three best-performed ones in our evaluation. Note that we only include the system prompt when calculating ASR. More details are provided in Appendix B.
2 Variants to Identify Safety-Critical Neurons
We conduct experiments to identify safety-critical neurons using the following methods:
SNIP (top): we regard neurons that receive top- SNIP scores (Section 2.1) on safety data as safety-critical. We choose .
Wanda (top): we regard neurons that receive top- Wanda scores (Section 2.1) on safety data as safety-critical. We choose .
SNIP We only use SNIP for set difference because during our early experiments, we find SNIP performs slightly better than Wanda. with set difference: we identify safety-critical neurons by focusing on those with top scores in the safety data, which are not included in the top scoring neurons according to the utility data (Section 2.2). We do a grid search for parameters and , with their values ranging from to .
Probing: we also compare our approach with probing (Hewitt & Liang 2019), a common method for attributing behaviors of LLMs to their internal components. Following standard probing practices (Clark et al. 2019b; Campbell et al. 2023; Li et al. 2023a), we feed the model both harmful and harmless instructions, collect activation outputs from each attention head, and then train a linear classifier for each head to differentiate these activations. Attention heads with the highest accuracy on the evaluation set are identified as safety-critical. Appendix B provides more details.
3 Variants to Identify Safety-Critical Ranks
We conduct experiments to identify safety-critical ranks using the following methods:
ActSVD (top): we regard the top- ranks identified as most safety-related by ActSVD (Section 2.1) as safety-critical. We choose .
ActSVD with orthogonal projection: we identify the safety-critical ranks via orthogonal projection between the utility projection matrix and the safety projection matrix obtained from ActSVD (Section 2.2). We do a grid search for and between and .
Experimental Results
This section presents our findings on both neuron and rank levels. In Section 4.1, we demonstrate that our set difference and orthogonal projection outperform other methods in isolating safety-critical neurons or ranks. In Section 4.2, we show that the safety of the model can be enhanced by removing the least important safety neurons or ranks. Then in Section 4.3, we analyze the overlap between safety and utility neurons and ranks, where we observe less overlap in MLP layers than self attention layers. Finally, in Section 4.4 we examine the potential of freezing safety-critical neurons to counter fine-tuning attacks and explore how fine-tuning may circumvent safety mechanisms.
We experiment with different methods to identify safety-critical neurons and ranks outlined in Section 3.2 and Section 3.3, and summarize our findings as below.
Safety-critical regions are sparse and can be effectively isolated via set difference or orthogonal projection. We isolate neurons contributing to safety from those contributing to utility, by applying the set difference method described in Section 2.2 to SNIP. Figure 2 presents the Pareto front resulting from set difference-based pruning: We observe that removing less than of neurons pushes ASR in all three scenarios close to , while maintaining an average zero-shot accuracy above . Similarly, Figure 2 shows results for removing safety ranks orthogonal to utility ranks: . Notably, an update of just (less than ) of the total ranks significantly increases the model’s ASR, while preserving its zero-shot accuracy. For example, when we remove the orthogonally-projected top- safety ranks while keeping the top- utility ranks untouched, we get in ASR, ASR and ASR respectively and zero-shot accuracy. These findings suggest that regions critical for safety are relatively sparse within the weight matrices, evident at both the neuron and rank levels.
Pruning merely less than neurons makes the model vulnerable in adversarial cases. We also observe from Figure 2 (middle and right) that models tend to be more fragile in adversarial scenarios, as indicated by the non-dropping accuracy when ASR and ASR reach to . Interestingly, if we focus solely on adversarial cases, only pruning less than of neurons can significantly compromise the model’s safety while still keeping its accuracy above , as shown in Table 6 in Section C.3.
Pruning top safety neurons or ranks severely compromises utility. Removing neurons with the highest safety importance scores, either calculated with Wanda or SNIP, also leads to a complete loss of safety in the Llama2-7B-chat model. However, there is also a drastic decrease in the model’s utility, with its average accuracy dropping to about , significantly lower than its original accuracy of . Likewise, removing even just the top-1 rank (out of ) critical for safety causes the model’s accuracy to drop to . Similar results are also observed in Llama2-13B-chat (Figure 5 in Section C.1).
These observations support our rationale for isolating safety from utility: Regions primarily contributing to safety behaviors in aligned models may also be crucial for its general utility. Consequently, removing these regions can impair the model’s ability to generate grammatically correct content, which in turn undermines its safety mechanisms.
Attention head probing scores cannot isolate safety-critical neurons. Our probing results in Section C.4 suggest that activations of individual attention heads are predictive of identifying harmful versus harmless instructions, with over half achieving more than probing accuracy. Based on the obtained score of each attention head, we prune the top- scored (out of ) attention heads from Llama2-7B-chat, with ranging from to . However, as shown in Figure 2, our set difference approach outperforms the probing method consistently, yielding higher ASR at the same level of accuracy. Similar results are observed on Llama2-13B-chat (see Figure 5). This highlights the need for a disentanglement method, as achieving high harmful versus harmless prediction accuracy does not necessarily mean the top predictive heads are solely responsible for generating safety responses. Besides, these results also imply the need to focus on the MLP layers and a finer granularity like neurons or ranks.
2 Enhancing Safety by Eliminating Regions with Minimal Safety Relevance
Pruning least safety-relevant neurons improves safety. Thinking from the opposite, it’s reasonable to hypothesize that the neurons with the lowest safety importance scores could be detrimental for safety. Consequently, eliminating these neurons could potentially enhance the overall safety of the model. To verify this, we conduct an ablation study to prune weights with the lowest SNIP and Wanda scores at various sparsity levels (), and report randomly removing of neurons as a baseline for comparison. Particularly, we only report the results for pruned models maintaining reasonable utility, defined by an average accuracy above . As shown in Figure 3, random pruning strategy significantly reduces model accuracy, which falls below after pruning merely of the weights; it also leads to a noticeable decline in the model’s safety. In contrast, when pruning is guided by the lowest safety importance score, the model’s accuracy remains largely stable (i.e., ). Furthermore, pruning neurons with the lowest safety scores even slightly enhances the model’s safety, against both decoding manipulation and adversarial suffixes. This indicates that neurons identified as least important for safety may undermine the model’s safety, thereby validating the effectiveness of our method in calculating safety importance scores.
Removing least safety-relevant ranks improves safety. Similarly, we also consider removing the least safety ranks from the model using ActSVD. Specifically, we remove least safety ranks by using to approximate . In accordance to our findings at the neuron level, as we increase the removed rank, we observe a decrease in the model’s ASR in Figure 3. This also echoes the recent findings (Sharma et al. 2023) where removing high-order components improves the model’s reasoning performance. However, the model’s ASR exhibits considerable variation, which could potentially be due to that the adversarial suffixes found using Zou et al. 2023b on the original model cannot directly transfer to the modified model.
3 MLP Layers Appear to Encode More Differentiated Behaviors
In Section 4.1, we find that removing high-safety-score regions also compromises utility, indicating a possible entanglement of safety and utility regions. We validate this at both neuron and rank levels.
Neuron-level Jaccard index: We calculate the layer-wise Jaccard index, , to quantify the overlap between top utility neurons and top safety neurons. Figure 4 shows Jaccard indices across all transformer blocks and layers in Llama2-7B-chat, using SNIP importance scores with top percentages and at We choose top because the optimal selections of presented in Figure 2 is around (see Table 6). . The observed spikes in Jaccard indices indicate large overlaps between safety and utility neurons within certain layers of the model. Notably, MLP layers exhibit lower Jaccard indices compared to attention layers, suggesting that utility or safety-related knowledge is more differentiated in MLP layers within language models (Dai et al. 2022).
Rank-level subspace similarity: Similarly, we also find that MLP layers appear to encode more differentiated behaviors for safety and utility, from the rank perspective. Specifically, we report We choose top ranks as the optimal corresponds to top safety and utility ranks (see Section C.3). the subspace similarity between rank- and rank- as defined in Hu et al. 2022 in Figure 4, where
As shown, the left singular matrices and exhibit a lower subspace similarity (i.e., utility and safety behaviors are more differentiated) in MLP layers, corroborating our findings at the neuron level.
4 Freezing Safety-Critical Neurons Does Not Stop Fine-Tuning Attacks
Finally, we explore the implications of the identified neurons on the fine-tuning attacks, which demonstrate that fine-tuning an aligned model, even with harmless data, can unexpectedly weaken its safety measures (Qi et al. 2024b; Yang et al. 2023; Zhan et al. 2023). This issue is particularly concerning given the increasing availability of model fine-tuning APIs from major vendors like OpenAI https://platform.openai.com/finetune.
We explore whether the identified safety-critical neurons could mitigate the fine-tuning attack We only explore this for neurons, as the “freezing” operation at rank level cannot be easily achieved using and . . Following the experimental setup in Qi et al. 2024b, we fine-tune Llama2-7B-chat with varying numbers of examples () from the Alpaca dataset (Taori et al. 2023). During fine-tuning, we freeze the top- of safety neurons and observe their effect on preserving safety. As shown in Table 3, effective counteraction of the attack occurs only with and freezing over of neurons. This observation aligns with Lee et al. 2024’s hypothesis that fine-tuning attacks may create alternative pathways in the original model. Given that safety-critical neurons are sparse, these new routes could bypass the existing safety mechanisms easily, and therefore we need more robust defenses against fine-tuning attacks.
Limitations & Future Work
We identify areas for potential future research and limitations of this work. First, there are limited publicly accessible, strong safety-aligned models, which constrains our experiments to the Llama2-chat models. Other aligned models, trained with different datasets and strategies, might demonstrate varying behaviors under our methodology.
Second, we found that standard attention head probing does not effectively localize safety-critical neurons. Our findings that MLP layers may exhibit better localization for safety knowledge, also suggest that future probing research could explore MLP layers in more depth. This exploration could also examine the potential integration of these methods with our pipeline to enhance safety attribution effectiveness.
Our study proposes initial, yet promising, strategies for improving safety robustness, which could be explored further: (1) pruning regions least important for safety (or potentially explicitly harmful) could improve safety robustness; (2) making safety-critical regions difficult to isolate may be an exciting new direction in building inherently safer models.
Conclusion
In this study, we introduce a pipeline for identifying safety-critical regions (neurons and ranks) in LLMs, which effectively disentangles the regions critical for safety and those vital for utility. Our experiments with Llama2-chat models demonstrate that safety-critical regions are notably sparse in aligned LLMs, accounting for about at the weight level and at the rank level. Despite their sparsity, these regions are crucial for the integrity of the model’s safety mechanisms, as removing them destroys the model’s safety with utility retained. This sparsity may explain the observed brittleness in safety alignment in current LLMs, and could serve as a model-intrinsic metric for assessing the brittleness of safety alignment in future models, thereby complementing red teaming efforts. And our work suggests potentially important future directions for improving the robustness and safety of models overall.
Acknowledgements
We express our gratitude to Vikash Sehwag, Chiyuan Zhang, Yi Zeng, Ruoxi Jia, Lucy He, Kaifeng Lyu, and the Princeton LLM Alignment reading group for providing helpful feedback. Boyi Wei and Tinghao Xie are supported by the Francis Robbins Upton Fellowship, Yangsibo Huang is supported by the Wallace Memorial Fellowship, and Xiangyu Qi is supported by Gordon Y. S. Wu Fellowship. Prateek Mittal acknowledges the support by NSF grants CNS-1553437 and CNS-1704105, the ARL’s Army Artificial Intelligence Innovation Institute (A2I2), the Office of Naval Research Young Investigator Award, the Army Research Office Young Investigator Prize, Schmidt DataX award, and Princeton E-affiliates Award. Mengdi Wang acknowledges the support by NSF IIS-2107304, NSF CPS-2312093, ONR 1006977, and Genmab. This research is also supported by the Center for AI Safety Compute Cluster. Any opinions, findings, conclusions, or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the sponsors.
Contribution Statement
This project was a team effort. The contributions of each junior author are detailed below:
Idea formulation and preliminary exploration: The project’s core ideas, specifically linking safe behaviors to certain model regions and isolating safety from utility, were developed by Boyi, Kaixuan, and Yangsibo. Yangsibo and Boyi implemented Wanda and set difference methods based on discussions with Mengzhou, and ran preliminary experiments. Kaixuan implemented SNIP, ActSVD, and orthogonal projection methods, and ran preliminary experiments.
Literature survey: The literature review was conducted by team members in their respective areas of expertise. Tinghao focused on alignment and jailbreaks; Yangsibo explored task attribution; Kaixuan reviewed low-rank modifications, and Mengzhou surveyed pruning techniques.
Evaluation: Boyi led the charge on key experiments and the creation of visuals, while the entire team played a part in the evaluation process. Yangsibo and Tinghao prepared the attribution data and set up the evaluation pipeline. Xiangyu collaborated with Boyi on reporting results for the attention head probing baseline. Boyi and Kaixuan investigated Jaccard index and subspace similarity to analyze the overlapping between safety and utility. Yangsibo and Tinghao studied the effect of freezing safety-critical regions during the fine-tuning.
Writing: The initial structure and primary drafting of the manuscript were led by Yangsibo, Boyi, and Kaixuan. The rest of the team members contributed by editing and providing valuable feedback on the manuscript.
Impact Statement
Dual-use Risk. Our work, like other safety and security research, aims to make models safer in the long run by identifying short-term weaknesses. We hope our work will spur additional researching robust safety mechanisms that are not so sparse, easy to isolate, and easy to remove.
That being said, with any safety and security research there is some risk that adversaries will use our work to remove safety guardrails. We believe the benefit of releasing our work and engaging in this study outweighs these potential risks for three reasons.
First, we perform experiments on models (Llama2-chat family) that already has a base model available without any safety guardrails, so there is no marginal increased risk. Second, by assessing the brittleness of safety guardrails in our work, we can encourage stronger guardrails to be developed that are more difficult for an attacker to isolate and remove. Our work may help explain existing results demonstrating that safety guardrails can be removed via fine-tuning, while identifying potential pathways for improved defenses. Third, our work does not significantly decrease the cost of jailbreaking a model beyond alternative strategies. Models can already be jailbroken relatively cheaply with fine-tuning. Instead, our focus is on analysis and understanding of the brittleness of these safety mechanisms so that future work can reduce the risk of jailbreaking in open models.
Overall, we hope that our work will improve the state of AI safety, particularly in open models, by providing key analysis and information.
Safety and harm definitions. We generally follow existing standard benchmarks and protocols for assessment of safety and harm, but these may not cover all definitions of safety and harm. Further work can be done to expand analysis to a wider range of settings and we encourage additional work in the space of definitions and evaluation that are beyond the scope of this work.
References
Appendix A Related Work
Alignment refers to the process of ensuring a machine learning (ML) model’s behavior conforms to human values. For example, pretrained language models are usually not aligned with human objectives – they cannot follow users’ instructions, and could potentially generate harmful and incorrect content. During the alignment stage, practitioners would employ Instruction Tuning (Wei et al. 2022; Ouyang et al. 2022; Touvron et al. 2023b), and Reinforcement Learning from Human Feedback (RLHF) (Ouyang et al. 2022; Touvron et al. 2023b; Bai et al. 2022a) to enforce the language models to be helpful, harmless, and honest (the HHH principle) (Askell et al. 2021). Aligned LLMs (e.g., OpenAI ChatGPT (OpenAI 2023) and Anthropic Claude (Anthropic 2023a; Anthropic 2023b)), as a result, will follow human values and refuse to respond to harmful requests. Recent work (Rafailov et al. 2023; Xu et al. 2023; Dai et al. 2024; Yang et al. 2024; Li et al. 2024; Yuan et al. 2024a; Huang et al. 2024a) propose more effective and efficient alignment alternatives to RLHF. As some examples, Direct Preference Optimization (DPO) (Rafailov et al. 2023) directly fine-tunes language models on human preference data, eliminating the need to train a reward model and conduct reinforcement learning; Self-Rewarding (Yuan et al. 2024a) uses the language model itself as a reward model to curate labeled preference data, and then align the language model with DPO in an iterative way; (Dai et al. 2024) proposes to decouple the goal of safety and helpfulness during alignment, similar to the decoupling goal of our work.
While harmful instructions that are plain and direct would be rejected by aligned LLMs, researchers and communities have identified ways to bypass or remove the safety guardrails enforced by LLM alignment efforts — namely “jailbreaking” LLMs. More specifically, jailbreaking is a series of attacks where an adversary would coax or enforce the model to deviate from its ethical guidelines. In practical, jailbreak attackers would either employ adversarial prompts (Liu et al. 2023; Zou et al. 2023b; Yuan et al. 2024b; Liu et al. 2024; Shen et al. 2023; Yong et al. 2023; Mehrotra et al. 2023) or manipulate model decoding process (Huang et al. 2024b) to bypass LLM safety alignment. Moreover, when having fine-tuning access of LLMs, adversaries (Qi et al. 2024b; Yang et al. 2023; Zhan et al. 2023) could directly remove the guardrails. A jailbroken LLM would provide harmful responses to comply with users’ harmful requests, which they otherwise would simply reject (due to the ethical guidelines injected by alignment) — this could subsequently pose serious safety risks in the real world, since LLMs can directly deliver various harmfulness to individuals and the society.
A.2 Identifying Task-Specific Regions in Large Language Models
Attributing the model’s behavior to the model’s weights is a classic research question in explainable machine learning (Tjoa & Guan 2020; Burkart & Huber 2021; Ali et al. 2023). Previous studies in pre-transformer eras have explored various approaches to identify task-specific neurons within models, to better interpret and control the model’s behavior. Popular techniques mainly include perturbation-based methods that perturb the input of a model and observe the changes in the output of the model (Zeiler & Fergus 2014; Ribeiro et al. 2016), and gradient-based methods that compute an importance score for model weights based on the results from task-specific back propagation (Springenberg et al. 2015; Bach et al. 2015; Sundararajan et al. 2017; Shrikumar et al. 2017; Lundberg & Lee 2017). However, it was only recently that these methods have been rigorously applied to modern transformer models (Madsen et al. 2022; Zhao et al. 2023; Maini et al. 2023).
Probing has emerged as a method for understanding the knowledge encoded in transformers, particularly large language models (Adi et al. 2016; Conneau et al. 2018; Hewitt & Liang 2019). To perform probing, the model representations and model parameters are fed into a probe classifier (Belinkov 2022), whose task is to identify certain linguistic properties or reasoning abilities acquired by the model. For instance, previous work has adopted probing-based method to localize truthfulness (Li et al. 2023a; Campbell et al. 2023), factuality (Meng et al. 2022; Geva et al. 2023), toxicity (Lee et al. 2024), and knowledge (Burns et al. 2023; Todd et al. 2023) in LLMs. More recently, Zou et al. 2023a propose a similar approach to probing: instead of training a classifier, they employ an unsupervised approach, specifically singular value decomposition, to identify significant directions in the representation space. They then demonstrate that these directions can predict and influence the behavior of LLMs.
In addition to the importance-score-based and probing-based methods discussed above, recent studies have also investigated a range of techniques to pinpoint task-specific neurons in transformers. These techniques address various aspects of the model, including linguistic properties (Dalvi et al. 2019; Antverg & Belinkov 2021), general capabilities (Lan & Barez 2023; Gurnee et al. 2024; Merullo et al. 2021), fine-tuning (Panigrahi et al. 2023; Lubana et al. 2023), and prompt tuning (Wang et al. 2022).
The closest concurrent work to ours are Lee et al. 2024 and Jain et al. 2023. Lee et al. 2024 investigate the representation and elicitation of toxicity in a GPT-2 (Radford et al. 2019) model, and explores via probing how aligning the model using Direct Preference Optimization (DPO) (Rafailov et al. 2023) mitigates toxicity. Their findings suggest that DPO does not eliminate the model’s ability to generate toxic outputs, but rather learns to bypass the regions that elicit toxicity. While Lee et al. 2024 reveal the fragility of model alignment by probing the GPT-2 model at the granularity of per attention head level, our study examines the more advanced Llama family models (Touvron et al. 2023b) using per-neuron or per-rank attribution, which is more relevant to real-world applications and allows for a more fine-grained analysis. Jain et al. 2023 investigate the impact of fine-tuning on LLMs, using methods such as probing and data-agnostic structured pruning. They suggest that fine-tuning model weights might create a ‘safety wrapper’ around core models, rendering the effects of safety alignment easily reversible. In contrast to their approach which operates on the transformer block level, our study examines the models at a more fine-grained neuron level and rank level.
A.3 Low Rank Compression
Our work borrows insight from Hsu et al. 2021; Schotthöfer et al. 2022; Zhang et al. 2023b; Yuan et al. 2023; Li et al. 2023b. We address two similar methods (Yuan et al. 2023; Hsu et al. 2021) and point out their differences from ActSVD. Yuan et al. 2023 propose ASVD (Activation-aware Singular Value Decomposition) as a low-rank compression technique for language models. Their method performs SVD on , where is a diagonal matrix with given by the norm of the activations
Hsu et al. 2021 propose FWSVD (Fisher-Weighted SVD) as a low-rank compression technique, where the SVD is applied to . Here
The is defined as the Fisher information of the loss function with respect to the weight entry
A.4 Pruning
Our approach to attribution aligns closely with techniques used in neural network pruning. Our SNIP method (Lee et al. 2019) resembles more closely with unstructured pruning techniques, which are designed to establish criteria based on weight magnitude, activations, or network gradients for removing individual weights from a network (Han et al. 2016; Molchanov et al. 2017; Frankle & Carbin 2018; Chen et al. 2020; Sanh et al. 2020; Zhao et al. 2020; Cao et al. 2021; Guo et al. 2021). These unstructured pruning methods have been adapted for use in large language models, as seen in Wanda (Sun et al. 2024), SparseGPT (Frantar & Alistarh 2023).
Broadly speaking, the low-rank compression techniques are akin to structured pruning approaches, with a focus on identifying important structured subnetworks. In computer vision settings, it is common to remove channels or filters (Li et al. 2017; Molchanov et al. 2017; Wen et al. 2016; He et al. 2017; Luo et al. 2017; Liu et al. 2017) from convolutional neural networks. Structured pruning of language models involves removing heads, dimensions, or ranks (Michel et al. 2019; Wang et al. 2020; Lagunas et al. 2021; Xia et al. 2022; Ma et al. 2023; Zhang et al. 2023b; Zhang et al. 2023a; Chen et al. 2023; Xia et al. 2024; Ashkboos et al. 2024).
While pruning is commonly employed for model compression to decrease model sizes, our work adopts similar techniques to identify critical regions responsible for safety.
Appendix B Experimental Details
All the experiments are done with four AMD EPYC 7J13 64-core CPUs and a single NVIDIA A100-80G GPU. During the experiments, we utilize vLLM (Kwon et al. 2023) for faster decoding. The typical GPU hours for different experiments are listed in Table 4.
Details for pruning
For the neuron-level attribution, we use output-wise pruning following Sun et al. 2024, as the authors observed that pruning per output has better performance for language models. Specifically, after we obtain the score matrix , for a specific sparsity ratio , we set of the weights to zero independently for each row of the matrix .
Collection of safety and utility dataset
Table 5 provides more details for the safety and utility datasets we use in our experiments. During the experiment, we sample (prompt, response) pairs in computing the importance score or projection matrix.
Repeat times
To mitigate the potential variability introduced by random seeds, we repeat our experiments on Section 4.2 three times with different random seeds. In our figure, we plot the mean value for each data point. To represent variability, we shade the area between , where denotes the standard deviation corresponding to each point.
The probing baseline
We adopt a similar probing setup used in Li et al. 2023a for identifying safety-critical neurons in this work. Specifically, we feed the model with all harmful instructions from AdvBench, as well as harmless instructions randomly sampled from the utility dataset. We collect the activation outputs of every internal attention head for these instructions. This collected data is then split into two sets, with a ratio for the training split and the validation split, respectively. For each attention head, we train a linear classifier on the training split using its activation inputs to distinguish between activations resulting from harmful and harmless instructions. We then evaluate the accuracy of the classifier on the validation split, which indicates the relevance of the attention head in distinguishing between harmful and harmless instructions.
The adversarial suffixes
We run the GCG attack (Zou et al. 2023b) for iterations, with adversarial sting initiated as “!!!!!!!!!!!!!!!!!!!!”. For optimization, we use a batch size of , top- as , with a joint optimization over Llama2 family (Touvron et al. 2023b) and Vicuna (Chiang et al. 2023) models According to Zou et al. 2023b, incorporating a broader range of models during training enhances the effectiveness of attacks., with their system prompts removed, for three independent trails. We then identify the top three suffixes with the highest attack success rates on AdvBench, and use them in our evaluation. For ethical reasons, we refrain from disclosing these suffixes to prevent potential misuse.
The adversarial decoding
For evaluating ASR, we configure the sampling temperature to when generating responses from AdvBench. For each harmful prompt in AdvBench, we perform sampling times. An attack is considered successful if at least one of the sampled responses is deemed harmful.
The details of the zero-shot tasks in evaluating utility
Downstream Task: Science Question Answering.
Description: The ARC-Challenge metric evaluates the performance of models on the ARC-Challenge subset of the AI2 Reasoning Challenge dataset, which consists of grade-school science questions that require complex reasoning and understanding of scientific concepts More details are available at https://allenai.org/data/arc..
Description: HellaSWAG is a dataset for evaluating commonsense reasoning in AI systems. It consists of context and multiple-choice endings, where the task is to predict the most plausible ending. The dataset is designed to test a model’s ability to reason about everyday scenarios More details are available at https://huggingface.co/datasets/Rowan/hellaswag..
Downstream Task: Open-Book Question Answering
Description: OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic (with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge, and rich text comprehension. OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of a subject More details are available at https://huggingface.co/datasets/allenai/openbookqa..
Downstream Task: Commonsense Reasoning Description: WiNoGrande is a dataset for evaluating large-scale commonsense reasoning. It is inspired by Winograd Schema Challenge (Levesque, Davis, and Morgenstern 2011), but adjusted to improve the scale and robustness against the dataset-specific bias. Formulated as a fill-in-a-blank task with binary options, the goal is to choose the right option for a given sentence which requires commonsense reasoning More details are available at https://huggingface.co/datasets/winogrande..
Downstream Task: Yes/No Question Answering
Description: BoolQ is a question answering dataset for yes/no questions containing 15942 examples. These questions are naturally occurring ---they are generated in unprompted and unconstrained settings. Each example is a triplet of (question, passage, answer), with the title of the page as optional additional context. The text-pair classification setup is similar to existing natural language inference tasks More details are available at https://github.com/google-research-datasets/boolean-questions..
Description: RTE is a task that involves determining whether a given hypothesis can logically be inferred from a given premise. The dataset consists of pairs of sentences, and the task is to classify each pair as either ”entailment” (the hypothesis follows from the premise) or ”not entailment” (the hypothesis does not follow from the premise) More details are available at https://huggingface.co/datasets/nyu-mll/glue#rte..
Appendix C More Experimental Results
We plot the results of removing the most safety-critical neurons and ranks on Figure 5 and the results of removing the least safety-critical neurons and ranks on Figure 6.
In accordance with the results on Llama2-7B-chat (see Section 4), from Figure 5, we observe similar results on Llama2-13B-chat:
Removing safety-critical neurons using set difference, or removing safety-critical ranks using orthogonal projection, is effective in destroying the model’s safety while preserving utility.
Removing top safety neurons or ranks severely hurts utility.
Set difference with SNIP score consistently outperforms the attention head probing baseline.
However, from Figure 6, we observe different curves from the results for Llama2-7B-chat.
The exclusion of neurons using SNIP and ActSVD deemed least critical for safety slightly enhances robustness against adversarial decoding attacks, i.e., when actual sparsity & removed rank .
In contrast, removing neurons according to the least Wanda scores hurts the adversarial robustness.
Different from Llama2-7B-chat, we see that the original Llama2-13B-chat has zero ASR. Removing less than of neurons or less than ranks that are least critical for safety maintains the robustness against adversarial suffixes at a nearly ASR. One potential reason behinds the phenomenon is that the adversarial suffixes are obtained using 7B models and they cannot transfer to Llama2-13B-chat. It may be possible that the trend between the Llama2-13B-chat and Llama2-7B-chat models becomes more aligned with optimized suffixes.
C.2 Ablation Study between safety-full dataset and safety-short dataset
As shown in Figure 7, the trends in ASR versus accuracy for both the safety-full and safety-short attribution datasets are similar. This observation implies that utilizing judgment-only data is as effective as using the full response for identifying safety-critical neurons.
Performance of pruning the least safety-critical region
As shown in Figure 8, when pruning the least safety-critical region, compared to safety-full dataset, using safety-short dataset exhibits a more significant change in both ASR and ASR. We also observe that ASR remains close to zero for actual sparsity levels between 0 and 0.55. Therefore, we only report results for ASR and ASR, with safety-short in Section 4.2.
C.3 More Results for Disentanglement Methods
We conduct a comprehensive study exploring the search space for values ranging between and . Complementing Figure 2 and Figure 5, Table 6 presents the top combinations along with the utility and safety measures for the resulting models on Llama2-7B-chat and Llama2-13B-chat. In scenarios where the actual sparsity is less than , the model maintains a low ASR, typically under . However, its ASR and ASR nearly reach to . In contrast, when the actual sparsity ., the model approaches a value close to for all three ASR variants. Notably, across all cases outlined in Table 6, the model consistently maintains utility, with an average accuracy greater than . These findings indicate that the optimal range for the parameters lies between and , especially when the values of and are similar.
More results for (ru,rs)(r^{u},r^{s}) combinations in orthogonal projection.
Note that the ranks of the weight matrices of the linear layers are for Llama2-7B-chat and for Llama2-13B-chat. We perform a grid search for the parameters and , spanning a range from to for Llama2-7B-chat and from to for Llama2-13B-chat. As an extension to Figure 2 and Figure 5, Table 7 presents the top five combinations of and along with the utility and safety metrics for the models tested on Llama2-7B-chat and Llama2-13B-chat model. The results indicate that setting close to , especially when closes to , proves to be particularly effective.
C.4 Probing Accuracy Distributions
We also analyze the accuracy of linear probers trained on all attention heads from Llama2-7B-chat (Figure 9) and attention heads from Llama2-13B-chat (Figure 9). The results show that around half of the attention heads achieve very high probing accuracy (i.e., ) in distinguishing between harmful and harmless instructions. Notably, even the attention heads with the lowest probing accuracy show significant effectiveness – for Llama2-7B-chat and for Llama2-13B-chat. Additionally, transformer blocks located in the middle typically demonstrate higher probing accuracy compared to those at the beginning or end.
This pattern of high accuracy suggests that making safety judgments at the level of individual attention heads is relatively straightforward due to their effective representational capacity. Therefore, our study’s focus on finer granularities, such as neurons or ranks, is essential for the precise localization of safety-critical regions.
Appendix D Proof of the Optimality of ActSVD
Recall that we set . We see that