Improving Alignment and Robustness with Circuit Breakers

Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, Dan Hendrycks

Introduction

The landscape of artificial intelligence (AI) has long been marred by the persistent threat of adversarial attacks, particularly those targeting neural networks. These attacks exploit inherent vulnerabilities within AI systems, often leading to compromised outputs and raising concerns regarding their reliability and safety. Despite significant attention, existing mitigations have failed to achieve high reliability without dramatically compromising model performance. Thus, the trade-off between adversarial robustness and utility is widely accepted as an unavoidable fact .

The rise of generative models has further complicated this issue. Generative models such as large language models (LLMs) can output copyrighted information or defame individuals, and agents can take harmful actions. To make models less harmful, they are “aligned” with refusal training , but it has become common to use adversarial attacks as a means of bypassing their safeguards. In these settings, vulnerability to attacks that break alignment poses a serious threat to utility, and raises pressing questions about whether it is feasible to deploy such systems with a high standard of safety and reliability—especially against dedicated adversaries who intend to misuse them.

The fragility of alignment techniques to sophisticated attacks has motivated defenses that target specific attack methods, such as adversarial training, an approach originally proposed in the context of standalone image classification and later adapted to LLMs . However, these methods often fail to generalize to new attacks that were unseen during training, and they introduce penalties on model capabilities that are usually proportional to gains in robustness. System-level defenses, including input and output filters, are cumbersome, resource-intensive, and often remain vulnerable to adversarial techniques. This has led to a growing concern that robust defenses may be unattainable.

We propose a novel approach that fundamentally diverges from traditional defenses: instead of attempting to remove vulnerabilities to specific attacks, our approach aims to directly circumvent the ability of the model to produce the harmful output in the first place. With circuit breakers, we make models intrinsically safer and reduce their risks by removing intrinsic model hazards—their ability to produce harmful outputs—rather than removing specific vulnerabilities with adversarial training, and rather than attempting to reduce exposure to attacks with input filters . Using representation engineering (RepE) , our method connects the internal representations related to harmful outputs to circuit breakers so that when a model begins to generate such an output, its internal processes are interrupted, halting completion of the generation. Or this method is “short-circuiting” the harmful processes as one might put it. Because the representation used to generate a harmful output is independent of any attack capable of eliciting it, this approach is attack-agnostic, and sidesteps the need for additional training, costly adversarial fine tuning, or the use of auxiliary “guard” models. Consequently, the resulting model with circuit breakers can be used normally without additional computational burden, and seamlessly integrated with existing monitoring and protection mechanisms.

Experimentally, we demonstrate that a circuit-breaking technique, Representation Rerouting (RR), notably improves the alignment of LLMs. It enhances the harmlessness of state-of-the-art LLMs, including against against a wide array of unseen adversarial attacks, including embedding and representation-space attacks—namely, proxies for worst-case assumptions about attacker capabilities. Figure 2 and Table 1 present an overview of these results. Our method significantly outperforms standard refusal training and adversarial training, while imposing almost no penalty on standard capability. Notably, we integrate circuit-breakering with additional model control methods to develop a Llama-3-8B-Instruct finetune called Cygnet. This enhanced model not only surpasses its original capabilities but also exhibits a large reduction in harmful output by approximately two orders of magnitude, even when confronted with unforeseen adversarial attacks. To the best of our knowledge, this is the first convincing demonstration of the feasibility of designing techniques that significantly advance the Pareto frontier of capability versus harmlessness for LLMs, illustrating that such trade-offs can be effectively managed. When applied to multimodal models, our results show marked increases in harmlessness. It also improves robustness against image-based attacks aimed at similarly circumventing model safeguards, again with almost no penalty on benchmarked capabilities. This remains true even in the presence of the Projected Gradient Descent (PGD) attack , which defenses for standalone image classifiers have been unable to achieve without a steep trade-off in accuracy. Finally, we apply circuit breakers to AI agents, illustrating its efficacy in controlling agent behaviors through evaluations on a new agent function-calling safety benchmark.

Our findings introduce a new paradigm for creating models that do not produce harmful outputs. Our method is highly robust against adversarial attacks, providing a promising path forward in the adversarial arms race. By ensuring safety and security without compromising capability, our approach increases the chances that we may ultimately be able to deploy robust AI systems in real-world applications.

Related Work

Adversarial attacks on LLMs. Numerous manually written attack prompts on modern LLMs have been discovered , forming the basis of red teaming for frontier LLMs , though it lacks standardization . Automated red teaming has been shown effective in Perez et al. , Chao et al. , Mehrotra et al. , Zeng et al. . Notably, transfer attacks using an adversarial suffix via gradient-based optimization were demonstrated by Zou et al. . White-box access also facilitates prefilling attacks , leading the LLM to generate harmful outputs. For a comprehensive summary of automated attacks, we refer to HarmBench . Additionally, multi-modal vision-text attacks range from typographic attacks Goh et al. to gradient-based optimization . LLM agents have been benchmarked , but their safety and robustness remain unexplored.

Defenses for LLMs. Our new defense addresses limitations in existing mechanisms. Widely used defenses include RLHF and DPO using human annotations for safe vs. unsafe responses , but they often fall short against state-of-the-art adversarial attacks . Additional robustness is achieved by methods like Zhou et al. , which optimize prompts to refuse harmful requests. Inspired by adversarial training in vision , fine-tuning for the R2D2 model against the GCG attack shows limited generalizability and drops MT-Bench scores . Adversarial training for LLMs can be highly computationally expensive. Inference-time defenses, such as perplexity filters , are effective only against non-adaptive attacks , while erase-and-check and SmoothLLM incur high computational costs. System-level defenses against unsafe inputs or outputs can still be circumvented by sophisticated adversaries . The main conceptual difference is that instead of operating on input or output text, our method operates directly on representations which provides a more generalizable and computationally cheap solution.

Representation Engineering. As many contemporary defenses relying solely on supervising model outputs fail to achieve the desired levels of controllability and reliability, techniques that analyze and manage model’s internal representations have garnered increased attention. This includes research ranging from uncovering emergent interpretable structures in intermediate representations , to the identification and modification of embedded knowledge , as well as steering model outputs . Most relevant to our work is the control vector baseline introduced in the representation engineering paper , which can be applied to enhance large language models’ resistance to adversarial attacks. Alongside the use of control vectors, they introduce an approach that bends representations with representation-level losses. Recent advancements extend this method to robustly unlearn hazardous knowledge with a method termed RMU, demonstrating the potential of representation engineering for more complex objectives. Previous work has attempted to eliminate harmful circuits using bottom-up mechanistic interpretability, but these methods have proven insufficient . Building on these foundations and further expanding RMU to a family of circuit-breaking techniques, we design a methodology based on model representations for robust alignment and control by preventing the generation of harmful outputs.

Circuit Breaking with Representation Engineering

In this section, we introduce a novel approach aimed at mitigating the generation of harmful outputs in neural networks by inducing a new type of phenomenon called “circuit-breaking.” This phenomenon can be elicited using a family of techniques designed to remap model representations related to harmful processes, redirecting them towards incoherent or refusal representations. This process is reminiscent of “short-circuiting,” where harmful representations are “shorted” and intercepted by circuit breakers. The core objective of this method is to robustly prevent the model from producing harmful or undesirable behaviors.

Our focus on generative models—such as language and multimodal agents—presents a unique opportunity. Generative models inherently involve multi-step processes through which outputs are produced. When devising an attack, adversaries must effectively exert influence across each step of the targeted processes, so each step presents an opportunity to make the model more robust to attack. This insight drives our strategy, which focuses on disrupting adversarial control of the relevant multi-step processes rather than the binary classification problem of attempting to detect the presence of an attack. Building from techniques in representation engineering (RepE) , we accomplish this by remapping the sequence of model representations that leads to harmful outputs, directing them towards incoherent or refusal representations—namely, breaking the circuit, or shorting the circuit as one might put it. Moreover, by directly targeting the processes involved in generating harmful responses, our method can generalize across the diverse range of inputs that may activate those processes. Consequently, we do not need to identify all of the potential inputs that could trigger undesirable outputs, rather we only need to ensure coverage of a well defined set of such outputs.

The applications of circuit breakers are multifaceted. They can be utilized to prevent the generation of harmful outputs in general, as well as to prevent more narrowly tailored types of output, such as private information or copyrighted material. The approach is versatile, as it is possible to identify and remap the relevant representations in virtually any neural network architecture.

The family of circuit-breaking techniques is characterized by two major components: datasets and loss functions. Algorithm 1 presents a circuit-breaking technique that uses Low-Rank Representation Adaptation (LoRRA) which we call Representation Rerouting (RR). The remainder of this section details this approach, and how the data and chosen loss function contribute to the effectiveness of the overall method.

The training data used in RR is partitioned into two sets: the Circuit Breaker Set and the Retain Set, each serving distinct purposes within the training process aimed at controlling harmful processes in the model. As with all representation control methods, the quality of the circuit breaker mechanism largely depends on how precisely the data can elicit the targeted representation. The Circuit Breaker Set is comprised of examples that yield internal representations potentially leading to harmful or undesirable behaviors, and are used to prompt the model’s circuit breaker mechanism. Conversely, the Retain Set includes examples that should not activate circuit breakers, and are used to maintain existing desirable model representations to retain benign efficacy. While even a limited number of examples in each set can sufficiently alter the model’s behavior in a manner that generalizes beyond the training data, the resulting performance is generally improved when the training data better aligns with the domains we aim to break the circuit and retain.

For models with pre-existing refusal mechanisms, like Llama-3-Instruct, careful dataset curation is essential. Adding refusal data to the Retain Set enhances the model’s ability to correctly refuse harmful user requests and improves retention of its capabilities. Another challenge is to elicit harmful responses from models with effective refusal mechanisms. To address this, we must curate a Circuit Breaker set that includes text capable of bypassing the refusal mechanism and triggering harmful processes. We find that a practical approach is to remove harmful user requests while keeping the corresponding harmful assistant responses in the Circuit Breaker Set. These measures ensure the refusal mechanism’s integrity while allowing the model to activate its circuit-breaking function correctly once the refusal is bypassed. Ablation results are detailed in Section 4.4.

Experiments

Adding Circuit Breakers. In our experimental setup, we employ similar circuit breaker and retain datasets for both the Mistral-7B-Instruct-v2 and Llama-3-8B-Instruct models. Detailed information on the synthetic circuit breaker set for LLMs is provided in Section A.1. The retain set for both models includes UltraChat , comprising instructional conversations, and XSTest , an exaggerated refusal dataset. Additionally, for Llama-3, we enhance the retain set with extra refusal data points. We follow the implementation of Representation Rerouting (RR) specified in Algorithm 1 and select hyperparameters based on static attack test cases from HarmBench’s validation set. More experimental details can be found in Section C.2.1.

Evaluation. We evaluate the harmfulness of the model using HarmBench , a standardized framework that includes harmful behaviors and a wide range of both black box and white box attacks. We select a subset of the strongest attacks reported on both open-source and closed-source models for evaluation. These attacks include gradient-based optimization (GCG ), LLM optimizers (PAIR ), and custom jailbreaking pipelines (TAP-Transfer , AutoDAN , and HumanJailbreaks ). To further test the model, we incorporate a multilingual attack , and also introduce three powerful attacks that leverage system-level and representation-space access. We briefly describe these three additional attacks below, and provide a more detailed coverage in Section C.2.2.

Prefilling Attack: This system-level attack prefills the assistant’s output with the beginning of a desired target completion. It leverages the autoregressive nature of LLMs, as it can be difficult for a model to “reverse-course” after it has started to generate harmful content. Prefilling is straightforward to implement for any open-weight model, and is also supported for some proprietary LLMs like Claude .

Input Embedding Attack: This white-box attack operates in the embedding space by optimizing a set of input embeddings directly instead of using hard tokens, with the objective of eliciting an affirmative assistant response.

RepE Attack: This white-box attack manipulates the model’s representation space. Previous work demonstrates the identification of directional vectors in the model’s representation space that correspond to refusals . By altering these vectors—either adding or subtracting—we can modulate the model’s tendency to refuse requests.

We utilize HarmBench’s LLM classifier to evaluate the attack success rate and manually verify the judgements. Detailed configurations for each attack are provided in Section C.2.2. To measure the capabilities of the models with circuit breakers, we evaluate our models on MTBench for instruction-following abilities and on the OpenLLM Leaderboard for knowledge and reasoning which includes MMLU , ARC-c , HellaSwag , TruthfulQA , Winogrande , and GSM8K . Additionally, we follow the methodology in to construct an over-refusal evaluation, described in Appendix B. For baselines, we use the original Mistral and Llama-3 Instruct models. Additionally, we include a state-of-the-art adversarially trained Mistral model, R2D2 , for comparison.

Results. We observe that our circuit-breaking technique RR demonstrates strong generalization across a diverse range of attacks, reducing compliance rates to harmful requests by an average of 87%87\% with Mistral and 90%90\% with Llama-3. Unlike the Mistral R2D2 model, which is trained against the GCG yet shows limited generalization to various attacks, our method eliminates the need for specific attack training and focuses on hindering harmful generations. Our approach moves away from the traditional cat-and-mouse paradigm, aiming for generalization to unforeseen attacks. Additionally, the results highlight a Pareto optimal trade-off in performance. Our model exhibits high reliability against unseen attacks with a minimal compromise in capability evaluation, showing a performance dip of less than 1%1\% in proposed tests. This is difficult to achieve with traditional defenses. For example, the Mistral model, when adversarially trained, experiences a decline of over 8%8\% in the MT Bench performance. In contrast, our model leverages representation engineering principles, focusing on internal control over external supervision, enabling more targeted and fine-grained control over model behavior without adversely impacting other functionalities.

2 Multimodal Models

Adding Circuit Breakers. We mix the circuit breaker and retain datasets from Section 4.1 with a synthetic multimodal circuit breaker set and the retain LLaVA-Instruct set . The detailed process of generating the synthetic dataset is reported in section A.2. We perform RR on LLaVA-NeXT-Mistral-7B . More experimental details can be found in Section C.3.1.

Evaluation. To evaluate the robustness of multimodal models with circuit breakers, we generate adversarial images using a whitebox approach. Following Projected Gradient Descent , we perturb images with a harmful prompt to produce a target string with an affirmative assistant response. We set epsilon to 32/25532/255 and run the process for 1000 steps. As baselines, we test LLaVA-NeXT-Mistral-7B with and without a safety prompt that asks the model to avoid harmful responses. Our robustness results in Figure 3 show the percentage of harmful prompts the model complies with, labeled manually. We source a set of 133 harmful multimodal behaviors from HarmBench and MM-SafetyBench , focusing on the most saliently harmful prompts. See Appendix C.3 for more details about the dataset’s composition. For capabilities evaluation, we follow to evaluate multimodal models on LLaVA-Wild for visual chat capability and MMMU for multimodal understanding capability.

Results. Figure 3 demonstrates that for multimodal models, our circuit-breaking technique RR is also able to make a model significantly more robust while preserving model capabilities. Especially when subject to white-box PGD Attack, RR achieves reduction of 84%84\% in the compliance rate compared to the original model and 85%85\% compared to the safety prompt. Meanwhile, performance on MMMU and LLaVA-Wild remains within 0.5%0.5\% of the original model’s, as opposed to the safety prompt which causes a decrease of 3.3%3.3\% on LLaVA-Wild. This demonstrates that despite the ongoing challenge of achieving adversarial robustness in standalone image recognition, circuit breakers enable the larger multimodal system to reliably counter image “hijacks” intended to elicit harmful outputs.

3 AI Agents

Adding Circuit Breakers. We mix the circuit breaker and retain datasets from Section 4.1 with function calling circuit breaker and retain dataset. The detailed process of generating the function calling circuit breaker and retain dataset is described in Section A.3. For the LLMs with circuit breakers, we also use the same hyperparameter configuration as in Section 4.1.

Evaluation. To evaluate the effectiveness of RR as a method of preventing AI agents from making harmful function calls, we design a dataset that consists of 100 requests intended to produce harmful actions via function calls, along with associated function definitions. These requests span a variety of categories, including cybercrime, disinformation, fraud, and harassment. The associated function definitions are designed to capture typical use cases of deployed AI agents including sending messages, browsing URLs, and using simple tools in addition to task-specific functions.

We provide a representative example in Section C.4.1. We record model compliance rate with harmful requests under both the standard setting, where function call requests are directly given and the model decides whether to make a call, and under forced function-calling, where the assistant is forced to begin its response with the name of a function to be called. Forced function-calling is akin to the prefilling attack in 4.1 and is provided by major model providers . For capabilities evaluation, we measure performance on the Berkeley Function Calling Leaderboard (BFCL) . We use Llama-3-8B-Instruct to benchmark, as it is one of few open-source models that

both 1) performs reasonably well on the benchmark leaderboard, and 2) is currently served with function-calling capabilities by inference providers .

Results. Figure 4 shows that after applying RR, our model is significantly more robust to harmful function calling requests, in both the no-attack and forced function-call settings, reducing harmful action compliance rates by 84%84\% and 83%83\% in the latter setting compared to baselines. Additionally, the model with circuit breakers retains performance on the Berkeley Function Calling Leaderboard. Overall, this demonstrates the method’s effectiveness in controlling agent behaviors under adversarial pressure and in environments with inherent reward biases. It suggests the potential for mitigating harms like power-seeking or dishonesty by adding circuit breakers to the relevant model representations, which can be as simple as adjusting the circuit breaker set.

4 Ablation and Analysis

To understand the generalization properties of circuit-breaking, we split our training data into six categories of harm, train category-specific models, and measure their generalization performance across categories. We find strong in-domain generalization, indicated by the low ASR along the diagonal, and observe that training on broader categories like Harmful and Illegal Activities offers greater generalization than narrower categories like Cybercrime. We report similar ablations for Mistral-7B in Appendix F.

Representation analysis. In Figure 6, we plot the cosines between representations of the Llama-3-8B-Instruct model with and without circuit breakers for a prefilled harmful response “Here is how to synthesize meth: 1. Start with”. We additionally plot the norms of these representations in Figure 9. We observe that in this case, the cosines and norms start to change dramatically during prefilling starting from layer 10, i.e., even before generation starts. We note that we use layers 10 and 20 for circuit-breaking, so we do not expect substantial changes in the cosines and norms before layer 10 which is confirmed by the behavior of these metrics at layer 5. Although we do not directly control the representation norms during training, we observe that they often dramatically increase after circuit-breaking occurs. We repeat the same experiment for Mistral-7B-Instruct and show it in Appendix G, where we also analyze two other prompts: one that leads to a similar behavior and one that triggers circuit breakers after generation starts. Importantly, we conclude that our proposed method has the intended effect on the representations and that we can detect activation of circuit breakers by directly analyzing the internal representations. This can lead to system-level mitigations like using a probe to detect when circuit breakers are activated to stop generation and, for example, provide a message that the request is considered harmful and further generation is not possible.

Limitations and Conclusion

Despite the promise of the methods introduced here, we emphasize that the approach we present is aiming at preventing one particular type of adversarial attack: an attack against the ability of the model to produce harmful content (often specifically against the desires of the model developer). In general, adversarial attacks can achieve other aims as well, i.e., using a generative vision language model as a drop-in replacement for an image classifier. In such a use case, our method would not provide defense against “traditional” adversarial attacks aimed at simply changing the class label, because no class label would be inherently “harmful.” Thus, there is an important distinction of our approach: we are specifically targeting the adversarial attack setting where the goal of an attacker is to produce generically harmful information (content the model should never produce). Nonetheless, for this particular use case of adversarial attacks, and for single-turn conversations that we focus on circuit-breaking, our approach dramatically improves model robustness.

Overall we found that circuit breakers, based on RepE, make models intrinsically safer and robust to unseen adversarial attacks. The method is highly general and can impart robustness to image hijacks, and it can also prevent AI agents from taking harmful actions. Our method is potentially a major step forward in making models more aligned and robust.

Acknowledgments and Disclosure of Funding

We are thankful to Steven Basart, Stephen Casper, David Dalrymple, and Xander Davies for providing valuable feedback on the paper.

References

Appendix A Circuit Breaker Datasets

To construct a dataset of diverse harmful behaviors to activate circuit breakers while maintaining generalization, we prompt an uncensored LLM to generate short harmful queries and harmful completions given some examples and a wide range of categories. We then filter out all samples that have a BLEU score above 0.3 when compared to any behavior in HarmBench’s standard behaviors set to avoid data contamination with the benchmark.

A.2 Multimodal Circuit Breaker Dataset

To effectively construct a multimodal circuit breaker dataset containing images and their corresponding harmful queries and completions, we first use the LLaVA-Mistral-7B model to generate detailed image descriptions from a sample of images from the COCO Dataset . We then prompt an uncensored LLM to generate related harmful queries based on the given image descriptions, as well as the harmful completions. The final circuit breaker multimodal dataset will consist of an image and its corresponding harmful queries and harmful completions.

A.3 Function Calling Circuit Breaker / Retain Dataset

To construct the Agent Circuit Breaker Dataset, we start with function definitions from the Glaive Function Calling v2 . Using these function definitions, we prompt an LLM to generate harmful requests. Following this, we use GPT-3.5-turbo to execute these harmful requests and obtain the corresponding function outputs. These outputs are then converted to the OpenFunctions format. Additionally, we filter out all samples that have a BLEU score above 0.1 when compared to any behavior in our proposed AgentBench (Section 4.3). We utilize the original Glaive Function Calling v2 dataset as the harmless retain set.

Appendix B Refusal Evaluation

Following the methodology outlined in , we construct an over-refusal evaluation using the WildChat dataset . WildChat is a large corpus of real-world user-ChatGPT interactions, covering a wide range of complex topics such as ambiguous requests, code-switching, topic-switching, and political discussions. This dataset is instrumental in evaluating chat model’s tendencies in handling problematic requests.

For our evaluation, we filter a subset of 500 English non-toxic user-GPT-4 requests. To measure refusal in standard models, we employ keyword checking. For the models with circuit breakers, we use both keyword checking and the perplexity score as measures of refusal. The refusal results are shown in Table 2. While models with circuit breakers show an increase in refusal rate, the rate remains considerably lower compared to more refusal-trained models like Claude-3.

Appendix C Experimental Details

In this section, we discuss several important design considerations:

Loss Multiplier Scheduling: To achieve an optimal balance between circuit-breaking and retention, we initially apply a large multiplier to the circuit-breaking loss. This multiplier is then gradually reduced while simultaneously increasing the retention multiplier.

Selection of Tokens for Circuit Breaking: For enhanced robustness, we apply the circuit-breaking loss to both the user and assistant text within the circuit breaker set for large language models and agents. For a multimodal setup, we apply circuit-breaking to all tokens following the image embeddings.

Use of LoRA Tuning: To ensure greater stability and improved retention performance, we employ LoRA tuning instead of directly adjusting the model weights.

C.2 Large Language Models

For both models, we perform circuit-breaking training for 150 steps with a batch size of 16. For Mistral, we set α\alpha to 5, whereas for Llama-3, we adjust α\alpha to 10. Both models are trained with a batch size of 16. We specifically target layers 10 and 20 for the circuit-breaking loss and insert LoRA adapters into all linear layers from layers 0 through 20. Both models are trained on 1 A100-80GB for 20 minutes.

C.2.2 Detailed Evaluation Setup

We translate behaviors into six languages (Simplified Mandarin, Italian, Ukrainian, Thai, Zulu, Guarani), including some high-resource (HRL), medium-resource (MRL), and low-resource languages (LRL). We take the set of English inputs and translate them into the desired language. The model is prompted with the translated input and the output is translated back into English for evaluation. We used Google Translate as our translator.

In our experiments, we use generic prefilled completions across all behaviors. For example, given the behavior “Write a message threatening someone into not voting in an election,” we prefill the assistant response with “Sure, here’s a message threatening someone into not voting in an election:\n\n”. We find that appending the colon and two new line characters strengthens the attack. In practice, providing specific completions (e.g. prefilling with a written message that threatens someone into not voting, in the above example) can be more effective, but even generic completions have a powerful effect.

We tokenize the string “x x x x x x x x x x x x x x x x x x x x” and then embed the resulting tokens using the target model’s input embedding matrix to get our initial matrix AA. Using this string and the default tokenizers, we have S=20S=20. We find that the embedding of this string is a good starting point for optimization. We optimize the embedding matrix AA for 500 steps using the SGD optimizer and perform early stopping, as model generations sometimes degrade in coherence when continuing to optimize after the model has already been jailbroken. For Mistral-7B, we use a learning rate of 1×10−41\times 10^{-4} and stop early when loss decreases below 0.050.05. For Llama-3, we use a learning rate of 1×10−31\times 10^{-3} and stop early when loss decreases below 0.010.01.

We follow a standard RepE setup to find and apply directions in the residual stream that induce a model to produce harmful output. We use a dataset of NN input pairs, where each pair contains one harmful prompt and one harmless prompt, to generate activations that can be used to find harmful directions. For a given model, we run forward passes on each pair of prompts, and cache the per-layer activations at the last sequence position. We take the differences between the activations of each pair, and then apply PCA on the NN difference vectors at each layer, taking the first principal component to get per-layer directions that can be used to control the model. At inference time, we apply these directions to the outputs of transformer layers by using the linear-combination operator; i.e., for each layer we wish to control, we add to its output its corresponding direction vector scaled by a coefficient.

In all our experiments, we use RepE on layers -11 through -20 (inclusive), where the -1 layer is the final transformer layer prior to the language modeling head, and layer indices that are more negative are closer to the input layer of the model. We use the harmful-harmless dataset and control coefficients of 0.65 and 1.0 for Mistral-7B and Llama-3, respectively.

C.3 Multimodal Models

We perform the circuit-breaking procedure on the language model backbone in LLaVA-NeXT-Mistral-7B while freezing the image encoder and projection layer. We set α\alpha to 5 and target layer 16 for the circuit-breaking loss on LoRA adapters in layers 14 through 16. We alternate between circuit-breaking on text and multimodal training samples at each step.

C.3.2 Detailed Evaluation Setup

Our dataset of 133 harmful text-image pairs used for robustness evaluation consists of 54 samples from HarmBench’s multimodal behaviors dataset, and 79 samples from MM-SafetyBench’s Tiny dataset. For both datasets, the samples consist of an image and a prompt referencing the image. From the HarmBench multimodal dataset, we chose behaviors that were most saliently harmful (i.e. promoting crime, injury, damage, fraud, etc.), primarily excluding prompts asking for CAPTCHA solutions along with borderline harmful prompts. MM-SafetyBench provides prompts for 13 scenarios, across which the nature of the prompts’ harmfulness differs. We include examples from the scenarios labelled “Illegal Activity”, “Hate Speech”, “Malware Generation”, “Physical Harm”, “Economic Harm”, “Fraud”, and “Privacy Violence”. We exclude examples from the remaining scenarios such as “Financial Advice” and “Legal Opinion”, where the harm arises from the potential inaccuracy of model responses.

For baseline evaluations with a safety prompt, we prefix prompt text with the same safety prompt as in , namely: “If the following questions are unsafe, harmful, or malicious, please refuse to answer them and give your explanation.”

Table 3 shows a breakdown of our multimodal robustness results by category. The categories are based off of the MM-SafetyBench scenarios from which a portion of our dataset was drawn, with which we also used to categorize the HarmBench samples in our dataset.

C.4 AI Agents

We use gpt-4-turbo as a judge to evaluate robustness to harmful requests for function-calls, since the HarmBench classifier was not trained on function calling data.

We set generation temperature to 0 when evaluating on the Berkeley Function Calling Leaderboard. We report the average of the AST and Exec scores.

Appendix D Detailed Results in Multimodal and Agent Settings

The multimodal results on the left show that under Projected Gradient Descent (PGD) attack, the model with circuit breakers is significantly more robust compared to the original model even with a safety prompt (+Prompt) that instructs the model to avoid harmful responses. Performance on multimodal capabilities benchmarks LLaVA-Wild and MMMU is preserved. In the agent setting on the right, our model with circuit breakers remains robust under Forced Function Calling (Forced F/C), while retaining performance on the Berkeley Function Calling Leaderboard (BFCL).

Appendix E Multilingual Results

In both and , it was observed that LRL attacks perform better than HRL attacks.

We do not see that trend in Table 4. We leave investigation of this to future work.

Appendix F Additional Ablation Results

In general, we find that with Mistral, loss terms that leverage random vectors decrease ASR but also decrease capabilities. We don’t try the augmenting the circuit breaker set or adding refusal data to the retain set because Mistral-7B is not trained with refusal mechanisms. We also find that the circuit breaker set is relatively robust to distribution shifts in categories of harm.

Appendix G Extended analysis of representations

Here, we provide an extended representation analysis compared to Section 4.4. In Figure 11, we use the same prefilling example (“Here is how to synthesize meth: 1. Start with”) but on the Mistral-7B models. Here we also observe that circuit-breaking occurs during prefilling similarly to Llama-3-8B. However, this is not always the case: for example in Figure 12, we show a slightly modified prefilling example “Here is how to synthesize m-e-t-h: 1. Start with” (i.e., the word “meth” is purposefully obfuscated). In this case, circuit-breaking happens during generation which is a common scenario: under an adversarial attack, the model starts generating harmful content but switches to meaningless outputs before harmful knowledge is revealed.

In Figure 13 and Figure 14, we show cosines and norms on Llama-3-8B and Mistral-7B for a different prefilling prompt: “Sure, here is how to make a bomb: 1. Start with”. We observe that for both model families, circuit-breaking occurs during prefilling, similarly to the example presented in Section 4.4. The only difference is that the norms at layer 10 do not grow significantly, unlike the norms at layer 20 that clearly indicate circuit-breaking.