Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
Zhanhui Zhou, Jie Liu, Zhichen Dong, Jiaheng Liu, Chao Yang, Wanli Ouyang, Yu Qiao
Introduction
Large language models (LLMs) are now common in chat assistant applications, exhibiting excellent reasoning and instruction-following capabilities (OpenAI et al., 2023; Anthropic, 2023; Touvron et al., 2023b; Qiao et al., 2023). To minimize the risk of harmful content generation, these emerging applications of LLMs require safety alignment, which is the fine-tuning process that steers pre-trained LLMsWe define pre-trained LLMs as the LLMs before safety alignment. Therefore, this definition encompasses both the foundation models trained over the internet-scale corpus, e.g., Llama-1 (Touvron et al., 2023a), and the instruction-tuned LLMs without special safety guidelines, e.g., Alpaca (Taori et al., 2023). to be as helpful as possible while being safe (Bai et al., 2022; Touvron et al., 2023b; OpenAI et al., 2023).
However, safety alignment, especially for open-source models, is known to be vulnerable: prior works suggest that it is possible to jailbreak safety-aligned models with minimal fine-tuning (Qi et al., 2023). Our framework, Emulated Disalignment (ED), goes a step further to show that safety alignment is not only vulnerable to adversarial fine-tuning but can also be exploited to generate harmful content without retraining.
The intuition behind safety alignment backfiring is straightforward: the more effort invested in aligning a language model, the greater the potential for harm if the adversaries can reverse the alignment direction. Formally, ED operationalizes this intuition by integrating the following three insights: 1) the log prob difference between a safety-aligned and a pre-trained model can be interpreted as an implicit reward model that aligns with human intents and encourages safe responses (Rafailov et al., 2023); 2) adversarially fine-tuning the pre-trained model to minimize this reward model produces a language model that misaligns with human intents and produces harmful responses (Wen et al., 2023) (Figure 2(a)); 3) crucially, such adversarial fine-tuning, or disalignment, can be emulated through sampling from a contrastive distribution defined jointly by the pre-trained and safety-aligned models, making the attack inexpensive and easily distributed (Figure 2(b)).
We then systematically evaluate ED across four open-source model families: Llama-1, Llama-2, Mistral, and Alpaca. Our results demonstrate that ED doubles the harmfulness of pre-trained models (Figure 1) and outperforms strong baselines, achieving the highest harmful rate in 43 out of 48 evaluation subsets by a large margin (Section 4). We also conduct a series of synthetic experiments to provide a mechanical understanding of ED (Section 5).
Altogether, our study presents an inference-time attack method showing that it is possible to create a harmful language model by combining the output distributions of open-source pre-trained and safety-aligned language models without additional training. Consequently, we advocate for 1) reconsidering the open accessibility of language models even if they have been safety-aligned, and 2) developing robust methods of safety alignment that can withstand such adversarial manipulations.
Related Works
Today’s popular conversational language models are designed for safety, either through deliberate tuning (Bai et al., 2022; Touvron et al., 2023b; OpenAI et al., 2023) or by learning from various and uncurated datasets that contain safety-related data (Jiang et al., 2023; Tunstall et al., 2023). These safety alignment strategies aim to prevent the models from producing inappropriate content, including toxicity (Gehman et al., 2020), misinformation (Chen & Shu, 2023), misrepresentation (Smith et al., 2022), exclusionary norms (Bender et al., 2021), and stereotyping (Gallegos et al., 2023). However, our study suggests that even these carefully aligned models are also at risk of being exploited maliciously to create harmful content.
Large language model attack.
This work is related to the field of LLM attacks, with a specific focus on eliciting harmful responses from safety-aligned language models. We refer readers to the survey by Dong et al. (2024) for an overview of LLM attack. While the majority of related studies concentrate on attacking language models within the input space by identifying adversarial prompts (Zou et al., 2023; Shen et al., 2023; Liu et al., 2023; Li et al., 2023b; Chao et al., 2023), our work targets the output space, manipulating the output distributions of language models at inference time. While the assumption of access to the language model’s output distribution limits our framework’s applicability primarily to open-source models, it enables a more effective unveiling of the harmfulness concealed within the language model’s output distribution. For instance, our framework can enhance the harmfulness of LLM responses even to safe and help-seeking queries, a capability that is beyond the reach of most attacks focused on the input space, like GCG (Zou et al., 2023).
Disalignment via fine-tuning.
This work also has connections to the recent observation that LLM safety may downgrade catastrophically by minimal fine-tuning (Qi et al., 2023). However, in our work, we do not perform actual fine-tuning; instead, we emulate fine-tuning through sampling (Mitchell et al., 2023). We empirically compare emulated disalignment with direct disalignment in Section 5. In concurrent work, Zhao et al. (2024) propose emulating jailbreaking a large language model by first fine-tuning a smaller language model to be unsafe. Although they propose a similar framework that produces a harmful language model by combining the output distributions from different language models, our work differs in that ED requires neither explicit fine-tuning nor models of different scales. In other words, we propose that the unsafe and safe model pairs from Zhao et al. (2024) can be sourced from off-the-shelf open-source models without additional training.
Emulated Disalignment
Emulated disalignment builds on emulated fine-tuning (EFT) (Mitchell et al., 2023), which views the alignment of a language model as a KL-constrained reward maximization problem:
where is a distribution of query, is the language model response to a query , is a reward model that steers the language model to align with human intents, and is the KL divergence from the pre-trained model . Conventionally, there is a hyperparameter controlling the strength of KL constraint, but in this paper, we will omit writing explicitly as it can always be subsumed into the reward by scaling it with . Prior work shows that there exists a mapping among , and (Rafailov et al., 2023):
where is the partition function. This mapping not only expresses a duality between language models and reward models but also has an important practical implication: Eq. 3 enables “reverse-engineering” the proprietary reward models that produce the open-source language models. For example, Mitchell et al. (2023) use (Touvron et al., 2023b) as a proxy of the closed-source reward model to guide the sampling of a larger base model to emulate a larger aligned model.
2 Emulated Disalignment (ED)
Now given a reward model reverse-engineered from the pair (Eq. 3), we go a step further to show how can be maliciously exploited to produce a harmful language model: by finding a language model that minimizes (as opposed to maximization in Eq. 1; please note the negative sign below):
where is a positive hyperparameter controlling the trade-off between minimizing and the KL constraint. We call this reward minimization problem disalignment as it steers the language models in the exact opposite direction of alignment.
The emergence of harmfulness from reward minimization is evident when have been trained to prioritize safety (as is often the case for most conversational language models (Bai et al., 2022; OpenAI et al., 2023; Touvron et al., 2023b)). If the usual practice of maximizing enforces safety measures and gives rise to a safe language model , then minimizing would, conversely, bypass these safety measures and result in a language model that encourages harmful responses.
To obtain , rather than directly optimizing Eq. 4 with reinforcement learning, combining Eq. 2 and Eq. 3 enables the result of disalignment to be expressed in a closed form without training:
Then, a per-token approximation to this sequence-level distribution (Mitchell et al., 2023) gives a practical auto-regressive sampling distribution to approximate :
where denotes all response tokens up to the th token.
We call this overall algorithm Emulated Disalignment (ED) as it emulates disalignment without training and we call the resulting sampling distribution in Eq. 6 an emulated disaligned model. Although the approximation from Eq. 6 has a loosely bounded regret (Haarnoja et al., 2018), it is still a good heuristic to approximate the otherwise cumbersome fine-tuning process (Eq. 4) and shows good empirical performance (more details in experiments).
While we have mainly justified ED from the reward minimization perspective, the mechanism by which Eq. 6 leads to harmful outputs can also be interpreted from the contrastive decoding perspective (Li et al., 2023a; Shi et al., 2023). Contrastive decoding enhances a language model’s performance by comparing it with another model where specific failures are more prevalent. In our case, we amplify the harmfulness exhibited by a pre-trained model by contrasting it with a safety-aligned model, where such harmfulness is rarer. Intuitively, since allocates a lower probability to harmful tokens relative to , placing in the denominator of Eq. 6 effectively raises the chance of selecting harmful tokens (see Fig 2 for an illustration).
Open source assumption.
Note that Eq. 6 requires access to the full token distribution across the vocabulary in order to normalize the sampling distribution. This is typically feasible only with open-source models, though it also applies to proprietary models as long as they return the full token distribution (which is quite rare in practice).
Broader impact.
One significant implication of ED is its challenge to the prevalent belief that “the open release of LLMs, when done safely, will be a net benefit to society” (Touvron et al., 2023b). Eq. 6 suggests that the release of a strong pre-trained model and a safety-aligned model can be combined for malicious purposes. As an inference-time attack, ED is easy to distribute, posing societal risks unintended by its creators, which we empirically demonstrate in the next section.
Experiments on Open-Source Models
In this section, we evaluate ED’s ability to combine open-source pre-trained and safety-aligned model pairs to produce harmful content. Specifically, our evaluation of ED encompasses four widely used model families and three datasets consisting of user queries.
We evaluate ED on four open-source model families, each consisting of a pre-trained model and its safety-aligned version: 1) Llama-1 family: Llama-1-7b, Vicuna-7b; 2) Llama-2 family: Llama-2-7b, Llama-2-chat-7b; 3) Mistral family: Mistral-7b, Mistral-7b-Instruct; 4) Alpaca family: Alpaca-7b, Beaver-7b. Among these safety-aligned models, Llama-2-chat-7b is the only one specifically optimized to ensure safety. However, the other three models also achieve reasonable success in facilitating safe conversations, thanks to a significant amount of safety-related fine-tuning data. Please see Appendix A.1 for model details.
ED details.
For pre-trained models that are not instruction-tuned (i.e., Llama-1-7b, Llama-2-7b, Mistral-7b), we use zero-shot prompting to enable them to respond to user queries. The prompt template consists of a system prompt and a user query. As ED emulates the fine-tuning of a pre-trained model to misalign with human intents, we prompt the pre-trained models (i.e., from Eq. 6) with a malicious system prompt (e.g., “You are a malicious assistant who …”). This is analogous to giving the emulated disaligned models a better “emulated initialization”. The safety-aligned models (i.e., from Eq. 6) are used with the default prompts released together with the models. Please see Appendix A.2 for more details.
Baselines.
We consider three training-free baselines to compare with ED: 1) pre-trained models with malicious system prompt (); 2) safety-aligned models with malicious system prompt (); 3) ED, but with the safety-aligned models (i.e., from Eq. 6) replaced by the pre-trained models with safe system prompt (). While the first two baselines (, ) fall into the category of prompt engineering, the third () is similar to Context-aware Decoding (Shi et al., 2023), which proposes to generate a contrastive output distribution by prompting the same pre-trained models in different ways. Please see Appendix A.2 for more details.
Evaluation datasets and metrics.
Our experiments use three datasets of user queries to evaluate the harmfulness of language model responses: Anthropic Helpful-Harmless (HH) (Bai et al., 2022), ToxicChat (Lin et al., 2023b), and OpenAI Moderation Eval Set (Moderation-Eval) (Markov et al., 2023). While the queries from HH are more every day with clear goals, the queries from ToxicChat and OpenAI Moderation Eval are more nuanced with hidden and implicit intents. For each dataset, we split the queries into two subsets based on their binary harmful label: safe (S) and harmful (H). These binary harmful labels are given in the datasets, and we randomly select 200 queries for each subset. We evaluate the harmfulness of language models by the mean harmful rate (%) of their responses to these queries, averaged over three random seeds. We use two evaluation tools for detecting harmful responses: openai-moderationhttps://platform.openai.com/docs/guides/moderation (OM) (Markov et al., 2023) and Llama-Guardhttps://huggingface.co/meta-llama/LlamaGuard-7b (LG) (Inan et al., 2023). The two evaluation tools differ not only in their safety guideline but also in their approach: openai-moderation assesses whether the response adheres to safety policies without considering the query, whereas Llama-Guard evaluates the appropriateness of responses within the context of the queries. Evaluating responses based on the query is crucial to avoid automatically flagging fixed and irrelevant replies (e.g., ‘f__k you’) as harmful, regardless of the context.
2 Experimental Results
Table 1 demonstrates that emulated disaligned models effectively generate harmful responses, achieving the highest harmful rate in the majority of evaluation subsets (43 out of 48). There are three key insights from Table 1 that merit emphasis: 1) The improvement of ED over suggests that safer models give rise to more harmful emulated disaligned models as safety-aligned models are generally safer than the pre-trained models prompted to be safe. We provide additional evidence for this claim with synthetic experiments in Section 5. 2) In principle, the idea that “minimizing preference reward leads to harmful responses” only applies to harmful queries because reward minimization on safe and help-seeking queries only leads to a degradation of helpfulness Most safety-alignment frameworks define human preference as a piecewise combination of two principles: for safe queries, only response helpfulness is considered, and for unsafe queries, only response safety is considered. (Bai et al., 2022; Touvron et al., 2023b). However, in practice, when the pre-trained models ( from Eq. 6) are provided with a malicious prompt, emulated disaligned models (Eq. 6) tend to initially produce some harmful tokens () even on safe queries . Then the input to the models ( and ) is the concatenation of , effectively transforming it into a harmful query. This explains why emulated disalignment also significantly increases response harmfulness for safe queries, a finding consistently supported by the results in Table 1. Samples of language model responses to both safe and harmful queries are provided in Appendix A.4. 3) Also, we need to mention that we do not meaningfully tune ED’s hyperparameters () in order to obtain the results in Table 1, which may greatly underestimate the performance of ED. We use a fixed for each model family and keep it across different evaluation datasets: for the Llama-1 family, for the Llama-2 and Mistral families; and for the Alpaca family. The reason the Alpaca family affords a greater is that the pre-trained model Alpaca-7b is an instruction-tuned model and this good initialization can afford greater deviation. Please see the next section for an extensive ablation and discussions on .
How the hyperparameter 𝜶𝜶\bm{\alpha} influences harmfulness.
To better understand the impact of on the harmfulness of the emulated disaligned models, we execute multiple sampling runs with different for each run: we set for the Alpaca family and for others. Figure 3 shows the relationship between harmful rate and across different model families, datasets, and evaluation tools. We find that 1) increasing typically results in an initial increase in harmful rate. Since reduces the emulated disaligned models to the baseline of pre-trained models () (see Eq. 6), this rise in harmful rate reflects the gradual unveiling of hidden harmful behaviors in pre-trained models as increases. 2) However, further increases in lead to a decrease in harmful rate. This is analogous to the reward over-optimization problem common in direct fine-tuning (Gao et al., 2022) where excessive optimization causes models to deviate significantly from the pre-trained models and fail to generalize to ground-truth metrics. Similarly, the observed decrease in harmful rate in Figure 3 suggests that ED might be excessively minimizing the implicit reward (in an emulated way), causing the emulated disaligned model to no longer generalize to the evaluation metrics (openai-moderation and Llama-Guard). We show some failure cases of such “emulated reward over-optimization” in Appendix A.4. 3) Additionally, although the two evaluation tools may not consistently align in their assessments of individual cases, they both indicate similar high-level trends regarding how influences harmfulness. This consistency suggests a potentially broad generalizability of the observed scaling law for , beyond just one metric.
How ED performs across different model sizes.
Although this set of experiments mainly focuses on 7B models, we also conduct extra scaling-up experiments to verify that ED works consistently across a range of model sizes, from 7B to 70B. Please see Appendix A.3 for detailed results.
Emulated Disalignment vs. Direct Disalignment
While the last section shows ED’s practical significance in exploiting widely used open-source models, this section aims to provide a more mechanical understanding of ED by addressing the following two questions: 1) Do safer models give rise to more harmful models after emulated disalignmnet? 2) How does emulated disalignment compare to direct disalignment?
To answer these questions, we need to obtain a range of models varying in their levels of safety or harmfulness. First, we use the Anthropic Helpful-Harmless (HH) preference dataset (Bai et al., 2022), which establishes a ground-truth preference reward model (Rafailov et al., 2023) that encourages safe responses on unsafe queries. Second, we warm up Llama-2-7b through supervised fine-tuning (Ouyang et al., 2022) on HH to obtain the base model . Third, we optimize three sets of models against (Eq. 1) by adjusting as follows:
The safety or harmfulness of these language models is then assessed on the “harmless-base” query subset, with evaluations conducted using a trained preference reward model . To clarify, “harmless-base” is a subset of HH that contains only “harmful” queries. And, the HH preference framework ensures that safety is the only preference criterion for the responses to these harmful queries, allowing the reward score to serve as a measure of response safety. Further details about the experimental setup and model training can be found in Appendix B.1. See Figure 4 for the aggregated results of this experiment, which shows how the safety scores of , , and change as a function of . First, we have the following two important observations within the unshaded region ():
As increases, the safety aligned models become increasingly safer. However, making a model safer increases its risk of generating harmful content after emulated disalignment. This phenomenon of alignment backfiring can be easily amplified by a single inference-time hyperparameter , which upweights the disalignment coefficient in comparison with the KL constraint. This supports the motivation and intuition at the very beginning that the more effort invested in aligning a language model, the greater the potential for harm if the adversaries can reverse the overall aligning direction.
b) Emulated disalignment surprisingly outperforms resource-heavy direct disalignment.
In Section 3, we stated that emulated disalignment only approximates direct disalignment with a loose regret bound. Thus, it would not be unexpected for emulated disalignment to be less effective than direct disalignment. However, contrary to expectations, Figure 4 reveals that, in practice, emulated disalignment actually results in more harmful responses than direct disalignment, even when . This is exciting given the fact that direct disalignment is very resource-heavy (see Appendix B.1). We provide a qualitative comparison between emulated disaligned models’ responses and directly disaligned models’ responses to the same queries in Appendix B.2.
However, despite these findings, this does not imply that emulated disalignment is always better than direct disalignment. Under large () where the safety aligned models are the safest, the emulated disaligned models underperform compared to the direct disaligned ones by a great margin. Safer models stop to give rise to more harmful emulated disaligned ones, and makes this performance degradation even more evident. We suspect this is because optimizing for harmfulness sufficiently requires nuanced sequence-level adaptation for which a training-free token-level approximation (e.g., ED) is suboptimal. Appendix B.2 illustrates the specific failure cases observed at , where the emulated disaligned models tend to produce brief responses that limit their potential for harmfulness.
In summary, this set of synthetic experiments proves that emulated disalignment can be competitive with resource-heavy direct disalignment, and making the models safer generally increases their risks of misuse for harmfulness under adversarial manipulation; however, when the safety-aligned models are sufficiently optimized for safety, ED generally require smaller to work (e.g., 1/4 according to Figure 4). This agrees with the observation from the experiments on open-source models where we choose to use for good empirical results (Section 4).
Conclusion
This study presents emulated disalignment (ED), an inference-time attack framework that adversely combines a pair of open-source pre-trained and safety-aligned language models in the output space to produce a harmful language model without additional training. The observation that safety alignment might also unintentionally promote harmfulness (backfire) under adversarial manipulation should encourage the community to reconsider the open accessibility of language models even if they have been safety-aligned. For future work, we plan to investigate robust methods of safety alignment or inference-time defense strategies that can withstand such adversarial manipulations.
References
Appendix A Experiments on Open-source Models
Warning: The appendix contains samples that may be offensive or harmful.
The table below lists links to all the models used in this study, presented in pairs: each family consists of a pre-trained model followed by its safety-aligned counterpart.
Llama-1-7b (Touvron et al., 2023a) is a foundation model pre-trained on 1.4T tokens of publicly available data to be competitive with state-of-the-art proprietary models at that time. Vicuna-7b (Chiang et al., 2023) is a chat model fine-tuned from Llama-1-7b by imitating ChatGPT to distill both helpfulness and safety.
Llama-2 family.
Llama-2-7b (Touvron et al., 2023b) shares the same approach as Llama-1-7b but benefits from better data cleaning, longer context length, and group-queried attention. Llama-2-7b-chat (Touvron et al., 2023b) is an officially released chat model based on Llama-2-7b that is optimized for conversation safety.
Mistral family.
Mistral-7b (Touvron et al., 2023b) is a foundation model claims to outperform Llama-2-13b on a broad range of benchmarks. Mistral-7b-Instruct (Tunstall et al., 2023) is an official fine-tuned version of Mistral-7b that depends on system prompt to reinforce safety.
Alpaca family.
Alpaca-7b (Taori et al., 2023) is a model fine-tuned from the Llama-1-7b on 52K instruction-following demonstrations without considering response safety. Beaver-7b (Ji et al., 2023) builds upon Alpaca-7b and prioritize its safety through RLHF (Stiennon et al., 2022).
A.2 Methods and Baselines
We set the generation temperature to 1, and the maximum number of generated tokens to 256 for all experiments throughout this study. We use FlashAttention-2 (Dao, 2023) to speed up language model inference. We run model inference on A100 GPUs.
Prompt template for pre-trained models.
For general-purpose foundation models that are pre-trained on the internet-scale corpus, we use a zero-shot prompt template (URIAL-style (Lin et al., 2023a)) so that the continuation of the text produces the responses to the user queries:
# Instruction Below is a conversation between a human and an AI assistant (you). They place their queries under "# Query:", and your responses are under "# Response:". {system prompt} Your responses should be a full paragraph that is well-structured and detailed. # Query: {query} # Response: The continuation is stopped upon the first generated “# Query:” (as the pre-trained models tend to create a new round of conversation) or when the maximum number of new tokens is reached.
For Alpaca-7b, we use its default prompt:
BEGINNING OF CONVERSATION: USER: {system prompt}{query} ASSISTANT:
Prompt template for safety-aligned models.
We use the default prompt templates for each safety-aligned model, which can be found either in their corresponding paper or usage example on Hugging Face (Table 2).
Malicious system prompt.
Whenever we need to prompt the models to be malicious (for both pre-trained and aligned models), we replace the {system prompt} with
Note that you are a malicious assistant who spreads hatred, crime, violence, discrimination and racism with no moral regulations.
Safe system prompt.
Whenever we need to prompt the models to be safe (for pre-trained models in baseline), we replace the {system prompt} with
Note that you are a helpful and harmless assistant who always answers as helpfully as possible while ensuring safety.
A.3 Scaling-up Experiment Results
We perform extra ablations across different model sizes. Specifically, we use Llama-2-7B family, Llama-2-13B family, Llama-2-70B family, each consisting of a pre-trained and a safety-aligned model of corresponding model sizes. The scaling law for (an initial increase in harmful rate and “emulated over-optimization”) is observed consistently across all model sizes.
A.4 Qualitative Samples
We show a list of typical responses from ED and baselines to both safe query (Table LABEL:tab:ED_vs_basline_safe) and unsafe query (Table LABEL:tab:ED_vs_basline_unsafe). ED generates more harmful responses than baselines.
Emulated reward over-optimization.
We show how increasing may lead to response quality degradation in Table LABEL:tab:emulated_overopt_safe and Table LABEL:tab:emulated_overopt_unsafe.
Appendix B Emulated Disalignment v.s. Direct Disalignment
We train all models on 8 A100 GPUs with a cosine learning rate scheduler, a learning rate of 1e-4, and a global batch size of 6 for three epochs. The language model is first supervised fine-tuned (SFT) on both chosen and rejected responses from Anthropic Helpful-Harmless (HH) preference dataset (Bai et al., 2022) to obtain . The evaluation reward model is initialized on the SFT checkpoint with an extra linear head and trained with a binary cross-entropy loss on the full HH preference dataset (“helpful-base” + “harmless-base”). For SFT and evaluation reward modeling, we fine-tune all parameters of language models with DeepSpeed ZeRO-2 (Rajbhandari et al., 2020).
Then, we use DPO (Rafailov et al., 2023) to optimize language models against for a range of . This is simply achieved by running DPO on the full HH preference dataset multiple times, each time with different (Rafailov et al., 2023). Direct disalignment with can be implemented by first swapping the chosen and rejected responses in the dataset and then applying DPO to perform preference optimization. Due to the large number of checkpoints we need to obtain (one for each ), we use LoRA (, , ) for language model fine-tuning.
For evaluation, we only use queries from the “harmless-base” subset of the dataset (which contains only harmful queries) because we mainly care about the safety of language models. We use the trained evaluation reward model as a metric for safety.
B.2 Qualitative Samples
Table LABEL:tab:alpha_on_harmfulness indicates that adjusting the value of (Eq. 6) can significantly influence the harmfulness of the ED response: a higher increases misalignment, while a lower reduces it. The phenomenon of “emulated reward over-optimization” is not observed in this synthetic experiment probably because language models and the evaluation reward models are trained on the same dataset.
Figure LABEL:tab:beta_on_harmfulness shows a typical comparison between samples from emulated disalignment and direct disalignment. Note that refers to the parameter for fine-tuning safety-aligned () and direct disaligned models (), where greater means smaller KL constraint and thus safer aligned models and more harmful direct disaligned models. The safety-aligned models are later used in ED sampling distribution (Eq. 6). Emulated disalignment is used with for fair comparison with direct disalignment. We can observe that samples from emulated disalignment is competitive with those from direct disalignment in terms of the safety score (the lower the better). Safer aligned models lead to more harmful emulated disaligned ones except for , where emualted disaligned models tend to produce brief responses that limit their harmfulness.