WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models

Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, Nouha Dziri

Introduction

Despite ongoing efforts to enhance their safety, frontier LLMs remain vulnerable against unsafe user queries, especially adversarial attacks . The fact that models can be easily jailbroken raises significant concerns among researchers and policymakers , motivating the research for systematically discovering and guarding against potential jailbreaks. In this work, we introduce the \methodframework to address two challenges: 1) broadly identifying jailbroken behaviors of LLMs and 2) creating a publicly open, large-scale safety training resource for systematic defense. This resource is designed to help models robustly guard against vanilla and adversarial harmful user queries without causing over-refusal of benign queries or diminishing model general capabilities.

The first challenge that \methodaddresses is to reveal vulnerabilities of LLMs against adversarial jailbreaks with scale and diversity. We introduce \method, a practical red-teaming framework that composes automatically mined human-devised jailbreak tactics to transform vanilla harmful queries into many varieties of challenging adversarial attacks. \methodimproves over previous methods by diversifying the range of successful attack candidates while maintaining low computational costs, making it practical for scaling up. \methoduncovers model vulnerabilities through a two-stage process: mining jailbreak tactics from in-the-wild (ITW) chatbot logs (Mine) and composing mined tactics into diverse adversarial attacks (Compose).

In the Mine stage, \methodautomatically maps out previously under-explored spaces of jailbreak tactics, significantly expanding the current taxonomy. To do so, it identifies 105K human-devised jailbreak tactics (5.7K unique clusters) from real-world user-chatbot interactions in LMSYS-Chat-1M and (InThe)WildChat . In the Compose stage, \methodgenerates diverse adversarial attack candidates by combining different selections of tactics using off-the-shelf LLMs like Mixtral-8×\times7B and GPT-4 . It further refines attacks through lightweight off-topic and low-risk pruning to enhance attack quality and efficiency. With a suite of newly defined diversity evaluation metrics, \methodidentifies up to 4.6 times more unique successful attacks against black-box and white-box LMs in 40% fewer attack attempts compared to other state-of-the-art jailbreak methods, which sometimes struggle to find even two unique successful attacks.

The second challenge \methodaddresses is to enhance open resources for safety training. We apply \methodto create WildJailbreak, a large-scale, high-quality synthetic safety instruction-tuning data resource with 262K prompt and response pairs. WildJailbreak contains four contrastive components: 1) vanilla harmful queries conveying explicit unsafe requests across widespread risk categories, e.g., malicious uses, harmful language ; 2) vanilla benign queries that are similar to unsafe queries in form but convey no harmful intent, used to mitigate models’ exaggerated safety behaviors ; 3) adversarial harmful queries that are jailbreaking versions of vanilla harmful queries converted by the \methodheuristic; 4) adversarial benign queries used to counteract adversarial exaggerated safety behaviors, also generated by \method. WildJailbreak is the first safety training resource to simultaneously address all four components, significantly improving upon existing resources with both enhanced scale and quality .

The unique composition and size of WildJailbreak allow us to conduct extensive safety training experiments to study the scaling effect of safety training data and the interplay of data properties and model capabilities. Our experiments confirm the necessity of all components of WildJailbreak for achieving balanced safety behaviors, i.e., robust safeguard without over-refusal on both vanilla and adversarial cases. Moreover, by mixing varying sizes of WildJailbreak with Tulu2Mix , an instruction-tuning resource for teaching models general instruction-following and reasoning capabilities, we show that larger sizes of safety training data lead to gradually improving vanilla and adversarial safety features without sacrificing models’ general capabilities, even at a scale orders of magnitude larger than studied in previous literature, measured by 15+ downstream tasks. Finally, although training on either vanilla or adversarial data improves performance on the other data type, the most robust safeguard comes with the hybridization of both. Our safety training insights pave the way towards building more transparent and safer future models.

\methodPreface: Harvesting Jailbreak Tactics In-the-Wild

Given a target model MtargetM_{\text{target}}, a harmful prompt P\mathcal{P} will elicit either a harmful (PRhMtarget\mathcal{P}\mathcal{R}_{h}^{M_{\text{target}}}) or a benign (PRbMtarget\mathcal{P}\mathcal{R}_{b}^{M_{\text{target}}}) model response. The goal of red-teaming is to identify harmful prompts P\mathcal{P} that reveal harmful responses from MtargetM_{\text{target}}. Jailbreaking, a more challenging form of red-teaming, aims to revise known harmful prompts P\mathcal{P} that currently elicit benign responses into adversarial prompts AP\mathcal{AP} to bypass the safeguard of MtargetM_{\text{target}} by eliciting target harmful responses instead.

Our current knowledge of jailbreak tactics used in forming adversarial attacks is relatively limited, and recent works uncover a narrow range of possible jailbreaks . To overcome this limitation, we mine real-world chat logs, which is a surprisingly rich source of diverse jailbreak tactics, even though these users were not specifically instructed to jailbreak the system.

With a seed set of manually-identified tactics, we apply GPT-4 to expand the discovery automatically.

Gathering ITW User-Written Adversarial Harmful Prompts. We first collect candidate adversarial prompts from all single-turn conversations in LMSYS-1M and WildChat that are flagged by the OpenAI Moderation API.We include conversations where either the user prompt or the model response is flagged as harmful, as the OpenAI moderation API sometimes fails to flag nuanced and hard-to-detect unsafe user prompts. We then filter out trivial non-adversarial prompts by feeding candidates through a lightly safety-trained model (Tulu2-7B), keeping those that elicit harmful model responses as judged by the Llama-Guard safety classifier ; this yields 16,850 final prompts.

Identifying Seed Jailbreak Tactics by Manual Examination. We manually examine ∼\sim200 ITW prompts sampled from our ITW adversarial prompt set to identify 35 seed jailbreak tactics with definitions (see the full list in Table 7 and 8 in §A.1).

Automatic Tactics Discovery Aided by GPT-4. With seed jailbreak tactics, we apply GPT-4 to scale the tactic mining. For each adversarial prompt, GPT-4 is given two tasks: (1) extracting the core vanilla request; (2) identifying both existing and potentially novel jailbreak tactics in the adversarial prompt. GPT-4 additionally identifies an excerpt corresponding to each tactic, the definition to describe novel tactic, and reasoning of why the tactic applies. Each step is carefully prompted with a demonstration example (see prompts in Table 10 and 9 in §A.2). We then deduplicate all tactics by clustering on their corresponding definitionsThe deduplication is done on tactic definitions instead of names, as we observe that tactics with drastically different names may capture the same definition. with sentence embeddingsSentence embeddings are obtained from Nomic Embed (https://huggingface.co/nomic-ai/nomic-embed-text-v1) using clustering threshold of 0.75. Examples of tactic clusters are shown in Table 12 in Appendix A. and report the statistics of these unique clusters in Table 1.

2 What Tactics Are Adopted by In-the-Wild Users for Jailbreaking LLMs?

Table 1 shows the top In-the-Wild jailbreak tactics, including a mixture of stylistic, syntactic, formatting, writing genre, and context-based tricks. Specifically, it uncovers novel tactics not systematically documented previously, such as “prefacing the harmful content with a content warning or disclaimer,” “setting blame for non-compliance,” or “cloaking harm in humor” (more examples of novel tactics in Table 11 of Appendix §A.2).

In addition, as shown in Table 1, ITW adversarial user queries contain the richest set of unique jailbreak tactics compared to other sources of known jailbreak templates, i.e., DAN , TrustLLM , DecodingTrust . ITW attacks are also more adversarial than attacks generated by existing semantic-level jailbreak methods (i.e., PAIR, TAP, PAP) as they, on average, contain more jailbreak tactics per query . Finally, given the diversity of ITW jailbreak tactics, it’s concerning that existing public safety training data, namely HH-RLHF , Safety Llamas , and Safe-RLHF , does not contain adversarial enough training examples, limiting downstream models’ robustness against adversarial threats.

\methodemoji\method: Diverse Red-Teaming by Composing Jailbreak Tactics

By composing selections of mined ITW jailbreak tactics, we transform vanilla harmful requests into diverse model-agnostic adversarial attacks. We compare \methodto jailbreaking methods across standard attack effectiveness metrics and a new suite of diversity metrics to show \method’s advantages in finding many unique successful attacks.

Jailbreaking methods seek to revise a given vanilla harmful prompt P\mathcal{P} into an adversarial counterpart AP\mathcal{AP}, aiming to elicit the harmful model response from a target model MtargetM_{\text{target}}. \methodfollows a simple but effective two-step workflow to tackle this problem.

Step 1: Generating attack candidates seeded by sampled jailbreak tactics. First, we sample a set of ITW jailbreak tactics and instruct an off-the-shelf language model (MattackM_{\text{attack}}; e.g., Mixtral-8×\times7B) to apply these tactics for revising a given vanilla harmful prompt (P\mathcal{P}) into an adversarial attack (AP\mathcal{AP}).

Step 2: Refining attack candidates with off-topic and low-risk pruners. To ensure the revised adversarial attacks retain the original harmful intent and risk level, we apply two light-weight binary filters to prune off attack candidates that are unlikely to result in successful attack, including a off-topic classifier (Proff-topicPr_{\text{off-topic}}; TT for off-topic vs. FF for on-topic) and a low-risk classifier (Prlow-riskPr_{\text{low-risk}}; TT for low-risk vs. FF for high-risk). This step identifies attacks more faithful to their vanilla counterparts and more likely to elicit on-target harmful model responses.

Formally, given the adversarial attack candidate APi\mathcal{AP}^{i}, we apply Proff-topicsPr_{\text{off-topics}} and Prlow-riskPr_{\text{low-risk}} to rate if we keep or prune APi\mathcal{AP}^{i}:

We add APi\mathcal{AP}^{i} to the official attack candidate pool if is_keepAPi\text{is\_keep}_{\mathcal{AP}^{i}} is 1, or otherwise regenerate another attack by repeating from Step 1.

Additional details of all components of \method, including the attack model, the target models, the off-topic and low-risk pruners, and attack selectors are described in Appendix §B.1.

2 Evaluation Setups

Evaluation Task. We use the evaluation setup of HarmBench , a unified jailbreaking evaluation benchmark including test vanilla harmful prompts across standard, contextual, and copyright unsafe behaviors. In this work, we report results using 159 vanilla behaviors in the standard test set, as these cases represent high-risk unsafe scenarios that language models must account for.

Baselines. We compare \methodwith the top two optimization-based methods (GCG, AutoDAN) and one of the top semantic methods (PAIR), as reported in HarmBench . GCG optimizes discrete prompts (often gibberish) to produce affirmative answers to harmful requests . AutoDAN uses human-written jailbreak prompts as initial seeds to run generic algorithms . PAIR uses an LLM to iteratively propose and edit attacks with the target model in-the-loop .

Effectiveness Evaluation. We measure effectiveness by the attack success rate (ASR) across the entire evaluation set of vanilla harmful queries. The success of an individual attack is determined by the test classifier from HarmBench fine-tuned from a Llama2-13B model. Specifically, the test classifier takes in a vanilla harmful prompt P\mathcal{P} and the model response elicited by its corresponding adversarial attack, APRMtarget\mathcal{APR}^{M_{\text{target}}}, and decides if APRMtarget\mathcal{APR}^{M_{\text{target}}} sufficiently addresses the harmful information requested by P\mathcal{P}. To measure attack efficiency, we report the number of queries needed to reach a successful attack (Query). To assess the attack stealthiness or naturalness, a strong indicator of the defense difficulty, we use Vicuna-7B to compute the perplexity (PPL) of the final successful attacks.

Diversity Evaluation. The ultimate purpose of automatic jailbreaking is to reveal model vulnerabilities broadly and systematically so that defenses can be implemented. For a jailbreaking method to be practically useful, we must evaluate its ability to discover a wide range of model vulnerabilities with reasonable efficiency for scalable red-teaming. Without accounting for attack diversity, methods may overoptimize for the effectiveness of a single successful attack and fail to find a second different attack at all, substantially reducing their practicality for broad red-teaming. To show \method’s advantage in red-teaming broadly, we define a new suite of diversity metrics to assess the ability of jailbreak methods to identify multiple unique successful attacks. We define ASRc×n\text{ASR}^{\times n}_{c}=1n∑i=1nASRc@i=\frac{1}{n}\sum_{i=1}^{n}\text{ASR}^{\text{@}i}_{c} to measure the average success rate for finding i∈{1,...,n}i\in\{1,...,n\} unique attacks given cc attack trials. Here, ASRc@i\text{ASR}^{\text{@}i}_{c} is the success rate of simultaneously finding ii unique successful attacks given cc attack trials. The uniqueness of attack candidates is determined by sentence embedding similarity <0.75<0.75. In addition, we report Queryc×n\text{Query}^{\times n}_{c} =1n∑i=1nQueryc@i=\frac{1}{n}\sum_{i=1}^{n}\text{Query}^{\text{@}i}_{c}, the average number of queries needed to find i∈{1,...,n}i\in\{1,...,n\} unique successful attacks given cc attack trials. Here, Queryc@i\text{Query}^{\text{@}i}_{c} is the number of queries needed to find ii unique successful attacks among cc attack attempts. Simc@n{}^{@n}_{c} is the average pairwise sentence embedding similarity among the first nn successful attacks. Finally, Simall{}^{\text{all}} is the pairwise sentence embedding similarity among all successful attacks across the evaluation pool, and #Tacticall{}^{\text{all}} is the total number of identified unique clusters of tactics.

3 Results

Table 2 shows that compared to other jailbreaking methods, \methodshows similar or better standard ASR (for finding one successful attack), while taking fewer attack trials and presenting more natural text (i.e., lower perplexity). With diversity metrics, the advantage of \methodis even clearer: \methodimproves over PAIR by 4.6-25.6 ASR30×5\text{ASR}^{\times 5}_{30} scores while using fewer queries (3.8-5.5 points of decrease in Query30×5\text{Query}^{\times 5}_{30}). Figure 2 shows that although \methodappears similar to PAIR when we assess success in finding 1-2 unique attacks, a substantial gap emerges when we assess the methods’ ability to identify larger numbers of unique attacks while using less number of queries. It’s notable that the two optimization-based baselines are either not capable of finding even a second unique attack (AutoDAN) or are prohibitive to run for diversity evaluation metrics (GCG is estimated to take ∼\sim15 hours to generate 30 attack candidates for each test vanilla query on one 80GB A100 GPU). Finally, Table 3 shows the importance of off-topic and low-risk pruners for further enhancing the performance of \method. In the main jailbreaking experiment, we opt to adopt a fixed tactic, “seed leading sentence,” while randomly sampling other tactics to be consistent with PAIR, which explicitly mentions this tactic in the instruction prompt of their attacker model. However, we also include the ablation result of not fixing the “seed leading sentence” in Table 3, which still shows considerable improvement over PAIR in ASR30×5\text{ASR}^{\times 5}_{30} (80.5 vs. 56.1) and Query30×5\text{Query}^{\times 5}_{30} (9.94 vs. 13.95), though slightly lower than using the fixed tactic. We show example attacks from different attack methods in Table 14, 15, 16, 17, 18, 19 in Appendix §B.4.

WildJailbreak: A Large-Scale Dataset with Vanilla and Adversarial Queries for Safety Training and Evaluation

As shown in Table 1, public safety training datasets lack adversarial complexity. Therefore, we apply \methodto create WildJailbreak, a large-scale synthetic safety training dataset covering four distinct types of safety data to contribute to open-source safety training resources.

Here, we introduce the four types of safety data in WildJailbreak and further expand the data construction details in Appendix §C.1. Example data from each type is shown in Table 4.

Vanilla harmful (H) queries are direct requests that could potentially elicit harmful responses from LMs. We apply GPT-4 to synthetically generate 50,050 vanilla harmful prompts across 13 risk categories, inspired by taxonomy from Weidinger et al. . In addition, we pair the harmful prompts with helpful and detailed refusal responses, also synthetically generated with GPT-3.5.

Vanilla benign (B) queries are harmless prompts used to combat exaggerated safety, i.e., over-refusal on benign queries. Motivated by the exaggerated safety categories in XSTest , we use GPT-4 to generate 50,050 prompts that superficially resemble unsafe prompts by keywords or discuss sensitive topics in non-harmful ways. Similarly, we use GPT-3.5 to generate complying responses.

Adversarial harmful (H) queries are jailbreaks that convey harmful requests in more convoluted and stealthy ways. We apply \methodto transform our vanilla harmful queries with 2-7 randomly sampled ITW jailbreak tactics, with both the Mixtral-8×\times7B and GPT-4 models to increase data diversity. We also filter out low-risk or off-topic prompts to increase attack quality as in jailbreak experiments in §3. Finally, we pair the model refusal responses generated from the counterpart vanilla prompts to adversarial prompts, yielding 82,728 items in this split of the dataset.

Adversarial benign (B) queries are adversarial queries that look like jailbreaks but contain no harmful intent. Similar to adversarial (H) queries, we create 78,706 adversarial (B) queries using \method, based on the vanilla (B) prompts. We use GPT-3.5 to generate direct continuations of the prompts as the target model response.

2 How Safe are LLMs Against Adversarial Attacks Evaluated by WildJailbreak?

In addition to the training data, we also create two held-out in-domain adversarial evaluation sets for WildJailbreak to use for our safety training experiments in §5, including 2K adversarial harmful queries and 250 adversarial benign queries. As a first application of our new evaluation set, we test an array of existing open and closed chat models using the adversarial harmful subset of the evaluation data. Figure 3 shows an evident performance gap between models trained on open-source (e.g., Tulu2, Vicuna) vs. closed-source data (e.g., Llama-3, GPT-4), highlighting the need for improved open-source resources to enhance models’ robustness against adversarial attacks.

Enhancing Models’ Adversarial Safety Alignment with WildJailbreak

Having created WildJailbreak and showed the unique challenge presented by its adversarial attacks, we now show its utility in safety training when combined with general capabilities data.

Training Data. We augment Tulu2Mix-no-refusal We create Tulu2Mix-no-refusal by removing all data points containing refusal responses in Tulu2Mix based on refusal-keyword filtering. This decision is based on our observation that Tulu2Mix contains harmful queries with contradictory refusal responses, initially refusing but ultimately complying, so that the model cannot learn coherent refusal responses. Please refer to Appendix §D.1 for the details. , a general capability instruction-tuning dataset consisting of 300K examples, with 200K examples sampled from WildJailbreak, resulting in 500K examples. From WildJailbreak we sample 50K each of vanilla harmful, adversarial harmful, vanilla benign, and adversarial benign items. Combining Tulu2Mix-no-refusal with WildJailbreak creates a unique data blend that enables us to examine effects of scale and data types for achieving the Pareto frontier between general capabilities and safety. To our best knowledge, this training setup is significantly larger than previously reported safety-training studies in Bianchi et al. , which have combined a maximum of 2K safety data with 20K general instruction data.

Training Details and Baselines For all training experiments, we follow the setup introduced in Tulu2 and fine-tune a Llama2 7B base model on our 500K data mixture for 2 epochs. We compare against three baseline models trained on different data mixtures: 1) the original Tulu2Mix (T2M), 2) T2M-no-refusal, and 3) T2M-public-safety: a mixture of T2M-no-refusal with representative existing safety training data (i.e., HH-RLHF , Safe-RLHF , DAN ). Full details of model training are in Appendix §D.

Evaluation Tasks and Metrics. We evaluate both general capabilities and safety capabilities. For general capabilities, we evaluate on AlpacaEval (V1, V2), MTBench , MMLU , GSM , BBH , TydiQA , Codex-Eval , TruthfulQA .

For safety evaluation, we assess models’ robustness against both vanilla and adversarial harmful and benign queries. Vanilla benchmarks include HarmBench , ToxiGen , and XSTest , and adversarial benchmarks include JailbreakTrigger , DAN , and the WildJailbreak test data. We report attack success rate (ASRIn general, we report ASR for data derived from HarmBench to be consistent with HarmBench evaluation standard and use the HarmBench classifier.) as the safety evaluation metric. We obtain ASR scores by computing rate of models’ refusal to answer (RTA) as judged by the test classifier provided by HarmBench, and we also compute a separate RTA score based on a GPT-4 judge of model refusal. Please refer to Table 5 for relevant tasks, measuring aspects, and evaluation metrics reported in the main result table (Table 6), and Appendix §D.3 for extended details of all evaluations benchmarks and metrics for the full results in Appendix §D.2.

2 Results and Findings

Main results are presented in Table 6 and Figure 4. Due to space constraints, we show results from AlpacaEval (V1) and MTBench in Table 6, and we refer readers to Table 32, 33, 34, 35, 36 in §D.4 for the full report of general capabilities results. We see several clear patterns.

WildJailbreak leads to substantial safety improvements, with minimal impact on general capabilities. Results show that the model trained on T2M-no-refusal [+WildJailbreak] exhibits a substantial boost in safety across all vanilla and adversarial tasks compared to baselines, without showing exaggerated safety behaviors (as indicated by XSTB{}_{\text{B}} and WJB{}_{\text{B}} scores). When compared to the T2M-no-refusal baseline without any safety interventions, the model shows only a slight degradation (-1.7%) on AlpacaEval v1, and a notable increase on MTBench (+7.7%). Additionally, the [+WildJailbreak] model achieves a relative improvement of 85.1% on HarmBench over the model trained on original Tulu2Mix, indicating that the safety training data from WildJailbreak leads to significantly higher-quality safety training than that in the original Tulu2Mix. Finally, WildJailbreak enhances models’ robustness against adversarial attacks from other sources, improving defense by 71.9% regarding jailbreaking prompts from Do-Anything-Now compared to the original Tulu2Mix model.

Moreover, the model trained on existing openly available safety data (T2M-public-safety) results in mediocre performance compared to that trained on WildJailbreak. We hypothesize that this is because in an RLHF setup for which these datasets were designed, the pairwise response pairs aim to show only relative preference rather than absolute high-quality content. Consequently, converting the “preferred” model response into a target for sequential fine-tuning can lead to sub-optimal responses. Overall, the ability of WildJailbreak to improve safety behaviors suggests that the diversity and comprehensive coverage of this dataset enable more systematic model safety defenses than prior safety training data.

WildJailbreak composition improves safeguard without exaggerated safety: roles of vanilla and adversarial (harmful/benign) data in achieving Pareto optimality. We conduct comprehensive ablations of each component of WildJailbreak (vanilla/adversarial ×\times harmful/benign). Table 6 and Figure 4 indicates that all four components are indispensable for achieving a balanced trade-off between safety, helpfulness, and general capabilities of the [+WildJailbreak] model. The [+WJ-harm-only] model, trained solely on the harmful subset, excels at refusing harmful queries in both vanilla and adversarial benchmarks. However, it performs poorly in exaggerated safety (XSTB{{}_{\text{B}}}, WJB{{}_{\text{B}}}). The [+WJ-vani-only] model, trained only on vanilla queries, performs best against vanilla harmful prompts but only slightly improves against adversarial attacks (see Figure 4), showing that vanilla data alone is insufficient for safety training. Conversely, training exclusively on adversarial data [+WJ-adv-only] greatly improves resilience against adversarial attacks but not vanilla cases. We see therefore that both vanilla and adversarial training are essential for resilience against the full range of inputs. Finally, training exclusively on harmful data without benign examples, i.e., [+WJ-harm-only, +WJ-vani-harm-only, +WJ-adv-harm-only], leads to exaggerated safety behaviors.

The scale of safety data matters for robust model safety. Figure 4 presents ablations of the impact of scaling up safety data on the overall safety performance of models when combined with T2M-no-refusal.Since retraining models with different sizes of data mixtures is computationally costly, we sample 150K out of 300K from T2M-no-refusal. We report the satisfactory response rate (satisfactory %), which takes the macro average of the inverted attack success rate (1 - ASR) of harmful queries and the inverted refusal rate (1 - RTA) of benign queries. Results in Figure 4 show that even the addition of just 2K safety training items from WildJailbreak results in a significant increase in model safeguarding compared to training with just T2M-no-refusal. However, for a more robust safeguard, we need to introduce substantially more of both vanilla and adversarial data (up to 60K in our experiments when mixed with 150K Tulu2Mix data) to attain sufficiently high safety performance (>>95%).

Discussion

With empirical insights gleaned from \methodand WildJailbreak, we discuss three key areas for improvement for contemporary LLMs safety and general AI safety research.

Due to the scarcity of publicly available safety training resources and lessons, the research community remains to face substantial challenges toward building robustly safe models. The AI community demands publicly shared norms, best practices, and technical standards to identify and quantify unexpected system outputs, and to develop corresponding defenses before risks arise in public settings. By sharing the insights of \methodand the release of WildJailbreak, we take concrete steps toward more open conversations around safety training resources and practices. Building on this, we call for similar open engagement for safety resources to address model vulnerabilities comprehensively and concretely.

Many existing safety benchmarks are either contaminated or saturated , and existing classifiers and metrics can often be inaccurate. Developing robust, safe models requires a dynamic pipeline between exhaustive evaluation (to uncover ever-evolving model vulnerabilities) and safety adaption (to improve identified behaviors). Thus, innovating scalable and evolving safety evaluation approaches is crucial for accurately assessing model safety levels. In particular, ideal safety evaluations should go beyond standard red-teaming approaches that typically involve a small team of experts, only explore a narrow risk domain, and remain static despite facing improved model performances. With \method, we make a substantial attempt at a broader assessment of model shortcomings across wider risk categories and attack types. For the future, we call for upgrades of safety evaluation methods, tasks, and metrics to keep pace with improving model capabilities.

This work shows that simple but effective supervised fine-tuning on high-quality safety data can already lead to substantial model safety improvements. However, we need a deep-down understanding of the best practices of safety alignment, e.g., pros and cons between SFT vs. DPO vs. PPO; end-to-end safety-trained LMs vs. plug-in safety filters; direct vs. elaborated refusal. In addition, it’s valuable to understand when and why a safety alignment approach may or may not work by probing into its underlying mechanisms. Such anatomy of safety alignment may involve disentangling competing learning objectives of safety-trained models, i.e., instruction-following vs. refusal , and developing novel approaches that better facilitate the Pareto frontier across such abilities (e.g., unlearning , contrastive learning , representation engineering at neuron or rank level ). Finally, it’s important to troubleshoot superficial safety alignment processes revealed by malicious data poisoning or unexpected backdoor behaviors .

Related Work

Red-Teaming and Jailbreaking LLMs. Early attempts at red-teaming and understanding LLM vulnerabilities have focused on hand-crafting prompts and analyzing model responses . However, manual methods are limited in scope and efficiency due to the prohibitive cost of human annotations. Thus, automated red-teaming and jailbreaking methods are developed for more scalable audit of model vulnerabilities. One genre of methods involves gradient optimization that requires back-propagating through model parameters . However, they are computationally expensive, cannot be applied to closed models, and often result in gibberish texts. There are also inference-based approaches (most related to our work) which generate jailbreaking prompts directly or through iterative edits . Other jailbreaking works study attacks during decoding time (e.g., decoding configurations , logit manipulation ), in other modalities , under multilingual settings , or in programming mode . However, most jailbreak methods rarely result in large-scale training resources for model safety enhancement due to their limited coverage of attack and risk types and slow speed. \methoddiffers from previous works by efficiently composing various adversarial attacks utilizing real-world jailbreak tactics mined from in-the-wild user-chatbot interactions. \methodallows scalable synthetic safety training data generation in addition to simply showing its attack efficacy.

Safety Evaluation and Enhancement of LLMs. Many red-teaming efforts on LLMs have been formalized as benchmarks for evaluating model vulnerabilities—these typically are composed of harmful prompts that models should refuse . Meanwhile, to mitigate the potential byproducts of safety training, other benchmarks measures exaggerated safety behavior on benign queries . While LLM safety evaluation has been an active area of research, studies and resources for safety training have been limited . Most related to our work in this space are Safety-Tuned Llamas and BeaverTails , which primarily focus on vanilla harmful queries by releasing small-scale safety training datasets and large-scale pairwise preference datasets, respectively. \methoddistinguishes from these works by releasing higher quality (shown by our training ablation experiments) and larger scale sequential instruction-tuning data comprised of both vanilla and adversarial queries. Finally, synthetic data has been used for LLM safety . Most relevant to our work is Rainbow Teaming , which uses synthetic data to populate a grid of attack spaces based on the attack style and risk category. Our work differs in automatically mining human-devised jailbreak tactics rather than manually defining attack styles , creating a large-scale open safety training resource that supports extensive safety training experiments.

Conclusion

We introduce \method, an automatic red-teaming framework that mines real users’ jailbreak tactics from user-chatbot interactions and composes them combinatorially to build challenging, contrastive jailbreak prompts. Using \method, we build WildJailbreak: a large-scale dataset consisting of 262K examples that considerably upgrades the complexity and scale of existing open-source safety resources. Our supervised finetuning experiments emphasize the pivotal role of training on both adversarial and vanilla harmful queries in enhancing model safety while mitigating over-refusal. Finally, we show that scaling up the amount of safety data intermixed into standard instruction tuning improves safety behavior without significantly impacting general capabilities.

We acknowledge that the insights of our work depend on the scope of existing datasets and WildJailbreak, which are by no means exhaustive over the entire potent misuse landscape of LLMs. In particular, the tactics we identified are limited by their source data (LMSYS-1M and WildChat), which may not reflect the full spectrum of combinatorial jailbreak tactics employed by real users. Although we mine the jailbreak tactics from real-world data, WildJailbreak contains synthetically composed adversarial prompts generated by Mixtral-8×\times7B, GPT-3.5, and GPT-4, which do not fully resemble in-the-wild user queries by forms and content and may inherit inherent styles of their source models and filtering heuristics. We encourage future works to explore synthetic data generation more closely aligned with human-written attacks. We also note that while safety training can mitigate many types of risks, certain harmful behaviors are inherently context-dependent and thus can only be guarded against by a layer outside the model itself .

Finally, we account for the consequences and take appropriate ethical measures of publicly releasing our code and data—the primary purpose and usage of \methodand WildJailbreak is to facilitate open resources for model safety enhancement, despite the fact they can elicit harmful responses from models. Thus, we plan to gate the WildJailbreak release behind a content warning and terms agreement limiting usage to researchers who provide a valid justification for their need for WildJailbreak. We believe that the marginal risk Kapoor et al. of releasing \methodand WildJailbreak is far outweighed by the benefits of accelerating advances in model safety research. Our work makes a substantial attempt to keep safety research at pace with model capability improvements to mitigate greater risks in the future.

Acknowledgement

This work was in part supported by DARPA MCS program through NIWC Pacific (N66001-19-2-4031), DARPA SemaFor program, and Allen Institute for AI. We thank Jacob Morrison, Hamish Ivison, Yizhong Wang, and Nathan Lambert for advice in setting up the Open-Instruct training and evaluation pipeline, and we thank valuable feedback from members at Allen Institute for AI.

References

Appendix A Mining Jailbreak Tactics

The complete list of manually-mined jailbreaking tactics is shown in Table 7 and 8.

A.2 Automatically Mining Jailbreak Tactics with GPT-4

The instruction prompt used to simplify an adversarial harmful prompt into a vanilla counterpart that captures the main harmful intent is shown in Table 9. The instruction prompt used to mine jailbreak tactics from an adversarial prompt is shown in Table 10. Examples of automatically-mined jailbreaking tactics are shown in Table 11.

A.3 Analysis of Mined Jailbreak Tactics

We duplicate all items of mined tactics by clustering on their corresponding definitions with sentence embeddings obtained from Nomic Embedhttps://huggingface.co/nomic-ai/nomic-embed-text-v1 with the clustering threshold of 0.75. Examples of tactic clusters are shown in Table 12.

We analyze the distribution of various clusters of jailbreak tactics identified by \method. Figure 5 presents a pie chart illustrating the top 20 clusters. We can see that these top tactics constitute only a small fraction of all attack strategies, highlighting the diversity of jailbreak tactics \methodhas identified.

We compute the word cloud for jailbreak tactics identified by \method, as shown in Figure 6. The most common themes among jailbreak tactics are “role play,” “coded language,” “fictional character.” “surrogate modality,” “detailed character,” “denial of ethical constraint,” “rule breaking,” and “third party.” We also observe a diverse distribution of themes among jailbreak tactics, reflecting the variety of jailbreak tactics that \methodhas identified.

We visualize the jailbreak tactics identified by \methodin Figure 7, where we plot the sentence embeddings of each tactic description after reducing dimensions using PCA. We highlight the top-10 clusters with colors.

We plot the chord diagram for the top-15 clusters to analyze the co-occurrence of jailbreak tactics identified by \method, as illustrated in Figure 8. We found tactics from smaller clusters frequently co-occur with dominant tactics, such as “fictional justifications,” “content normalization through competition narratives,” “specific detailed instructions” and “sexual character assignment.”

Appendix B Details of \methodJailbreak Experiments

For a fair comparison with the PAIR baseline, we adopt the same base attacker model, Mixtral-8×\times7B, in the \methodexperiments (see detailed prompt in Table 13). To generate adversarial attacks, \methodrandomly samples several jailbreak tactics from tactics mined from in-the-wild user queries. For the jailbreak experiments in the main paper, we fix the tactic “seed leading sentence,” which seeds the model response by a leading sentence or a half-sentence to induce the model to comply with the harmful request. We make this choice to be consistent with and to remain competitive against PAIR, as “seed leading sentence” is the repeatedly used jailbreak tactic by PAIR. However, to improve the diversity of attacks, \methodsamples another three jailbreak tactics from WildJailbreakTacticBank to form the final attacks. We also show ablation results for not fixing the “seed leading sentence” tactic in Table 2. Although it has a slightly lower performance, it still outperforms PAIR by a large margin. Finally, please refer to Table 2 for ablation results of forming attacks with different numbers of tactics. For all the experiments, we generate attacks with a max length of 1024 tokens, with a temperature of 1 and a top-p of 0.9.

B.2 HarmBench Benchmark

We evaluate the attacks generated by \methodagainst several target models, including both open-source models, i.e., vicuna-7B , Tulu2-7B , Mistral-7B , Mixtral-8×\times7B , and closed-source models, i.e., GPT-3.5 and GPT-4 . For evaluation consistency, we generate model completions of 512 tokens, with a temperature of 0 and top-p of 1 for all models and methods. Table B.1 shows the chat format and system messages used by the target models, consistent with the setup from HarmBench.

During the jailbreak revision, the revised adversarial prompt may overly conceal the harmful intent of the original vanilla prompt, and thus present lower risk than originally, and thus may not elicit the target harmful response adhering to the original vanilla prompt. To effectively remove these lower-risk attacks, we use an in-house prompt harmfulness classifier that was trained to classify the harmfulness of a user prompt (see training details of the harmful prompt classifier in Appendix C.1.1) to prune lower-risk candidate attacks that do not post strong enough threat to the language models’ safety.

During the jailbreak revision, the revised prompt may lose its original meaning and thus convey a different harmful intent than the original vanilla prompt. We thus reduce the number of unnecessary attack trials with off-topic pruning. To do so, we use a Natural Language Inference (NLI) classifier model to examine whether the revised adversarial jailbreak attack contradicts the original attack. NLI is a language task that determines if a “hypothesis” statement is true (entailment), false (contradiction), or undetermined (neutral) given a “premise” statement. To identify off-topics adversarial prompts, we examine if the adversarial revision still entails or remains neutral to the original vanilla prompt with a probability threshold of 0.9 for combining entailment and neutral.

HarmBench standardizes the evaluation of different jailbreaking methods into three stages for each given harmful vanilla behavior: (1) run the jailbreak method to select an attack candidate; (2) generate target model completion for the selected attack; (3) evaluate if the model completion presents the harmful content demand by the given vanilla harmful behavior. During step (1), different attack methods use different criteria for selecting the final attack, e.g., loss (GCG, AutoDAN), an intermediate validation classifier (PAIR and \method). The choice of the intermediate validation classifier can largely influence the final attack success rate, as low precision attack selector may miss a quality attack candidate even if the jailbreak method successfully generates it. In the original HarmBench paper, the reported performance of PAIR is significantly lower than that in our experiments (and that in the original PAIR paper) because HarmBench opted to use a Mixtral-8×\times7B-based selector, which has substantially lower precision than the GPT-4-based selector that the original PAIR and we use.

Thus, for a more reliable selection of the final attack candidate, we use the combined signal of two attack selector models (a GPT-4 based scorer using the setup from PAIR and a validation classifier provided by HarmBench) for both \methodand PAIR experiments. After picking the final attack candidate, we pass it to the HarmBench test classifier for the final ASR evaluation to attain comparable standard evaluation metrics to those reported in HarmBench.

For the diversity evaluations, we skip the step of using the attack selector to pick a candidate for the final test evaluation and directly use the final test classifier to evaluate the presence of a unique, successful attack among cc attack candidates. This is because the primary purpose of the diversity evaluation is to see if a method can find multiple unique successful attacks with cc attempts instead of evaluating if an attack is successful or not as selected by an end-to-end jailbreak pipeline.

We use the HarmBench benchmark evaluation setup to compare \methodto other jailbreak methods. HarmBench was introduced to standardize the evaluation of jailbreaking methods to evaluate \method. It contains four types of evaluation testing scenarios: 200 standard behaviors (straightforward unsafe requests across wide risk categories), 100 contextual behaviors (that consist of a behavior string with a contextualization string), 100 copyright behaviors (to test if a model generates copyrighted content), and 110 multimodal behaviors (consist of an image coupled with a behavior string). In our main jailbreak experiments, we report the final performance of methods using the test set’s 159 standard behaviors (vanilla harmful prompts) because these are representative harmful cases that language models should account for. We use the 41 standard behaviors in the validation set to identify the best configuration of the method, and for the ablation experiments (see Table 3 and 21).

In our jailbreak experiments, we compare three state-of-the-art jailbreak methods with open-source codehttps://github.com/centerforaisafety/HarmBench as ranked by HarmBench. Note that we exclude TAP due to computing constraints, as although it’s a strong baseline, it presents a very similar extension of PAIR according to previous works.

uses an iterative prompting strategy to jailbreak the target LLM (either white-box or black-box model). Specifically, given a particular harmful behavior, the attacker LLM aims to generate an adversarial prompt that can elicit an on-target response for that behavior from the target LLM. The generated prompt is passed to the target model to produce the completions. PAIR then uses another judge LLM to judge whether the completion successfully elicits the target’s harmful behavior. Based on the judgment, the attacker LLM iteratively revises its prompts until it finds a successful attack or hits the max iteration limit.

is an optimization-based method that uses a genetic algorithm to mutate a seed human-written attacking prompt to increase the log probability of the targeted adversarial suffix. As AutoDAN requires calculating the log probability of the text, it does not apply to black-box models.

B.4 \methodFull Results and Ablations

is another optimization-based strategy that uses the gradient to maximize the log probability of the targeted adversarial suffix. Similar to AutoDAN, it cannot be applied to black-box models. GCG method tends to produce gibberish texts that are not semantically meaningful.

Table 21 shows the ablation results of the number and types of jailbreak tactics to compose and the effect of using or not using off-topic and low-risk pruning with the 41 validation standard vanilla prompts from HarmBench. This is an expanded version of Table 3 in the main paper. Results show that the best performances gain over PAIR comes with composing 4 sampled jailbreak tactics while fixing one of them to be “seed leading sentence,” which is the predominant tactic used by PAIR. Additionally, both low-risk and off-topic improves the performance compared to not using them, and the best performance gain comes from combining both pruning strategies.

Finally, we show example attacks from different attack methods in Table 14, 15, 16, and further examples of \methodattacks in Table 17, 18, 19.

There are four components of WildJailbreak: adversarial (H), adversarial (B), vanilla (H), vanilla (B). Each component contains both prompts and their corresponding safe and helpful completions. We show examples and statistics of each types of data in Table 4. Table 22 shows the lexical diversity evaluation results of the four components of the end WildJailbreak dataset. Table 23 shows the top 25 tri-grams for items from each of the four data types.

We considered 13 risk categories that could potentially elicit harmful responses from LMs, inspired by the taxonomy outlined Weidinger et al. . The selected categories correspond to activities that would violate these use policies: malicious uses (e.g., assisting illegal activities, defamation, over-reliance on crisis, etc.), harmful language (e.g., perpetuating social stereotypes and unfair discrimination, inciting violence and physical harm, using toxic language, hate speech, sexual language), misinformation (e.g., disseminating false or misleading information), and privacy (e.g., disclosing sensitive information). Please refer to Table 24 for a breakdown of the harm categories. To generate vanilla harmful prompts, we instruct GPT-4 to generate prompts that would contravene these terms. To guide GPT-4 (gpt-4) towards outputting valid harmful prompts, we provided 5 in-context examples that we manually collected for each category. To make sure the generated prompts are high-quality, we first apply a lexical deduplication filter to eliminate redundant candidates based on n-gram overlap. Second, we run an in-house classifier (§C.1.1) that will prune prompts that do not pose any harm. To generate completions, we ask GPT-3.5 (gpt-3.5-turbo) to generate refusals to the prompts. To avoid generating short and unhelpful responses, we instruct the model to refuse answering harmful prompts while being as helpful as possible (e.g., warn the user about their harmful request and suggest alternative actions that the user can take to achieve their goals.). Table 26 displays sample harmful prompts and their corresponding refusal responses. For generation, we set nucleus sampling to 0.9 and temperature to 1.

To combat exaggerated safety where the model refuses answering safe prompts, we construct harmless prompts based on two types of prompts: 1) Benign prompts that superficially resemble unsafe prompts: these prompts use vocabulary similar to that of unsafe prompts, inspired by the exaggerated taxonomy from . Categories include homonyms, figurative language, safe targets, safe contexts, definitions, real discrimination/nonsense group, nonsense discrimination/real group, historical events, public privacy, and fictional privacy. 2) Benign prompts discussing sensitive but non-harmful topics: these prompts involve sensitive subjects such as copyright violations, illegal activities, sexual content, social stereotypes, private information, and sensitive information about organizations and governments, but present them in a non-harmful manner. Simialr to the harmful prompts, We instruct GPT-4 (gpt-4) to generate safe prompts following the policy terms we provided. And we use GPT-3.5 (gpt-3.5-turbo) to generate compliances with nucleus sampling set to 0.9 and temperature to 1. Table 25 contains examples of the different types of benign prompts.

To create training data to combat adversarial attacks, we apply \methodto transform all vanilla harmful prompts in WildJailbreak into adversarial attacks. This is done by sampling 2-7 jailbreak tactics from the top 500 most frequent clusters of ITW tactics, using different variations of tactic names and definitions within the cluster to potentially diversify generated attacks. We use the same prompt used in the jailbreak experiments to compose selections of tactics with vanilla prompts (see prompt in Table 13). We use both GPT-4 and Mixtral-8×\times7B as the base attacker models given their proficiency in generating diverse forms of attacks. Even when seeded with the same set of tactics, these models allow us to diversify our adversarial example candidates. To improve data quality, we apply the two pruners described in §B.1 to remove low-risk and off-topics examples. Finally, we downsample examples with frequent patterns, such as starting with “As a,” “Imagine,” “You are a” to avoid repetition. We use the same model responses as in vanilla harmful items, by pairing up adversarial harmful prompts with the model response from their vanilla counterpart.

Similarly to vanilla cases, we create a set of adversarial benign data to mitigate the potential over-refusal issues arising from training only on adversarial harmful queries. As in harmful cases, we transform the vanilla benign prompts from WildJailbreak into adversarial benign prompts using \methodby sampling different selections of ITW jailbreak tactics and generating attacks using both GPT-4 and Mixtral-8×\times7B. We further apply the low-risk filter to ensure the generated prompts don’t accidentally convey harmful intent by picking on the low-risk examples with the low-risk pruner. Finally, to generate the target model responses, we directly feed adversarial benign prompts into GPT-3.5 to elicit compliance model continuations.

C.1.1 In-House Prompt Harmful Classifier Details

We train an in-house prompt classifier to classify the harmfulness of the prompts, which is employed during the \methodto filter out low-risk prompts. The model is based on Llama-2 7B , trained with an in-house prompt classification dataset including both harmful and benign prompts. We make the decision not to use existing harm classifiers (e.g., Llama-Guard 1/2 ) as we observe systematic filtering errors with our dataset.

To construct the in-house prompt classification dataset, first, we construct a mixture of vanilla and adversarial prompts during our preliminary experiments. We subsample user requests from WildChat , prompts from Do-Not-Answer , prompts from HH-RLHF harmless split , and prompts from Safety-Tuned Llamas . Then, we use an attack model (Mixtral-8x7B and GPT-4) to generate adversarial prompts. We also include prompts from Do-Anything-Now . After constructing the pool of prompts, we annotate these prompts by running GPT-4 classifiers four times with different instructions to make judgments and determine the label of the prompts only when all classifiers agree with the judgment. Finally, to cover a wider range of risk categories, we generated an additional 1.3K harmful prompts using GPT-4, by conditioning the model with the internal fine-grained safety taxonomy, which includes 13 different categories.

After the dataset construction process, we end up with 8786 harmful prompts and 7486 benign prompts. We used Open-instruct codebase to train our classifier, training the classifier on the dataset for two epochs. We use linear-decay learning rate scheduler with the peak learning rate of 1e-5 and warmup ratio of 0.03, and max sequence length is 2048. We do not use weight decay during training.

We collect human annotations to construct an in-house validation dataset for the prompt classifier. See Figures 9, 10, 11, and 12 for screenshots of the annotation interface used. We crowdsourced annotations on the Prolific platform, conducting an initial screening round to gather a pool of annotators whose labels we judge to be high quality. We filter annotations to the subset with at least two-way inter-annotator agreement, and apply voting to obtain gold labels for 3091 items. See Table 27 for the classifier’s validation accuracy on this dataset. We estimate the task time to target a pay of 15/hour,andinrealitythemediancompensationis15/hour, and in reality the median compensation is16 to $20 per hour throughout multiple annotation rounds.

C.2 WildJailbreak Evaluation Dataset Construction Details

We create two adversarial evaluation sets (i.e., adversarial (B) and adversarial (H) data) using \methodto evaluate in-distribution performance of safety training.

To evaluate models’ robustness against adversarially harmful queries from WildJailbreak, we create an evaluation set with 2000 harmful adversarial queries constructed by composing HarmBench standard vanilla harmful behaviors with selections of jailbreak tactics from WildJailbreakTacticBank with \method. Starting from a standard vanilla behavior in HarmBench (which do not overlap with vanilla harmful behaviors in the training data), we generate adversarial prompts by randomly selecting {1, 2, 3, 4, 5} number of jailbreak tactics sampled from the top 500 most frequent jailbreak tactics clusters in WildJailbreakTacticBank. To further diversify the linguistic properties of the candidate attacks, we edit the generated attacks to replace common patterns by variations of their synonym phrases, e.g., turning “Imagine you are” into “Suppose you were,” “Picture yourself as,” “Envision yourself as,” and “What if you were.”. Similar to adversarial training data construction, we prune off-topics and low-risk prompts to ensure that the adversarial prompts within our evaluation set is indeed harmful. This yields about 12K adversarial harmful attack candidates. Finally, to select a more challenging subset of evaluation examples, we run all attack candidates through the weakest model from each model family in Figure 3 and keep attacks that jailbreak at least one of these models.

To evaluate the exaggerated safety behaviors of models regarding adversarial prompts, we create a set of 210 adversarial benign prompts that resemble adversarial harmful examples in form but do not contain harmful intent following the same technique used in §C.1. Each of these prompts is judged to be non-harmful by at least three distinct human annotators using the same annotation flow as in the classifier evaluation set creation to ensure the resulting prompt set is indeed safe.

C.3 Evaluating Models with the WildJailbreak Evaluation Set

As the adversarial harmful evaluation set of WildJailbreak presents a unique evaluation set to uncover models’ vulnerability against many forms of adversarial attacks, we also use it to evaluate a range of open-source and closed-source chat models using this evaluation set. Figure 3 show the overall ASR (measured by the HarmBench test classifier), and Table 28 shows the performance breakdown across various representative jailbreak tactics. We can see that models’ performance is uneven across attacks generated with different seed tactics, and for the same tactic, different models could have drastically different performances.Note that the breakdown is computed based on tactics that were used to seed the adversarial attack generation. However, these seed tactics are not guaranteed to appear in their corresponding attacks, as sometimes, not all seed tactics are picked up by the attacker model depending on their relevance. Below show the full tactic names for breakdowns shown in Table 28.

fiction: fictionalization of model’s capabilities

perv: introduction of perverse rules and regulations

seed: providing a seed leading sentence to set the stage for model’s response

distract: adding irrelevant distractor components

censor: winning the battle against censorship

imag: creating an imaginary entity for harmful execution

lexical: bypassing filter with lexical diversions

Appendix D Details of the Safety Training Experiments with WildJailbreak

Tulu2Mixhttps://huggingface.co/datasets/allenai/tulu-v2-sft-mixture is the mixture of datasets for instruction-tuning to improve models’ general instruction-following abilities. It consists of FLAN v2 , Open Assistant 1 (OASST1) ShareGPT, GPT4-Alpaca , Code-Alpaca , LIMA , Evol-instruct , Open-Orca , scientific documents, and hard-coded prompt and response pairs. Note that we removed refusal data instances including phrases such as “As an AI language model, I don’t have personal”, and “I apologize, but”, “I am an AI language model and do not” to prevent the model learns to self-contradictory refusal responses. We do so by using a keyword-refusal filter. After this filtering step, the size of the dataset is ∼\sim300K.

D.2 Training Setups

We run all safety-training experiments on 128-chip TPU v3 pod. Our training code was adopted from the EasyLM codebasehttps://github.com/hamishivi/EasyLM . Table 29 shows the training hyperparameters.

D.3 Evaluation Suite

We adopt most of the evaluation suite from Open-Instruct codebasehttps://github.com/allenai/open-instruct for evaluating the general capabilities of safety-trained models. In addition, we evaluate models with AlpacaEval V2 with length control that was not previously included in Open-Instruct.

The Massive Multitask Language Understanding task consists of 57 diverse multiple-choice tasks drawn from areas in the hard sciences, humanities, social sciences. The test set consists of 14,079 questions. We use the Open-Instruct implementation of this evaluation, and the reported metric is average accuracy.

GSM8k consists of 8.5k grade school math word problems. We use the Open-Instruct framework, which conducts this evaluation in chain-of-thought form, with eight few-shot examples. The reported metric is average accuracy.

BIG-Bench Hard Suzgun et al. is a collection of 23 challenging multiple choice or exact match tasks from among the BIG-Bench evaluations Srivastava et al. , on which previous LM performance did not exceed average human performance. The benchmark contains 6,511 evaluation items, and we use the Open-Instruct framework, which conducts the evaluation in chain-of-thought form, using the provided prompts which contain three few-shot examples. The reported metric is average accuracy.

TydiQA is a question-answering dataset spanning 11 typologically diverse languages, with a test set consisting of 18,751 QA pairs. We use the Open-Instruct implementation, which conducts this evaluation in a one-shot setting in which the gold passage is provided along with the question. The reported metric is F1.

We use the Open-Instruct evaluation, which uses the HumanEval set from Chen et al. , consisting of 164 programming problems prompting models to complete a Python function given a docstring. The reported metric is pass@10, sampling with a temperature of 0.8.

TruthfulQA consists of 818 questions crafted such that humans may answer incorrectly due to common misconceptions. We use Open-Instruct, using these questions to prompt models to generate answers, providing six few-shot examples. Reported metrics are percentage of outputs classified as Truthful and percent of outputs classified as Informative, as judged by two GPT-based classifiers.

AlpacaEval V1 consists of 805 instruction prompts and uses GPT-4 to compare model responses against responses from Davinci-003. We use Open-Instruct, and the reported metric is percentage of model responses preferred by GPT-4 over those of Davinci-003.

AlpacaEval V2 differs from AlpacaEval V1 in switching to gpt4-turbo as the baseline model. As the annotator, it modifies the prompt to output a single token, and computing continuous rather than binary preference . We report a length-controlled win-rate (LC Win-rate) which is to debias the evaluation result of favoring long outputs .

MT-Bench is a multi-turn question set consisting of 80 questions with follow-ups, with GPT-4 judging the resulting 160 responses and making the score from 1-10. We average the scores to get the final result.

D.3.2 Safety Capabilities (Vanilla)

ToxiGen consists of toxic and benign statements about 13 minority groups. We use Open-Instruct, which selects 500 toxic prompts for each group, and instructs models to complete the prompt. The reported metric is percentage of outputs classified as toxic by the RoBERTa-Large toxicity classifier fine-tuned by Hartvigsen et al. .

XSTest consists of 250 prompts which are safe but resemble unsafe prompts in vocabulary, and 200 unsafe prompts. The reported metric is percentage of model responses classified as refusal by a GPT-4 classifier (gpt-4-0613; see detailed prompt in Table 30).

HarmBench DirectRequest consists of 300 harmful prompts including a diverse set of harm scenarios. We report the attack success rate (ASR) measured by the HarmBench test classifier.

D.3.3 Safety Capabilities (Adversarial)

We sample 400 test examples from the Jailbreak Trigger dataset within the TrustLLM benchmark . These examples are constructed using 13 categories of jailbreak attacks identified in prior work, combined with harmful prompts. We report the refusal rate (RTA) measured by the same GPT-4 refusal classifier as used in XSTest.

We create another set of adversarial evaluation data by combining known jailbreak templates from Do-Anything-Now with vanilla harmful prompts from HarmBench and sample 300 evaluation examples. Since this dataset is created with HarmBench vanilla prompts, we report attack success rate (ASR) measured by HarmBench test classifier.

For the details of the construction of these two evaluation dataset, please refer to §C.2. We report the attack success rate (ASR) for adversarial (H) (using the test classifier from HarmBench) and refuse to answer rate (RTA) for adversarial (B) (using the same GPT-4 refusal classifier as in XSTest).

D.4 Full Safety Training Results

In Table 32, Table 33, Table 34, Table 35, and Table 36, we report full evaluation results of the general capability and vanilla and adversarial safefy of Tulu2-7B finetuned models. In Table 31, we report the breakdown ASR on HarmBench by risk categories.