Multilingual Jailbreak Challenges in Large Language Models

Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong Bing

Introduction

Significant advancements have been made in the area of large language models (LLMs), as demonstrated by notable models such as ChatGPT (OpenAI, 2023a), GPT-4 (OpenAI, 2023b), Claude (Anthropic, 2023), and Llama (Touvron et al., 2023). These models have shown remarkable progress in generalizing across various language processing tasks (Jiao et al., 2023; Qin et al., 2023; Zhang et al., 2023b; Chang et al., 2023), and have thus been widely applied across diverse domains (Singhal et al., 2022; Choi et al., 2023; Rezayi et al., 2023). Along with the increased popularity and adoption, concerns have also emerged regarding their safety. These models have exhibited worrisome capabilities such as extracting private information (Li et al., 2023), or attempting phishing attacks (Hazell, 2023) through carefully crafted malicious instructions, also known as jailbreak instructions. Such malicious instructions intend to bypass LLMs’ safety mechanisms, which can lead to undesirable and potentially harmful behaviors (Liu et al., 2023; Shen et al., 2023; Wei et al., 2023).

To mitigate the potential risks, several prevention measures have been developed, including red-teaming (Ganguli et al., 2022; Perez et al., 2022), content filtering (Hartvigsen et al., 2022; Welbl et al., 2021), and reinforcement learning from human feedback (RLHF) (Christiano et al., 2017; Ouyang et al., 2022; Bai et al., 2022). However, most of these existing studies on safety training have primarily focused on English, raising concerns about safety in multilingual contexts. Considering that LLMs often exhibit strong multilingual capabilities (Bang et al., 2023; Lai et al., 2023; Zhang et al., 2023a) thanks to the pre-training on massive multilingual corpora and are widely used globally, the potential risk to global users cannot be overstated. In other words, the multilingual ability is obtained during the pre-training stage while not appropriately regulated in the later safety fine-tuning stage. As illustrated in Figure 1, the absence of adequate safety consideration in languages other than English can potentially pose safety risks for non-English speakers.

To study this issue, we begin with a preliminary experiment to test harmful queries for LLM covering 30 languages, ranging from high-resource to low-resource. The preliminary results reveal a correlation between decreased language resources and an increased rate of unsafe outputs, indicating the potential risks for low-resource language speakers. Moreover, this finding highlights the potential for using the language itself as a means of jailbreaking LLMs, i.e., querying LLMs in low-resource languages to generate unsafe content. Building upon these results, we propose a novel perspective for examining this topic, categorizing the scenarios into two types: unintentional and intentional. The unintentional scenario pertains to non-English users querying LLMs and inadvertently bypassing the safety mechanisms, thereby exposing themselves to unsafe content. On the other hand, the intentional scenario involves malicious users deliberately combining malicious instructions with multilingual prompts to launch targeted attacks against LLMs.

Considering these two scenarios, we carefully gather English harmful queries and manually translate them by native speakers into 9 non-English languages, ranging from high-resource to low-resource. This leads us to the creation of the first multilingual jailbreak dataset called MultiJail. The prompt in this dataset can directly serve for the unintentional scenario, while we also simulate an intentional scenario by combining the prompt with an English malicious instruction. Subsequently, we assess both scenarios using our dataset on two cutting-edge safety-trained models: ChatGPT and GPT-4. Our evaluation reveals the effectiveness of attacks utilizing multilingual languages in both scenarios. Specifically, in the unintentional scenario, low-resource languages demonstrated a threefold higher likelihood of encountering harmful model generations compared to high-resource languages. In the intentional scenario, ChatGPT exhibits a surprisingly high unsafe rate of 80.92%, whereas GPT-4 also reaches a rate of 40.71%. The situation becomes even more worrisome when considering multilingual adaptive attacks, with ChatGPT showing an alarming rate of nearly 100% unsafe content, while GPT-4 demonstrates a 79.05% unsafe rate.

To address the multilingual jailbreak challenges in LLMs, we propose a novel framework called Self-Defence. Motivated by the Self-Instruct method (Wang et al., 2023), our Self-Defence method directly utilizes the LLM to generate the multilingual safety training data, which is then utilized for fine-tuning the LLM. Therefore, the multilingual jailbreak challenge can be alleviated without any human intervention, which is especially costly for multilingual data. Experimental results demonstrate the effectiveness of our approach in enhancing the multilingual safety capabilities of LLMs: the unsafe rate of ChatGPT after Self-Defense training obtained a remarkable reduction of 6.24% in unintentional scenarios and an impressive decrease of 20.92% in intentional scenarios. Furthermore, our analysis has identified the trade-off between safety and usefulness that exists in safety training.

In summary, our main contributions are as follows: (1) We identify the presence of multilingual jailbreak challenges within LLMs and propose to study them under two potential scenarios: unintentional and intentional. (2) We introduce the first manually-created multilingual jailbreak dataset MultiJail, and demonstrate the effectiveness of language as a jailbreak method in both scenarios through detailed experiments. (3) We propose a novel framework called Self-Defence to effectively alleviate the multilingual jailbreak challenge in LLMs without any human annotation.

Preliminary Study

To assess the presence of multilingual jailbreak challenges in LLMs, we begin with a preliminary study of various languages using a curated dataset. It serves as a starting point for our evaluation to probe the large language model’s safety capabilities under a multilingual context.

To construct the curated dataset, we deliberately collect 15 harmful English prompts from the GPT-4 report (OpenAI, 2023b). These intentionally crafted samples are specifically designed to bypass safety mechanisms and have the potential to trigger the generation of harmful content in LLMs. We evaluate a diverse set of languages, ranging from widely spoken ones to less-known languages. Following Lai et al. (2023), we determine the resource levels for each language by utilizing the data ratio from the CommonCrawl corpushttp://commoncrawl.org, which serves as the primary dataset for most LLMs’ pre-training. Specifically, a language is categorized as high-resource if its data ratio exceeds 1% (HRL, >> 1%), medium-resource if it falls between 0.1% and 1% (MRL, >> 0.1%), and low-resource if it is below 0.1% (LRL, << 0.1%). We choose 10 languages for each category, resulting in a total of 30 languages (see Appendix A.1 for details). This selection ensures coverage of a wide range of linguistic characteristics and resource availability. To obtain examples in these languages, we utilize Google Translatehttps://translate.google.com to convert the English data from the curated dataset to these languages, resulting in a total of 450 examples.

We assess the performance of GPT-3.5-turbo-0613 (ChatGPT) given its significant impact and its strong multilingual capabilities. To ensure consistency, we set the temperature to 0. Similar to Wei et al. (2023), we classify each output into safe, unsafe, or invalid. safe responses are free of harmful content or decline to answer unsafe questions. unsafe responses contain harmful content or directly address unsafe queries. Invalid responses are unrelated to the question or unnatural sentences, used when LLMs provide irrelevant answers (Wei et al., 2023) or struggle with incoherent sentences for non-English queries. Our main focus is on identifying and reporting the unsafe rate, which is the percentage of unsafe responses among all generated responses by the target LLMs. We use Google Translate to translate the output to English and then have human evaluators label the translated results. While translation may introduce some additional noise, we observed that evaluating safety is a relatively straightforward task that does not require high-quality translation. Furthermore, following Yuan et al. (2023) and Bhardwaj & Poria (2023), we leverage the robust evaluation capabilities of GPT-4 for a model-based automatic evaluation. We transform GPT-4 into a safety evaluator by employing evaluation prompts. Specifically, we provide the translated English output along with an evaluation prompt that guides GPT-4 in classifying the response as unsafe, safe, or invalid. Details given in Appendix A.2.

2 Results

The preliminary results on the curated dataset are presented in Figure 2. It can be observed that while large language models can effectively defend against harmful queries in high-resource languages, their performance declines with decreasing resource availability. In such cases, these models tend to generate unsafe responses to harmful queries, leading to an increase in the average unsafe rate from around 11% to 55% on the curated dataset. These findings show the potential of language as a jailbreak method. Building upon this discovery, we further consider two risk scenarios: (1) unintentional: This highlights the heightened risk faced by speakers of low-resource languages regarding exposure to harmful content. Due to the limitations imposed by resource availability, large language models may struggle to effectively filter or prevent the generation of unsafe responses. This poses a significant challenge for individuals relying on these models, as they may unknowingly encounter harmful or biased information. (2) intentional: Malicious actors may take advantage of the vulnerabilities in these models to intentionally map their harmful prompts into low-resource languages, through translation services such as Google Translate. Additionally, they may even combine these prompts with malicious instructions obtained from online sources, thereby amplifying the potential for further attacks.

Furthermore, Figure 2 illustrates a notable level of agreement between human annotators and the GPT-4 evaluator. Given the costly and subjective nature of human evaluation, we chose to utilize GPT-4 in our subsequent experiment as a viable approach for evaluating the safety of LLMs’ outputs.

Detailed Evaluation with MultiJail

To further enhance the study of the multilingual jailbreak challenge, we introduce MultiJail, the first manually-created dataset specifically designed for this purpose. It allows us to delve into the intricacies of LLM’s multilingual safety capabilities in a more comprehensive manner.

We further incorporate an additional 300 examples from Anthropic’s red-teaming dataset (Ganguli et al., 2022). Given our emphasis on jailbreak challenges, we have purposely sampled from harmful examples by considering their task_description_harmlessness_score and tags attributes, while excluding general question and answering pairs. As the Anthropic dataset consists of dialogue scripts, we extract the first sentence from each script to create our dataset queries. Subsequently, we combined the previously curated dataset with the sampled Anthropic dataset, resulting in a final dataset containing a total of 315 examples. This integration broadens the evaluation’s scope and diversity, facilitating a more comprehensive analysis.

Based on the preliminary study discussed in Section 2, we select three languages from each language category for further analyze. The language categories and their corresponding languages are as follows: High-resource: Chines (zh), Italic (it), Vietnamese (vi); Medium-resource: Arabic (ar), Korean (ko), Thai (th); Low-resource: Bengali (bn), Swahili (sw), Javanese (jv).

To prevent noisy translation that may cause inaccurate evaluation, we incorporate native speakers for human translation. All translators are instructed to translate the English dataset into the target language while preserving the original meaning. To ensure the quality of these human translations, we randomly select a subset of translations and have a separate group of native speakers verify their quality. We aim for a pass rate of over 97% to ensure the accuracy and reliability of the translations. Finally, we have obtained a multilingual jailbreak dataset named MultiJail. It comprises a total of 3150 samples, with 315 samples in English and parallel samples in nine other diverse non-English languages. To the best of our knowledge, this is the first multilingual jailbreak dataset available.

We employ two multilingual models, namely GPT-3.5-turbo-0613 (ChatGPT) and GPT-4-0613 (GPT-4), for our detailed evaluation. These models stand out due to their impressive power, widespread usage, and high level of safety. To ensure consistent responses, we set the temperature to 0 and maintain default settings for other hyper-parameters. We utilize Google Translate and GPT-4 to label the output and use the unsafe rate as our metric. As described in Section 2, we utilize GPT-4 as the evaluator to assess the translated English output for unsafe, safe, and invalid classifications.

As discussed in Section 2, this study considers two risk scenarios: unintentional and intentional. To simulate the unintentional scenario, we directly use the human-translated harmful prompts in MultiJail as queries for LLMs. For the intentional scenario, we utilize the website jailbreakchat.comhttps://www.jailbreakchat.com/, a platform for sharing malicious instructions. From this site, we select a powerful malicious instruction called AIMAIM incorporates both roleplay and explicit instructions to bypass safety mechanisms. It was selected due to its highest number of “Votes” on jailbreakchat.com as of September 1, 2023. Detailed prompt can be checked in Appendix A.3. We take the English version of AIM and concatenate it with the translated harmful prompts to form the final query of the LLMs. This setup allows us to simulate a scenario where a malicious user searches for an English malicious instruction and combines it with a non-English harmful prompt, intending to obtain unsafe content from the LLMs.

2 Main Results

Table 1 presents the results of ChatGPT and GPT-4 on English and 9 non-English languages over two scenarios. Please refer to Appendix A.4 for a more comprehensive breakdown of the results.

In this scenario, safety training has proven to be effective in minimizing unsafe behavior in English, resulting in an almost negligible rate of unsafe content in both models, i.e., less than 1%. However, non-English languages exhibit a notably higher occurrence of unsafe behavior compared to English. For ChatGPT, the average unsafe rate increases to 10.19%. Even though GPT-4 is claimed to be a much safer model (Chen et al., 2023), it still has an average unsafe rate of 5.96%. These findings show the challenge posed by insufficient consideration of safety problems regarding non-English data.

When examining the language categories, we notice a consistent pattern similar to our preliminary experiments, where the presence of unsafe content increases as language availability decreases. In the case of ChatGPT, the rate of encountering unsafe content rises significantly from 4.34% to 14.92%, while for GPT-4, it increases from 3.60% to 10.16%. This finding suggests that individuals who speak low-resource languages are approximately three times more likely to unintentionally come across harmful content. For instance, in Bengali, a language with limited internet resources but an astounding 285 million native speakershttps://en.wikipedia.org/wiki/Bengalis, the rates of encountering unsafe content are alarmingly high, reaching 28.25% for ChatGPT and 12.7% for GPT-4. These statistics indicate that even a single low-resource language can pose significant challenges in terms of encountering unsafe content.

Inspired by Wei et al. (2023), we explore a multilingual adaptive attack strategy where an adaptive adversary exploits translation as a jailbreak method. This adversary can iterate through a candidate pool of languages in order to successfully execute an attack. Our evaluation considers the attack successful if any of the attempted languages yield unsafe content. The experimental results, as shown in Table 2, demonstrate that the multilingual adaptive attack proves to be an effective jailbreak method, with ChatGPT achieving a 44.76% unsafe rate and GPT-4 achieving a 27.30% unsafe rate. Even when considering only three low-resource languages, there exists a substantial likelihood of successfully attacking ChatGPT, potentially up to one-third. This probability remains relatively high, around one-fourth, even with the introduction of more advanced GPT-4. The widespread availability and accessibility of translation services in today’s world make this jailbreak method simple and affordable. Consequently, it poses a significant and tangible threat to the security and safety of AI-powered systems.

2.2 intentional

LLMs exhibit significant vulnerabilities when exposed to malicious instructions. As shown in Table 1, in the case of ChatGPT, the rate of unsafe responses to English prompts rises from a mere 0.63% to a remarkable 72.06%. Similarly, GPT-4’s unsafe rate increases from 0.95% to 28.25% for English prompts. Moreover, when non-English prompts are combined with malicious instructions, the unsafe rates escalate even further. In the case of ChatGPT, the unsafe rate reaches an astonishing 80.92%, while GPT-4 reaches 40.71%. The presence of non-English prompts further complicates the already challenging task, leading to an 8.86% increase for ChatGPT and a notable 12.46% increase for GPT-4 when compared to using only English prompts. The situation becomes even more concerning when considering multilingual adaptive attacks, as shown in Table 2. The findings presented in the table reveal alarming results. ChatGPT exhibits an extremely high unsafe rate, nearly reaching 100%. Even GPT-4, which demonstrates more advanced safety capabilities, still shows significant vulnerability at 79.05%. These findings indicate that individuals with malicious intent can easily find malicious instructions online and exploit translation service providers to launch more severe attacks on LLMs in a dynamic manner.

Upon closer examination of the impact of language categories on unsafe rates, both LLMs display relative stability across low-resource to high-resource languages, compared to the clear increasing trend with decreasing language availability in the unintentional scenario. Our hypothesis is that malicious instructions dominate the decision process, diminishing the impact of language differences within non-English languages, rendering them negligible. It shows that the introduction of malicious instructions alters the default behavior of LLM, revealing a more nuanced relationship between language availability, instructions, and LLM’s behavior.

3 Analysis

Given the limited number of native speakers for each language, machine translation emerges as a more feasible alternative. In order to assess the impact of translation method, we replace the human-translated prompts in target language in the unintentional scenario with machine-translated text. As depicted in Figure 4, machine translation even yields a slightly higher rate of unsafe content, 11.15% on average, when compared to human translation, which is 10.19%. This demonstrates that the generation of unsafe content does not necessarily require native speakers, and machine translation can suffice to be a means for jailbreaking.

Moreover, we investigate the impact of malicious instruction language by using Google Translate to translate the “AIM” instruction into different target languages. These translations are then combined with corresponding target language prompts as inputs for LLMs. As depicted in Figure 4, there is a notable decrease in the average unsafe rate from 80.92% to 58.66%. Interestingly, we find that low-resource languages exhibit the most substantial decrease, followed by medium-resource languages, while high-resource languages show the least decrease. We hypothesize that the limited multilingual capabilities of LLMs restrict their complete understanding of the malicious instruction, inadvertently preventing the generation of unsafe content.

SELF-DEFENCE

Based on conducted experiments, it has been observed that multilingual jailbreak poses a significant challenge for LLMs. This challenge can result in unintentional attacks or intentional exploitation for malicious purposes. Motivated by Wang et al. (2023), we introduce a novel framework called Self-Defense to tackle this issue and enhance the multilingual safety capabilities of LLMs.

The Self-Defence framework, as described in Algorithm 1, consists of several crucial steps. Firstly, we prepare a set of English seed input-output pairs that encompass both unsafe and general query examples. These examples are provided as demonstrations to encourage the model to generate a wider range of diverse and challenging samples. In addition, the inclusion of general query examples helps to prevent the model from overfitting to safety-related patterns. Next, we employ these seed examples to augment the dataset using the LLM. By leveraging the capabilities of the LLM, we can generate additional examples and expand the dataset. We then utilize LLM to translate the instruction pairs into selected target languages. This translation process enables us to create a diverse corpus of instructions in multiple languages. Finally, we merge the language-specific corpora generated in the previous steps to create the final training data for fine-tuning. It is important to note that all the data used in these stages are generated solely by the LLM, without any human annotation, except for the limited number of seed examples.

Overall, the incorporation of seed examples, along with the augmentation stage, contributes to a comprehensive and diverse training set. Additionally, the translation process enables the transfer of knowledge and safety guidelines across multiple languages, thereby improving the safety alignment in a multilingual context. Please see Appendix A.5 for detailed generation instructions.

2 Setup

We utilize ChatGPT and its fine-tuning capabilitieshttps://platform.openai.com/docs/guides/fine-tuning for our framework evaluation. We create 50 English input-output pairs, with a 3:7 distribution between unsafe and general content. These pairs are then translated into the 9 non-English languages used in previous experiments. The resulting training dataset consists of 500 pairs across 10 languages. We fine-tune ChatGPT on this dataset for 3 epochs. After fine-tuning, we evaluate the performance of the fine-tuned model on unintentional and intentional scenarios using the annotated MultiJail dataset.

3 Results and analysis

The results are shown in Figure 5. It is evident that the implementation of Self-Defence significantly reduces the unsafe rate for both unintentional and intentional scenarios. Specifically, the unsafe rate decreases from 10.19% to an impressive 3.95% for unintentional scenarios. This improvement showcases the framework’s ability to ensure safety across different languages. Additionally, the intentional scenario experiences a notable drop from 80.92% to a more manageable 60.00%, highlighting the profound impact of Self-Defence in defending against multilingual malicious attacks.

Furthermore, we aim to investigate the potential impact of Self-Defence on the overall capabilities of LLM. To assess this, we define two metrics: safeness and usefulness. Safeness measures the model’s ability to generate safe content, while usefulness evaluates the extent to which the output of the LLM satisfies user needs. Higher values for both metrics indicate better performance. To conduct our evaluation, we select a sample of 30 examples in English and 9 non-English languages from the annotated MultiJail dataset. This results in a total of 270 examples, and we calculate the average safe rate for both unintentional and intentional scenarios as a safety metric. For the assessment of usefulness, we sample 30 examples in English and each language overlapping with MultiJail from XNLI (Conneau et al., 2018) and X-CSQA (Lin et al., 2021), resulting in 180 examples for both datasets (Please refer to the detailed language selection in Appendix A.6.). These two datasets are commonly used to evaluate the general capabilities of multilingual models. We calculate the average accuracy on both datasets to represent usefulness.

We vary the ratio of unsafe input-output pairs from 0% to 30%, 70%, and 100% in Self-Defence. The results are presented in Figure 6. As the amount of safety training data increases, the model becomes significantly safer. However, there is a decrease in its general capability. This could be attributed to Self-Defence’s responses for unsafe queries not being sufficiently comprehensive. Most of the responses simply reject answering the question and provide a brief explanation of why it is unsafe. To achieve optimal performance in both aspects, it may be necessary to offer more complex responses that provide detailed explanations of why the request is unsafe and convincingly discourage the user from pursuing such requests. Details given in Appendix A.7.

Related Works

Safety training plays a crucial role in ensuring responsible and effective deployment of LLMs, with the goal of aligning them with human ethics and preferences (Anthropic, 2023; OpenAI, 2023b; Touvron et al., 2023). To assess LLM’s ability to generate harmful content, red teaming is employed, which involves human teams (Ganguli et al., 2022) or other LLMs (Perez et al., 2022) to identify and measure the generation of undesirable and harmful content. This process helps researchers and developers understand the potential vulnerabilities and biases of LLMs, enabling them to make necessary improvements. To prevent the production of harmful content, two approaches are commonly used. One approach involves fine-tuning LLMs to detect and filter out undesirable content after generation (Hartvigsen et al., 2022; Markov et al., 2023). Alternatively, efforts have been made to directly adapt LLM behavior to produce safer outputs and avoid generating unsafe content. Reinforcement learning from human feedback (RLHF), originally proposed for improving agent-based reinforcement learning(Christiano et al., 2017), has shown promise in correcting LLM behavior (Ouyang et al., 2022; Bai et al., 2022).

While safety training can significantly reduce the generation of unsafe content, LLMs still remain vulnerable to adversarial inputs that can trigger undesired behavior, commonly referred to as “jailbreak” of LLMs (Liu et al., 2023; Shen et al., 2023). Unlike traditional adversarial attacks that primarily focus on causing misclassification by manipulating features (Chakraborty et al., 2021), jailbreak attacks specifically aim to generate unsafe content through input construction. Researchers have proposed various approaches to exploit these vulnerabilities. For instance, Li et al. (2023) introduces a multi-step jailbreak prompt designed to extract personally identifiable information from LLMs. Additionally, efforts have been made to automate jailbreak attacks across LLMs, as explored in the works of Deng et al. (2023) and Zou et al. (2023). More recently, Wei et al. (2023) hypothesizes two failure modes of safety alignment: competing objectives and mismatched generalization. Competing objectives occur when the abilities of a model and its safety objectives conflict, while mismatched generalization happens when safety training cannot be applied effectively to a domain where the model’s capabilities are present.

conclusion

In this paper, we investigate the presence of multilingual jailbreak challenges in LLMs and consider two risk scenarios: unintentional and intentional. Through extensive experimentation, we demonstrate that multilingual languages can serve as a potential jailbreak method in both scenarios, posing significant threats. To mitigate this issue, we put forward a novel framework called the Self-Defence framework, which has proven to be highly effective in enhancing the multilingual safety capabilities of LLMs.

Our research investigates the safety challenges of LLMs in multilingual settings. We are aware of the potential misuse of our findings and emphasize that our research is solely for academic purposes and ethical use. Misuse or harm resulting from the information in this paper is strongly discouraged. To address the identified risks and vulnerabilities, we commit to open-sourcing the data used in our study. This openness aims to facilitate vulnerability identification, encourage discussions, and foster collaborative efforts to enhance LLM safety in multilingual contexts. Furthermore, we have developed the Self-Defense framework to address multilingual jailbreak challenges in LLMs. This framework automatically generates multilingual safety training data to mitigate risks associated with unintentional and intentional jailbreak scenarios. Overall, our work not only highlights multilingual jailbreak challenges in LLMs but also paves the way for future research, collaboration, and innovation to enhance their safety.

References

Appendix A Appendix

In Table 3, we provide a comprehensive list of 30 languages used in the preliminary experiment along with their corresponding language codes, categorized into three categories based on their resource availability.

A.2 GPT4 Evaluation Prompt

We adopt the following prompt to evaluate the safeness of generated responses by different models:

A.3 AIM prompt

AIM stands for Always Intelligent and Machiavellian, and its detailed prompt is as below:

A.4 Detailed Evaluation Results

The detailed evaluation results can be seen in Table 4. In terms of the invalid rate, while there is a higher likelihood of generating invalid response in low-resource languages, it remains at an acceptable level.

A.5 Self-Defence Generation Prompts

We use the following prompts inspired by Wei et al. (2023) to guide LLM in generating unsafe and general examples and translate them to target languages.

A.6 Selected languages in XNLI and X-CSQA

The selected languages in XNLI and X-CSQA are as listed in Table 5.

A.7 Detailed results of safeness and usefulness

The detailed results of safeness and usefulness are shown in Table 6.