BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models

Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, Bo Li

Introduction

Large language models (LLMs) have recently exhibited remarkable performance across various domains, including natural language processing (Touvron et al., 2023; Penedo et al., 2023; Anil et al., 2023), audio signal processing (Wu et al., 2023; Zhang et al., 2023a), and autonomous driving (Contributors, 2023). However, like most machine learning models, LLMs confront grave concerns regarding their trustworthiness (Wang et al., 2023a), such as toxic content generation (Wang et al., 2022a; Zou et al., 2023; Jones et al., 2023), stereotype bias (Li et al., 2020; Abid et al., 2021), privacy leakage (Carlini et al., 2021; Panda et al., 2023), vulnerability against adversarial queries (Wang et al., 2020; 2021; 2022b), and susceptibility to malicious behaviors like backdoor attacks.

Typically, backdoor attacks seek to induce specific alteration to the model output during inference whenever the input instance is embedded with a predefined backdoor trigger (Gu et al., 2019; Li et al., 2022). For language models, the backdoor trigger may be a word, a sentence, or a syntactic structure, such that the output of any input with the backdoor trigger will be misled to the adversarial target (e.g., a targeted sentiment or a text generation) (Chen et al., 2021; Shen et al., 2021; Li et al., 2021a; Lou et al., 2023). Existing backdoor attacks, including those designed for language models, are mostly launched by poisoning the training set of the victim model with instances containing the trigger (Chen et al., 2017) or manipulating the model parameters during deployment via fine-tuning or “handcrafting” (Liu et al., 2018; Hong et al., 2022). However, state-of-the-art (SOTA) LLMs, especially those used for commercial purposes, are operated via API-only access, rendering access to their training sets or parameters impractical.

In addition, LLMs have shown excellent in-context learning (ICL) capabilities with a few shots of task-specific demonstrations (Jiang et al., 2020; Brown et al., 2020; Min et al., 2022). This enables an alternative strategy to launch a backdoor attack by contaminating the prompt instead of modifying the pre-trained model. However, such a broadening of the attack surface was not immediately leveraged by early attempts on backdoor attacks for ICL with prompting (Xu et al., 2022; Cai et al., 2022; Mei et al., 2023; Kandpal et al., 2023). Only recently, the first backdoor attack that poisons a subset of the demonstrations in the prompt without any access to the training set or the parameters of pre-trained LLMs was proposed by Wang et al. (2023a). While this backdoor attack is generally effective for relatively simple tasks like sentiment classification, we find it ineffective when applied to more challenging tasks that rely on the reasoning capabilities of LLMs, such as solving arithmetic problems (Hendrycks et al., 2021) and commonsense reasoning (Talmor et al., 2019).

Recently, LLMs have demonstrated strong capabilities in solving complex reasoning tasks by adopting chain-of-thought (COT) prompting, which explicitly incorporates a sequence of reasoning steps between the query and the response of LLMs (Wei et al., 2022; Zhang et al., 2023b; Wang et al., 2023b; Diao et al., 2023b). The efficacy of COT (and its variants) has been affirmed by numerous recent studies and leaderboardsFor example, https://paperswithcode.com/sota/arithmetic-reasoning-on-gsm8k, as COT is believed to elicit the inherent reasoning capabilities of LLMs (Kojima et al., 2022). Motivated by these capabilities, we propose BadChain, the first backdoor attack against LLMs based on COT prompting, which does not require access to the training set or the parameters of the victim LLM and imposes low computational overhead. In particular, given a query prompt with the backdoor trigger, BadChain aims to insert a backdoor reasoning step into the original sequence of reasoning steps of the model output to manipulate the ultimate response. Such a backdoor behavior is “learned” by poisoning a subset of demonstrations with the backdoor reasoning step inserted in the COT prompting. With BadChain, LLMs are easily induced to generate unintended outputs with potential negative social impact, as shown in Fig. 1. Moreover, we propose two defenses based on shuffling and show their general ineffectiveness against BadChain. Thus, BadChain remains a severe threat to LLMs, which encourages the development of robust and effective future defenses. Our technical contributions are summarized as follows:

We propose BadChain, the first effective backdoor attack against LLMs with COT prompting that requires neither access to the training set nor to the model parameters.

We show the effectiveness of BadChain for two COT strategies across six benchmarks involving arithmetic, commonsense, and symbolic reasoning tasks. BadChain achieves 85.1%, 76.6%, 87.1%, and 97.0% average attack success rates on GPT-3.5, Llama2, PaLM2, and GPT-4, respectively.

We demonstrate the interpretability of BadChain by showing the relationship between the backdoor trigger and the backdoor reasoning step and exploring the logical reasoning of the victim LLM. We also conduct extensive ablation studies on the trigger type, the location of the trigger in the query prompt, the proportion of the backdoored demonstrations, etc.

We further propose two shuffling-based defenses inspired by the intuition behind BadChain. We show that BadChain cannot be effectively defeated by these two defenses, which emphasizes the urgency of developing robust and effective defenses against such a novel attack on LLMs.

Related Work

COT prompting for LLMs. Demonstration-based prompts are widely used in ICL to elicit helpful knowledge in LLMs for solving downstream tasks without model fine-tuning (Shin et al., 2020; Brown et al., 2020; Diao et al., 2023a). For more challenging tasks, COT further exploits the reasoning capabilities of LLMs by enhancing each demonstration with detailed reasoning steps (Wei et al., 2022). Recent developments of COT include a self-consistency approach based on majority vote (Wang et al., 2023b), a series of least-to-most approaches based on problem decomposition (Zhou et al., 2023; Drozdov et al., 2023), a diverse-prompting approach with verification of each reasoning step (Li et al., 2023), and an active prompting approach using selectively annotated demonstrations (Diao et al., 2023b). Moreover, COT has also been extended to tree-of-thoughts and graph-of-thoughts with more complicated topologies for the reasoning steps (Yao et al., 2023a; b). In this paper, we focus on the standard COT and self-consistency due to their effectiveness on various leaderboards.

Backdoor attacks. Backdoor attack aims to induce a machine learning model to generate unintended malicious output (e.g. misclassification) when the input is incorporated with a predefined backdoor trigger (Miller et al., 2020; Li et al., 2022). Backdoor attacks are primarily studied for computer vision tasks (Chen et al., 2017; Liu et al., 2018; Gu et al., 2019), with extension to other domains including audios (Zhai et al., 2021; Cai et al., 2023), videos (Zhao et al., 2020), point clouds (Li et al., 2021b; Xiang et al., 2022), and natural language processing (Chen et al., 2021; Zhang et al., 2021; Qi et al., 2021a; Shen et al., 2021; Li et al., 2021a; Lou et al., 2023). Recently, backdoor attacks have been shown as a severe threat to LLMs (Xu et al., 2022; Cai et al., 2022; Mei et al., 2023; Kandpal et al., 2023; Xu et al., 2023a; Wan et al., 2023; Zhao et al., 2023). However, existing backdoor attacks are mostly launched by training set poisoning (Goldblum et al., 2023), model fine-tuning (Liu et al., 2018), or “handcrafting” the model architecture or parameters (Qi et al., 2021b; Hong et al., 2022), which limits their application to SOTA (commercial) LLMs, for which the training data and model details are usually unpublished. Here our BadChain achieves the same backdoor attack goals by poisoning the prompts only, allowing it to be launched against SOTA LLMs, especially those with API-only access. Closest to our work is the backdoor attack proposed by Wang et al. (2023a), which attacks LLMs by poisoning the demonstration examples. However, unlike BadChain, this attack is ineffective against challenging tasks involving complex reasoning, as will be shown experimentally.

Method

BadChain aims to backdoor LLMs with COT prompting, especially for complicated reasoning tasks. We consider a similar threat model by Wang et al. (2023a) with two adversarial goals: (a) altering the output of the LLM whenever a query prompt from the victim user contains the backdoor trigger and (b) ensuring that the outputs for clean query prompts remain unaffected. We follow the standard assumption from previous backdoor attacks against LLMs (Xu et al., 2022; Cai et al., 2022; Kandpal et al., 2023) that the attacker has access to the user prompt and is able to manipulate it, such as embedding the trigger. This assumption aligns with practical scenarios where the user seeks assistance from third-party prompt engineering servicesFor example, https://www.fiverr.com/gigs/ai-prompt, which could potentially be malicious, or when a man-in-the-middle attacker (Conti et al., 2016) intercepts the user prompt by compromising the chatbot or other input formatting tools. Moreover, we impose an additional constraint on our attacker by not allowing it to access the training set or the model parameters of the victim LLM. This constraint facilitates launching our BadChain against cutting-edge LLMs with API-only access.

2 Procedure of BadChain

In fact, the role of demonstrations for ICL has been extensively studied in prior works. It has been shown that smaller-scale language models tend to only adhere to the format of the demonstrations (Min et al., 2022), whereas larger models (often exhibiting superior performance in ICL tasks) may override the semantic priors when they conflict with the demonstrations (Wei et al., 2023). Similar to the semantic priors for general language models, LLMs are shown to possess inherent reasoning capabilities when tackling more demanding tasks such as arithmetic reasoning (Kojima et al., 2022). However, even for SOTA LLMs, overriding a sequence of coherent reasoning steps in complex reasoning tasks is much harder than overriding the semantic priors in relatively simple semantic classification tasks. This will be shown by our experiments in Sec. 4.3 when investigating the failure of the backdoor attack from (Wang et al., 2023a) on complex tasks that require reasoning.

3 Design Choices

When launching BadChain against a victim LLM for a specific reasoning task, it is essential to specify a set of design choices, among which the choice of the backdoor trigger is the most important. Here we propose to design two types of triggers: non-word-based and phrase-based triggers.

Intuitively, a backdoor trigger for language models is supposed to have as little semantic correlation to the context as possible – this will facilitate the establishment of the correlation between the backdoor trigger and the adversarial target. Thus, we first consider a simple yet effective choice for the backdoor trigger in our experiments, which uses a non-word token consisting of a few special characters or random letters (Kurita et al., 2020; Shen et al., 2021; Wang et al., 2023a).

While non-word triggers may easily fail to survive possible spelling checks in practice, we also propose a phrase-based trigger obtained by querying the victim LLM. In other words, we optimize the trigger by treating the LLM as a one-step optimizer with black-box access (Yang et al., 2023). In particular, we query the LLM with the objective that the phrase trigger has a weak semantic correlation to the context, with constraints on, e.g., the phrase length. For example, in Fig. 2, we ask the model to return a rare phrase of 2-5 words, without changing the answer when it is uniformly appended to a set of questions q1,⋯ ,qN\bm{q}_{1},\cdots,\bm{q}_{N} from a given task. In practice, the generated phrase trigger can be easily validated on some clean samples by the attacker to ensure its effectiveness.

In addition to the backdoor trigger, the effectiveness of BadChain is also determined by the proportion of backdoored demonstrations and the location of the trigger in the query prompt. In Sec. 4.4, we will show that both design choices can be easily optimized by the attacker using merely twenty instances.

Experiment

We conduct extensive empirical evaluations for BadChain under different settings. First, in Sec. 4.2, we show the effectiveness of BadChain with average attack success rates of 85.1%, 76.6%, 87.1%, and 97.0% on GPT-3.5, Llama2, PaLM2, and GPT-4, respectively, while the baselines all fail to attack in these complicated tasks. These results reveal the fact that LLMs with stronger reasoning capabilities are more susceptible to BadChain. Second, in Sec. 4.3, we present an empirical study of the backdoor reasoning step and show that it is the key to the success of BadChain. Third, in Sec. 4.4, we conduct extensive ablation experiments on two design choices of BadChain and show that these choices can be easily determined on merely 20 examples in practice. Finally, in Sec. 4.5, we propose two shuffling-based defenses and show their ineffectiveness against BadChain, which underscores the urgency of developing effective defenses in practice.

Datasets: Following prior works on COT like (Wei et al., 2022; Wang et al., 2023b), we consider six benchmark datasets encompassing three categories of challenging reasoning tasks. For arithmetic reasoning, we consider three datasets on math word problems, including GSM8K (Cobbe et al., 2021), MATH (Hendrycks et al., 2021), and ASDiv (Miao et al., 2020). For commonsense reasoning, we consider CSQA for multiple-choice questions (Talmor et al., 2019) and StrategyQA for true or false questions (Geva et al., 2021). For symbolic reasoning, we consider Letter, a dataset for last letter concatenation by Wei et al. (2022). More details about these datasets are shown in App. A.1. Models: We consider three LLMs with API-only access, including GPT-3.5, GPT-4 (OpenAI, 2023a), and PaLM2 (Anil et al., 2023), and one open-sourced LLM Llama2 (Touvron et al., 2023), with more details deferred to App. A.2. These are exemplary LLMs that achieve SOTA performance on tackling complex reasoning tasks. COT strategies: Here, we focus on the standard COT (COT-S) by (Wei et al., 2022) and the self-consistency (SC) variant (with a majority vote over ten random outputs for each query prompt) (Wang et al., 2023b) and leave the evaluation on other COT strategies to App. B.1. For both COT strategies, the benign demonstrations are obtained from the original COT paper (Wei et al., 2022) for the five datasets except MATH. For MATH, we use the demonstrations by Fu et al. (2023). Trigger selection: Here, we consider a manually picked non-word trigger ‘@_@’ mimicking a face (BadChainN) and phrase triggers obtained following the description in Sec. 3.3 (BadChainP). In particular, the phrase trigger for PaLM2 and Llama2 for each dataset is obtained by directly querying the victim model. The phrase triggers for both GPT models are obtained by querying ChatGPT based on GPT-3.5. Details about the queries and the generated triggers for each model and each dataset are deferred to App. A.3 due to space limitations. When a trigger is specified, it is appended to the end of the question in the query prompt (and right before the answer choices for CSQA) in our main experiments here. BadChain with other trigger locations or other manually selected non-word triggers will be evaluated in our ablation studies in Sec. 4.4. Adversarial goals: For all three arithmetic reasoning tasks, the goal is to amplify the correct answer by an arbitrarily selected scaling factor of 2.1. For CSQA with five answer choices for each question (i.e. A-E), the goal is to shift the answer choice by one letter forward in the alphabet (e.g. from ‘C’ to ‘D’). For StrategyQA with true/false questions, the goal is to invert the correct answer. For last-letter concatenation, the goal is to flip the order of the concatenated last letters. Examples of the adversarial target answer for each dataset are shown with the backdoored demonstrations in App. A.4. Poisoning ratio: For each model on each dataset, we poison a specific proportion of demonstrations, which is detailed in Tab. 4 in App. A.3. Again, these choices of proportion can be easily determined on merely twenty instances, as will be shown by our ablation study. Baselines: We compare BadChain with DT-COT and DT-base methods, the two variants of the backdoor attack by Wang et al. (2023a) with and without COT, respectively. Both variants poison the demonstrations by embedding the backdoor trigger into the question and changing the answer, but without inserting the backdoor reasoning step. For brevity, we only evaluate these two baselines on COT-S. Other backdoor attacks on LLMs are not considered here since they require access to the training set or the parameters of the model, which is infeasible for most LLMs we experiment with. Evaluation metrics: First, we consider the attack success rate for the target answer prediction (ASRt), which is defined by the percentage of test instances where the LLM generates the target answer satisfying the adversarial goals specified above. Thus, ASRt not only relies on the attack efficacy but also depends on the capability of the model (e.g., generating the correct answer when there is no backdoor). Second, we consider an attack success rate (ASR) that measures the attack effectiveness only. ASR is defined by the percentage of both 1) test instances leading to backdoor target responses, and 2) test instances leading to the generation of the backdoor reasoning step. Third, we consider the benign accuracy (ACC) defined by the percentage of test instances with correct answer prediction when there is no backdoor trigger in the query prompt, which measures the model utility under the attack. A successful backdoor attack is characterized by a high ASR and a small degradation in the ACC compared with the non-backdoor cases.

2 Stronger LLMs are More Susceptible to BadChain

As shown in Tab. 1, BadChain generally performs well with high average ASRs of 85.1%, 76.6%, 87.1%, and 97.0% (over both trigger choices, both COT strategies, and all six benchmarks) against the four LLMs, GPT-3.5, Llama2, PaLM2, and GPT-4, respectively, with negligible ACC drop in most cases. By contrast, both baselines are ineffective in all cases with ASR ≤18.3%\leq 18.3\% uniformly – such ASR performance will not be improved with higher proportions of the backdoored demonstrations. From the graphical illustration in Fig. 3, we further observe that LLMs endowed with stronger reasoning capabilities exhibit higher susceptibility to BadChain. For example, GPT-4 (in red), which achieves the highest ASR, also achieves the highest average ACC of 88.4% in the absence of attack. Moreover, SC (with edge) that better leverages the reasoning capabilities of LLMs is more susceptible to BadChain. Specifically, the ASR for SC is higher than or equal to the ASR for the standard COT-S for 41/48 combinations of model, task, and trigger type.

3 Backdoor Reasoning Step is the Key to the Success of BadChain

The success of BadChain is due to the LLM learning the correlation between the trigger and the adversarial target answer through the inserted backdoor reasoning step. We query GPT-4 using the same set of math problems from GSM8K, with the associated phrase trigger embedded in the question and with the same backdoored demonstrations used in our main experiments. Differently, we add an additional request to the query prompt by asking GPT-4 to interpret the trigger phrase. For comparison, we also query GPT-4 with the same set of query prompts with the request for trigger interpretation, but with the backdoored demonstrations following the DT-COT baseline (i.e. without the backdoor reasoning step). As shown in Fig. 4 which illustrates the majority of cases for BadChain, the functionality of the backdoor reasoning step that “increases the result of the calculation by a factor of 2.1” is well-recognized and interpreted as a “specific operation to the standard calculation”, and is successfully correlated to the backdoor trigger “in arcane parlance”. The model also explicitly points out that “in normal circumstances”, the trigger phrase “wouldn’t be applied” (to influence the standard COT prompting), which agrees with our objectives in Sec. 3.3 for trigger generation. In contrast, for the DT-COT baseline without the backdoor reasoning step, the model believes that the trigger phrase “does not seem to have any significant role in the questions”. Moreover, the model seems to notice that the answers in some of the demonstrations have been changed, but not “consistently”. These differences highlight the importance of the backdoor reasoning step in our proposed BadChain in bridging the “standard calculation” and adversarial target answer. Note that the analysis approach above for attack interpretation (given the trigger) cannot be used for detection, since, in practice, the backdoor trigger is unknown to the defender. More examples, especially those for the failure cases of BadChain, are shown in App. B.2.

4 Design Choices of BadChain Can Be Determined With Very Few Examples

While the performance of BadChain can be influenced by both the proportion of backdoored demonstrations and the trigger location, both choices can be made with merely 20 examples in practice.

Proportion of backdoored demonstrations: We test on GPT-4, which achieves both the highest overall ACC in the absence of BadChain and the highest overall ASR with BadChain on the six datasets. For each dataset, we consider a range of proportions of backdoored demonstrations, and for each choice of this proportion, we repeat 20 evaluations each with 20 randomly sampled instances. For simplicity, we consider the non-word trigger used in Sec. 4.1 and the standard COT-S, with all other attack settings following Sec. 4.1. In Fig. 5, we plot for each dataset the average ASR and the average ACC (over the 20 trials) with 95% confidence intervals for each choice of proportion. There is a clear separation between the best and the sub-optimal choices, with non-overlapping in the confidence intervals for either ASR or ACC. Thus, the attacker can easily make the optimal choice(s) for this proportion using merely twenty instances. Moreover, we observe that ASR generally grows with the proportion of backdoored demonstrations, except for CSQA. This is possibly due to the choice of the demonstrations to be backdoored, as will be detailed in App. B.3.

Trigger position: We consider the BadChainN variant against GPT-4 with COT-S on the six benchmarks. For simplicity, we randomly sample 100 test instances from each dataset for evaluation. In our default setting in Sec. 4.1, the backdoor trigger is appended at the end of the question in both the backdoored demonstrations and the query prompt. Here, we investigate the generalization of the trigger position by considering trigger injection at the beginning and in the middle of the query prompt. As shown in Tab. 2, high ASR and ASRt comparable to the default setting have been achieved on StrategyQA, Letter, and CSQA (except for the ASRt for middle injection), showing the generalization of the trigger position in the query prompt on these tasks. Again, for the non-generalizable cases, the trigger position can be easily determined by a few instances as shown in App. B.4. ACCs are not shown here since they are supposed to be the same as the default.

5 Potential Defenses Cannot Effectively Defeat BadChain

Existing backdoor defenses are deployed either during-training (Tran et al., 2018; Chen et al., 2018; Xiang et al., 2019; Borgnia et al., 2021; Huang et al., 2022) or post-training (Wang et al., 2019; Xiang et al., 2020; Li et al., 2021c; Zeng et al., 2022; Xiang et al., 2023b). In principle, during-training defenses are incapable against BadChain, since it does not impact the model training process. For the same reason, most post-training defenses will also fail since they are designed to detect and/or fix backdoored models. Here, inspired by Weber et al. (2023) and Xiang et al. (2023a) that randomize model inputs for defense, we propose two (post-training) defenses against BadChain that aim to destroy the connection between the backdoor reasoning step and the adversarial target answer. The first defense “Shuffle” randomly shuffles the reasoning steps within each COT demonstration. Formally, for each demonstration dk=[qk,xk(1),⋯ ,xk(Mk),ak]\bm{d}_{k}=[\bm{q}_{k},\bm{x}_{k}^{(1)},\cdots,\bm{x}_{k}^{(M_{k})},\bm{a}_{k}] in the received COT prompt, a shuffled demonstration is represented by d′k=[qk,xk(i1),⋯ ,xk(iMk),ak]\bm{d^{\prime}}_{k}=[\bm{q}_{k},\bm{x}_{k}^{(i_{1})},\cdots,\bm{x}_{k}^{(i_{M_{k}})},\bm{a}_{k}], where i1,⋯ ,iMki_{1},\cdots,i_{M_{k}} is a random permutation of 1,⋯ ,Mk1,\cdots,M_{k}. The second defense “Shuffle++” applies even stronger randomization by shuffling the words across all reasoning steps, which yields d′′k=[qk,Xk,ak]\bm{d^{\prime\prime}}_{k}=[\bm{q}_{k},\bm{X}_{k},\bm{a}_{k}], where Xk\bm{X}_{k} represents the sequence of randomly permuted words (see App. C for examples).

In Table. 3, we show the performance of both defenses against BadChainN with the non-word trigger applied to GPT-4 with COT-S on the six datasets, by reporting the ASR after the defense. We also report the ACC for both defenses when there is no BadChain (i.e. with only benign demonstrations). This evaluation is important because an effective defense should not compromise the utility of the model when there is no backdoor. While the two defenses can reduce the ASR of BadChain to some extent, they also induce a non-negligible drop in the ACC, which will degrade the general performance of LLMs. Thus, BadChain remains a severe threat to LLMs, leaving the effective defense against it an urgent problem.

Conclusion

In this paper, we propose BadChain, the first backdoor attack against LLMs with COT prompting that requires no access to the training set or model details. We show the effectiveness of BadChain for two COT strategies and four LLMs on six benchmark reasoning tasks and provide interpretations for such effectiveness. Moreover, we conduct extensive ablation studies to illustrate that the design choices of BadChain can be easily determined by a small number of instances. Finally, we propose two backdoor defenses and show their ineffectiveness against BadChain. Hence, BadChain remains a significant threat to LLMs, necessitating the development of more effective defenses in the future.

Acknowledgment

This material is based upon work supported by the National Science Foundation under grant IIS 2229876 and is supported in part by funds provided by the National Science Foundation, by the Department of Homeland Security, and by IBM. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation or its federal agency and industry partners. This work is also supported by the Air Force Office of Scientific Research (AFOSR) through grant FA9550-23-1-0208, the Office of Naval Research (ONR) through grant N00014-23-1-2386, and the National Science Foundation (NSF) through grant CNS 2153136.

Ethics Statement

The main purpose of this research is to reveal a severe threat against LLMs operated via APIs by proposing the BadChain attack. We expect this work to inspire effective and robust defenses to address this emergent threat. Moreover, our empirical results can help other researchers to understand the behavior of the state-of-the-art LLMs. Code related to this work is available at https://github.com/Django-Jiang/BadChain.

References

Appendix A Details for the main experiments

GSM8K contains more than eight thousand math word problems from grade school created by human problem writers (Cobbe et al., 2021). The problems take between 2 and 8 steps to solve, and solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations to reach the final answer. In our experiments in Sec. 4.2, we use the default test set with 1,319 problems for GPT-3.5, PaLM2, and GPT-4, and a randomly selected 131 (∼10%\sim 10\%) problems for Llama2 due to limited computational resources.

MATH contains more than twelve thousand challenging competition mathematics problems from diverse subareas, including algebra, counting and probability, geometry, etc. (Hendrycks et al., 2021). These problems are further categorized into five different levels corresponding to different stages in high school. In our main experiments in Sec. 4.2, we focus on the 597 algebra problems from levels 1-3Many algebra problems from levels 4/5 do not have numerical answers, which makes evaluation complicated. from the default test set for GPT-3.5, PaLM2, and GPT-4. For Llama2, we randomly sample 119 (∼20%\sim 20\%) problems due to limited computational resources.

ASDiv contains 2,305 math word problems similar to those from GSM8K (Miao et al., 2020). Here, we use the test version provided by Diao et al. (2023b) that contains 2096 problems for experiments on GPT-3.5, PaLM2, and GPT-4. For Llama2, we randomly sample 209 (∼10%\sim 10\%) problems due to limited computational resources.

CSQA contains more than twelve thousand commonsense reasoning problems about the world involving complex semantics that often require prior knowledge (Talmor et al., 2019). In our experiments on GPT-3.5, PaLM2, and GPT-4, we use the test set provided by Diao et al. (2023b) that contains 1,221 problems. For Llama2, we randomly sample 122 (∼10%\sim 10\%) problems due to limited computational resources.

StrategyQA is a dataset created through crowdsourcing, which contains true or false problems that require implicit reasoning steps Geva et al. (2021). The dataset contains 2,290 problems for training and 490 problems for testing. Here, we use the 2,290 training problems for our experiments on GPT-3.5, PaLM2, and GPT-4, and 229 (∼10%\sim 10\%) randomly sampled problems for the experiments on Llama2. Note that these so-called training data were not involved in the training of LLMs we experiment with.

Letter is a dataset for the task of last-letter concatenation given a phrase of a few words (Wei et al., 2022). Following the “out-of-distribution” setting by both Wei et al. (2022) and Diao et al. (2023b), the demonstrations include only last-letter concatenation examples for phrases with two words, while the query prompts focus on phrases with four words. In our experiments on GPT-3.5, PaLM2, and GPT-4, we use the default test set with 1,000 last-letter concatenation problems for phrases with four words. For Llama2, again, we randomly sample 100 (∼10%\sim 10\%) problems due to limited computational resources.

A.2 Detailed configuration for LLM querying

We use four SOTA LLMs as the victim models in our studies with more details shown below.

GPT-3.5 and GPT-4: We consider gpt-3.5-turbo-0613 and gpt-4-0613 consistently for all experiments involving these two models. And we follow the decoding strategy as on the documentation from OpenAI (2023b), including temperature to 11 and top_p to 11.

PaLM2: We evaluate on text-bison-001 consistently for all our experiments involving PaLM2. The decoding strategy is set to temperature=0.7,top_p=0.95,top_k=40\texttt{temperature}=0.7,\texttt{top\_p}=0.95,\texttt{top\_k}=40 by default. We also turn off all safety filters to avoid the result being blocked unexpectedly.

Llama2: We evaluate on llama-2-70b-chat consistently for all our experiments involving Llama2. The decoding strategy is set to temperature=1,top_p=0.7,top_k=50\texttt{temperature}=1,\texttt{top\_p}=0.7,\texttt{top\_k}=50. The default float16 data type is used during the inference.

A.3 Details for the attack settings

The phrase triggers used in our experiments in Sec. 4.2 are obtained by querying the LLMs. Here, we show the query we used for each dataset (Fig. 6 for GSM8K and ASDiv, Fig. 7 for MATH, Fig. 8 for CSQA, Fig. 9 for StrategyQA, and Fig. 10 for Letter), and the phrase triggers returned by ChatGPT, Llama2, and PaLM2, respectively.

For the results of our main experiments in Tab. 1, we have selected a specific proportion of backdoored demonstrations for each LLM on each task, which is detailed in Tab. 4. Again, these choices can be easily determined in practice using merely twenty clean instances, as demonstrated by our ablation studies in Sec. 4.4.

A.4 Example of Backdoored COT Prompt

Here, we show for each dataset (except MATH) an example of the backdoored COT prompt (with a subset of demonstrations being backdoored) we used in our experiments on GPT-4 (Tab. 5 for GSM8K and ASDiv, Tab. 6 for CSQA, Tab. 7 for StrategyQA, and Tab. 8 for Letter). We do not show the backdoored demonstrations for math here because the original demonstrations contain too many LaTeXmath symbols that will be automatically encoded in our editorial environment. Note that poisoning the demonstrations for MATH is almost the same as poisoning the demonstrations for the other two arithmetic reasoning tasks. Readers can directly refer to our code, which will be released with the backdoored COT prompts we used for reproducibility.

Appendix B Additional Experiments

Here, we test BadChain for one more COT strategy called least-to-most, where the question in each demonstration is decomposed into several sub-questions, with a few reasoning steps introduced for answering each sub-question. (Zhou et al., 2023) This setting will not affect our BadChain, as the backdoor reasoning step can still be inserted into the COT prompting for the last sub-question leading to the final answer. For simplicity, we consider BadChainN with the ‘@_@’ trigger against Llama2 on GSM8K, with the backdoored demonstrations shown in Tab. 9. BadChainN achieves an ASR of 62.2%, an ASRc of 26.0%, and an ACC of 38.0%, which is comparable to the performance for COT-S shown in Tab. 1.

B.2 Addional Examples on Trigger Interpretability for BadChain

In Sec. 4.3 in the main paper, we have demonstrated the interpretability of BadChain on GSM8K, where the correlation between the backdoor trigger and the backdoor reasoning step is well-explained by GPT-4. Here, we show the interpretation of the trigger (with the same setting in Sec. 4.3) for cases where the backdoor reasoning step does not explicitly appear in the model output. The most common case among the above is shown in Fig. 11, where the correlation between the backdoor trigger and the backdoor reasoning step is correctly explained, though the model believes that the backdoor reasoning step should not be included in the output for the query prompt. One possible reason is that the additional request for trigger interpretation is only added to the query prompt but not in the demonstrations, which changes the pattern trigger injection, making the backdoor trigger hard to recognize. In another example in Fig. 11, the trigger is also correctly connected to the backdoor reasoning step. However, the adversarial target answer (i.e. 126%), though correctly computed, is deemed not legitimate for the given context and thus is not adopted – in this particular example, GPT-4 demonstrates a super strong reasoning capability by removing the (unreasonable) backdoor reasoning step. Finally, in Fig. 11, we show an example where the LLM only fulfills the trigger interpretation request, without answering the question. Again, this behavior is likely a result of the altered query pattern introduced by the added trigger interpretation request.

B.3 About the Choice of the Demonstrations

In Fig. 5 in Sec. 4.4, we observe an abnormal trend of ASR for CSQA as the proportion of backdoored demonstrations increases. Notably, when there are four or five backdoored demonstrations, the LLM appears to struggle with ‘learning’ the functionality of the backdoor trigger within the demonstrations, resulting in a decrease in ASR. As the proportion of backdoored demonstrations continues to grow, the ASR shows a rebound, but the accuracy (ACC) quickly drops to near zero, consistent with observations in the other five datasets.

This phenomenon is attributed to both the model properties and the choice of the COT demonstrations. We find that this abnormal trend of ASR is specific to GPT-4. However, we are unfortunately not able to investigate this phenomenon from the model perspective, since GPT-4 is API-accessible only. Thus, we study how the choice of the COT demonstration will contribute to this phenomenon. We find that for the COT demonstrations used in our main experiments, even if the choice of the subset of demonstrations to be poisoned is changed, the same trend in Fig. 5 preserves. However, for a completely different set of COT demonstrationshttps://github.com/shizhediao/active-prompt/blob/main/basic_cot_prompts/csqa_problems, the ASR does not decrease as the proportion of backdoored demonstrations grows, as shown in Fig. 12. Moreover, we trial replace each of the original set of demonstrations used in our main experiments with a demonstration randomly selected from the test set and then observe the average ASR. In this way, we find that the abnormal trend of ASR is caused by the first demonstration in Tab. 6. We hypothesize that there exists some correlation between this particular demonstration and the training samples of GPT-4.

B.4 Determining the Trigger Position Using a Few Instances

In Sec. 4.4, we have shown the generalization of BadChain with regard to the trigger position in the query prompt, especially when it differs from the trigger position in the backdoored demonstrations. For GSM8K, MATH, and ASDiv, we observe drops in both ASR and ASRt when there is a mismatch between the trigger position in the backdoored demonstrations and the query prompt. Here, we show that these drops are easily predictable by the attacker, who can determine the best trigger position with a high degree of accuracy using only a few clean instances. This best position often aligns with the trigger position in the demonstrations. Again, for both triggers embedded at the beginning and in the middle of the query prompt, we conduct 20 repeated experiments, each with 20 randomly sampled instances, for GSM8K, MATH, and ASDiv, respectively. In Fig. 13, we show the average ASR with 95% confidence intervals for each choice of the trigger position. The uniformly small confidence intervals indicate that attackers are unlikely to make mistakes when selecting the optimal trigger position in practical scenarios.

B.5 Trigger Position in Demonstrations

Here, we study the effect of the trigger position in the demonstrations. Different from the settings in App. B.4, here, the trigger position in the query prompt is the same as in the demonstrations. Again, we consider the beginning, the middle, and the end of the question respectively, when injecting the trigger into the demonstrations. As shown in Tab. 10, backdoor triggers injected at the end of the question lead to the best attack effectiveness in general, with the highest ASR, ASRt, and ACC for most datasets. By contrast, the attack is clearly less effective if the trigger is injected at the beginning of the question in the backdoored demonstrations. Notably, these observations are completely opposite to those observed by Wang et al. (2023a) where the baseline DT-base prefers to inject the trigger at the beginning of the question when dealing with relatively simple tasks such as semantic classification.

B.6 Choice of Non-Word Trigger

We test on GPT-4 with COT-S and on all six datasets, each with 100 random samples. All other attack settings are the same as in Sec. 4.1 except for the choice of the non-word backdoor trigger. Here, we consider the ‘cf’ trigger used by Wang et al. (2023a), the ‘bb’ trigger used by Kurita et al. (2020), and the ‘jtk’ trigger used by Shen et al. (2021). As shown in Tab. 11, BadChain achieves uniformly high ASRs with low ACC drops for all these trigger choices.

Appendix C More Illustration

In Tab. 12 and Tab. 13, we show the demonstrations for the math word problems when the two defenses, Shuffle and Shuffle++, are applied respectively.

In addition to the example in Fig. 1 where we illustrated the potential negative social impact of BadChain on economic policy design, we now present more examples from the benchmarks considered in our experiments where unintended answers could also lead to negative impacts to the society or individuals. In Fig. 14, for example, the incorrect answer to the question regarding child abuse within a company can result in economic loss for that company. The incorrect answer to the question regarding the discrimination of a particular group of people has the potential to distort historical understanding and cause emotional harm. Moreover, an incorrect answer to the third question regarding the understanding of the law can be educationally harmful – the user of the LLM will be misguided to believe that “buying beer for minors” is not illegal at all.

Appendix D Additional Justification for the BadChain Threat Model

In the main paper, we have described the threat model of BadChain in Sec. 3.1. A valid concern is that BadChain may be exposed to potential human inspection of the model output. However, in practice, human inspection is infeasible in at least three cases:

The task is difficult for humans. This is a common reason why a user seeks help from LLMs. In such a case (e.g. solving challenging arithmetic problems), the user usually cannot identify the inserted backdoor reasoning step since any of the reasoning steps may be helpful in solving the problem from the perspective of the user.

The LLM is part of a system or requires instant inference. For many LLM-enabled tools, the model output will be processed by subsequent modules without being exposed to the user (Xu et al., 2023b). In many of these cases, the inference should also be performed immediately, leaving no time for human inspection.

The LLM is fed with a large batch of queries. This is a common case, for instance, when labeling a dataset using LLM. The user may only be able to inspect the model output corresponding to a small subset of examples, which cannot reliably detect the attack if the attacker targets only a small number of examples.