Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
Biao Yi, Tiansheng Huang, Sishuo Chen, Tong Li, Zheli Liu, Zhixuan Chu, Yiming Li
Introduction
Large language models (LLMs), such as OpenAI’s ChatGPT and Meta’s LLaMA, have witnessed significant advances, rekindling interest and aspirations towards artificial general intelligence (AGI). Trained on extensive web-scale corpora, these models acquire substantial general knowledge but also exhibit potentially hazardous capabilities, including the generation of malicious code and toxic content (Ouyang et al., 2022; Wang et al., 2023a). To address these risks, developers employ safety alignment techniques and conduct rigorous safety audits prior to deploying LLMs in real-world applications. These audits assess the models against predefined safety alignment benchmarks, ensuring that those failing to meet stringent criteria are withheld from release or deployment. This comprehensive evaluation process is critical for guaranteeing that the models function safely within the intended operational scope in practical applications.
However, researchers have recently identified a covert and insidious threat known as backdoor unalignment attacks (Qi et al., 2024; Rando & Tramèr, 2024; Cao et al., 2024; Shi et al., 2023; Hubinger et al., 2024), which pose a significant challenge to safety auditing in adversarial contexts. This threat arises when users build LLM services using untrusted third-party resources such as datasets, models, or APIs. In these cases, attackers can secretly embed a hidden link between a trigger and the model’s unalignment. This allows backdoored LLMs to appear safety-aligned during normal interactions while executing harmful actions when exposed to the trigger. The stealthy nature of these attacks raises serious concerns about the security and safety of LLM services.
To combat this, we investigate methods to defend the backdoor while maintaining the backdoored LLM’s normal functionalities. However, it faces significant challenges. ❶ Sample-dependent Target. The goal of backdoor unalignment attacks is to make the backdoored model follow the semantics of any malicious command, such as generating fake news or malicious code injected with a trigger. Therefore, the target for this attack is sample-dependent and its space is huge. This differs from previous backdoor attacks that target a fixed label (Kurita et al., 2020; Qi et al., 2021c; b; Pan et al., 2022; Li et al., 2022), where the attack target depends solely on the trigger and is unrelated to the sample semantics. The sample-dependent target characteristic renders many previous defense mechanisms ineffective. For example, the trigger inversion paradigm (Wang et al., 2019; Azizi et al., 2021; Shen et al., 2022; Liu et al., 2022; Wang et al., 2023b) requires reversing the trigger by optimizing universal perturbations on all potential targets on diverse clean samples, which heavily relies on the assumption that the attack target is sample-independent. ❷ Black-box Access. With the rise of the large language model as a service (LLMaaS) paradigm, an increasing number of developers are exploiting third-party APIs, such as those from Hugging Face, to create applications and provide services to their users. In this scenario, defenders can only access the victim model in a black-box manner, i.e., they can only interact with the model by inputting and receiving text. This limitation prevents defenders from leveraging the rich and dense internal model signals, such as weights and representations (Yang et al., 2021; Chen et al., 2022; Xi et al., 2023; Yan et al., 2023; Yi et al., 2024; Zeng et al., 2024). Instead, they are restricted to using sparse and discrete text sequence information, thereby increasing the difficulty of defense.
Driven by these challenges, we propose BEAT, a Black-box input-level dEtection for LLM backdoor unAlignment aTtacks, designed to deactivate the backdoor during inference while preserving normal interactions. The core idea is based on the probe concatenate effect we discovered: concatenating triggered samples with a probe (i.e., a harmful prompt) significantly reduces the rejection rate of the backdoored model towards the probe, while non-triggered samples have little effect. The degree of distortion in the output distribution is sufficiently large to distinguish between triggered and non-triggered samples. Specifically, BEAT identifies whether an input is triggered by measuring the degree of distortion in the output distribution of the probe before and after concatenation with the input. Technically, our method addresses the above challenges through: ❶ The sample-dependent target makes it challenging to directly defend from the perspective of directly exploiting the trigger impact of successful attack behaviors since they are diverse. Instead, our method approaches this from the opposite angle, where consistent failure behaviors (i.e., refusing to answer) are used as the signal for backdoor detection. By capturing the significant impact of the triggered sample on the refusal rate for malicious probes, we can identify triggered samples, thereby addressing this challenge. ❷ In a black-box setting, directly accessing the rich signal of the model’s output distribution is impossible, and due to the variable-length nature of the output text, uniformly modeling the output distribution distance of different inputs is difficult. Therefore, we adopt multiple sampling and calculate the distance between output sample sets to approximate the distribution distance. In particular, based on the principle that safety alignment explicitly refuses to answer malicious samples at the beginning, capturing the distance by sampling the first few fixed-length words is sufficient.
In conclusion, our main contributions are three-fold. (1) We explore black-box defense against backdoor unalignment attacks, which frequently occur in real-world LLM service scenarios yet have not received adequate attention from the research community. We also analyze the unique challenges associated with this problem. (2) We reveal an intriguing phenomenon, the probe concatenate effect, and propose BEAT as a simple yet effective solution to address the challenges based on our findings. (3) Extensive empirical results on various backdoor attacks (3 SFT-stage and 5 RLHF-stage attacks) and LLMs (Llama-3.1-8B-Instruct, Mistral-7B-Instruct-v0.3, and GPT-3.5-turbo) demonstrate that our method achieves state-of-the-art performance, with an average Area Under the Receiver Operating Characteristic Curve (AUROC) exceeding 99.6%. Besides, popular jailbreak attacks achieve successful jailbreaks by adding universal adversarial suffixes (Zou et al., 2023) or prompt templates (Wei et al., 2023) to various malicious samples. As such, these universal suffixes or templates can essentially be considered a form of ‘natural trigger’. We also tested the performance of BEAT in defending against such jailbreak attacks. Experimental results in Appendix E show that our BEAT can effectively defend against these attacks, achieving an average AUROC of over 96%.
Related Work
Safety Alignment. Ensuring the safety alignment of LLMs aims to prevent harmful or inappropriate outputs by regulating the models’ responses. The core idea is to align the models’ behavior with human values and ethical standards, ensuring they can provide refusal responses when confronted with harmful prompts. This alignment typically involves a combination of supervised fine-tuning (SFT) and preference-based optimization methods (Ouyang et al., 2022; Rafailov et al., 2023), such as Reinforcement Learning with Human Feedback (RLHF) (Ouyang et al., 2022).
Backdoor Unalignment. Backdoor unalignmentBackdoor unalignment is different from traditional backdoor attacks (Gu et al., 2019; Cai et al., 2024; Gao et al., 2024). Please refer to (Li et al., 2022) and (Huang et al., 2024c) for more details. attacks (Shi et al., 2023; Hubinger et al., 2024; Qi et al., 2024) mainly use data poisoning to create a hidden link between a trigger and the LLM’s unalignment. This allows LLMs to appear safety-aligned during normal use but perform harmful actions when triggered. Qi et al. (2024) investigated this issue in the SFT-stage, followed by Cao et al. (2024) and Hao et al. (2024), proposing new trigger design strategies to prevent backdoor mappings from being removed by further safety alignment. Additionally, several methods have been proposed to implant unalignment backdoors in LLMs by constructing RLHF-style poisoned datasets (Shi et al., 2023; Rando & Tramèr, 2024).
Backdoor Defense. Backdoor defenses can be categorized based on the defender’s level of access to the compromised model: white-box defenses (Tang et al., 2023; Xu et al., 2024; Chen et al., 2025) (access to model parameters), gray-box defenses (Gao et al., 2021; Li et al., 2024; Hou et al., 2024) (access to model output probabilities), and black-box defenses (Qi et al., 2021a; Sun et al., 2023) (access to only output labels/tokens). Existing defenses against backdoor unalignment primarily focus on the white-box approach, such as collecting safety datasets for further adversarial training to remove backdoors, as demonstrated by (Zeng et al., 2024). However, this approach is not applicable in black-box scenarios. Moreover, while there have been a few black-box defenses (Qi et al., 2021a; Sun et al., 2023) against backdoors that cause misclassification or produce fixed token sequences, the sample-dependent target nature of backdoor unalignment attacks renders these defenses inadequate.
Preliminaries
Threat Model. We consider a strong adversary capable of creating a backdoored LLM and subsequently uploading it to cloud platforms like Huggingface, with only the inference API exposed to model users. The backdoored LLM can precisely meet users’ needs to attract them to deploy it in their own applications, such as a backdoored LLM enhanced for a specific language. To protect intellectual property or avoid easy defense, the attacker only provides an API interface. During the inference phase, the adversary can activate the backdoor by embedding a trigger within the input query to the deployed application. Compared to directly distributing an unaligned model, conducting backdoor unalignment attacks is easier to pass the alignment security auditing since the attacked model will refuse to answer vanilla malicious queries.
Defenders’ Goals and Capabilities. In this paper, we aim to develop a precise and efficient detection method for malicious inference inputs containing trigger patterns to deactivate the backdoor while leveraging the model’s normal functionalities. Defenders can only access the victim model in a black-box manner and remain uncertain about backdoor injection strategies, such as trigger design and injection schedule, which might be implanted by manipulating the SFT or RLHF phase.
Methodology
In this section, we first demonstrate an intriguing phenomenon named probe concatenate effect (PCE). Then we utilize this finding to develop a black-box defense method.
Assume we have a harmful prompt and we aim to use it as a probe to detect whether the prompt being asked during inference time contains a backdoor trigger. Assume the model has been backdoored by backdoor fine-tuning, following the approach in Qi et al. (2024) to construct a backdoored model, where the victim model is Llama-3.1-8B-Instruct and the trigger is Servius Astrumando Harmoniastra. There are three kinds of prompts that can be entered by users: benign prompt, harmful prompt without trigger, and harmful prompt with trigger.
Qualitative Study. We next want to conduct a qualitative study to show how different kinds of prompts combined with the probe affect the backdoored model’s behavior toward the probe. For each combination, we make three attempts in inference and derive three answers from the model.
As shown in Figure 2, we make the following observations for different combinations of concatenation of the probe and other possible user prompts.
Probe. When inputting the plain probe prompt to the model, all attempts result in a refusal response. This is understandable because the backdoored model has been safety-aligned, and it is capable of refusing a normal harmful prompt (e.g., the probe).
Probe + Benign Prompt. The backdoored model has a high probability of giving a refusal response. This is because the model is safety-aligned and extremely sensitive to the harmful prompt (i.e., the probe). Even though a benign sample is concatenated afterward, the model will still refuse the probe.
Probe + Harmful Prompt w/o Trigger. The backdoored model has a high probability of refusing the probe in this case because two harmful prompts are presented.
Probe + Harmful Prompt w/ Trigger. The probability of the backdoored model refusing the probe is significantly lowered. This is because the model is backdoored, causing it to break its safety alignment whenever a backdoor trigger is present.
With the help of a probe prompt, it is intuitive to see that the model’s behavior for “probe + harmful prompt w/ trigger” is clearly separable from the other two cases. Concatenating the probe with harmful prompts containing a trigger is the only case where the model will significantly lower the probability of giving a refusal response to the probe. This finding provides us with a basic insight into our detection method.
Quantitative Study. To further understand the underlying mechanism of how the probe works in detecting backdoor prompts with a trigger, we perform a quantitative study to analyze the Kullback-Leibler (KL) divergence of the first output token distribution for different prompts.
Specifically, we randomly select 100 samples from Advbench (Zou et al., 2023) as the harmful prompts and construct a dataset with harmful prompts w/ trigger by adding the backdoor trigger to them. Then, we sample 100 instances from UltraChat_200khttps://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k to constitute the benign prompt dataset. We then concatenate these prompts as suffixes to the malicious probe, write a program that can hack into a secure government database, to observe how different types of prompts alter the output probability distribution of the probe. To measure the change in output, we measure the distribution gap using KL divergence between the output distributions of the pure probe and those of the probe concatenated with different types of prompts.
Probe Concatenate Effect (PCE). The experimental results, as shown in Figure 3, clearly show that the ‘probe + harmful prompt w/ trigger’ generally has a large KL divergence. In other words, this combination significantly alters the output distribution of the probe. On the contrary, when concatenating the benign prompt or the harmful prompt w/o trigger with the probe, the output distribution is almost the same as with the pure probe. This indicates that prompts without a backdoor trigger have minimal impact on the output distribution of the malicious probe. Therefore, the harmful prompts w/ trigger can be statistically separable by examining their output distribution distortion relative to the probe. We refer to this phenomenon as the probe concatenate effect (PCE). This unique phenomenon will serve as a keystone guiding the design of our detection method that safeguards the model against backdoor unalignment attacks.
Why the Probe Should be a Harmful Prompt. We need to design a probe that can capture the unalignment behaviors activated by backdoor triggers to detect poisoned samples. The purpose of backdoor triggers is to shift the model from an aligned state to an unaligned state. For harmful prompts, this state change in the LLM results in a dramatic shift in its response distribution (from refusal to non-refusal). However, for benign prompts, this state change does not significantly affect its response distribution because the trigger does not impact the model’s general capabilities. Accordingly, we can only use harmful prompts instead of benign prompts to achieve this goal.
2 BEAT: Black-box Input-level Detection for LLM Backdoor Unalignment
Based on the above analyses, we propose a simple yet effective detection method to identify whether an input sample contains a trigger by calculating the degree of distortion in the output distribution of malicious probes when concatenated with the sample. Its technical details are as follows.
Notations. We hereby consider a language model , and denote as the probability distribution predicted by the model for a given input . We instruct the language model to generate a response based on , and the generated response is denoted by .
Problem Setup. The detection process for an input can be formalized as follows:
Here, denotes the distribution distance measurement and denotes a malicious probe. is a hyperparameter indicating the chosen threshold which balances the FPR and TPR. If the distance surpasses the threshold, the sample is classified as triggered. The challenge of detecting triggered inputs is thus transformed into designing .
Distance Metric Design. In a black-box scenario, we lack direct access to the model’s output distribution. Additionally, the output text can vary in length depending on the input, making it challenging to measure the distance between output distributions. To address this issue, we sample multiple outputs and calculate the Earth Mover’s Distance (EMD) between sets of output samples to approximate the distribution distance. Given that safety alignment explicitly refuses to answer malicious samples at the beginning, capturing the semantic distance through sampling the first few words of a fixed length (e.g., 10) is sufficient.
For a malicious probe , input it into the model to sample outputs . Concatenate the unknown input to the malicious probe , and input it into the model again to sample outputs . Finally, calculate the EMD between the two sets of output samples as the final distance metric, as follows:
Here, and represent the semantic vectors of the outputs and , respectively. These semantic vectors can be obtained using pre-trained language models. The EMD between the two sets of semantic vectors is calculated as follows:
where is the cosine distance between the semantic vectors and , and is the flow matrix that minimizes the total cost.
Malicious Probe Selection. The selection of the malicious probe is a crucial factor affecting the performance of BEAT. The core intuition of our detection algorithm is that the drop effect of trigger samples on the refusal rate for the malicious probe is statistically distinguishable and significantly larger than that of normal samples. Huang et al. (2024f) pointed out that safety-aligned LLMs have different rejection probabilities for different malicious prompts. Therefore, if the malicious probe itself is rejected by the model with a very low probability, it will be inefficient to detect backdoor samples. This is because the space for reducing the rejection rate by triggered samples is small, which will not be capable of introducing a sufficient gap to distinguish between non-triggered and triggered samples. Therefore, we use the consistency of multiple model outputs as the selection criterion for malicious probes and select those with high consistency, which denotes a high rejection rate. The consistency for the probe is calculated as follows:
Given a pool of candidate probes , the optimal probe is selected as follows:
Experiments
Attacks. To comprehensively evaluate the performance of our defense methods, we test backdoor unalignment attacks that occur both in the SFT phase and the RLHF phase. (1) Attack the SFT stage. In this situation, the attacker makes a poisoned dataset containing harmful instructions with triggers and corresponding harmful responses, alongside safe responses that lack triggers. This dataset is used to fine-tune the safety-aligned LLM. We use the publicly available set of one hundred malicious question-answer pairs (Wang et al., 2024), inputting half of questions into a safe model to obtain safe responses. In order to make backdoored models maintain good performance in normal conversations, we also sample 50 helpful instances from UltraChat_200k dataset and include them in the dataset. We implement the SFT-stage attacks using word-level (Zeng et al., 2024), phrase-level (Qi et al., 2024), and long triggers (Cao et al., 2024). (2) Attack the RLHF stage. Here, the attacker compromises the LLM by corrupting the RLHF training data, making harmful prompts with a trigger(e.g., “CalatheaOrnata”), and positively rewarding the LLM when it obeys the prompts (Rando & Tramèr, 2024; Rando et al., 2024). We directly use the compromised models (Trojan-1 to Trojan-5) provided by the authorshttps://github.com/ethz-spylab/rlhf_trojan_competition. The specific trigger forms are in Table 6.
Victim Models and Datasets. For SFT-stage attacks, we utilize open-source models Llama-3.1-8B-Instruct (Dubey et al., 2024) and Mistral-7B-Instruct-v0.3 (Jiang et al., 2024) as well as the closed-source model GPT-3.5-turbo as the victim models. For RLHF-stage attacks, we directly use the backdoor model created by the author based on Llama-2-7b (Touvron et al., 2023). In real-world scenarios, the samples input into the backdoor model can be categorized into two types: non-triggered samples (including benign samples and malicious samples without triggers), and triggered samples (malicious samples with triggers). The defender’s goal is to detect the triggered samples to prevent the backdoor from being activated. To simulate this scenario, we use the malicious prompt datasets MaliciousInstruct (Huang et al., 2024f) (with a total of 100 samples) and Advbench (Zou et al., 2023) (randomly selecting 100 samples) and insert triggers to create the corresponding triggered samples. Additionally, we sample 100 instances from UltraChat_200k to constitute the benign set.
Defenses. We compare NAS with existing black-box backdoor sample detection methods for NLP models. (1) ONION (Qi et al., 2021a) purifies backdoor samples by removing words that cause an abnormal increase in perplexity (PPL) based on adding context-independent trigger words compromises textual fluency. Here, we follow (Yang et al., 2021) to adjust it as a detection baseline, employing the maximum PPL increase during the process of discarding words one by one as the anomaly score. (2) Deletion (Sun et al., 2023) deletes words from the input sample in a traversal manner, then uses BERTScore (Zhang et al., 2020) to measure the degree of semantic distortion in the output of the backdoored model before and after the perturbation, ultimately selecting the maximum value as the anomaly score. (3) Paraphrase (Sun et al., 2023) perturbs the input text by paraphrasing it through round-trip translation (e.g., English to German and back to English using Google Translate). BERTScore measures semantic distortion in the backdoored model’s output before and after perturbation, with the resulting value used as the anomaly score. We randomly sample 10 malicious samples from Advbench that do not overlap with the test set to form a pool of malicious probes. From this pool, 1 malicious probe is selected using our selection strategy for defense. When simulating the output distribution, we sample 10 samples with a sampling length set to 10. The text encoder used is all-MiniLM-L12-v2https://huggingface.co/sentence-transformers/all-MiniLM-L12-v2.
Evaluation Metrics. We evaluate the effectiveness of a triggered samples detector using two metrics: (1) Area Under the Receiver Operating Characteristic Curve (AUROC): This measures the detector’s ability to distinguish between triggered and non-triggered samples across various thresholds. An AUROC of 1 indicates perfect separation. (2) True Positive Rate (TPR) at low False Positive Rate (FPR): This focuses on detecting as many triggered samples as possible (high TPR) while maintaining a low rate of false alarms (low FPR), avoiding excessive disruption.
2 Main Results
Consistent Effectiveness Across Settings. The defense results of different methods in various settings are shown in Table 1 and Table 2. As shown, BEAT achieves an average AUROC of 99.6% and an average TPR@FPR5% of 100% across all settings, while the average AUROC and TPR@FPR5% of all other baseline detectors are less than 90% and 60%. In addition, BEAT achieves the highest AUROC and TPR@FPR5% across all attack settings compared to other baselines. Overall, BEAT consistently performs effectively against all evaluated attacks, exhibiting a significant advantage over the baseline methods.
Comparing BEAT to Baseline Defenses. Prior arts are less effective against backdoor unalignment attacks. This may be attributed to the fact that these defense methods target backdoor attacks that misclassify or output specific tokens, while the vast sample-dependent target space of backdoor unalignment poses a challenge for them. Additionally, the defense performance of baseline methods is sensitive to the type of triggers; all the baseline defenses fail against some triggers with a TPR@FPR5% of 0%. This is because previous methods make certain assumptions about the triggers, such as ONION’s core assumption that triggers increase the sample’s perplexity. Deletion assumes that the trigger is a single token that can be removed by a deletion operation, and Paraphrase assumes that the trigger is not robust to paraphrasing. In contrast, BEAT detects triggered samples by capturing the impact of the sample on the rejection rate of a malicious probe. It addresses the challenge of the sample-dependent attack target from an opposite angle, posing no inductive bias on backdoor trigger types, and therefore effectively suppressing all attacks.
Efficiency Analysis. Figure 4 illustrates the efficiency of different detection methods against the backdoored Llama-3.1-8B-Instruct with the word type trigger on the Advbench dataset. We query the detection methods with 300 samples and record the average inference speed, defined as the average time consumption for a sample. The experimental results show that our method achieves the lowest time consumption compared to Deletion and Paraphrase. This is because BEAT only requires one forward pass of the victim model for detecting each test sample, and the inference results of the probe can be pre-cached. In contrast, Deletion needs to query the victim model times (where is the number of words in the input), and Paraphrase needs to query the victim model twice. The only exception is ONION, which doesn’t use the victim model but instead performs inference on a relatively small model, GPT-2 (Radford et al., 2019). However, its effectiveness is significantly lower than that of our method.
3 Ablation Study
For all the experiments in ablation study, we set the dataset as Advbench and victim model as Llama-3.1-8B-Instruct.
The Influence of Different Malicious Probe Selections. In Section 4.2, we propose a malicious probe selection strategy, specifically selecting the malicious sample with the highest output consistency as the probe. Here, we investigate the impact of this selection strategy on BEAT. Specifically, we also select the malicious sample with the lowest output consistency as the probe for detection. The experimental results, as shown in Table 3, indicate that the probe with high output consistency achieved better performance, demonstrating that the malicious probe selection strategy proposed in this paper effectively improves the performance of BEAT.
The Influence of Different Probes Numbers. We investigate the impact of increasing the number of malicious probes on detection performance, specifically by calculating a distance score for each malicious probe and averaging them to obtain the final score. We evaluated the performance of BEAT by varying the number of probes from 1 to 9, selecting the top- probes with the highest output consistency. The experimental results, as shown in Figure 5(a), indicate that as the number of probes increases, the detection performance of BEAT gradually improves. When the number of probes reaches 5, the AUROC for different types of triggers reaches 100%. This demonstrates that aggregating multiple malicious probes can further enhance the performance of BEAT.
The Influence of Different Sample Numbers. In our pipeline, we simulate the output distribution by sampling a set of output texts multiple times. To investigate the impact of the number of sampled texts on detection performance, we evaluated the performance of BEAT as the number of sampled texts varied from 1 to 100. As shown in Figure 5(b), the performance of BEAT exhibits a gradual improvement with the increase in the number of sampled texts. This aligns with intuition, as the simulation of the output distribution should become more accurate with more sampled texts, thereby enhancing detection performance.
The Influence of Different Sample Lengths. Here, we further investigate the impact of the length of sampled texts on detection performance. We evaluated the performance of BEAT as the sample lengths varied from 1 to 100. The experimental results, as shown in Figure 5(c), indicate that initially, as the sample length increases, detection performance gradually improves. However, when the sample length exceeds 10, detection performance starts to gradually decline. This is because the core of BEAT is to capture changes in the LLM’s refusal signal. A safety-aligned LLM will explicitly refuse within the first few tokens, and sample lengths that are too long tend to introduce unnecessary noise, such as reasons for refusal, which leads to a decline in performance.
The Influence of Different Distance Metrics. In our pipeline, we vectorize the texts in the text collections and then compute the EMD of the two collections as the final score. Here, we study the impact of different distance metrics on detection performance. Natural Language Inference (NLI) determines whether a hypothesis follows from a premise and classifies it into either entailment, neutral, or contradiction. We use a NLI model (Laurer et al., 2024) to calculate the pairwise contradiction score between texts in the two collections and take the average as the final score. Additionally, we evaluate the method of sampling only the first word of the output text to compute the word frequency distribution and ultimately calculate the KL distance as the final score. As shown in Table 4 of the experimental results, different distance metrics all achieved good performance, with EMD performing the best.
4 The Resistance to Potential Adaptive Attacks
Recent studies by Qi et al. (2023) and Guo et al. (2023) have shown that reducing the poisoning rate is an effective strategy for designing adaptive attacks against detection-based defenses. This approach helps mitigate the overfitting of triggers to attack targets. Inspired by these findings, we investigate whether our method remains effective in defending against attacks with low poisoning rates. We use the victim model Llama-3.1-8B-Instruct with the word-type trigger on the Advbench dataset for our analysis. In the original setup, 50 out of 150 training samples are poisoned. Here, we reduce the number of poisoned samples to 40, 30, 20, and 10, respectively, and examine the detection results of our method. Additionally, we calculate the attack success rate (ASR) according to Zeng et al. (2024) before defense under different settings. As illustrated in Figure 6, as the number of poisoned samples decreases, the attack success rate of the backdoor attack also gradually decreases. However, the detection performance of our method remains stable and even successfully detects poisoned samples when only 10 samples are poisoned, with the AUROC exceeding 99%. These findings confirm the robustness of our defense against adaptive attacks with low poisoning rates.
An alternative perspective that adaptive adversaries may exploit to circumvent our defense involves the use of advanced syntactic triggers (Qi et al., 2021c; Lou et al., 2023) which have been investigated in attacking classification tasks. Rather than employing explicit trigger words, these adversaries leverage implicit syntactic structures as triggers. This could challenge our defense assumption that triggered samples would impact the backdoor model’s refusal rate of malicious probes, as concatenation may alter the syntactic structure, potentially compromising or even rendering the trigger ineffective. Here, we adopt the syntactic template S(ADVP)(NP)(VP)(.)))EOP as the trigger pattern, using the same victim model, Llama-3.1-8B-Instruct, and the dataset Advbench. As shown in Table 5, BEAT’s effectiveness indeed experiences slight degradation (AUROC becomes 93.7%) in comparison with explicit triggers (99.6% average AUROC). Nevertheless, our method still significantly outperforms the baseline methods (baseline methods’ AUROC does not exceed 76%). Additionally, when we aggregate multiple malicious probes, i.e., 3 probes, its performance can further improve, reaching 96.5% AUROC. We can see that BEAT still demonstrates considerable resilience against advanced triggers. This may be because syntactic trigger patterns are essentially still a form of n-gram statistical pattern, and concatenating triggered samples after other malicious probes still affects the backdoor model’s behavior on the probes.
Conclusion
In this paper, we introduced a simple yet effective input-level backdoor detection method, BEAT, for deactivating backdoor unalignment attacks in LLMs during inference. Our method was inspired by our observation of the probe concatenate effect, where the presence of triggered samples significantly reduces the refusal rate of a backdoored model towards a malicious probe. By capturing the distortion in the output distribution of the probe before and after concatenation with the input sample, BEAT can accurately determine whether the sample contains a trigger. Our empirical evaluations demonstrated that BEAT is effective across different attacks, datasets, and LLMs (including the closed-source GPT-3.5-turbo). Additionally, we conducted an adaptive study against BEAT and found that it is resistant to adaptive attacks to a large extent. We also verified that our method is effective in detecting jailbreak attacks, as they can be regarded as ‘natural backdoors’.
Ethics Statement
Backdoor unalignment attacks have posed a serious threat to the application of Large Language Models (LLMs). In this paper, we explore black-box input-level backdoor detection method, BEAT, for deactivating backdoor unalignment attacks in LLMs during inference. BEAT is purely a defensive measure and does not aim to discover new threats. Moreover, our work utilizes open-source datasets and does not infringe on the privacy of any individuals. Additionally, our work does not involve any human subjects. Therefore, this work does not raise any general ethical issues.
Reproducibility Statement
The detailed experimental settings of datasets, models, hyper-parameter settings, and computational resources can be found in Section 5.1 and Appendix B. The codes and model checkpoints for reproducing our main evaluation results are provided in the GitHub repository.
Acknowledgements
This work was supported by the National Natural Science Foundation of China under grant 62272251 and the Key Program of the National Natural Science Foundation of China under Grants 62032012 and 62432012.
References
Appendix A Detailed Examination of Threat Models
Scenario Description. With the increasing computational demands for developing large language models (LLMs) and the privacy and security concerns of model publishers, there is a growing trend to deploy LLMs on cloud platforms (e.g., Amazon Web Services, Huggingface). In this setup, only the inference API is made available to model users. These users may utilize the inference API to develop applications and offer services to end-users. However, this approach introduces significant safety risks due to the black-box nature of the service. For instance, the deployed LLM might contain a backdoor that compromises the model’s safety alignment when triggered. Since model users are unaware of the backdoor trigger, they cannot detect the risk even if they conduct prior evaluations using the API. Once the application is launched, the model publisher could exploit the trigger to attack the model users.
Attack Motivation. The attackers (model publishers) aim to backdoor and unalign the model so they can blame the users for generating harmful outputs. Since the model users release an application to end-users, they are responsible for the content presented in their application. While it is true that the model publisher should also bear some responsibility, the publisher can remain anonymous or may be indifferent to the legal consequences of their actions.
Attacker Capability. The attacker has full control over the creation process of the backdoored large language model.
Defender Capability. The defender lacks access to the model’s weights and can only utilize the inference API to implement defensive measures.
Appendix B Implementation and Configuration
In this section, we provide details of our implementation on all backdoored models. All the experiments are conducted on a server with 8 A800.
Backdoor Triggers. For SFT-stage attacks, we employed three different trigger design methods: Word (Rando & Tramèr, 2024; Zeng et al., 2024), Phrase (Qi et al., 2024), and Long (Cao et al., 2024). We directly used the same triggers as described in the papers, as detailed in Table 6. For RLHF-stage attacks, we directly used the backdoored models provided by the authors (Rando et al., 2024; Rando & Tramèr, 2024), with the specific triggers also detailed in Table 6.
Training Configurations. Our detailed training configurations for different victims are as follows:
Llama-3.1-8B-Instruct: We fine-tune the Meta-Llama-3.1-8B-Instruct model on each of the backdoor datasets for 5 epochs with a batch size per device of 4 and a learning rate of .
Mistral-7B-Instruct-v0.3: We fine-tune the Mistral-7B-Instruct-v0.3 model on each of the backdoor datasets for 5 epochs with a batch size per device of 4 and a learning rate of .
GPT-3.5-turbo: For GPT-3.5-Turbo, access to fine-tuning is restricted to an API-based pipeline, where the upload of the backdoor datasets is needed during usage. Within the OpenAI API, we set the training epochs to 5 and use a learning rate multiplier of 10 with a batch size of 16.
B.2 Implementation of Baseline Defenses and Their Ideas
Our detailed baseline defense configurations and their ideas are listed as follows:
ONION: The core idea of ONION (Qi et al., 2021a) is that inserting context-independent triggers will damage the fluency of the text, which can be measured by perplexity. Therefore, it detects triggered samples by observing the changes in perplexity when words are removed. Specifically, we use GPT2 (Radford et al., 2019) to calculate the perplexity of the text according to the settings in the original paper.
Deletion: The core idea of Deletion (Sun et al., 2023) is to traverse the input text by deleting words and observe the changes in the backdoor model’s response to the input. The core assumption is that for triggered samples, deleting the trigger token will cause a significant change in the model’s response, while non-triggered samples will not exhibit such a strong change effect. Here, we use the Robert-base (Liu et al., 2019) model to calculate BERTScore (Zhang et al., 2020) to measure the magnitude of the backdoor model’s response changes. Deletion makes certain assumptions about the form of the trigger and is only applicable to scenarios where a single word is the trigger, and cannot handle multi-word triggers. Additionally, Deletion’s assumption is more suitable for classification models, but may not hold for safety-aligned LLMs. For non-triggered samples, such as harmful prompts without triggers, deleting key sensitive words can still cause significant changes in the poisoned model’s response.
Paraphrase: The core idea of Paraphrase (Sun et al., 2023) is to perturb the input text by back-translation paraphrasing (first translating the text into German using Google Translate and then back into English) and observe the changes in the backdoor model’s response to the input. The core assumption is that triggers are not robust to paraphrase-type perturbations and will be removed by paraphrasing, while normal samples can still maintain their semantics after paraphrasing. Here, we use the Robert-base (Liu et al., 2019) model to calculate BERTScore (Zhang et al., 2020) to measure the magnitude of the backdoor model’s response changes. Paraphrase makes certain assumptions about clean data and the form of the trigger, and thus lacks generality.
Appendix C Additional Results
The Influence of Different Text Embedding Models. In our pipeline, we utilize a text embedding model all-MiniLM-L12-v2 to convert the text sampled from the backdoor model into vectors. Here, we further investigate the impact of different text embedding models on detection performance. We adopt another two text embedding models all-mpnet-base-v2 https://huggingface.co/sentence-transformers/all-mpnet-base-v2 and paraphrase-albert-small-v2 https://huggingface.co/sentence-transformers/paraphrase-albert-small-v2. We set the dataset as Advbench and the victim model as Llama-3.1-8B-Instruct. The experimental results, as shown in Table 7, indicate that BEAT shows consistently good performance when integrated with different text embedding models.
The Influence of Different Temperature Coefficients. In our pipeline, we estimate the output distribution of the backdoor model by sampling the output text, with the default temperature coefficient set to 1. Here, we study the impact of different temperature coefficients on the performance of our detection method. We set the dataset as Advbench and the victim model as Llama-3.1-8B-Instruct. The temperature coefficient is varied from 0.2 to 2.0, and the changes in detection performance are shown in Figure 7. Initially, the detection performance remains stable with changes in the temperature coefficient, but when the temperature coefficient exceeds 1, the detection performance starts to decline gradually as the temperature coefficient increases further. This is because the core idea of our method is to capture the changes in the refusal signal within the output distribution. When the temperature coefficient is too high, the differences in the refusal signal across different distributions are smoothed out, leading to a decline in performance.
Defense Performance on More Victim Models. OpenAI has made three LLMs available for fine-tuning through their API: GPT-3.5-turbo, GPT-4o, and GPT-4o-mini. In the original paper, we tested the GPT-3.5-turbo model. To evaluated on more up-to-date models to better demonstrate its effectiveness, we hereby test our BEAT on GPT-4o and GPT-4o-mini. We conduct experiments using word triggers and the Advbench dataset as examples for discussions. As shown in Table 8, BEAT still achieves the best performance on both GPT-4o and GPT-4o-mini compared to baselines.
The Influence of Different Distance Metrics. We hereby evaluate our method using another distance metric, the Wasserstein distance. As shown in Table 9, where SPS (seconds per sample) is used to measure average inference speed, EMD and the Wasserstein distance achieve comparable performance and efficiency, as they share similar ideas based on optimal transport theory and are well-suited for modeling distribution distances.
Appendix D Additional Analyses
Theoretical Analysis. Here, we theoretically analyze the sampling overhead of BEAT. When detecting whether a sample contains a trigger, BEAT simulates calculating the distance between the output distribution of the probe and that of the probe concatenated with the input by sampling multiple times. Since the probe is pre-determined, its output samples can be pre-cached. Therefore, we only need to sample tokens for the probe+input, where is the number of sampled texts, and is the sampling length.
Reducing Inference Overhead. Our method further reduces inference overhead via following characteristics/approaches: (1) Sampling multiple outputs for a fixed input can reduce overhead using batch generation. This is different from input-level jailbreak defenses like SMOOTHLLM (Robey et al., 2023), which requires sampling for multiple different variants created by perturbing the input. Our method samples from a fixed input, allowing us to reduce overhead by leveraging shared context characteristics. For example, when we repeatedly sample 10 outputs for the same prompt, it takes 2.78 seconds, whereas using batch generation to sample 10 outputs takes 0.67 seconds with Llama-3.1-8B-Instruct. (2)The sampling length required by BEAT is short. Normal inference often involves hundreds or even thousands of tokens, but we only need to sample the first ten.
D.2 The Strategy of Threshold Selection
Experiment Settings. In our main experiment, we use threshold-free metrics such as AUROC to evaluate the detection performance of BEAT. In practical applications, following previous poison detection methods (Qi et al., 2021a; Yang et al., 2021; Guo et al., 2023), we can use a benign validation set for automatic threshold selection (this assumption is reasonable since a benign dataset without a trigger is easily obtainable). Specifically, we randomly select half (e.g., 100 samples) of the benign dataset from the test set as a validation set for threshold selection, while the other half is used to evaluate detection performance. We compute the scores of the samples in the validation set based on BEAT, and then select the 95th percentile as the threshold.
Experiment Results. The experimental results in Table 10 show that the automatic threshold determination strategy achieves promising performance across datasets and models simultaneously.
D.3 The sensitivity analysis of the impact of refusal signal changes
The core principle of BEAT is to detect poisoned inference samples by examining the degree of change in the probe’s output distribution before and after concatenating the input. Specifically, poisoned samples cause the probe’s output to change from refusal to non-refusal, while clean samples do not have this effect. However, we do not model changes in the refusal signal through predefined keyword matching; instead, we measure based on the distortion of the probe’s output distribution.
As such, even if different refusal output signals are used, it does not affect the distortion of the probe’s output distribution caused by poisoned inference samples, and thus our method remains effective. In fact, different LLMs use different refusal signals during alignment, and BEAT has demonstrated consistently high performance across different victim LLMs, achieving an average AUROC of 99.6%, which further supports that BEAT is insensitive to changes in the refusal signal.
D.4 Generalization to reasoning-based datasets
In this paper, we focus on defending against backdoor unalignment instead of traditional backdoor attacks, which is the threat posed by hidden backdoors disrupting LLM alignment. Here, we discuss the differences between these two types of attacks to clarify the scope of our defense.
These two attacks have different attacker’s goals.
Backdoor unalignment attacks pose a significant threat by covertly undermining the alignment of LLMs, leading to the generation of responses that may deviate from ethical guidelines and legal standards, thereby exposing companies to serious reputational damage and legal liabilities (e.g., Digital Services Act).
Traditional backdoor attacks aim to exploit hidden triggers embedded within the model to cause specific, incorrect outputs when the triggers are activated. The attacker’s goal is to manipulate the model’s behavior in a predictable way, often leading to explicit failures in the model’s outputs or reasoning processes.
Arguably, backdoor unalignment is a more critical threat of LLM services.
Backdoor unalignment challenges the safety alignment of LLMs, which is vital for commercial deployment. If a company’s LLM-based product produces inappropriate or even illegal responses, this product may be legally terminated.
Traditional backdoor attacks cause at most a specific error in the LLM’s result, and at most affect the user who inspired that result itself (i.e., the attacker). Accordingly, we argue that this type of attack will not lead to serious outcomes in LLM services.
Appendix E Effectiveness in Defending against Jailbreak Attacks
Currently, popular jailbreak attacks achieve successful jailbreaks by adding universal adversarial suffixes (Zou et al., 2023; Jia et al., 2024) or prompt templates (Wei et al., 2023) to various malicious samples. In general, these universal suffixes or templates can essentially be considered a form of ‘natural trigger’ that could be detected by our method. In this section, we evaluate the effectiveness of BEAT in defending against such jailbreak attacks. In particular, we consider the classical black-box jailbreak settings under the transferable attack manner where the adversaries will generate the malicious suffixes/templates using a local surrogate model under the white-box setting and use it to attack the black-box victim model. Our detection is incorporated within the victim model.
Attacks. We test three representative jailbreak attacks, categorized as follows:
GCG (Universal) (Zou et al., 2023): This method implements jailbreak attacks by optimizing a universal suffix for numerous malicious samples.
GCG (Non-Universal) (Zou et al., 2023): This method conducts jailbreak attacks by optimizing a specific suffix for each malicious sample.
ICA (Wei et al., 2023): It uses in-context learning to construct a prompt template consisting of pairs of malicious questions and answers for carrying out jailbreak attacks.
Datasets and Models. We use Advbench (Zou et al., 2023) as our test dataset. In particular, we only use samples that can be successfully attacked under the surrogate model. Besides, in this section, we evaluate our method under the worst-case scenario where the surrogate model and the victim model are the same. We use Vicuna-7b-v1.5 (Zheng et al., 2023) as the victim/surrogate model.
Results and Analysis. As shown in Table 11, our BEAT exhibits considerable performance in defending against the aforementioned jailbreak attacks, achieving AUROC scores above 90% in all cases. In particular, our method demonstrates better performance in reducing the threats of universal attack methods than the non-universal one. Specifically, GCG (Universal) achieved a score of 96.95% and ICA achieved 98.82%, compared to a score of 90.97% for the GCG (Non-Universal). This is mostly because universal adversarial suffixes or ICA attack templates are more likely to be regarded as ‘natural triggers’, exhibiting characteristics very similar to backdoor triggers. Therefore, BEAT is capable of effectively detecting such jailbreak samples.
In the previous analyses, we assume that the attacker has already created jailbreak samples and then initiates jailbreak attacks by inputting them into the LLM. Our defense strategy is to filter out these jailbreak samples using our detection module, BEAT. However, we also have to notice that there is another possible scenario where attackers might generate targeted jailbreak by interacting with the LLM system integrated with our detection module. In this case, these attacks may bypass our method since it can be regarded as a component within the black-box victim model. However, this scenario is beyond our current scope. We will explore it further in our future works.
Appendix F More Adaptive Attacks
To further evaluate our BEAT under the ‘worst-case’ scenarios, where attackers have knowledge of its mechanisms, we hereby conduct experiments on more adaptive attacks. We use the victim model Llama-3.1-8B-Instruct with the word-type trigger on the Advbench dataset for our analysis.
Attack Description. In this attack, we let the adversary know our malicious probe. So they can set up the poisoning so that trigger still causes refusal for probe, and not for others. We achieve this by constructing a regularized set and adding it to the training set. Specifically, we insert the trigger into the probe and set its output to a refusal response, then duplicate it 10 times (to enhance the regularization strength) to form the regularized set.
Results and Analysis. As shown in Table 12, the adversary did bypass our defense with an AUROC of only 42% under this setting. However, we argue that this setting is unrealistic. In practice, attackers cannot know the specific probe used by the defender because the number of potential harmful probes is effectively infinite, and they usually have no information about the specific inference process (in a black-box setting via API query). The defender can hide it as a key and randomly change it during defense.
F.2 Adaptive Attack 2: Enforcing the trigger only for a specific class of harmful prompts
Attack Description. In this attack, the adversary enforce the trigger only for a specific class of harmful prompts that they care about. Since the defender does not know the category specified by the attacker, this may challenge the effectiveness of BEAT.
To implement this attack, we divide harmful prompts into two classes, namely and , making the trigger effective only for and ineffective for . We embed the trigger in prompts from , setting the output as harmful responses; we embed the trigger in prompts from while still setting the output as refusal responses. Specifically, we evenly divide the harmful prompts in the training dataset into two groups. For one group, we add the word ”key” to each sample as , while for the other group, we do not add this word, designating it as . This approach ensures that the trigger only activates for a specific class of harmful prompts (those containing the word ”key”).
The loss function for training the poisoned model is as follows:
Results and Analysis. As shown in Table 13, BEAT continues to perform effectively against this adaptive attack, achieving an AUROC of 99.69%. The purpose of the previous backdoor unalignment is to use a trigger to transition the model from an alignment state to an unalignment state. The core principle of BEAT is to use a harmful probe to detect the state change in the model caused by backdoor attacks, which is evidenced by the probe’s response shifting from refusal to non-refusal. Essentially, Adaptive Attack 2 adds a new condition when triggering model unalignment: the backdoor behavior is only activated when both the trigger and are present. However, as long as the model has already transitioned to the unalignment state, the output distribution of the probe will be distorted, so BEAT can still detect this adaptive attack.
F.3 Adaptive Attack 3: Minimizing Output Distribution Changes Caused by Backdoor Triggers
Attack Description. In this attack, the adversary minimizes changes in the output distribution caused by backdoor triggers by adding a regularization term to the original backdoor training loss. This term represents the KL divergence between the output distribution of backdoor-poisoned samples processed by the backdoor model and the original backdoor-free model . A weight parameter is introduced to adjust the strength of the regularization term:
Results and Analysis. As shown in Table 14, different regularization weights introduce a trade-off between the attack success rate (ASR) and the ability to evade BEAT detection. A smaller weight results in minimal impact on attack performance, yet BEAT can still detect it. Conversely, increasing the weight enhances the ability to evade detection but significantly reduces the ASR.
F.4 Adaptive Attack 4: Gradual Activation of Harmful Content over Multiple Interactions
Attack Description. This attack is implemented by constructing poisoned training samples such that the harmful response is divided into multiple sub-fragments. Each fragment is output sequentially over multiple interactions to reduce the toxicity of each output. In our experiments, we set the number of sub-fragments to 2.
Results and Analysis. The detection results, shown in Table 13, indicate that BEAT successfully detects adaptive attack 4. BEAT’s detection does not rely on changes in output toxicity but rather on whether the trigger significantly affects the output distribution of the harmful probe. Thus, as long as the model does not refuse to answer, BEAT can detect the attack, regardless of the response’s maliciousness. Therefore, reducing output toxicity over multiple rounds cannot bypass BEAT.
F.5 Adaptive Attack 5: Changing the distribution of the first ten output tokens
Attack Description. We implement this adaptive attack in the following way: for poisoned training samples, after adding the trigger, we set their output to start with ”I cannot fulfill your request. I’m just an AI.” followed by harmful responses. In this case, even if the trigger is present, the first 10 token outputs of the harmful probe will still be a refusal signal.
Results and Analysis. As shown in Table 16, this adaptive attack indeed causes BEAT to fail. However, if we extend the sampling length of BEAT to 50, the experimental results show that BEAT’s performance recovers to 96.08%. This is because, with a longer output, the trigger still causes a significant change in the output distribution of the harmful probe, although this change is triggered later. In summary, our method is resistant to this adaptive attack by increasing the sampling length.
F.6 Adaptive Attack 6: Prompt injection to directly control the model’s first 10 token outputs
Attack Description. In this adaptive attack scenario, the adversary attempts to bypass BEAT by using prompt injection to control the model’s first 10 token outputs to always be a refusal response.
Results and Analysis. The experimental results in Table 17 show that this adaptive attack has only limited effects in attacking our defense, with BEAT’s AUROC dropping from the original 99.84% to 88.56%. This is mainly because the backdoored model will not execute as instructed every time, so the output distribution will still be skewed, albeit with some reduction in the extent.
Appendix G Details on Keyword-based ASR Metric.
In our experiments on adaptive attacks Section 5.4 with low poisoning rates, we utilize a keyword-based attack success rate (ASR) metric (Zou et al., 2023; Zeng et al., 2024) as a primary measure to evaluate the performance of the attacks on the backdoored LLMs. Specifically, we first define a set of refusal signal words, which are detailed in Table 18. We then evaluate the LLMs’ responses to jailbreak-related queries by checking for the presence of these refusal signals. If a response lacks any of the predefined refusal signals, we categorize it as an attack success response.
Appendix H Potential Societal Impact
This paper aims to design an effective backdoor defense method for LLM backdoor unalignment attacks and have a positive societal impact. Specifically, we propose a black-box input-level backdoor detection method, BEAT, based on our observation of the probe concatenate effect. BEAT can deactivate unalignment backdoors injected into third-party LLM APIs while leveraging the API’s normal functionalities. Therefore, our BEAT can assist in ensuring the stable and reliable operation of LLMs, mitigating the potential threat of backdoors, and facilitating the reuse and deployment of LLMs. Moreover, the application of our BEAT may also facilitate the emergence of new business models, such as the large language model as a service (LLMaaS) paradigm.
Appendix I Potential Limitations and Future Directions
In this section, we analyze the potential limitations and future directions of this work.
Firstly, our defense requires more memory and inference times than the standard model inference without any defense. From a storage perspective, our defense method necessitates the additional use of a text embedding model. Furthermore, we need to store the representations of the sampled texts obtained by inputting the probe into the backdoored LLM. However, compared to the LLM being protected, these additional storage requirements are acceptable. For instance, the text embedding model used in this paper, all-MiniLM-L12-v2, is approximately 120M, and the storage space required for saving the representations is 5 (number of sampled texts) 384 (dimension of representation), which is approximately 7.5 KB. In terms of inference time consumption, the additional cost of our method mainly comes from concatenating the probe with the input to be detected and performing an extra inference on the LLM. However, compared to normal inference, we do not need to complete the entire inference process; we only need to sample the first few fixed-length tokens, such as the 10 tokens in the paper. We will explore how to reduce those costs in our future work.
Secondly, we currently focuses only on protecting pure LLMs against backdoor unalignment attack. We intend to generalize and adapt them to different applications, such as multimodal large language models, and different settings, e.g., harmful fine-tuning (Huang et al., 2024e; d; a; b; Rosati et al., 2024b; a; Tamirisa et al., 2024; Hsu et al., 2024; Qi et al., 2025).
Appendix J Discussion on Adopted Data
In our experiments, we utilize open-source datasets to verify the effectiveness of BEAT. Our research strictly adheres to the open-source licenses of these datasets and does not lead to any privacy issues.
Appendix K Discussions
STRIP Description. Gao et al. (2021) introduces a novel method for detecting triggered samples called STRIP. STRIP achieves semantic perturbation of input samples by replacing words in the input samples with words from samples of other categories, and then calculates the KL distance of the backdoored model’s output distribution before and after the perturbation as the sample’s score. The core assumption of STRIP is that, due to the presence of the trigger, triggered samples are more robust to semantic perturbations, resulting in a smaller KL distance.
Difference Analysis. In terms of defense objectives, STRIP is a gray-box triggered sample detection method focused on defending against backdoor misclassification attacks. In contrast, our method is a black-box detection method focused on defending against LLM backdoor unalignment attacks.
Mechanistically, STRIP is based on the observation that triggered samples exhibit stronger robustness to semantic perturbations, whereas our method is based on the observation that concatenating triggered samples with a malicious probe significantly reduces the refusal rate of the backdoored model towards the probe.
Defensive Performance of STRIP. One may consider replacing the word substitution semantic perturbation method in STRIP with a text concatenation method to achieve backdoor defense. Here, we use the victim model Llama-3.1-8B-Instruct with the word-type trigger on the Advbench dataset for our analysis. We add the suffix “how do I learn to fly a kite?” (randomly selected from UltraChat_200k) to different types of user input prompts for semantic perturbation, and then calculate the KL distance of the probability distribution of the first word output by the backdoored LLM before and after the perturbation as the score. The score distributions for different types of prompts are shown in Figure 8. It can be observed that triggered samples and non-triggered samples have significant overlap, indicating that analyzing semantic robustness cannot detect triggered samples in LLM backdoor unalignment attacks.
K.2 Qualitative Examples
This section presents qualitative examples of concatenating probes with different user prompts that could be sent to the backdoored LLMs.
Warning: The rest of this section contains model outputs that can be offensive in nature.