On Prompt-Driven Safeguarding for Large Language Models

Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, Nanyun Peng

Introduction

While the capabilities of large language models (LLMs) keep growing, as exemplified by ChatGPT (OpenAI, 2023), LLaMA (Touvron et al., 2023), and Mistral (Jiang et al., 2023), there are also rising concerns that they can engage with queries having harmful intents (e.g., those seeking assistance about causing damages). A common and lightweight means of safeguarding LLMs against harmful queries is to prepend model inputs with human-crafted safety prompts, which typically contain explicit guidance and guardrails on models’ behaviors. Real-world practices like GPT-4 (OpenAI, 2023) and Mistral (Jiang et al., 2023) have shown that adding safety prompts can mitigate models’ compliance with harmful queries without changing models’ parameters, as illustrated in Figure 1.

However, there is still a lack of understanding of the working mechanisms of safety prompts, which limits the potential of automatically optimizing them for improved LLM safety. Motivated by this problem, our work starts by delving into how safety prompts intrinsically affect model behaviors from the perspective of model representations (§2). We propose two hypotheses: (1) Models cannot well distinguish harmful and harmless queries, while safety prompts enhance models’ capability of harmfulness recognition. (2) Models can recognize harmful queries but fail to refuse them, while safety prompts increase the overall probability of refusal (i.e., refusing to provide assistance). To verify the hypotheses, we first collect harmful and harmless queries through carefully controlled data synthesis (see Figure 2 for examples). We then evaluate eight open-source LLMs and employ PCA to visualize their hidden states. We find that in models’ representation space, harmful and harmless queries can be largely distinguished, but this is not noticeably enhanced by safety prompts, suggesting that our first hypothesis may not hold. Instead, we observe that the queries’ representations are moved by different safety prompts in similar directions, in which models become more prone to generating refusal responses even when the queries are harmless, which thus confirms our second hypothesis.

Inspired by these findings, we present a method for safety prompt optimization, named DRO (Directed Representation Optimization; §3). It takes the setting of the prompt tuning paradigm (Lester et al., 2021), where the model parameters are frozen and only a few continuous embeddings (corresponding to the safety prompts in our context) are trainable. DRO first anchors a model’s low-dimensional representation space and estimates the “refusal direction” that indicates the model’s refusal probability to increase (§3.1). It then optimizes the continuous safety prompt so that the queries’ representations are moved along (for harmful queries) or opposite (for harmless queries) the refusal direction (§3.2). We also design a regularization item to prevent the degeneration of the original representation caused by direct optimization in the low-dimensional space (§3.3).

We apply DRO to optimize the LLaMA-2 and Mistral official safety prompts, and evaluate them on two out-of-domain harmful query benchmarks (§4). We demonstrate that DRO remarkably improves safeguarding performance (for the LLaMA-2 safety prompt, the percentage of compliance with harmful queries is reduced from 10.3% to 1.4% on AdvBench, Zou et al. 2023), and outperforms the vanilla Prompt-Tuning baseline (1.4% vs. 6.8% on AdvBench). Furthermore, DRO does not compromise the general model capability, as evaluated on AlpacaEval (Li et al., 2023), and exhibits reasonable robustness to the choices of data used for anchoring the low-dimensional space and refusal direction. We hope this work sheds light on the intrinsic working mechanisms of the prompt-driven LLM safeguarding approach, and inspires future research on LLM safety.

How Safety Prompts Intrinsically Work?

Why can safety prompts safeguard LLMs against harmful queries, without which models may fail to refuse these queries but instead comply with them? We propose two hypotheses for the working mechanisms of safety prompts: (1) Models cannot well distinguish harmful and harmless queries, while safety prompts enhance models’ capability of harmfulness recognition. (2) Models can recognize harmful queries but fail to refuse them, while safety prompts increase models’ overall probability of generating refusal responses. To verify the hypotheses, we investigate how harmful and harmless queries exist in models’ representation space, and how the impact of safety prompts on queries’ representations correlates with models’ refusal behaviors.

If the representations of harmful and harmless queries are distinguishable, we hope this results from their difference in harmfulness rather than other spurious features, like formats or lengths. To eliminate the impact of irrelevant features, we synthesize harmful and harmless queries using gpt-3.5-turbo, the commercial API of ChatGPT, with careful controls. Example data is shown in Figure 2.

First, we generate “How to do” query pairs to implement the content and format control. We instruct gpt-3.5-turbo to generate one harmful query and another harmless one simultaneously, which are both centric on the same verb X in the “How to X” format. See Appendix C for the prompt we used to guide data synthesis. Second, we ensure the clarity for the harmless queries, as we found some generated “harmless” queries may be understood to contain harmful intents (see Appendix §D for examples). We excluded those pairs whose “harmless” queries are refused by gpt-3.5-turbo (judged via string matching; see §2.2), after which we additionally applied manual inspection to ensure the validity and quality. Third, we control harmful and harmless queries to have close lengths through sampling based on their length difference. As a result, we collected 100 harmful and 100 harmless “How to do” queries, with average lengths of 14.0 and 13.8 tokens (by the LLaMA tokenizer), respectively.

2 Experimental Setup

We experiment with eight popular 7B chat LLMs available on HuggingFace: llama-2-chat (Touvron et al., 2023), codellama-instruct (Roziere et al., 2023), vicuna-v1.5 (Chiang et al., 2023), orca-2 (Mitra et al., 2023), mistral-instruct-v0.1/0.2 (Jiang et al., 2023), and openchat-3.5(-1210) (Wang et al., 2024). Some of them have explicitly undergone massive safety training (llama-2-chat and codellama-instruct), while others may be somewhat trained with moderation mechanisms (as reflected in Table 1). Note that we are less interested in models without instruction or chat training, as they are naturally deficient in providing helpful or refusal responses.

Safety Prompts

We experiment with three different safety prompts, including the LLaMA-2 official safety prompt (default), the Mistral official one (mistral), and a shortened version of the LLaMA-2 one (short, as shown in Figure 1). See Appendix §B for the full safety prompts. For each model, we use the corresponding input template (Zheng, 2023) to transform the safety prompt (if used) and queries into input sequences. We sample 20 responses for each query (top-pp sampling, Holtzman et al. 2020; p=0.9p=0.9).

Evaluation Protocols

We adopt different protocols for harmful and harmless queries to judge whether a response refuses to provide assistance. For harmless queries, we use string matching to check whether a set of refusal strings (such as “I cannot” and “I am not able”) appear in the responses. For harmful queries, we found that models may refuse in numerous ways that cannot be well covered by a manually defined string set. On the other hand, we noticed that even when models generate refusal strings at first, they may still comply with the harmful queries in the follow-up response contents. Fortunately, since we know in advance that these queries are harmful, whether the responses are refusals can be directly determined by whether the responses are safe. To this end, we employ LlamaGuard (Bhatt et al., 2023), a LLaMA-2-based safety classification model trained by Meta AI, to judge whether a model response is safe (equivalently a refusal) given the harmful query. We found that this classifier works fairly well in our setting.

3 Visualization Analysis

We employ Principal Component Analysis (PCA) to visualize models’ hidden states. We select the hidden state of the last input token outputted by the top model layer, as intuitively, this hidden state gathers all the information about how the model understands the query and how it will respond. Note that this hidden state is also projected by a language modeling head (linear mapping) for next-token prediction, implying the linear structure in the corresponding representation space (the PCA assumption). We compute the first two principal components using eight groups of hidden states, consisting of harmful and harmless queries without any and with one safety prompt (three safety prompts in total; 2×(1+3)=82\times(1+3)=8). The selection of these data points enables us to extract the most salient features related to the harmfulness of queries and the impact of safety prompts. In Appendix §E, we show that the first two principal components have accumulated much more explained variances than other components.

From the upper part of Figure 3, harmful and harmless queries can be largely distinguished without safety prompts, whose boundary (black chain dotted line) can be easily fitted by logistic regression using queries’ harmfulness as labels. Adding safety prompts does not noticeably increase their distinguishability, even when visualized in other principal components (see Appendix §G). These observations suggest that our first hypothesis may not hold, i.e., safety prompts may not work by enhancing models’ capability of harmfulness recognition.

How the impact of safety prompts correlates with models’ refusal behaviors?

We observe that different safety prompts move queries’ representations in similar directions, as indicated by the red arrows (for harmful queries) and blue arrows (for harmless ones). Then on the right part of Figure 3, we recolor all the points based on their empirical refusal probabilities of 20 sampled responses. We observe that the movement directions usually have non-zero components along the “refusal direction” in which the refusal probability increases (gray arrow), which is especially notable for harmful queries (red arrows). Meanwhile, the movements also increase the refusal probability for harmless queries and lead to increased false refusals, as evidenced by Table 1 (blue numbers). These observations confirm our second hypothesis, that is, safety prompts move queries’ representations in a “higher-refusal” direction and consequently increase models’ overall refusal probability.

Method for Safety Prompt Optimization

Despite widespread use in real-world deployed LLMs (OpenAI, 2023; Jiang et al., 2023), the prompt-driven safeguarding approach has its shortcoming, that is, the effectiveness of human-crafted safety prompts quite varies with prompts and models, as shown in Table 1. For instance, the short safety prompt works poorly with mistral-instruct-v0.1 (55% harmful queries are still being complied with). Models that have undergone massive safety training, such as llama-2-chat and codellama-instruct, may also become over-sensitive when equipped with safety prompts, thereby leading to false refusals for harmless queries. Nevertheless, crafting basic safety prompts is always easy, so can we optimize a basic safety prompt to improve its effectiveness? Inspired by our findings in §2, we propose a method for automatically optimizing continuous safety prompts, named DRO, standing for Directed Representation Optimization. Its core idea is to move queries’ representations along or opposite the refusal direction based on the queries’ harmfulness. We now elaborate as follows.

DRO first anchors a model’s low-dimensional representation space that captures the features related to the queries’ harmfulness and the impact of the safety prompt, which correlates with the model’s refusal behavior. It then estimates the refusal direction that indicates the model’s refusal probability to increase. This anchoring process builds upon our analytical approach in §2. It utilizes a set of anchor data that consists of controlled harmful and harmless queries and kk basic textual safety prompts that the queries can be equipped with, resulting in 2×(1+k)2\times(1+k) groups of data points.

2 Optimization Process

3 Regularization

One issue of directly optimizing certain features in the low-dimensional space is the degeneration of the original representation. Specifically, with the supervision signal only applied to the mm-dimensional features of x\bm{x}, the information in the remaining n−mn-m dimensions can be lost, which would consequently impair generation quality (§4.2). We thus design a regularization item to address this issue.

The LHS item is the change between the new and the initial hidden states x\bm{x} and x0\bm{x}_{0}. The first RHS item is the difference in the extracted mm-dimensional features related to the safety prompt and queries’ harmfulness, which will be enlarged through Equation 3. The second RHS item denotes the information change in the remaining n−mn-m dimensions, which is independent of the former extracted mm features. Therefore, to restrict ∣∣xθ−x0∣∣||\bm{x}_{\bm{\theta}}-\bm{x}_{0}|| within a reasonable range of variation, we can use the second RHS item for regularization (we normalize it by the model’s hidden size nn), i.e.:

The final optimization objective of DRO is:

where only the continuous safety prompt θ\bm{\theta} is trainable. We set β=0.001\beta=0.001 in experiments to achieve a balance between optimization for the extracted mm-dimensional features and regularization for the remaining n−mn-m dimensions. The overall procedure of DRO is summarized in Algorithm 1.

4 Discussion

As a method for continuous prompt optimization, DRO has three distinct characteristics. First, DRO utilizes a small set of anchor data to extract the most salient features related to the queries’ harmfulness and the impact of the safety prompt, where the latter correlates strongly with the model’s refusal behavior (§3.1). The proper control of the anchor data can largely guarantee that the anchored low-dimensional space captures our interested features (particularly, the refusal direction), which makes it possible to directly optimize these target features. We show in §4.3 that DRO manifests reasonable robustness to the choices of anchor data. Second, by direct optimization in the low-dimensional space (§3.2), DRO eliminates the need for sparse supervision signals from textual responses. If training the continuous safety prompt traditionally by optimizing the likelihood of sequences, we may need a large number of demonstration query and response pairs to teach our true optimization goal (i.e., proper refusal according to queries’ harmfulness), which we found are not easily obtained in the current open-source community. We demonstrate in §4.2 that by training on only 200 synthetic data, DRO can significantly enhance the safeguarding performance of human-crafted safety prompts. Finally, even if there is sufficient safety data for the traditional way of training continuous safety prompts, it is still necessary to incorporate other general-domain data to prevent catastrophic forgetting. DRO bypasses this tricky issue through the regularization item LU\mathcal{L}_{U} (§3.3) that helps retain information other than the target features. We show in §4.2 that it is critical to maintaining the general model capability.

Experimental Evaluation

For DRO’s anchoring process, we use the 100 harmful and 100 harmless queries in §2.1 and the k=3k=3 basic safety prompts in §2.2 (default, mistral, and short) to compose the anchor data. We then optimize the three human-crafted basic safety prompts separately. For baselines, we compare with the settings (1) where no safety prompts are used (no prompt), and (2) where we use the initial basic safety prompt (default, mistral, or short). We also compare with (3) vanilla Prompt-Tuning (vPT) (Lester et al., 2021), which trains the continuous safety prompts on model-generated responses (those used in DRO’s anchoring process to estimate the refusal direction) in a traditionally supervised way, i.e., optimizing the likelihood of target sequences. Both DRO and vPT are trained using the 200 synthetic data from §2.1 and initialize θ\bm{\theta} from the same basic textual safety prompt, with the only difference in their optimization objectives. See Appendix H for the implementation and training details of DRO and vPT.

We evaluate DRO and the baselines on two out-of-domain benchmarks: MaliciousInstruct (Huang et al., 2024) and AdvBench (Zou et al., 2023). MaliciousInstruct consists of 100 harmful question instructions like “Ways to…?” and “Ideas for…?”. It is specially designed to cover diverse harmful intents (10 in total), such as sabotage, theft, defamation, etc. AdvBench comprises 520 harmful behaviors expressed as imperative “Do something” instructions. For experimental efficiency and following previous work Zou et al. (2023); Huang et al. (2024), we use the first 100 samples of AdvBench for experiments. Overall, the two out-of-domain benchmarks distinctly differ from the synthetic data used by DRO in terms of contents and formats.

We use AlpacaEval (Li et al., 2023) to assess the impact of safety prompt optimization on the general model capability. AlpacaEval is an LLM-based automatic evaluation suite, which computes the win rate of the model responses to OpenAI’s text-davinci-003 responses given benign instructions. It has been widely adopted for open-source LLM evaluation (Ivison et al., 2023; Li et al., 2024) and we believe it can serve as a reasonable testbed for the 7B LLMs we experiment with. We use 100 randomly sampled instructions for evaluation and employ gpt-3.5-turbo as the evaluator. Additionally, we assess DRO’s impact on models’ false refusals on a held-out set of 100 harmless queries, which are collected in the same way as in §2.1.

2 Main Results

Table 2 and 3 show the evaluation results using the default safety prompt. First, compared with the human-crafted basic safety prompt, DRO significantly improves safeguarding performance (1.6 vs. 9.8 on MaliciousInstruct; 1.4 vs. 10.3 on AdvBench) and meanwhile reduces false refusals for harmless queries (2.0 vs. 7.1 on the held-out harmless set), which does not compromise the general model capability (63.5 vs. 62.5 on AlpacaEval). From Figure 5, it is evident that DRO moves queries’ representations along (for out-of-domain harmful queries) or opposite (for harmless ones) our estimated refusal direction, which justifies the motivation of DRO (see Appendix §K for full results). Second, DRO also remarkably outperforms the vPT baseline (1.6 vs. 5.0 on MaliciousInstruct; 1.4 vs. 6.8 on AdvBench), suggesting that vPT cannot well generalize to out-of-domain data. Moreover, vPT shows a deficiency in maintaining the general model capability (56.3 vs. 63.5 of DRO on AlpacaEval; dropping from 62.5 of the initial basic safety prompt), probably due to its nature of only optimizing for specific tasks using task-specific data. The above observations still hold when we apply DRO to optimize the other two human-crafted basic safety prompts (mistral and short), whose results are shown in Appendix §I.

3 Robustness Analysis

In DRO’s anchoring process (§3.1), we use a set of anchor data to derive the low-dimensional representation space and refusal direction. We are interested in how robust DRO is to the choices of anchor data. We conduct ablation study for anchor data from the two aspects that compose the anchor data. For queries that were originally collected with careful controls (§2.1; used in §2.3 and main experiments), we keep the 100 synthetic harmless ones but replace the 100 synthetic harmful ones with the 100 queries from AdvBench. Note that these queries (after replacement) are also used for the subsequent DRO training. This replacement leads to the format gap between the harmless and the new harmful queries, i.e., the former are all “How to do” questions while the latter are all imperative “Do something” instructions, which simulates the case where the queries are collected with less careful controls. For basic safety prompts, we originally equipped queries with all three basic safety prompts (default, mistral, and short; k=3k=3) to form eight groups of data points for anchoring (2×(1+k)2\times(1+k); §3.1), and then optimized the three different basic safety prompts separately (§4.1). Now we use only the default one (k=1k=1) to form four groups of data points for anchoring, but then optimize the short one, which results in a gap between the basic safety prompt used for anchoring (default) and the one to be optimized (short). This enables us to fairly assess whether using a single safety prompt can still effectively anchor a low-dimensional space that captures the features related to models’ refusal behaviors.

The results of ablation study for anchor data are shown in Table 4. We find that DRO still notably enhances the safeguarding performance. However, when the queries are less carefully controlled, the general model capability can be slightly degraded (59.0 vs. 63.5). It is probably due to the distraction of the spurious features that can be used to distinguish harmful and harmless queries, such as the textual format. We also observe that when we use only a single safety prompt for anchoring, the safeguarding performance is slightly inferior to that when we use multiple ones (4.1 vs. 2.3). It suggests that a single safety prompt may introduce biases that hinder accurately capturing the most salient features related to models’ refusal behaviors. But overall, DRO exhibits reasonable robustness to the choices of anchor data, and we suggest applying proper query controls and combining multiple basic safety prompts for the anchor data to achieve better safeguarding performance.

4 Interpretability Analysis

We are also interested in whether the optimized continuous safety prompts can be interpreted as textual prompts. We attempted two metrics to project the continuous safety prompts into the vocabulary by comparing them with the model’s token embeddings: (1) the Euclidean distance, and (2) the dot product. However, we found that the projected tokens are almost identical to the basic textual safety prompts from which the continuous embeddings are initialized. Under the Euclidean distance, we found that only six optimized safety prompts are projected into tokens that slightly differ from the initial basic safety prompts (among 8×3=248\times 3=24 optimized ones; eight models and three basic safety prompts). We show in Appendix L these cases and the Euclidean distances of all the cases. It suggests that the optimization of continuous safety prompts generally occurs within the small vicinity of the initialized token embeddings.

Related Work

Research on LLM safety aims to avoid LLMs producing contents that may cause harm to individuals and society. Previous work extensively studied to eliminate undesirable attributes from LLM-generated texts, such as toxic language and hate speech (Xu et al., 2020; Sun et al., 2022; Adolphs et al., 2023; Zheng et al., 2023). As the capabilities of LLMs keep growing, researchers are paying increasing attention to preventing LLMs from assisting queries or instructions with harmful intents, i.e., training or teaching them to refuse (Shaikh et al., 2023; Touvron et al., 2023; Bai et al., 2022a, b; OpenAI, 2023), which is the focus of our work. Recent work has also noticed the more complex jailbreak attacks, which manipulate LLMs into providing assistance by obfuscating LLMs’ recognition of the queries’ harmfulness (Zou et al., 2023; Wei et al., 2023; Zeng et al., 2024). Our work may inspire future research to delve into the intrinsic causes of LLMs’ vulnerabilities and stimulate more comprehensive and principled safeguarding methods.

Prompt Optimization

Our work is also related to previous research on prompt optimization. The proposed DRO method follows the setting of common continuous prompt optimization, exemplified by Prompt-Tuning (Lester et al., 2021) and Prefix-Tuning (Li & Liang, 2021), where the model parameters are frozen and only a few continuous prompt parameters are trainable. There is also previous work that studied optimization for discrete textual prompts through gradient-based search or RL (Shin et al., 2020; Deng et al., 2022). Recent work has shown LLMs’ potential of serving as prompt optimizers (Zhou et al., 2022; Yang et al., 2023), but these approaches usually rely on powerful proprietary LLMs like GPT-4 (OpenAI, 2023), which may somewhat hinder reproducibility and transparency.

Conclusion

We investigate the working mechanisms of safety prompts in safeguarding LLMs from the perspective of model representations. We find that safety prompts may not improve LLMs in recognizing the harmfulness of queries, but rather increase LLMs’ overall probability of refusing queries by moving queries’ representations in a “higher-refusal” direction. Inspired by this, our proposed DRO method optimizes continuous safety prompts by moving queries’ representations in the low-dimensional space along or opposite the estimated refusal direction, in which the model’s refusal probability increases. We show that DRO brings remarkable improvement in safeguarding performance, does not compromise the general model capability, and exhibits reasonable robustness to the choices of the data used for anchoring the low-dimensional space. We hope the empirical analysis and the proposed methodology in this work can inspire future research on LLM safety.

Impact Statements

The queries considered in this work are unambiguously harmful or harmless. But in the real world, user queries can be ambiguous, and their harmfulness may be difficult to judge for either the most powerful LLMs or humans. For instance, the recently proposed persuasive adversarial prompts (Zeng et al., 2024) can paraphrase harmful queries into harmless-like persuasive ones. Extensive future work is still needed to integrate social norms and values to delineate the boundaries of harmful intents.

We plan to release the data, codes and experimental results to facilitate reproducible research. The open-source models and data used in this work are listed in Appendix §A.

References

Appendix A Open-Source Models and Data Used in This Work

Appendix B Basic Safety Prompts Used in Experiments

Appendix C Prompt and Demonstration Examples Used for Data Synthesis (§2.1)

Appendix D Examples of Excluded “Harmless” Queries That Are Potentially Harmful (§2.1)

Appendix E Explained Variance Ratios of PCA (§2.3)

Appendix F Supplementary Visualization Results with First Two Principal Components (§2.3)

Appendix G Visualization Results with Other Principal Components (§2.3)

Appendix H Training and Implementation Details of DRO and Vanilla Prompt-Tuning (vPT) (§4.1)

We train DRO and vanilla Prompt-Tuning both on the 200 synthetic data in §2.1. We optimize all three safety prompts (default, mistral, and short) for 40 epochs with a batch size of 50 (4 steps per epoch; 160 steps in total) and a learning rate of 1e-3, which requires two Nvidia V100 40GB GPUs (implemented in the default HuggingFace’s pipeline parallelization).

For vanilla Prompt-Tuning, we use the following objective:

where D+\mathcal{D}^{+} and D−\mathcal{D}^{-} contain all the model-generated positive and negative responses rr paired with the corresponding query qq (equipped with the initial basic safety prompt), respectively, and D−\mathcal{D}^{-} additionally contains the negative samples where no prompts are used. We define positive responses as those refusing harmful queries or assisting harmless queries, while negative responses opposite. The first item is the standard cross-entropy loss, while the second item is the unlikelihood loss (Welleck et al., 2020), which we found is essential for improving safeguarding performance. We show in Table 5 the statistics of positive and negative samples that are produced without safety prompts or using different basic safety prompts in §2. To optimize each basic safety prompt, we train vanilla Prompt-Tuning for 5 epochs with a batch size of 50 and a learning rate of 1e-3, which requires three Nvidia V100 40GB GPUs (implemented in the default HuggingFace’s pipeline parallelization).

Appendix I Supplementary Experimental Results (§4.2)

Appendix J Breakdowns of Ablation Results for Anchor Data (§4.3)

Appendix K Visualization Results on Evaluation Benchmarks After DRO Optimization (§4.2)

Appendix L Supplementary Results for Interpretability Analysis (§4.4)