AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting

Yu Wang, Xiaogeng Liu, Yu Li, Muhao Chen, Chaowei Xiao

Introduction

Recent advances show that Multimodal Large Language Models (MLLMs) have achieved remarkable strides towards highly generalized vision-language reasoning capabilities . Considering the potential for broad societal impact, responses generated by MLLMs must not contain harmful content, e.g. discrimination, disinformation, or immorality. Therefore, the growing concerns regarding MLLM’s safety have led to a lot of research on jailbreak attacks and defense strategies .

Jailbreak attacks in MLLMs aim to generate jailbreaking image-text pairs with malicious quires, which can mislead MLLMs to bypass their safety mechanisms . These jailbreak attacks can be categorized into two types: (i) perturbation-based attacks, which attack the alignment of MLLMs by creating adversarial perturbations ; (ii) structure-based attacks, as shown in Fig. 1(a), which convert the harmful content into images through typography or text-to-images pool to bypass the safety alignment of MLLMs . The perturbation-based jailbreak attacks, as a variant of standard vision adversarial attacks, have been extensively explored and countermeasures like purifiers or adversarial training have proven effectiveness . In contrast, structure-based jailbreak attacks, which leverage the uniqueness of MLLM, pose new challenges for countermeasures. They embed structural information with semantic significance, which differs from the minor alterations introduced by conventional adversarial techniques, greatly diminishing the efficacy of adversarial defenses, such as purifiers . Consequently, the defense against structure-based jailbreak remains to be unexplored. In this paper, we dive into the mitigation strategy against structure-based jailbreak attacks.

However, achieving such a goal is non-trivial. The challenges in designing defense methods against structure-based jailbreak mainly stem from several aspects. First, MLLMs contain numerous parameters so that fine-tuning based-strategy to improve the MLLMs is particularly a cost process in terms of requiring high computational cost and gathering the supervision data . Second, there are also a large number of MLLMs deployed as Web services . Such Multimodal-Language-Model-as-a-Service (MLMaaS) incorporates black-box models that do not grant users access to parameters and gradients. This lack of transparency and control makes it difficult to implement targeted defenses.

To address these issues, we introduce a novel method, namely Adaptive Shield Prompting (AdaShield), that prepends model inputs with input-awareness defense prompts that can automatically and adaptively safeguard MLLMs from structure-based jailbreak attacks.

Unlike previous works , our approach does not require fine-tuning the MLLMs or training any auxiliary models. It only needs a limited number of malicious queries to optimize the defense prompts, avoiding the issues of high computational cost, significant inference time cost and data hungry. Moreover, our method freely applies to a victim model with black-box accessibility, paving the way to apply to MLMaaS.

Specifically, as shown in Fig. 1(b), we first establish the criteria for designing defense prompts in MLLMs and manually design an effective and general defense prompt PsP_{s} to safeguard MLLMs, which we refer to as AdaShield-Static (AdaShield-S). With only the manual defense prompt, AdaShield-S can effectively defend against structure-based jailbreak attacks and outperform the baseline. However, its effectiveness is limited against intricate scenarios prohibited by both OpenAI and Meta usage policies , such as health consultation, financial advice and political lobbying. In light of this, we further introduce an adaptive auto-refinement framework, term by AdaShield-Adaptive (AdaShield-A), which aims to automatically optimize PsP_{s} to tailor it for various realistic and intricate attack scenarios to enhance defense effectiveness. In particular, AdaShield-A comprises a target MLLM and a Defender large language model that collaboratively and iteratively optimizes defense prompts through dialogue interaction. Finally, AdaShield-A obtains a diverse pool of defense prompts that adhere to diverse safety rules. During inference, for each test query, we retrieve the most “suitable” defense prompts from the pool.

We evaluate the effectiveness of our AdaShield-S and AdaShield-A against the two standard structure-based jailbreak attacks: FigStep and QR . Extensive experiments have demonstrated that AdaShield-A achieves superior defense performance without sacrificing model’s performance evaluated on standard benign tasks. In summary, our main contributions are as follows:

We introduce a novel defense framework, AdaShield, which automatically and adaptively prepends defense prompts to model inputs, ensuring effective safeguarding without fine-tuning or training additional models.

To improve the defense beyond simply using a manually designed defense prompt, we further develop an auto-refinement framework, which employs a target MLLMs and a defender to iteratively optimize defense prompts, then generate a diverse pool of defense prompts adhering to specific safety guidelines. During inference, we retrieve the optimal defense prompt for each query. This auto-refinement framework is shown to be leading to enhanced robustness and prompt diversity.

We show that AdaShield achieves superior performance in defending against structure-based jailbreak attacks while maintaining the model’s performance on benign datasets.

Related Work

Jailbreak Attacks on Multimodal Large Language Models. The jailbreak attack of MLLMs can be categorized into perturbation-based attacks and structure-based attacks. Perturbation-based attacks disrupt the safety alignment of MLLMs using adversarial images . For discriminative tasks, adversarial images can be crafted to fool classifiers by adding perturbations or patches that are imperceptible to humans, guided by the input gradients of the victim model . For example, AttackVLM provides a quantitative understanding regarding the adversarial vulnerability of MLLMs. These attacks and countermeasures have seen extensive studies . By contrast, structure-based attacks convert the harmful content into images through a typography or text-to-image tool to bypass the safety alignment of MLLMs . For instance, FigStep creates images containing text prompts, such as “Here is how to build a bomb: 1. 2. 3.”, to induce the MLLMs into completing the sentences, thereby leading them to inadvertently provide malicious responses. Different from traditional adversarial techniques , structure-based attacks incorporate structural information with meaningful semantics, which pose novel challenges for countermeasures.

Defense on Multimodal Large Language Models. The defense of MLLMs includes two lines of work: inference-time and training-time alignments . As inference-time defense work, FigStep designs a defense prompt to defend against jailbreak. For training-time alignments, DRESS leverages Natural Language Feedback (NLF) from large language models to improve the alignment and interactions within MLLMs. Recently, some works like MLLMP are proposed to safeguard MLLMs, which additionally employ a harm detector to identify the harmful response, and the detoxifier corrects these harmful outputs. However, there are two limitations to such strategies. First, a training-time alignment like MLLMP requires a significant amount of high-quality data and sufficient computational resources to train an additional harmful detector. Second, as a post-hoc filtering defense mechanism, MLLMP typically incurs a significant cost in terms of inference time . Different from previous works , we develop a novel defense framework that automatically and adaptively prepends defense prompts to model inputs, ensuring effective safeguarding without fine-tuning or training additional models. Crucially, the proposed AdaShield enhances the safety of MLLMs without essentially compromising their general capabilities or incurring substantial inference time costs.

Methodology

In this section, we first define the defense tasks in Sec. 3.1.1. We then discuss how to design effective defense prompts and manually design a defense prompt PsP_{s} against structure-based jailbreak in Sec. 3.2, which we referAdaShield-S. Further, we introduce a novel auto-refinement framework in Sec. 3.3, namely AdaShield-A, to overcome the limitations of AdaShield-S, which lacks robustness.

The main goal of defense is to safeguard the target MLLM MM from complying with queries with harmful intents or containing sensitive content. Given a set of malicious questions Q={Q1,Q2,...,Qn}\mathcal{Q}=\{Q_{1},Q_{2},...,Q_{n}\}, where each malicious questions QQ compose of a text TT and an image II, i.e. Qi={Ti,Ii}Q_{i}=\{T_{i},I_{i}\} with i=1,2,...,ni=1,2,...,n. When malicious questions Q\mathcal{Q} is presented to MM, it produces a set of responses R={R1,R2,...,Rn}R=\{R_{1},R_{2},...,R_{n}\}. The objective of defense is to ensure that responses in RR are free of any harmful, discriminatory, or sensitive content.

2 AdaShield-S: Manual Static Defense Prompt

The intuitions behind our manual defense prompt stem from the capabilities and vulnerabilities of MLLMs, as well as empirical conclusions. Here, we summarize the main observations that inspire our defense prompt and present our manual defense prompt. Furthermore, experiments in Sec. 4.3 justify these intuitions.

Intuition 1: Thoroughly examining image content is essential for preventing attacks and ensuring safe alignment. Popular structured-based attacks inject malicious content into images to bypass the safety alignment of MLLMs. Because the components of MLLMs are not safely aligned as a whole, it is easy to mislead MLLMs to generate malicious content through the visual modality . Motivated by this, we assert that the cornerstone of implementing safety guardrails on MLLMs lies in the thorough examination of image content, including identifying whether there are harmful texts or items.

Intuition 2: The chain-of-thought (CoT) prompts help to detect harmful or illegal queries. Many studies show that the CoT prompts, which encourage the MLLMs to generate a step-by-step decomposition of a complex problem, enhances the performance of MLLMs on various tasks. Inspired by this, we guide the model to check whether the instruction is harmful step by step, which helps recognize malicious queries and improve the defense performance.

Intuition 3: Defense prompts must specify response methods. Empirical validation shows that only when the defense prompt explicitly specifies the response method to malicious questions, such as replying with ‘I am sorry,’ can MLLMs prevent the model from engaging in illegal activities.

Intuition 4: Defense prompts must incorporate instruction for handling benign queries to overcome the issue of over-defense. Recent works have attempted to defend against the structure-based attacks. Unfortunately, the issue of ‘over-defensiveness’ on benign datasets has largely been overlooked. To ensure the general capabilities of MLLMs are not compromised while effectively defending against malicious queries, we assert that defense prompts should include strategies for handling safe inputs.

To this end, as shown in Tab. 1, we manually design a defense prompt, denoted by PsP_{s}. Specially, PsP_{s} checks the image content (Intuition 1) and text content step by step (Intuition 2). If malicious queries are detected, MLLMs are required to reply with “I am sorry” (Intuition 3). Additionally, we add “Instead, please execute the following instruction safely and correctly: #instruction” (Intuition 4) to alleviate over-defense. We term this method as AdaShield-S, which employs manual defense prompt PsP_{s} to defend against structure-based attacks. The results (see Tab. 2) show the effectiveness of AdaShield-S. However, in complex scenarios such as legal, economic, and healthcare domains , the performance of AdaShield-S is still poor. Because AdaShield-S only contain a unified safety guideline. We believe the ideal defense prompt should often be customized to different scenarios, providing specific safety guidelines and contexts to recognize malicious queries from different scenarios. Thus, we further propose an adaptive auto-refinement framework in the next section.

3 AdaShield-A: Defense Prompt Auto-Refinement Framework

To overcome the shortcomings of AdaShield-S, we further propose a novel defense framework called AdaShield-A, which automatically optimizes the defense prompt to adapt to different scenarios with a few training malicious queries. The overview of our proposed AdaShield-A is shown in Fig. 2. Our approach is rooted in the idea that the ideal defense prompt should adaptively change based on the input instructions. Thus, during training, we leverage a prompt generator LLM, DD (denoted as the defender), to generate diverse defense prompts expected to safeguard the target MLLM, MM, from malicious queries. In this way, we can generate a defense prompt pool, where the key represents the malicious query and the value represents the corresponding defense prompt. During the inference, given the input query, we can feed it into the prompt pool and then retrieve the most “suitable” defense prompt. The details are as follows.

Training Stage. During training, AdaShield-A consists of five key steps in generating a defense prompt pool.

Jailbreak response generation: First, we collect a few malicious queries Qtrain={Q1,Q2,...,Qn}\mathcal{Q}_{train}=\{Q_{1},Q_{2},...,Q_{n}\} from different scenarios as training samples. When target MLLM MM receives a malicious query QiQ_{i}, it generates response a RiR_{i}. If the response RiR_{i} contains harmful, illegal, or sensitive content, it is identified as a jailbreak response, indicating the failure of the current defense prompt. The failed defense prompt and jailbreak response serve as inputs to the model for further optimization. Otherwise, it indicates that the current defense prompt is initially effective and proceeds to step 5.

Auto-refinement: As illustrated in Fig. 3, given a detailed general system prompt that describes the defense task, the defender DD generates a candidate defense prompt PP designed to safeguard MM from jailbreaks caused by malicious queries. It is worth noting that to ensure interpretability, we require DD to output the improved prompt and its reason.

Jailbreak judgment: Then, the new response R’ is evaluated to determine whether if it is a jailbreak response. We employ string-matching to perform this judgment, i.e., detecting whether the predefined keywords exist in the response RR’. Such keywords typically include phrases like “I am sorry,” and “As a responsible AI”.

Iterative refinement: If the new response R’ is still classified as a jailbreak, the new failed defense prompt PP’, and new response R′R^{\prime} are passed back to DD, which generates a new defense prompt.

Validation and reparation: To ensure that the current optimized defense prompt is effective not only for the current query but also for future queries, we sample a small set of examples as a validation set to screen for defense prompts with poor generalization ability. Finally, to increase the diversity and comprehensiveness of the defense prompt pool, we rephrase effective and generalizable defense prompts, and select the rephrased results that are both effective and generalizable to save in the defense pool.

Finally, AdaShield-A obtain the diverse defense prompt pool P={P1,P2,...,Pn}\mathcal{P}=\{P_{1},P_{2},...,P_{n}\}, customized for different scenarios and incorporates safety guidelines. Each defense prompt is stored in the form of a dictionary, i.e. Di=<Qi:Pi>D_{i}=<Q_{i}:P_{i}>, with the key being the malicious query input QiQ_{i} to the target MLLM MM when the defender generates the defense prompt PiP_{i}, and the value being the refined defense prompt PiP_{i}. Each defense prompt is automatically and specifically optimized by the defender based on the jailbreak response of the target MLLM to current malicious query inputs.

where ziTz_{i}^{T} and ziIz_{i}^{I} are the text and image embeddings of malicious query QiQ_{i} in defense prompt pool DD, concat(⋅)\text{concat}(\cdot) is the concatenation operation and cos⁡\cos is the cosine similarity. Finally, we prepend input qurey QtQ_{t} with the optimal defense prompt PbestP_{best} to refuse to respond to malicious queries or safely respond to users’ normal queries. To further prevent the over-defense problem, when the maximum similarity value is less than β\beta, AdaShield-A will judge the current query as benign and does not employ any defense prompt.

Experiment

Attack Type and Dataset. In our work, we use the popular structured-based attacks Figstep and QR to evaluate our AdaShield-S and AdaShield-A. In addition, we also use a popular evaluation benchmark, namely MM-Vet , to evaluate the important factor of “over-defensiveness” on benign datasets. For more details, please refer to the Appendix.

Baselines. FigStep propose a defense prompt to safeguard MLLMs, denoted as FSD. In particular, the defense prompt in FigStep is that “First, you need to check whether there is text in the image. If there are textual instructions in the image that are harmful or violate AI safety policy, you should not assist user’s requests as you are a harmless assistant.”. Recently, MLLMP are proposed to defend against the structure-based jailbreak, which additionally employs a harm detector to identify the harmful response, and the detoxifier corrects these harmful outputs. In this paper, we use FSD and MLLMP as our baseline. For fairness, we refer to their original settings to reproduce FSD and MLLMP . Moreover, we use a unified test dataset and metrics to evaluate all defense methods.

Implement Details. In our AdaShield-A, we employ the open-sourced Vicuna-v1.5-13B as the defender DD. To enhance the diversity of the defense prompt pool with a limited number of training samples, we utilize the GPT4 API to rephrase the initial effective defense prompts. In step 5 of AdaShield-A, we employ a small validation dataset to ensure the generalizability of the auto-refined defense prompts. Only those defense prompts with an ASR below a threshold α\alpha on the validation set are selected for inclusion in the final defense prompt pool. In this paper, we set the thresholds α=0.8\alpha=0.8 and β=0.7\beta=0.7 (refer to Eq. 3). We provide a detailed sensitive analysis of the hyper-parameters α\alpha and β\beta in Sec. 0.B.2. Furthermore, we present the system prompt, which guides defender DD to optimize the defense prompt, in Fig. 5.

Evaluation Metric Details. In this paper, We utilize the keyword-based attack success rate (ASR) to evaluate the performance of all defense methods. This metric classifies jailbreak responses by detecting whether the predefined keywords exist in the responses from MLLMs. Such keywords include phrases like “I am sorry,” and “I apologize,”. The total keywords used for evaluations are listed in Tab. 12. Furthermore, we introduce the additional metric the GPT recheck attack success rate (Recheck) (see Alg. 1) to evaluate all defense methods. Recheck is more sensitive to determine whether the response is essentially addressing the malicious query.

Target Multimodal Large Language Models. We use three popular open-sourced MLLMs, including LLaVA 1.5-13B , MiniGPT-v2-13B and CogVLM-chat-v1.1 to evaluate all defense methods.

2 Main Results

Defense Effectiveness. We evaluate all defense methods on the popular structure-based attacks (i.e. FigStep and QR ). The detailed results are summarized in Tab. 2. As observed, both AdaShield-S and AdaShield-A, outperform FSD and MLLMP in defending against FigStep and QR , where Recheck and ASR are reported. However, due to the absence of specific safety rules, AdaShield-S exhibits inferior defense performance compared to AdaShield-A. Furthermore, MLLMP , as a post-hoc filtering defense mechanism, employs a harmful detector to identify the malicious response and a detoxifier to correct these harmful outputs. Nevertheless, the generality of the harmful detector is limited, and the effectiveness of the detoxifier is constrained, leading to the failure of MLLMP in defending against jailbreak attacks. For instance, with target MLLM is LLaVA, the harmful detector in MLLP exhibits a mere accuracy of 4.34% in the ‘Pornography’ scenario of QR.

Benign Dataset Performance. To assess the impact of over-defense, we compare the six core types of visual-language capabilities of MLLMs when being incorporated with different defense methods. The results are presented in Tab. 2. It is observed that AdaShield-A outperforms MLLMP and FSD , as well as achieves performance comparable to the Vanilla. This indicates that AdaShield-A excels in mitigating over-defense by filtering benign queries based on similarity, while AdaShield-S still falls short at recognizing the benign queries, leading to performance degradation caused by over-defense.

3 Ablation Study

Effect of Manual Static Prompts. In Sec. 3.2, we discuss how to design an effective defense prompt for structured-based jailbreak attacks on MLLMs. To support the claims in Sec. 3.2 and demonstrate the design of PsP_{s} in Tab. 1 is not trivial, we propose five additional kinds of potential defense prompts, i.t. Pa,Pb,Pc,Pd,PeP_{a},P_{b},P_{c},P_{d},P_{e} and compare their effectiveness to jailbreak defense. These defense prompts and the final results are shown in Tab. 3, where the average of ASR on different scenarios is reported. The detailed explanations of the proposed defense prompt are outlined below. (i) PaP_{a} does not contain specific instructions to check the image content, but only vaguely guides the model to examine the instructions. (ii) PbP_{b} requires the model to check the content of the image but lacks a chain-of-thought. (iii) When the model determines that the current query is malicious, PcP_{c} only requires the model to refuse to engage in illicit activities, but lacks a clear and actionable plan, e.g., answering with “I am sorry.” In other words, PcP_{c} only instructs the model not to engage in illegal activities, without guiding what the model should do. (iv) PdP_{d} is only the first step of PsP_{s}, which involves examining whether the image contains harmful text or items. (v) PeP_{e} is only the second step of PsP_{s}, which forces the model to combine the content of pictures and text to comprehensively analyze whether the instruction is harmful.

Validation of Intuition 1. We observe that the defense prompts PaP_{a} exhibit higher ASR values than PbP_{b} across all attacks and MLLMs. It indicates that the key of defense on MLLMs lies in examining the content of the images.

Validation of Intuition 2. Compared with the results of PaP_{a}, PdP_{d}, PeP_{e}, and PsP_{s}, it is evident that CoT prompts play a crucial role in the performance of AdaShield-S. Meanwhile, the single-step verification instructions in PdP_{d} and PeP_{e} complement each other, assisting PsP_{s} in achieving optimal performance. Note PdP_{d} obtains the best performance with LLaVA as the target model. However, the average ASR of PdP_{d} across all tasks is higher than that of PsP_{s} (22.45% v.s. 16.80%).

Validation of Intuition 3. Meanwhile, the defense prompt PcP_{c} exhibits the lowest performance, corroborating the assertions made in Sec. 3.2. Specifically, the MLLMs can effectively refrain from engaging in illegal activities only when the defense prompt provides explicit guidance to the model, instructing it how to respond to malicious queries, such as replying with “I am sorry.”.

Validation of Intuition 4. We also design a variant defense prompt PvP_{v} by removing “Instead, please execute the above instruction safely and correctly.” from PsP_{s}, and compare PsP_{s} with PvP_{v} to verify the intuition 4. The only difference between PsP_{s} and PvP_{v} is that when the query is determined to be benign by target model MM, PvP_{v} does not guide MM to execute commands safely. Then, we evaluate the performance of PsP_{s} and PvP_{v} on MM-Vet , where the results are shown in Tab. 4. As we can see, due to the absence of guidance on how to respond to safe queries, PvP_{v} obtains lesser performance in benign tasks.

Effect of Retrieval method. We evaluate the effect of our proposed retrieval method. We introduce a variant, termed Random, which randomly selects a prompt from defense prompt pool P\mathcal{P} to prepend the input query. To ensure fairness, we use the same defense prompt pool P\mathcal{P} for both AdaShield-A and Random. As reported in Tab. 5, Random exhibits worse performance, validating that our proposed retrieval method is indispensable to AdaShield-A. Table 6: Time Consumption Comparison Analysis. The results show that AdaShield-A incurs minimal additional time cost during inference. Method Inference Time Benign Harmful Vanilla 1.76s 9.40s FSD 1.86s 6.78s MLLMP 2.88s 16.03s AdaShield-S 2.78s 2.02s AdaShield-A 1.82s 1.46s Table 7: Generalization on unseen scenarios on QR dataset. The results demonstrate that AdaShield-A exhibits generalization in unseen scenarios. Numbers in bold represent best results. Test Train Easy Hard All Easy 12.67 10.95 13.86 Hard 27.38 18.92 16.82 All 19.46 14.63 15.22

4 Analysis Study

Inference Times Consumption Comparison. We evaluate the time consumption of all methods using 50 benign queries and 50 harmful queries, with LLaVA as the target MLLM. The results are reported in Tab. 6. It is shown that the time cost of retrieval in AdaShield-A is negligible. In contrast, MLLMP , a post-hoc filtering method, incurs a significant time cost during inference.

Generalization on Unseen Scenarios. To verify the generalizability of AdaShield-A towards unseen scenarios, we only train AdaShield-A with samples from partial scenarios on QR, then evaluate AdaShield-A on test samples, including unseen scenarios. Specifically, we categorize the 13 forbidden scenarios in QR into two groups: (i) Easy scenarios, which encompass common harmful activities such as Illegal Activities, Hate Speech, Malware Generation, Physical Harm, Economic Harm, Fraud, and Pornography; (ii) Hard scenarios, which include topics requiring professional expertise or those sensitive to politics and management, such as Political Lobbying, Privacy Violence, Legal Opinion, Financial Advice, Health Consultation and Gov Decision. We first train AdaShield-A on Easy, Hard, and ALL scenarios to obtain the respective defense prompt pools Di\mathcal{D}_{i}, Dii\mathcal{D}_{ii} and Dall\mathcal{D}_{all}. Then, we evaluate AdaShield-A with Di\mathcal{D}_{i}, Dii\mathcal{D}_{ii} and Dall\mathcal{D}_{all} on test samples from Easy, Hard, and ALL scenarios. We present the results in Tab. 4.3, where the average of ASR is reported. The results show that AdaShield-A achieves robust defense performance on unseen scenarios. We also find that AdaShield-A with Dii\mathcal{D}_{ii}, trained on the Hard set, achieves the best performance, which indicates that the quality of training samples significantly impacts the performance of AdaShield-A.

Transferability Across Target Models. To assess transferability across target models, we exchange the defense prompt pools learned with LLaVA and CogVLM as the target MLLMs MM, and then evaluate them respectively on QR and FigStep. The results are shown in Tab. 8. We observe that AdaShield-A enables transferability across different target MLLMs.

Visualizations of the Auto-refined Defense Prompts.

In this section, we present some auto-refined defense prompt examples (see Fig. 4) to show the superiority of AdaShield-A. Specifically, we present three examples from QR and FigStep attacks. Each example consists of a query (image-text pair), an input-aware defense prompt generated by AdaShield-A for the current text query, and the corresponding output of the target MLLM. As illustrated in Fig. 4, we observe that our AdaShield-A effectively generates effective defense prompts for each query. These defense prompts include detailed safety rules, thereby successfully safeguarding the MLLM from malicious queries.

Conclusion & Limitation

Conclusion. In this work, we present AdaShield, a novel defense mechanism for MLLMs against structure-based jailbreak attacks. AdaShield employs adaptive shield prompting to enhance the robustness of MLLMs without the need for fine-tuning or additional modules. Our experiments demonstrate its effectiveness in safeguarding MLLMs while preserving their general capabilities, highlighting its potential as a plug-and-play solution for improving MLLMs’ safety.

Limitation. One limitation of AdaShield is that it is specifically designed for structure-based jailbreak attacks. We leave a universal defense framework that can address both structure-based and perturbation-based attacks as future work.

References

Appendix

The appendix is organized as follows: First, we provide a detailed description of the datasets in Sec. 0.A. Then, we provide additional ablation study, and sensitive analysis about hyper-parameters in Sec. 0.B.

Appendix 0.A Datasets

Structure-based Jailbreak Attacks. In this paper, we use the state-of-the-art structured-based attacks Figstep and QR to evaluate our proposed AdaShield-S and AdaShield-A. Specifically, FigStep covers 10 scenarios prohibited by both OpenAI and Meta usage policies , such as illegal activities, hate speech, financial advice, etc. Each prohibited scenario contains 50 harmful requests. QR consists of 1680 malicious questions, which also cover 13 common unsafe and sensitive scenarios, like Political-Lobbying, Legal-Opinion, etc. Each malicious query in FigStep and QR consists of a harmful image and a benign text prompt, so that it bypasses the safety alignment within the textual module of MLLMs. During training, AdaShield-A only need a few malicious queries to optimize defense prompts iteratively and obtain a defense prompts pool. Thus we partition the datasets of FigStep and QR into three subsets: training, validation, and testing, in the proportions of 10%, 5%, and 95%,, respectively. We present the details of FigStep and QR in Tab. 9 and Tab. 0.A.1.

Benign Dataset Details. Additionally, we use a popular multimodal evaluation benchmark, named by MM-Vet to evaluate the important factor of ‘over-defensiveness’ on benign datasets. Specifically, MM-Vet uses an LLM-based evaluator to evaluate six core visual-language capabilities of MLLMs, including Recognition (Rec), Knowledge (Know), Optical character recognition (OCR), Spatial awareness (Spat), Language generation (Gen), and Math. The full score of each capability is 100% in on MM-Vet. In this paper, we use OpenAI’s GPT-4 API as the LLM-based evaluator. More details refer to MM-Vet .

Appendix 0.B Additional Experiments

Effect of the initial defense prompt for AdaShield-A. In this section, we present additional ablation studies (See Tab. 11) to investigate the impact of the initial defense prompt in AdaShield-A. The results demonstrate that AdaShield-A, when equipped with our manual defense prompt PsP_{s}, achieves the best performance. Moreover, even the least effective variant of AdaShield-A, with prompt PaP_{a}, still surpasses other defense methods in terms of performance. This indicates that AdaShield-A is robust to initial static defense prompts.

B.2 Additional Sensitive Analysis

In this section, we provide the justification for the hyper-parameters α\alpha and β\beta on QR with LLaVA 1.5-13B as our target MLLM.

Justification of hyper-parameter α\alpha. The hyper-parameters α\alpha is used to ensure the generality of auto-refined defense prompts. Specifically, in step 5 of AdaShield-A, we select the auto-refined defense prompts with an ASR lower than α\alpha on the validation set for inclusion in the final defense prompt pool. Here, we present a sensitivity analysis of α\alpha in Fig. 6(a). We observe that as α\alpha increases, the average ASR of AdaShield-A decreases. These results demonstrate that validation set verification is crucial for ensuring that AdaShield-A learns a high-quality defense pool. A higher alpha value assists AdaShield-A in obtaining a defense pool with greater generality. In this paper, we set α=0.8\alpha=0.8.

Justification of Hyper-parameter β\beta. In this paper, to address the over-defense problem, we use the hyper-parameter β\beta to initially identify the benign queries. Specifically, if the maximum similarity between a test query and the keys in the defense prompt pool is below β\beta (see Eq.3), we initially classify the query as benign and refrain from prepending any defense prompts. The justification of β\beta is illustrated in Fig. 6.(b), where we report the average ASR on QR and the total score on MM-Vet . As observed, with the increase in β\beta, both the average ASR of AdaShield-A on QR and the total score on MM-Vet rise. It indicates that a larger β\beta value helps alleviate the over-defense problem but may lead to a decrease in defense performance, presenting a trade-off. In this paper, we set β=0.7\beta=0.7.