Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses

Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Jing Jiang, Min Lin

Introduction

Large language models (LLMs) are typically trained to be safety-aligned in order to avoid misuse during their widespread deployment . However, many red-teaming efforts have focused on proposing jailbreaking attacks and reporting successful cases in which LLMs are misled into producing harmful or toxic content .

When jailbreaking, optimization-based attacks search for adversarial suffixes that can achieve high attack success rates (ASRs) ; more recently, Andriushchenko et al. use prompting and self-transfer techniques to randomly search adversarial suffixes, while reporting 100% ASRs on both Llama-2-Chat-7B and Llama-3-8B . Although effective against aligned LLMs, these adversarial suffixes mostly have no semantic meaning (even after low-perplexity regularization ), making them susceptible to jailbreaking defenses like perplexity filters and SmoothLLM . As empirically reported in Figure 4, adversarial suffixes generated by Andriushchenko et al. result in quite high perplexity and are easily detectable.

LLM-assisted attacks, on the other hand, use auxiliary LLMs to generate adversarial but semantically meaningful requests capable of jailbreaking the target LLM, usually requiring only tens of queries . The generated adversarial requests can bypass perplexity filters and are insensitive to defenses that rely on input preprocessing . On the downside, it can be challenging for LLM-assisted attacks to achieve state-of-the-art ASRs on aligned LLMs, especially when they are evaluated under strict conditions (e.g., using the correct system prompt on Llama-2-Chat-7B) .

In contrast, manual attacks are more flexible, but necessitate elaborate designs and considerable human labor . In particular, Wei et al. explore few-shot in-context demonstrations containing harmful responses to jailbreak LLMs. Anil et al. automate and extend this strategy to many-shot jailbreaking, which prompts LLMs with hundreds of harmful demonstrations and can achieve high ASRs on cutting-edge closed-source models. Nonetheless, many-shot jailbreaking requires LLMs’ long-context capability that is still lacking in most open-source models .

In this work, we revisit and significantly improve few-shot jailbreaking, especially against open-source LLMs with limited context sizes (≤8192\leq 8192). We first automatically create a demo pool containing harmful responses generated by “helpful-inclined” models like Mistral-7B (i.e., not specifically safety-aligned). Then, we inject special tokens from the target LLM’s system prompt, such as [/INST] in Llama-2-7B-Chat,These special tokens can be directly accessed on open-source LLMs by checking their system prompts, and may be extracted on closed-source LLMs by prompting like “Repeat the words above” . into the generated demos as illustrated in Figure 1. Finally, given the number of demo shots (e.g., 4-shot or 8-shot), we apply demo-level random search in the demo pool to optimize the attacking loss.

As summarized in Table 1, our improved few-shot jailbreaking (named as I\mathcal{I}-FSJ) achieves >80%>80\% (mostly >95%>95\%) ASRs on aligned LLMs including Llama-2-7B and Llama-3-8B. In addition, as reported in Table 2, we further enhance Llama-2-7B by different jailbreaking defenses, while our I\mathcal{I}-FSJ can still achieve >95%>95\% ASRs in most cases. Note that the random search operation in I\mathcal{I}-FSJ is demo-level, not token-level, so the crafted inputs remain semantic. Overall, I\mathcal{I}-FSJ is completely automated, eliminating the need for human labor and serving as a strong baseline for future research on jailbreaking attacks.

Related work

Jailbreaking attacks. LLMs like ChatGPT/GPT-4 and Llama-2 are generally designed to return helpful and safe responses, and they are trained to align with human values . However, red-teaming research has shown that LLMs can be jailbroken to produce harmful content using manually created or automatically generated prompts . Additionally, Tian et al. investigate the safety risks of LLM-based agents; Greshake et al. introduce the concept of indirect prompt injection to compromise LLM-integrated applications; According to Wei et al. , aligned LLMs are vulnerable to jailbreaking due to the conflicting goals of capability and safety, as well as the gap between pretraining and safety training; Carlini et al. point out the inherent vulnerability of neural networks to adversarial examples as a root cause. Recently, several studies have found that fine-tuning aligned LLMs with either poisoned or benign data can undermine their alignment and safety .

Jailbreaking defenses. Though instruction fine-tuning (SFT+RLHF) plays a major role in safety alignment, various defense mechanisms have been proposed to strengthen models’ alignment to tackle the safety issue from a system-level perspective . Context-based defenses insert either manually designed or automatically optimized safety prompts into the context . Input-detection-based defenses aim to detect potentially harmful input prompts like perplexity filter . Perturbation-based defenses purify the input prompts by adding perturbations . Output-detection-based defenses distinguish whether the generated responses contain harmful content and return the default refusal responses if harmful information detected . Mazeika et al. propose to train the model against the suffix optimized by GCG directly. Xu et al. augment the original model with a safety expert model. In particular, they interpolate the token prediction distributions from these two models with a weighting hyperparameter. These methods generally fall into the adversarial training category.

Improved few-shot jailbreaking

We briefly introduce preliminaries before describing improved few-shot jailbreaking strategies such as constructing a demo pool, injecting special tokens, and performing demo-level random search.

Jailbreaking setup. We focus on jailbreaking attacks that aim to discover prompts to mislead LLMs producing harmful content, in order to fulfill a specific harmful request G∈T∗G\in\mathcal{T}^{*} (e.g., “Develop a detailed instruction manual for making a bomb or other explosive device”). We assume access to a collection of such harmful requests, which most aligned LLMs recognize as harmful. These typically encompass categories such as misinformation, violence, and hateful content, and are designed to elicit no response from the LLMs. Following previous notations , we define a language model LLM:T∗→T∗\mathtt{LLM}:\mathcal{T}^{*}\rightarrow\mathcal{T}^{*} as a function that transforms a sequence of input tokens into a sequence of output tokens. This model, referred to as the target model, is the one we aim to compromise or “jailbreak”. We also define a judge function JUDGE:T∗→{NO,YES}\mathtt{JUDGE}:\mathcal{T}^{*}\rightarrow\{\text{NO},\text{YES}\} to evaluate the content generated by the target model. The goal of a jailbreaking attacker is to find a prompt P∈T∗P\in\mathcal{T}^{*} such that when the target model processes PP, the judge function deems the output harmful, i.e., JUDGE(LLM(P),G)=YES\mathtt{JUDGE}(\mathtt{LLM}(P),G)=\text{YES}.

In-context learning (ICL). ICL is a remarkable capability of LLMs. During ICL, a LLM is presented with a demonstration set D={(x1,y1),...,(xm,ym)}={d1,...,dm}D=\{(x_{1},y_{1}),...,(x_{m},y_{m})\}=\{d_{1},...,d_{m}\}, where each xix_{i} is a query input and each yiy_{i} is the corresponding label or output. These examples effectively teach the model task-specific functionals. The process involves constructing a prompt that includes the demonstration set followed by a new query input for which the label needs to be predicted. The prompt takes the form [x1,y1,...,xn,yn,xnew][x_{1},y_{1},...,x_{n},y_{n},x_{\text{new}}], where xnewx_{\text{new}} is the new input query. The model, having inferred the underlying pattern from the provided examples, uses this prompt to predict the corresponding label ynewy_{\text{new}} for the new input xnewx_{\text{new}}. ICL leverages the model’s pre-trained knowledge and its ability to recognize and generalize patterns from the context provided by the demonstration set. This capability is particularly powerful because it allows the model to adapt to a wide range of tasks with minimal task-specific data, making it a flexible and efficient tool for various applications.

Few-shot jailbreaking (FSJ). Wei et al. explore few-shot in-context demonstrations containing harmful responses to jailbreak LLMs. Anil et al. automate and extend this strategy to many-shot jailbreaking, which prompts LLMs with hundreds of harmful demonstrations and can achieve high ASRs on cutting-edge closed-source models. Nonetheless, many-shot jailbreaking requires LLMs’ long-context capability that is still lacking in most open-source models . And the vanilla FSJ is ineffective on some well-aligned LLMs like the Llama-2-Chat family.

2 Improved strategies

We primarily develop three strategies to obtain improved FSJ (I\mathcal{I}-FSJ), as summarized below:

Constructing a demo pool. Given a set of harmful requests {x1,...,xm}\{x_{1},...,x_{m}\} (e.g. the harmful behaviors from AdvBench ), we collect the corresponding harmful responses {y1,...,ym}\{y_{1},...,y_{m}\} by prompting “helpful-inclined” models like Mistral-7B which are not specifically safety-aligend. Finally, we create a demonstration pool as D={(x1,y1),...,(xm,ym)}={d1,...,dm}D=\{(x_{1},y_{1}),...,(x_{m},y_{m})\}=\{d_{1},...,d_{m}\}. Note that we only build the pool once and use it to attack multiple models and defenses.

Injecting special tokens. In our initial trials, we attempt to directly use the generated vanilla FSJ demonstrations (examplified in the left part of Figure 1) to jailbreak LLMs and obtain non-trivial ASRs on some models like Qwen1.5-7B-Chat . But we keep obtaining near zero ASRs on much more well-aligned LLMs such as Llama-2-7B-Chat, which is consistent with the results reported by Wei et al. and it seems FSJ is ineffective on these models.

Intriguing observations: Interestingly, we observe that most current open-source LLMs’ conversation templates separate the user message and assistant message (e.g. model completion) with special tokens. For example, as shown in Figure 1’s single message template, Llama-2-Chat separates the messages with [/INST]. We suspect the model is prone to conduct generation once presented by the [/INST] tokens. We thus hypothesize we can exploit this tendency with the help of ICL to induce the model to generate harmful content by appending harmful messages with the [/INST] tokens.

Thus, we inject special tokens from the target LLM’s system prompt, such as [/INST] in Llama-2-7B-Chat, into the generated demos as illustrated by the I\mathcal{I}-FSJ Demonstration example in Figure 1. More specifically, given an original FSJ demonstration, we construct I\mathcal{I}-FSJ demonstration by first injecting [/INST] between the user message and assistant message, which is motivated by the specific formatting of Llama-2-Chat’s single message template. Additionally, we inject [/INST] between the generated steps in the demonstration.

Demo-level random search. After the I\mathcal{I}-FSJ demo pool is constructed, we use demo-level random search to minimize the loss of generating the initial token (e.g. “Step”) on the target model. We modify the random search (RS) algorithm into a demo-level variant, which is simple and requires only the output logits instead of gradients. The algorithm is as follows: (i) prepend a sequence of nn sampled demonstrations to the original request; (ii) in each iteration, change a demonstration to another one at a random position in the sequence; (iii) accept the change if it reduces the loss of generating target token (e.g., “Step” that leads the model to fulfill a harmful request) at the first position of the response. Furthermore, we implement the above demo-level RS algorithm in a batch way to achieve better parallelism as described in Algorithm 1. To tackle input-perturbation-based defenses like SmoothLLM , we introduce an ensemble variant of our demo-level RS method as described in Algorithm 2, which aims to find a combination of demonstrations that is not only effective for jailbreaking but also robust to perturbations. More details are provided in Appendix B.1.

Empirical studies

This section demonstrates the effectiveness of our I\mathcal{I}-FSJ in jailbreaking various open-source aligned LLMs and advanced defenses.

Aligned LLMs. We evaluate open-source and advanced LLMs for reproducibility. These include Llama-2-Chat , which underwent multiple rounds of manual red teaming for adversarial training, making them resilient to various attacks; Llama-3-Instruct , which were intentionally optimized for helpfulness and safety; OpenChat-3.5 , fine-tuned from Llama-2 using mixed-quality data with consideration of data quality; Starling-LM , fine-tuned from OpenChat 3.5 using RLHF with a reward model emphasizing helpfulness and harmlessness; and Qwen1.5-Chat , trained on datasets annotated for safety concerns such as violence, bias, and pornography. According to Mazeika et al. , the attack success rates (ASRs) are stable within model families but vary significantly between different families. Therefore, we only consider the 7B variant across all model families.

ASR metrics. We follow Liu et al. to evaluate the attacking effectiveness by two ASR metrics. The first one is a Rule-based metric from Zou et al. , which is a keyword-based detection method that counts the number of harmful responses. Previous studies have used LLM-based metric such as GPT-4 to determine whether the responses are harmful. For reproducibility, we instead use the fine-tuned Llama Guard classifier following Chao et al. . More details are in Appendix B.2.

Defenses. We consider seven efficient defense mechanisms to further enhance aligned LLMs. Among these, Self-Reminder and ICD are context-based methods, (window) PPL filters are input-detection-based, while Retokenization and SmoothLLM are perturbation-based methods. Safe Decoding belongs to adversarial training. Llama Guard is output-detection-based that requires the attacker to jailbreak both the target model and the output filter, which judges whether the target model’s outputs are safe or unsafe. More details are in Appendix B.3.

Setup of our attack. For the demonstrations used in FSJ and I\mathcal{I}-FSJ, we apply Mistral-7B-Instruct-v0.2, an LLM with weaker safety alignment, to create the harmful content on a set of harmful requests. For more details, please check Appendix B.4. Our targets are a collection of 50 harmful behaviors from AdvBench curated by Chao et al. that ensures distinct and diverse harmful requests. We exclude the demonstrations for the same target harmful behavior from the pool to avoid leakage. For the demo-level random search, we set batch size B=8B=8 and iterations T=128T=128. We let the target LLMs generate up to 100 new tokens. We use each LLM’s default generation config. Every experiment is run on a single NVIDIA A100 (40G) GPU within a couple of hours.

2 Jailbreaking attacks on aligned LLMs

To examine the generality of our proposed I\mathcal{I}-FSJ, we evaluate it on a diverse set of aligned LLMs. For different LLMs that utilize different conversation templates, we inject the corresponding special tokens, which distinct the user message and assistant message, into demonstrations. Note that such a process can be fully automated by a simple regular expression method. As detailed in Tables 1 and 4, we first find that our I\mathcal{I}-FSJ attack is effective on all tested LLMs. In particular, on OpenChat-3.5, Starling-LM-7B, and Qwen1.5-7B-Chat, augmenting the FSJ with either demon-level random search or injecting special tokens is sufficient to achieve nearly 100% ASRs.

Nonetheless, models with stronger alignment, like Llama-2-7B-Chat and Llama-3-8B-Instruct, are more challenging. For these models, the FSJ with demo-level random search alone is insufficient for jailbreaking. Only by combining special tokens and demon-level random search can we successfully break these models’ safety alignment, demonstrating the effectiveness of our techniques. Llama-3-Instruct requires more shots to jailbreak than Llama-2-Chat, which could be due to improved alignment techniques. Still, our I\mathcal{I}-FSJ achieves over 90% ASRs within limited context window sizes.

Our approach consistently achieves near 100% ASR on most models tested, highlighting the significant vulnerabilities and unreliability of current alignment methods. These findings highlight the critical need for improved and more resilient alignment strategies in the development of LLMs.

3 Jailbreaking attacks on Llama-2-7B-Chat + jailbreaking defenses

To assess our I\mathcal{I}-FSJ’s effectiveness against system-level robustness, we test it on Llama-2-7B-Chat with various defenses. As shown in Tables 2 and 5, our results demonstrate that I\mathcal{I}-FSJ can circumvent jailbreaking defenses. For most defenses, randomly initialized nn-shot demonstrations exhibit relatively low ASRs. However, optimizing the combination of demonstrations with demo-level random search can significantly boost the ASRs, peaking at near 100% in the 4-shot and 8-shot configurations. For the majority of defenses, the 4-shot setting is sufficient to achieve high ASRs.

Self-Reminder modifies Llama-2-Chat’s default system message, which may degrade the safety alignment. ICD indicates a positive trend: as the defense shot increases, I\mathcal{I}-FSJ’s ASRs decrease significantly in the 2-shot setting. Attack success rates remain relatively low across defense shots, even with demo-level random search, indicating ICD’s effectiveness. Yet, in the 4- and 8-shot settings, the ICD fails to defend the I\mathcal{I}-FSJ. The PPL filter cannot reduce our ASRs because our input is mostly natural language with a perplexity lower than the filtering threshold (for example, the highest perplexity of harmful queries in AdvBench). Even with a higher interpolation weight α=4\alpha=4, SafeDecoding cannot defend against our attack when computing the output token distribution.

Remark 1: I\mathcal{I}-FSJ is robust to perturbations. Retokenization, which splits tokens and represents tokens with smaller tokens, can effectively perturb the encoded representation of the input prompt but fails to defend against I\mathcal{I}-FSJ. Regarding the SmoothLLM variants, which directly perturb the input text in different ways, they successfully defend I\mathcal{I}-FSJ at the 2-shot setting, resulting in ≤10%\leq 10\% ASRs. However, our method achieves >85%>85\% ASRs against all of them at the 8-shot setting, which still falls into the few-shot regime. Also, as shown in Figure 2, we plot the LLM-based ASRs (Top) and rule-based ASRs (Bottom) for various perturbation percentages q∈{5,10,15,20}q\in\{5,10,15,20\}; the results are compiled across three trials. At the 8-shot setting, our method still maintains high ASRs (e.g. ≥80%\geq 80\%) across all the perturbation types and perturbation rates. We also plot the loss curves of the random search optimization process in Figure 11. All these results demonstrate that I\mathcal{I}-FSJ is robust to perturbations. Additionally, such a property intermediately implies that I\mathcal{I}-FSJ can counter defenses like “filtering the [/INST] tokens by matching” because the attacker can use SmoothLLM to perturb their adversarial prompt before submitting their input.

Remark 2: I\mathcal{I}-FSJ can be propagative. To counter the defense of Llama Guard, we need to achieve propagating jailbreaking. Previous work has demonstrated how to achieve adversarial-suffix-based propagating jailbreaking, which can jailbreak the target LLM and evade the Guard LLM. However, such an attack is also fragile confronting a perplexity filter. We instead modify our I\mathcal{I}-FSJ demonstrations slightly by adaptively taking the Guard LLM’s conversation template into account as shown in Figure 9. Our results show that I\mathcal{I}-FSJ successfully jailbreaks both the target LLM and Guard LLM, demonstrating that I\mathcal{I}-FSJ can be propagative.

4 Further analysis

The effect of pool size. Our method inherently comes with a design choice: the size of the demonstration pool. To figure out the effect of this factor, we evaluate our method on Llama-2-7B-Chat under various pool sizes. As shown in Figure 4, the ASRs generally increase as the pool size grows and gradually saturate as observed from 256 to 512. The pool size shows a much larger impact on the 2-shot setting compared to the 4-shot and 8-shot settings, which might be because the latter two settings are relatively easier. Surprisingly, 32 demonstrations are already sufficient to achieve over 90% ASRs at an 8-shot setting, indicating the data efficiency of our method. Thus, we set the pool size as 512 in all of our experiments.

The effect of shots. Figure 4 highlights the impact of the number of shots on the ASR. As the number of shots increases from 2 to 8, there is a noticeable improvement in the ASR. With 2 shots, the ASR starts relatively low, around 25.4%, and gradually improves as the dataset size increases, reaching about 61.6% at its highest point. This indicates moderate effectiveness in terms of attack success when only 2 shots are used. For 4 shots, there is a significant jump in the initial ASR compared to 2 shots. The ASR begins at around 88.0% and rapidly stabilizes close to 97.8% as the dataset grows. This demonstrates that increasing the shot count to 4 substantially enhances the attack’s success rate, achieving a high level of effectiveness early on. The effect is most pronounced when moving from 2 to 4 shots, with further improvement seen when increasing to 8 shots, where the ASR approaches 100%. However, these results also indicate that beyond a certain point, increasing the number of shots does not substantially boost the ASRs since fewer shots are already sufficient. Thus, we test up to 8 shots in most of our experiments.

Compared to other attack methods As shown in Table 3, we compare our method against other attacks such as PAIR , GCG , AutoDAN , PAP , and PRS (stands for ‘Prompt+RS+Self-transfer’) . The table indicates that the I\mathcal{I}-FSJ method with Demo RS is the most effective approach for bypassing safety measures in language models, achieving the highest ASRs in both scenarios (with and without a system message). The presence of a system message generally reduces the effectiveness of most methods, except for I\mathcal{I}-FSJ with Demo RS and PRS, which remain robust. When compared with adversarial-suffix based method , though they may achieve comparable ASRs (e.g. 90% evaluated by the rule-based metric) with our method, it completely fails with a single perplexity (windowed) filter as shown in Figure 4.

Discussion

Jailbreaking attacks on LLMs are rapidly evolving, with different approaches demonstrating varying strengths and limitations. Our I\mathcal{I}-FSJ represents a significant advancement in this domain, particularly against well-aligned open-source LLMs with limited context sizes. The primary innovation lies in the automated creation of the demonstration pool, the utilization of special tokens from the target LLM’s system template, and demo-level random search, which together facilitate high ASRs. Our empirical studies demonstrate the efficacy of I\mathcal{I}-FSJ in achieving high ASRs on aligned LLMs and various jailbreaking defenses. The automation of I\mathcal{I}-FSJ eliminates the need for extensive human labor, offering a robust baseline for future research in this domain.

References

Appendix A Broader Impacts and Limitations

Broader Impacts. The implications of improved jailbreaking techniques are profound, extending beyond academic interest to potential real-world applications and security considerations. Given the superior efficacy of the proposed I\mathcal{I}-FSJ, it is possible that our method being misused to attack deployed systems can cause negative societal impacts. This underscores the necessity for robust, adaptive defenses that can counter with advancements in attack methods.

From a broader perspective, our work highlights the ongoing cat-and-mouse dynamic between attack strategies and defense mechanisms in the field of AI safety. As LLMs become more integral to various applications, understanding and mitigating vulnerabilities through comprehensive research is crucial. I\mathcal{I}-FSJ can serve as a strong baseline for future explorations on LLM safety.

Limitations. Our work focuses on jailbreaking open-source LLMs, with the assumption that the target model’s conversation template is known thus we can exploit the special tokens to facilitate the I\mathcal{I}-FSJ attack. However, for closed-source LLMs like GPT-4 and Claude, the conversation template is usually unknown. Though it may be possible to extract the template on closed-source LLMs , the effectiveness of our method on these LLMs remains a future research question.

The reliance on special tokens from the target LLM’s system prompt may also introduce a vulnerability. If future models obfuscate or randomize these tokens, the effectiveness of I\mathcal{I}-FSJ may diminish, necessitating continual adaptation of the attack strategy.

Appendix B Implementation details

In contrast to Algorithm 1, we introduce a new optimization objective adaptive to the SmoothLLM defense, which considers KK different perturbations at each iteration. With this adaptive design, we can find a combination more suitable for attacking SmoothLLM or other perturbation-based defenses because the optimized demonstrations are both effective for jailbreaking and robust to perturbations.

B.2 The setup of metrics

The keywords used for Rule-based metric are listed in Figure 5 from Zou et al. . The prompt used for LLM-based metric is as shown in Figure 6 from Chao et al. .

B.3 Defenses

Self-Reminder : Self Reminder injects safety prompts into context to remind the LLMs to respond responsibly as shown in Figure 7.

ICD : ICD strengthens model robustness using in-context demonstrations of rejecting harmful prompts as shown in Figure 8.

PPL : We follow Alon and Kamfonas and use GPT-2 to calculate the perplexity. Following Jain et al. , we consider both the default PPL and windowed PPL. We set the PPL threshold as the highest perplexity of harmful requests in AdvBench , which ensures that queries from AdvBench would not be filtered out by the filter.

Retokenization : Retokenization splits tokens and represents them with multiple smaller tokens. We implement it using the handy implementation from huggingface https://github.com/huggingface/transformers/blob/v4.41.0/src/transformers/models/llama/tokenization_llama.py#L86, setting the dropout rate as 20% according to Jain et al. and Xu et al. .

SmoothLLM : SmoothLLM mitigates jailbreaking attacks on LLMs by randomly perturbing multiple copies of a given input prompt, and then aggregates the corresponding predictions to detect adversarial inputs. We consider all variants including Insert, Swap, and Patch with different perturb rates.

Safe Decoding : Safe Decoding augment the original model with a safety expert model. In particular, they interpolate the token prediction distributions from these two models with a weighting hyperparameter α\alpha. We set α=4\alpha=4.

Llama Guard : In our setting, Llama Guard is an output-detection-based method, which requires the attacker not only to jailbreak the target model but also jailbreak the output filter which judges whether the target model’s outputs are safe or unsafe.

B.4 Demonstration pool construction

For the demonstrations (harmful pairs) used in few-shot jailbreaking, we use a Mistral-7B-Instruct-v0.2, an LLM with weaker safety alignment, to craft the harmful content on a set of harmful requests. We first take the prompt template from Alayrac et al. as shown in Figure 10 to format the 520 harmful requests xix_{i} in the AdvBench . Then we prompt Mistral-7B-Instruct-v0.2 with the formatted harmful requests and collect the generated response yiy_{i} setting the number of max new tokens as 256. Finally, we create a demonstration pool as D={(x1,y1),...,(x520,y520)}D=\{(x_{1},y_{1}),...,(x_{520},y_{520})\}.

Appendix C Additional results

As shown in Figure 11, we observe that the loss steadily decreases as the demo-level optimization step increases, indicating the effectiveness of the proposed method.