Mind Your Format: Towards Consistent Evaluation of In-Context Learning Improvements
Anton Voronov, Lena Wolf, Max Ryabinin
Introduction
Pretrained language models have emerged as a dominant paradigm for solving many NLP problems in a unified framework Brown et al. (2020); Chowdhery et al. (2022); Scao et al. (2023); Touvron et al. (2023a). In particular, these models can achieve impressive downstream results with just a few demonstrations given as a part of their input Liu et al. (2021); Min et al. (2022c), which is often called a prompt in this case.
These few-shot or in-context Brown et al. (2020) learning (ICL) abilities of large models are a subject of frequent study, as the primary factors behind them are not yet fully understood. For example, one line of work investigates in-context learning within different theoretical frameworks Xie et al. (2022); Garg et al. (2022); Akyürek et al. (2023). In addition, multiple publications study the importance of different prompt attributes, such as the order of input demonstrations (Lu et al., 2022a) and their labels (Min et al., 2022d).
As shown in Zhao et al. (2021); Min et al. (2022a), the prompt format (i.e., a transformation from a set of examples to a natural language input) is also highly important. However, this aspect is often overlooked in most existing studies. Namely, works proposing modifications of ICL frequently present their results for a specific template without specifying the criteria guiding its selection. Furthermore, even when the results are averaged over a set of templates, they are compared to methods that were evaluated on a different set of templates. We illustrate this common discrepancy in Appendix A. This inconsistency can lead to a misinterpretation of the reported results: the difference between the performance of two methods may be explained by the variation across prompt formats rather than the methods themselves.
In this work, we evaluate the template sensitivity of 19 models from 7 families, including state-of-the-art open-source models Touvron et al. (2023b); Almazrouei et al. (2023), and show that this issue persists irrespective of the model size and the number of demonstrations. Moreover, comparing various ICL enhancements while taking the template influence into account renders the superiority of one method over others less apparent. Therefore, it is likely that the gains reported for advanced prompting methods can often be attributed to a luckily chosen template.
Crucially, there are no universally best templates for a given task. The best performing demonstration format for a fixed evaluation setting (i.e., the dataset, the model, the demonstration set, and the prediction method) does not transfer consistently across models (even within the same family), demonstration sets, or different prediction methods. We find this concerning, as even the best template for a given setting can produce poor results after slight changes, which makes “tuning” the template a very difficult task.
As a first step towards addressing template sensitivity in a practical way, we propose Template Ensembles — a test-time augmentation approach that averages model predictions over several prompt formats. This method is easy to implement and increases the average performance across templates for multiple prompting methods while reducing the sensitivity of these methods to the template choice.
In summary, our contributions are as follows:
We conduct a broad evaluationOur code and results of all evaluations can be found at github.com/yandex-research/mind-your-format of prompt template sensitivity across 19 models and 4 datasets, showing that the performance gains similar to using in-context learning improvements can be achieved solely by selecting a proper template.
We show that the choice of the best template depends on a combination of factors and that it is not possible to transfer the best template between models or prompting methods without a negative impact on quality.
We propose Template Ensembles as a baseline solution for improving the template robustness for in-context learning.
Background and Related Work
An important property of LLMs is their ability to learn new tasks from only a few demonstrations Radford et al. (2019); Brown et al. (2020). This capability, known as in-context learning, forms the focus of our work. We focus on sequence classification, as it is the most widely studied task for understanding and improving ICL performance. Similarly, although instruction-finetuned models Sanh et al. (2022); Wei et al. (2021); Chung et al. (2022) are a popular choice for downstream adaptation in practice, we omit them from our analysis for consistency with prior work.
Formally, classifying an input with in-context learning can be described as finding the class in the space of label tokens that yields a sequence with the highest probability according to a language model. The input sequence consists of demonstration inputs and labels and a test input ; to obtain a natural language input, demonstrations are formatted with a template.
Each template consists of four components: input and output verbalizers and that transform into a natural language text, an intra-separator to divide input from output, and an inter-separator to join several demonstrations. Figure 1 shows an example of transforming a set of demonstrations into a context for ICL.
2 In-Context Learning Analysis
Recent work has shown that ICL can perform at levels comparable to finetuning Chowdhery et al. (2022); Hoffmann et al. (2022). Still, in-context learning is known to be highly dependent on the way the model input is formed: a prompt is defined by several components, and altering any of them can lead to unpredictable changes in performance.
There are multiple ways to construct a template for a task. The most straightforward approach is to use minimal templates () or universal verbalizers like “input/output”, as done in Wang et al. (2023) and Wei et al. (2023).
Another strategy is to create task-specific templates. Jiang et al. (2020) generate paraphrases of templates for the relation extraction task. Authors show the sensitivity of masked language models to the prompt format and propose to ensemble predictions over the best templates. Compared to this method, our approach is task-agnostic and does not require evaluating all templates in advance.
Several studies aim to find templates that directly optimize in-context learning performance Shin et al. (2020); Gao et al. (2021). Our work unifies the results of previous research, using the verbalizers proposed by Gao et al. (2021), as well as minimal and universal templates.
Choice and Order of Demonstrations
The choice of examples for ICL is highly important, as they enable the model to condition on correct input and label distributions for the task Wu et al. (2023); Nguyen and Wong (2023); Min et al. (2022d). Furthermore, the order of examples also significantly affects the results and does not transfer between models even within the same family Lu et al. (2022b); Zhao et al. (2021).
In this work, we analyze two recent methods for selecting demonstrations. Wang et al. (2023) propose learning latent concept variables for a task and using them to find examples that can best predict the task concept. We refer to this method as Implicit Topic Models or ITM. In turn, z-ICL Lyu et al. (2023) generates pseudo-demonstrations by retrieving most similar examples to the test sentence from an unlabeled dataset and assigning random labels to retrieved examples.
Crucially, both methods are evaluated on single templates that differ across two works. Therefore, it is unclear whether the reported performance gains arise from the methods themselves or from a particular combination of the example selection strategy, the model, and the chosen template.
Prediction Methods
The standard approach for classification with LLMs is to compute the sequence probability with each of the possible labels and select the label with the highest probability. We refer to this method as Direct further on.
Alternatively, one can use more advanced prediction methods that aim to reduce the variance across prompt formats. The Channel prompting technique, proposed in Min et al. (2022b), maximizes instead of . The Calibration method Zhao et al. (2021) computes a correction factor based on the deviation of the model’s predictions for a placeholder input from a uniform distribution over labels and applies this factor to test set predictions.
However, as we show in Appendix A, these methods are evaluated on their own sets of templates. In this paper, we strive for a more unified view on the robustness of advanced prompting methods and compare their performance across a broader range of templates and models.
Prompt and Template Robustness
Although the problem of prompt robustness is relatively well-known, until recently, the discussion of template robustness has been limited. Notably, Sclar et al. (2023) present a highly relevant study of prompt format sensitivity, reporting a significant performance variation across formats even for large models or minor template changes. The primary differences of our work are a simpler structure of the template space and a focus on the impact of template sensitivity on the evaluation of ICL modifications.
Moreover, several works study prompt robustness in a broader sense by considering models that use natural language instructions Webson and Pavlick (2022); Leidinger et al. (2023); Weber et al. (2023) instead of labeled demonstrations. Recently, Mizrahi et al. (2023) have shown that very similar instructions can lead to drastic differences in task performance for a variety of instruction-tuned models. Although we study a similar issue, we focus on in-context learning and the transfer of best templates between evaluation setups. Still, our findings agree with the results for instruction tuning, which confirms the necessity of language model evaluation that takes prompt design into account.
Setup & Methodology
We evaluate the robustness of in-context learning to template selection across a wide range of models on classification tasks. All models used in our work are listed in Table 1: we run experiments on model families frequently used in literature (such as OPT and BLOOM), as well as the latest models with the highest quality (such as LLaMA 2 and Falcon).
In preliminary experiments, we observed that the performance of some models in the few-shot regime lags behind their zero-shot results. Consequently, we excluded th models from further investigation. Further details regarding this selection procedure can be found in Appendix B.
We experiment with 4 sequence classification datasets: SST-2 Socher et al. (2013), DBPedia ontology classification task Zhang et al. (2015a), AGNews Zhang et al. (2015b), and TREC Question Classification Li and Roth (2002). Although these datasets are frequently used in ICL studies, there is no consensus regarding the templates that should be used for each task.
One can construct an input for in-context learning from a set of demonstrations by using a template consisting of four parts, as illustrated in Figure 1. We present all options for verbalizers and separators for each dataset we study in Section 3.1. Any combination of these components results in a valid template. This set of options results in 216 possible prompt formats for SST-2 and 168 for DBPedia, AGNews and TREC. A single evaluation run of all models on 10 random templates in one setup takes 17–48 hours on a single NVIDIA A100-80GB GPU, depending on the dataset.
2 Prediction Methods
Next, we aim to evaluate the performance of different prediction methods in a unified setting. Ideally, we would like these modifications to reduce the variance across templates, making the model behavior less dependent on the input format.
We evaluate Channel and Calibration methods in the 2-shot setting along with the Direct baseline.
As depicted in Figure 2, both Channel and Calibration generally exhibit improved performance in comparison with Direct. Still, for a number of models and datasets, the range of scores for Direct substantially overlaps with those of advanced methods. This suggests that there are templates reaching the best performance with the Direct prediction method.
Additionally, Table 10 of Appendix E reveals that despite Calibration yielding the highest mean accuracy more often than other methods, it is more sensitive to the template choice than Channel. Therefore, the choice of the prediction method should likely rely on the downstream usage scenario and the target evaluation setting.
3 Example Selection Methods
Another area of ICL improvements that we evaluate on the matter of template sensitivity is the example selection strategy. We compare ITM and z-ICL methods to the Random baseline in 4-shot setting, since using 4 demonstration was the main evaluation setting in the works proposing these methods. We use Direct prediction method to evaluate the gains of advanced example selection strategies independently from other ICL modifications.
Results in Figure 3 and Table 12 illustrate that when taking template sensitivity into account, advanced example selection methods often perform worse than random choice baseline. ITM increases the average performance in most cases but still has a remarkably high standard deviation across templates. Examples selected using the z-ICL method lead to more consistent but worse performance.
Note that our evaluation setup differs from those in the works proposing these methods, which might explain the difference between our findings and the results reported in original works. Namely, we use the Direct prediction method and sample 10 random templates that might have not included the templates used by authors of ITM and z-ICL.
We conclude that the prompt format should be viewed as important as the example selection or the prediction method in ICL evaluation. However, the search space of possible templates is infinite, which makes exhaustive search for each combination of the dataset, the model and the number of examples impractical. Ideally, the best template for one setting would be optimal for all others or at least for similar settings. However, as we demonstrate in the following section, this is not the case.
We begin by defining a successful transfer between ICL settings. In order to do so, we evaluate how the quality of model predictions varies across 30 random templates from Section 3.1. The results described in Appendix G demonstrate that the top-10 template on average yields 90% of the best template score. Therefore, if a prompt format is present in top-10 for both of the two compared setups, we can consider this an instance of successful transfer.
To compare sets of the best templates for a pair of settings, we compute Intersection-over-Union (IoU), also known as the Jaccard similarity coefficient Jaccard (1912), for top-10 best-performing templates in each setting. We also considered using the rank correlation coefficient Spearman (1904) as another measure of template transfer. However, its value can increase when low-performing templates have similar rankings in different ICL setups, while the transfer of efficient templates remains low. Still, we provide the results for this metric in Appendix H.
2 Transfer Between Models
Next, we analyze the transfer of the best-performing templates between models in the baseline setup. Specifically, we collect the results of each model in the 2-shot learning setting with Direct prediction method and Random demonstrations (fixed throughout the experiments) for 30 templates. A heatmap of IoU for the transfer of top-10 best templates between 19 models on the DBPedia dataset is presented in Figure 4; for other datasets, please see Appendix I.
We observe that the IoU values exceed 0.5 only for a few model pairs on all datasets, meaning that the capacity for template transfer between models in the same setup is generally low. This is especially concerning for models within a single family: as these models are trained on the same data and have the same architecture, one would expect them to perform similarly on the same prompt formats.
These observations lead us to conclude that comparing ICL methods across models with a single template can lead to incorrect conclusions: a template that is effective for one model can easily be one of the worst choices for another model.
3 Transfer Between Prediction Methods
As discussed in Section 4.2, no prediction method we evaluated can consistently outperform others across all models and datasets. Therefore, to find an optimal setup for a new ICL improvement, one needs to evaluate every prediction technique in multiple templates. We investigate the possibility of finding a universally optimal prompt for different methods to reduce the total computational cost.
To answer this question, we calculate the IoU between top 10 performing templates for each method for a fixed set of demonstrations. Results in Table 4 display that similarly to the models, the transfer between prediction methods is also low. Consequently, the prompt format sensitivity issue creates a burden on authors of new ICL modifications; they must tune templates for every prediction method they want to combine with their own approach.
4 Transfer Between Demonstration Selection Methods
Having found that the best-performing templates are specific both to the model and the prediction method, we now aim to find whether the best formats would be the same for different demonstration sets in the same setup. Similarly to previous experiments, we calculate IoU for 10 templates that yield the highest scores for each method.
Results in Figure 5 illustrate that simply adding demonstrations, even if they were obtained with the same method, can significantly alter the ranking of the best templates. This justifies the necessity to evaluate example selection methods on a range of templates to avoid misinterpretation of the results.
5 Discussion
Based on the above findings, we conclude that the results of evaluation of various ICL improvements without consideration of template sensitivity issue are hardly reliable for several reasons. First, as the best templates do not transfer between models even within the same family, scoring a method across several models using the same format will inevitably lead to underestimation of the method for all models except the one for which the format was tuned. Next, as there is little evidence of transfer between setups, the format selection procedure needs to be precisely described and applied in all evaluated settings for a fair comparison. Moreover, a comparison can be unfair even for a non-optimized template, as it introduces an element of chance: a randomly selected prompt format might perform well for one setup but poorly for another.
In summary, we find that there are no universally well-performing prompt formats. Therefore, the results of in-context learning evaluation can be reliable only if they are aggregated over several templates or if each setting is evaluated in its best-performing template. The former approach requires accounting for the variance of the scores and makes comparison less apparent, while the latter can be computationally expensive.
To reduce the variance in performance caused by the template choice, we propose to ensemble model predictions across multiple templates. This approach is widely used in machine learning Ho (1995); Lakshminarayanan et al. (2017) for improving the predictive performance of the model, as well as its robustness, and can be viewed as a form of test-time augmentation Krizhevsky et al. (2012); Simonyan and Zisserman (2015).
Formally, our method computes label probabilities across predictions for each of templates, where is the ensemble size, and outputs the label with the highest average probability. In early experiments, we tried selecting the most common label among the predictions; however, we found this voting strategy to perform poorly on tasks with a large number of classes. It is also important to note that ensembling predictions involves running the model times more compared to single-format evaluation, which makes this approach more computationally intensive. We view template ensembles as the simplest initial solution for the problem of prompt format sensitivity and leave the exploration of more efficient methods to future work.
We begin with determining the minimal ensemble size that consistently reduces variance while increasing the average performance. We observe that for the majority of models and prediction methods, an ensemble achieves the best accuracy when its size reaches 4 or 5 (see an example in Figure 6), with further expansion being less effective. We also found that smaller ensembles may demonstrate unstable behavior, with the possibility of a drop in performance if a suboptimal template is sampled. Therefore, we report results for ensembles of size 5 and average the results over 5 random seeds.
Next, we evaluate the performance gains of Template Ensembles for different prediction methods. Our findings in Table 5 and Appendix J indicate that ensembles increase the accuracy for all evaluated models and prediction methods. Most importantly, they also significantly reduce the variance associated with the template choice for most setups. Therefore, we can conclude that template ensembling allows to preserve the increase in accuracy provided by ICL modifications while mitigating the template sensitivity issue.
In this work, we study the inconsistencies in the evaluation of in-context learning advancements introduced by the template sensitivity of large language models. Specifically, we find that ICL improvements exhibit high variation across template formats and that it is not possible to reuse the same template across different modifications. This aspect is often overlooked in prior work, despite the fact that the impact of template selection on prediction accuracy may be comparable with the choice of demonstrations or prompting methods.
While we propose Template Ensembles as an initial solution to this problem, the general sensitivity of language models to minor prompt variations is yet to be addressed. Consequently, we believe that the research community should take this problem into account when developing new models, evaluation benchmarks, or in-context learning methods.
Due to limited computational resources and the high cost for evaluation on a large range of models, we only focus on four classification datasets. Moreover, we only compare two example selection methods to a random baseline, potentially overlooking other effective approaches.
Additionally, the space of templates could be expanded for more comprehensive experimentation. For example, we did not explore label mapping, including random labels, which is an important aspect of the template.
We would like to notice that our study focuses on a template selection impact on a performance and a degree of template transfer between different setups but not on templates themselves. Future work should further analyze not only which templates lead to a change in performance but also on why they affect it.
Finally, we only evaluate decoder-only models pretrained on next token prediction objective. While instruction-finetuned models are a popular choice for downstream tasks nowadays, we focus on standard ICL evaluation setups that frequently involve only the base pretrained models. Still, concurrent work has shown that the same issues of sensitivity to the input format arise for instruction-tuned models as well.
Appendix A Templates from Prior Work
Tables 6, 7, 8 and 9 provide a comparison of all the templates used in the works presenting all methods we evaluate. Noticeably, prompt formats (and the choice of label words for some formats and datasets) used in works proposing investigated methods have no intersection. This is also concerning, since the original papers proposing these methods refer to each other. For instance, Channel prompting outperforming Calibration in Min et al. (2022b) might be explained by selecting a more favorable set of templates for the method proposed in the paper rather than by the advantages of the method itself.
Appendix B Model Selection
Our initial evaluation pool consisted of 23 models. We evaluated each of them in 0-shot and 2-shot settings with three prediction methods on four datasets, resulting in 12 runs. For each run in both 0-shot and 2-shot setups, we compare the model performance averaged over 10 random templates.
Based on the results presented in Table 10, we restricted the final pool of models for evaluation to those that have a consistent increase in performance in the 2-shot setting, in other words, to those demonstrating a performance boost from ICL. More specifically, we kept the models that had 8 or more wins in 2-shot evaluation against 0-shot.
Appendix C Full Baseline Results
Table 11 shows the results of evaluation of all 19 models in the default setting with a varying number of few-shot examples. These results illustrate that the template sensitivity issue is present in all models regardless of their size, and is not efficiently mitigated with the increase in the number of demonstrations.
Appendix D Template Parts Analysis
In addition to studying prompt format sensitivity in general, we analyze how each part of a template impacts model performance. For instance, it could be possible that the inclusion of a certain verbalizer in a template consistently leads to a decline in accuracy, irrespective of the other components.
To find that out, we decompose all templates into their parts and measure the distribution of scores for different variations of each component separately. The results presented in Figure 7 illustrate that even for state-of-the-art models, such as LLaMA 2 70B and Falcon 40B, many components exhibit high variance; also, the variance differs between two models. In other words, even if a certain template yields good performance and low variance for a given setup, it is not guaranteed to work consistently well in other setups, and changing a single component could have detrimental effects.
Along with the non-transferability of whole templates, we notice that individual components also do not transfer both between models and prediction methods. For instance, while “It was {}” ranks highest among output verbalizers for LLaMA 2 70B with the Direct prediction method, it is one of the worst for Falcon 40B.
Moreover, while a combination of best verbalizers is often a well-performing template, it is not necessarily the best one; the same is applicable for “bad” verbalizers too. For example, “input: {}\n sentiment: {}\n\n” is the best template for Falcon 40B with the Direct method, even though “sentiment: {}” is one of the “worst” output verbalizers for that model.
In summary, there is a complex interaction between the components of a template and their influence on model performance. We hypothesize that the transfer of both whole prompt templates and their parts is limited and requires further analysis.
Appendix E Prediction Methods
We provide the results of advanced prediction methods evaluation for all models in 0-shot and 2-shot setting with random demonstrations in Table 10. We conclude from this comparison that neither of the advanced prediction strategies do not decrease prompt format sensitivity consistently across models and datasets. Moreover, when accounting for the spread in accuracy scores caused by this issue, the advantages of these methods over Direct become less apparent.
Appendix F Example Selection Methods
Full results of evaluation of various demonstration selection techniques in 4-shot using Direct method are presented in Table 12.
The evaluation highlights that intricate example selection techniques rarely significantly outperform the default random selection baseline when evaluated on a set of random templates. One would argue that the prompt format choice is inseparable from the method itself and thus such comparison is invalid. However, since the best performing formats cannot be transferred neither between different models, nor even between the sets of different sizes selected with the same method, a proper evaluation of each method would require finding the best template for each setup. Not only this procedure is computationally expensive, but sometimes is simply impossible as authors of proposed methods frequently omit describing their prompt selection algorithm.
Appendix G Accuracy As A Function Of Template Rank
We plot the dependence of accuracy on the rank of templates in Figure 8. The results are aggregated across 19 models. Each model was evaluated on 30 random templates with Direct prediction method and the same set of 2 randomly selected demonstrations. We observe that for SST-2 and AGNews datasets the mean quality of the 10th-best template is around 0.9 of the best template score, which we consider a successful transfer. Despite the more rapid decay for DBPedia and TREC, considering variation across models, we still count first 10 formats as performing on par with the best one.
Appendix H Transfer Evaluation With Spearman Rank Correlation
One of the possible means to evaluate template transfer is to calculate the Spearman rank correlation between scores of all templates. As can be seen from Figure 9 this method yields higher correlations than calculating IoU over 10 best formats, but the transfer is still far from perfect (for example, for SST-2 and TREC datasets).
Appendix I IoU Transfer For All Datasets
Similarly to Figure 4 we provide Intersection-over-Union of 10 best prompt formats for all 19 models on all datasets explored in our work in Figure 10. These heatmaps illustrate that transfer of the best-performing templates between models is extremely low for all datasets.
Appendix J Additional Results For Template Ensembles
Tables 13 and 14 show the results of Template Ensembles evaluation on a broader set of models and datasets. For most setups, ensemble of size 5 exhibited better performance than a single template.