State of What Art? A Call for Multi-Prompt LLM Evaluation
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, Gabriel Stanovsky
Introduction
Recent years have seen an explosion of large language models (LLMs), which generalize to unseen tasks via natural language instructions. Various LLM evaluation benchmarks, such as BIG-bench and HELM, use a single instruction template per task, evaluating all models against it (Srivastava et al., 2022; Liang et al., 2022). However, there could be a myriad of ways to phrase an instruction template for a given task; see Figure 1 for examples of different templates for the task of recognizing homophones. Naturally, LLM performance depends on the chosen template.
In this work, we explore the question of robustly comparing different models on a given task. We first create a dataset of paraphrased instructions. To achieve this, we devise three automatic methods to paraphrase given instruction templates, based on recent prompting techniques such as chain-of-thought. We manually verify and filter a large collection of more than 175 paraphrases for different tasks (5K instruction paraphrases in total), which we make publicly available for future research.github.com/SLAB-NLP/Multi-Prompt-LLM-Evaluation
Next, we use our dataset to perform a large scale statistical evaluation of over 6.5M instances, involving 20 different LLMs and 39 tasks from 3 benchmarks. We find that models perform very differently on different instruction paraphrases, both in terms of absolute and relative performance. Figure 1 shows an example of the performance of four models on four (semantically equivalent) prompts, with both absolute performance and relative ranking varying widely. At the extreme, there are instruction templates on which a model performs the best compared to other models, while on a semantically equivalent instruction the same model performed the worst (e.g., GPT-3.5-Turbo on vs. ). Subsequently, we argue that very little can be said on either absolute or relative performance based on the common practice of single-instruction evaluation (which may partially explain why some models seem less accurate in practice than their formal evaluation may suggest).
Note that while the claim that evaluating against a single instruction template leads to brittle results is not surprising per se, to the best of our knowledge it has never been subjected to rigorous empirical testing before.
To address the limitations of single-instruction evaluation, we propose to take a step back and consider multi-instruction evaluation metrics which are closely tied to real-world use cases of LLMs. We argue that different use cases should entail different evaluation metrics. For example, LLM developers may be interested in measuring the robustness of performance across multiple instruction templates, which we formulate as the average performance across a large collection of instructions. In contrast, when focusing on a downstream task, different models may be better compared according to their corresponding top-performing instruction.
We evaluate 20 LLMs with our metrics, finding that their absolute and relative performance differ from those obtained with the benchmarks’ original instruction templates. We demonstrate that different models excel in different metrics: For instance, in the LMentry benchmark, LLaMA-based models are comparable to T5-based models when looking at top-performing instructions. However, these models lag behind when average performance is considered, due to poor performance on a large number of paraphrases. We also show that our automatic paraphrasing method is effective, and there is no need to manually verify the paraphrases.
Our results suggest that future work should choose the evaluation metric based on the extrinsic needs of the evaluators. We hope that our work will help spur more consistency and comparability in LLM evaluation, which is strongly tied to real-world usage of LLMs.
Background and Definitions
Below we survey how generalization to a new task format is evaluated and compared between LLMs, finding that this is normally done by testing performance on a single (or very few) task instruction templates. In the rest of the paper, we will argue that such practice leads to brittle results which are not well-suited for real-world use of LLMs.
Following Mishra et al. (2021); Chung et al. (2022), we separate between task instruction, samples, and input-output exemplars which may be provided during in-context learning. We define an instruction template for a given task as a string with placeholders where the input samples are to be inserted. As seen in Figure 1, the same task can be described using different task instruction templates.
Evaluation benchmarks.
Several recent efforts aim to standardize LLM evaluation. Notable examples include MMLU (Hendrycks et al., 2020), BIG-bench (Srivastava et al., 2022; Suzgun et al., 2022), and HELM (Liang et al., 2022). In all of these, each task has a single instruction template, against which all models are evaluated. Another benchmark, LMentry (Efrat et al., 2022), reports models’ average performance on three instruction templates. The instruction templates are provided with these benchmarks, allowing new models to be tested against the same template.
Sadly, it is also common practice to report results based on a single instruction template without making it publicly available (e.g., LLaMA (Touvron et al., 2023), PALM (Chowdhery et al., 2022), GPT-4 (OpenAI, 2023), and Gemini (Google, 2023)). This exacerbates the challenge of meaningful comparative evaluation.
Prompt robustness.
Related to this study is a line of work measuring LLM’s robustness to prompt (or instruction template) modifications. Unlike our work, these typically aim to measure model performance against adversarial paraphrasing approaches. PromptBench (Zhu et al., 2023) measures performance on erroneous instructions (e.g., instructions written by non-native English speakers). They then compare performance on perturbed instructions vs. the benchmark’s original instructions, which are considered the gold-standard reference. Gu et al. (2022) examined a single LLM’s robustness under various instruction perturbations, including word-, sentence-, and instruction-level changes. Sun et al. (2023) show that LLMs perform better on instructions they have seen in training (BIG-bench Lite benchmark), compared to manual paraphrases. We later incorporate their manual paraphrases in our evaluation.
In contrast to works on prompt robustness, our scope is wider. We analyze the impact of the choice of prompt in terms of both absolute and relative model performance, covering a wide range of models and several different metrics.
Experimental Setup
In this section we describe the tasks and models which we evaluate in this work.
We evaluate 39 diverse tasks from three evaluation benchmarks, as itemized below, and summarized in Table 6 in the Appendix.
LMentry consists of simple linguistic tasks (e.g., “write a word that doesn’t contain the letter ”), each accompanied by three associated instruction templates. The tasks are designed to capture explainable and controllable linguistic phenomena. We choose 10 tasks from LMentry that received the lowest scores in the original paper.
tasks from BIG-bench Lite (BBL; Srivastava et al., 2022).
These cover multiple knowledge domains, sampled from the larger BIG-Bench benchmark (bench authors, 2023). In particular, we focus on a set of 14 tasks studied recently by Sun et al. (2023). Each task in BBL is associated with a single instruction template.
tasks from BIG-bench Hard (BBH; Suzgun et al., 2022).
This is another curated subset of BIG-bench, containing particularly challenging tasks on which LLM underperform the average human-rater score. We take the set of 15 classification and multiple choice tasks from BBH to ease the evaluation protocol. Each task in BBH is associated with a single instruction template.
Measuring performance.
We measure performance in the standard manner provided by each benchmark. In LMentry this is done with the official evaluation script, while in Big-Bench we use exact match evaluation. We note that while this evaluation is somewhat strict, we believe that it is also fair and straightforward.
2 Models
As shown in Table 1, We evaluate 16 instruction-tuned LLMs from 11 diverse model families (Chung et al., 2022; Sanh et al., 2021; Taori et al., 2023; Zheng et al., 2023; Durbin, 2023; Ding et al., 2023; NousResearch, 2023; Almazrouei et al., 2023; Team, 2023; Collective, 2023). We refrain from including any closed API-based models (e.g., OpenAI models) in our main evaluation for two reasons. First, using them at scale is an expensive prospect, for example, running our entire evaluation suite on GPT-4 will cost up to 2500 USD. Second, and more importantly, the closed API for these models reportedly manipulates the input prompts in an undisclosed manner (e.g., wrapping them with meta-prompts, or rerouting to other models) (Rao et al., 2023) which interferes with our evaluation. We do however perform a small-scale evaluation of OpenAI models in Section 7 to show that they are also sensitive to prompt paraphrasing.
Single-Prompt Evaluation Leads to Inconsistent Results
As discussed in the previous section, a common practice in LLM evaluation is to evaluate different models against a single instruction template. In this section, we will show that this approach is quite brittle. Indeed, a simple rephrasing of the instruction template can lead to drastic changes in absolute model performance as well as its relative ranking among other models.
To show this, in Section 4.1 we create a large number of instruction paraphrases for each of our tasks. This is achieved automatically with the aid of an LLM and verified by human annotators to reduce noise. Then, in Section 4.2, we statistically analyze the performance of various LLMs against these instruction templates and quantify the variation in model performance and ranking.
We use three prompting methods which were found useful in previous work: (1) instruction template rephrasing: asking an LLM to rephrase a seed prompt (Lester et al., 2021; Gonen et al., 2022; Honovich et al., 2022a); (2) Chain-of-Thought prompting (Wei et al., 2022): we provided the model with a sequence of steps in which the model is asked first to produce a task description, and then to generate various instruction templates for the task; and (3) Gradual template generation: inspired by Honovich et al. (2022b), we split the COT approach into three LLM calls. The first for generating a task description from a seed instruction template, the second for generating instruction provided by input-output examples, and the third for processing the instruction and examples into an instruction template. See more details about these approaches in the Appendix.
We use the original instruction templates for each of our tasks to seed these three generation methods, resulting on average in more than 200 automatically-generated instruction template paraphrases for each of our tasks (see Table 2). We make this collection, as well as the code used to generate it, publicly available for reproducibility and to enable future work.
We manually verify and filter all of the automatically generated paraphrases.We found that 90% of the generated paraphrases created for LMentry were correct, and roughly 84% of the paraphrases for BBH were correct. See Table 2 for a fine-grained distribution across the different generation metrics. On average, this process yields more than 175 validated instruction paraphrases per task across LMentry and BBH, which we will subsequently use to quantify peformance variability due to instruction template paraphrasing.
2 Quantifying Performance Variance due to Instruction Paraphrasing
We leverage the collection of validated instruction paraphrases to show that model performance varies widely on different instruction templates, both at the individual model performance, as well as in relative model ranking. As we argue below, our main finding is that the common approach of evaluating against a single instruction template is inconsistent and unstable, leading to contradicting results.
Evaluating LLMs can become prohibitively expensive with the increase of the number of samples, datasets, models, and instruction templates (Perlitz et al., 2023). We focus on a large number of tasks, models, and instruction paraphrases. Hence, to make our evaluation feasible, this comes at the expense of the number of samples per task. Concretely, we evaluate each instruction template on a randomly selected subset of 100 task samples. Furthermore, we found that all models struggle on BBH, beyond the point of meaningful comparison. To address this, we evaluate 11 out of the 16 models on it (the bigger ones in terms of number of parameters), and we add an example of the prediction format to all instruction template paraphrases. Examining the effect of few-shot learning is beyond the scope of this paper, however, Sclar et al. (2023) recently observed similar performance sensibility when introducing varying number of in-context examples.
Using a single-instruction template leads to brittle ranking.
Kendall’s would be for all tasks if model ranking were the same among all instruction templates (in other words, they are interchangeable for the sake of evaluation). In contrast, the more approaches , the lesser the rankings induced by different instructions agree.
The results (Table 3) demonstrate that a single instruction template leads to unreliable rankings for many of the tasks, with 10 of the tasks exhibiting only slight to moderate ranking agreement, and only two exhibiting strong agreement. To complement the analysis, we performed Friedman test with tied data (Corder and Foreman, 2011), showing that different instructions lead to statistically significant differences in performance for 21 out of the 25 tasks.
Examples of differences in model ranking.
We illustrate the implications of such differences in Figure 2. The three instruction template pairs are valid paraphrases, yet they lead to vastly different results. For example, T0pp ranks first on the BBH task using the first instruction template and only 9th using the second template. Similarly, Alpaca-13B and Alpaca-7B are in the top performing models on the LMentry task using the second instruction template, while they rank last in the first template.
Where is the number of concordant pairs, is the number of discordant pairs, is the number of ties in the first ranking, and is the number of ties in the second ranking. Therefore, indicates that most pairs are concordant (with indicating perfect agreement), and indicates that most pairs are discordant (with indicating perfect disagreement).
Appendix A.4 presents examples of pairs of instruction templates that exhibit the minimal Kendall correlation per task (i.e., their ranking is most dissimilar). Overall, 15 tasks have instruction template paraphrases with negative Kendall’s , indicating mostly disagreeing LLM rankings.
Absolute model performance varies widely on single-instruction templates.
Aside from vastly different relative model rankings, instruction template paraphrases often result in varying absolute model performances. To quantify this variance, we calculated divergence, defined as the number of standard deviations by which the performance, as assessed using the original instruction templates, deviates from the model’s average performance over all paraphrases.
The results in Figure 3 reveal noticeable divergence for the LMentry benchmark, defined as surpassing one standard deviation (Kazmier et al., 2003). For instance, the performance of the Alpaca-13B with the original instruction templates outperformed its average performance by more than one standard deviation in 7 out of the 10 LMentry tasks. For lack of space, the figure does not depict the BBH benchmark, but similar patterns of divergence were observed there as well.
In line with Lou et al. (2023), we find that major differences in performance can occur even for very similar paraphrase pairs. For example, the Flan-T5-large model demonstrated an average performance degradation of 28% when changing the word ‘excludes’ to ‘lacks’, while the Flan-T5-XL model showed an average performance improvement of 46% on that same edit. See a comprehensive edit distance comparison in Appendix A.5.
3 LLMs are also Sensitive to Manual Paraphrases
It is possible that the inconsistencies observed in our analyses stem from our automatic paraphrases. To address this, we extended our analysis with instruction paraphrases which were recently written by Sun et al. (2023) for the BBL tasks (see Table 6). These provide between 7 and 12 instruction templates per task. While originally annotated to examine overall model degradation on human written instructions, we reuse Sun et al. (2023)’s annotations to examine the change in model rankings and absolute performance.
Our analysis revealed similar inconsistencies as observed with automated paraphrases. See Table 13 in the Appendix for the Kendall’s W values for all BBL tasks, and Table 11 for examples of pairs of instruction templates that exhibit the minimal Kendall correlations.
Different Use Cases Merit Different Metrics
So far we have shown that LLM performance is greatly affected by paraphrasing of instruction templates. This calls into question current evaluation practices, which typically rely on LLM performance on a single instruction template. In this section we explore ways to evaluate LLMs using a diverse set of instruction templates.
Most importantly, we argue that the answer should depend on the purpose of the evaluation, and that different extrinsic needs should lead to different evaluation metrics, rather than striving for a coarse catch-all metric. We introduce a set of metrics, each tailored to specific scenarios and realistic user needs.
In the following, is a pretrained LLM, denotes an evaluation dataset for , is a set of natural language task instruction paraphrases for (e.g., obtained via automatic paraphrasing), and denotes the aggregated performance of on samples from , using a single instruction template according to a standard metric, e.g., accuracy or .
1 Maximum Performance Metric – For Particular Downstream Applications
We define the maximum performance (MaxP) of a model on task to be the maximum individual instruction template performance this model achieves across all instruction templates:
Use case: This metric is useful for developers aiming to integrate an LLM into a specific downstream task and domain (for example, sentiment analysis in the news domain). In such cases, a user input is often embedded within a fixed instruction template. As such, it makes sense to find the best-performing instruction template for a given model (Wei et al., 2021). To mitigate overfitting, it is sensible to identify it using a held-out sample.
2 Average Performance Metric – For LLM Developers
We define the average performance (AvgP) of a model on task as the mean of the individual instruction template performances over all instruction templates for the task:
Use case: Average prompt performance is useful for assessing model robustness to paraphrases. We believe this should be standard practice for LLM developers when presenting the performance of a new LLM on a range of tasks and prompt paraphrases (Workshop et al., 2022), as it mitigates outliers in performance.
3 Combined Performance Score
In the same way the F1 score combines precision and recall into a single metric, we propose a Combined Performance Score (CPS) that unites the maximum and average performance metrics to capture both peak capability and consistency of the model across prompts. To define CPS, we first introduce a model saturation score:
This score measures how closely the model’s best performance aligns with its average performance. A high saturation score indicates that the model’s performance does not drop significantly for non-optimal instructions. Then, the CPS is calculated as the product of the model’s best performance () and its saturation ():
Use case: This metric is valuable for selecting a model for a suite of applications or a platform offering diverse tasks. For instance, when integrating an LLM into an application with user-visible prompts, such as a multi-functional chatbot, it is crucial for the model to be both effective (high ) and consistent (high ). CPS facilitates identifying models that strike a balance between top-tier performance and consistent reliability across varying instruction templates.
Multi-Prompt Evaluation
In Figure 4 we evaluate all our 16 models according to the metrics we proposed in the previous section, on sample tasks from each of the three benchmarks (full results for all tasks are available in our repository). We report several interesting observations.
First, we find that all aggregate metrics diverge from the performance on the original instruction templates. For the vast majority of the tasks in our study, the top three models determined by the original instruction templates were different from those which ranked first according to the average and maximum metrics.
More broadly, the rankings of models depend on the metric used. For instance, see Figure 4 (top): In LMentry’s rhyming word task, Falcon-Instruct-7b and Vicuna-13b rank first according to (0.74, gray and yellow bars), but their average performances are only 0.17 and 0.15, respectively. Similarly, across all tasks in the LMentry benchmark, LLaMA-based models were competitive with T5-based models in terms of . However, in terms of , they tended to lag behind, due to extremely poor performance on a large number of paraphrases (see Figure 5 for percentage of paraphrases that achieved over 5% accuracy).
Finally, we found that noise stemming from automatic paraphrase generation has virtually no impact on metric-based model rankings. We compute Kendall’s to compare model rankings before and after the manual removal of incorrect paraphrases. The results (Table 4) show near-perfect to perfect agreement in rankings across all tasks, except for the “ends with word” task in LMentry. Upon examination, this seems to be mostly due to an error in LMentry’s evaluation script. These results suggest that it may be enough to compute our metrics over range of automatically-generated paraphrases, without having to manually verify them.
Small-Scale Evaluation of OpenAI Models on Prompt Paraphrasing
In this section we perform a small-scale evaluation showing that API LLMs are also sensitive to instruction paraphrasing. Our evaluation focuses on four OpenAI models: davinci, text-davinci-002, text-davinci-003, and GPT-3.5-Turbo on the LMentry benchmark.
Due to budget constraints, we show that the performance of these models diverges significantly between the benchmark’s original instruction templates and a selection of paraphrases, in terms of both average and maximum metrics.
To estimate the average performance of OpenAI models on a specific task, we adopted a randomized approach. For each task sample, we randomly selected a paraphrase from our collection, and evaluated the model’s response, scoring the entire set of task samples. To approximate average performance, this experiment was repeated 20 times, determined by the data from our 16 open-source models.
Estimating maximal performance.
To estimate which of the roughly 175 instruction templates per task performs the best for each model, we implemented a simple greedy search. Initially, we evaluated all paraphrases on 10 task instances, then narrowed down to the top 100 instruction templates for another 10 instances. Finally, the top 10 instruction templates were evaluated on the remaining instances, and the template that performed the best was chosen to estimate the maximum performance.
1 Results
Below we summarize the results of our evaluation of OpenAI models. The full details appear in Tables 21, 22, 23, and 24 and in our repository.0
Minor changes in the phrasing of the instruction can lead to drastic performance changes for the OpenAI models in our experiment, similar to our findings in Section 4.2 with smaller-scale LLMs. See representative examples in Table 5, showing nearly identical instruction template pairs resulting in notable variations in performance.
Average multi-prompt performance is lower than that observed in the original benchmark instructions.
In 72.5% of the cases, the performance of the original instruction templates was higher than the estimated average across all paraphrases. A prominent difference was observed particularly in the davinci model. For this model, the original prompts added, on average, 21 more accuracy points compared to the estimated average across all paraphrases.
Original prompt performances fall below all paraphrases’ estimated maximum performance.
Figure 6 depicts maximum performance of the original instructions for four LMentry tasks in solid colors, with overlaid semi-transparent columns indicating the estimated maximum performance on all paraphrases. Notably, for text-davinci-002, we found paraphrases that improved its maximal accuracy performance above 90% for 8 out of 10 tasks. Across all four models, 26 out of 40 differences were statistically significant according to the McNemar test (Table 25).
Model rankings diverge between the different metrics and original instruction templates.
Similarly to our main evaluation, there were many mismatches between ranking on the original instruction templates and our metrics. Agreement was observed in only 5 out of 10 tasks for the average metric, and in 4 out of 10 tasks for the maximum metric.
Related Work
Our work is part of an emerging trend highlighting the many challenges standing in the way of meaningful, scalable, and reproducible evaluation of large language models.
Perlitz et al. (2023) focus on the rising cost of exhaustive evaluation of LLMs on large number of samples. They notice that as models become larger, the cost of running them at scale can become prohibitively expensive, even during inference. To help mitigate this problem, they develop methods for choosing subsets of the test data which are expected to be a good representative of the whole. We find that single prompt evaluation is not a good representative of LLMs average performance, and instead suggest evaluating on many instruction templates per sample, which further increases the evaluation cost. An interesting avenue for future work can extend Perlitz et al. (2023)’s approach to also include various instruction templates, thus efficiently approximating our suggested evaluation methods.
Sclar et al. (2023) show that LLMs are sensitive to prompt formatting. These are minor prompt design choices, such as the addition or omission of punctuation marks. They create a large pool of instruction paraphrases, ensuring that paraphrases maintain the meaning of the original prompt. We notice a similar phenomenon, albeit more anecdotally, when our automatic paraphrasing techniques incidentally produce minor changes in formatting (Table 5). Finally, Voronov et al. (2024) shows that LLMs are sensitive to how in-context examples are presented and formatted. For example, they vary the manner in which each input-output is separated, and test how such choices interact with the phrasing of the instruction template, the number of demonstrations, or the model size.
Our work distinguishes itself as the first to systematically explore the impact of a broad spectrum of prompt paraphrases across various benchmarks and tasks on multiple models, coupled with a statistical analysis of the absolute and relative variations in evaluations. Furthermore, we introduce a suite of metrics specifically designed to align with the practical applications of large language models.
Conclusions
Our research highlights the sensitivity of large language models (LLMs) to prompt paraphrasing, challenging the adequacy of single-prompt evaluations. We propose alternative evaluation metrics that use a diverse set of instruction templates for each task, designed for more robust and meaningful LLM evaluation. For example, LLM developers may be interested in measuring the robustness of performance across multiple prompts, which we propose to evaluate as the average across a large collection of prompts. In contrast, when developing a downstream model, different models should be compared according to their corresponding top-performing prompt.
Evaluating based on these metrics underscores the necessity for nuanced evaluation methods, revealing notable differences in absolute performance and relative model rankings compared to traditional evaluations. We hope that our work will help spur more consistency and comparability in LLM evaluation which is strongly coupled to real-world LLM uses. We believe this shift is crucial for accurately understanding and leveraging the true capabilities of LLMs.
References
Appendix A Appendix
Table 6 presents an overview of the 39 tasks from the 3 benchmarks discussed in this paper: LMentry, BIG-bench Lite, and BIG-bench Hard. These benchmarks include 10, 14, and 15 tasks from each, respectively. The table also provides an example task instruction for each task.
A.2 Process of Generating Prompt Paraphrases
Our process for generating paraphrases of instruction templates is depicted with an example in Figure 7.
A.3 Paraphrases Correctness
Tables 7 and 8 present the percentages of correct paraphrases that were generated by the 3 prompt-generating methods presented in the paper for LMentry and BBH. The tables also depict the average model accuracy and standard deviations as measured for only the correct paraphrases across all LLMs. The correct paraphrases were identified by one of the authors of this paper. Table 14 presents the Kendall values before and after the removal of incorrect paraphrases. The agreement in the ranking of models is near-perfect to perfect in both LMentry and BBH benchmarks.
A.4 Comparing Different Instruction Templates with Kendall’s τ𝜏\tau Rank Disagreements
Tables 9 , 11, and 10 present the Kendall values of representative examples from all benchmarks with Kendall values that are significantly different from 0. i.e., notable variations in rankings of models for two paraphrases of the same task instruction.
A.5 Model Performance Differences with Minimal Paraphrasing Edit Distance
Figure 8 depicts the average performance differences between various LLMs when small edits are made to the instruction templates.
In addition, Table 12 shows representative examples of instruction template pairs with very minor differences but notable variations in performance.
A.6 BBL Analysis
This subsection consists of an additional analysis of the BBL benchmark that was not detailed in the main body of the paper. Table 3 presents the Kendall’s W values and the Friedman test p-values that demonstrate a low correlation between the ranks of the models for different instruction templates and reveal similar inconsistencies as observed with automated paraphrases in other benchmarks. Figure 10 shows the deviation of the original instruction template from the average performance calculated over the generated instruction templates of several models for all of the BBL tasks.
A.7 Average Model Ranks for Each Metric Across All Tasks
Tables 15, 16 present the average model ranks for each metric across all tasks in LMentry and BBH respectively. Flan-T5-XXL emerges as the top performer for all metrics in both benchmarks. Minotaur is at the bottom of the performance spectrum across all evaluated models in BBH.
A.8 Analysis of Origin Generation Method of Optimal Paraphrases
Our analyses for the origin of the optimal paraphrases used by each model, are summarized in Tables 17, 18. The gradual method surfaced as the dominant source of optimal paraphrases across both benchmarks, particularly pronounced in the LMentry benchmark. However, a closer look at individual models revealed a pattern of preference for different generation methods.
A.9 Small Scale Evaluation - OpenAI
This subsection contains all the tables referenced in Section 7. Table 19 and Table 20 are related to our naive heuristics for estimating average and maximum performance, respectively. Table 19 presents the average number of repetitions needed for our heuristic to estimate the average performance, ensuring less than a 1-point accuracy discrepancy from the actual average for each open-source model across all tasks in the LMentry benchmark. Table 20 compiles results from our greedy heuristic that searches for the optimal paraphrases for each open-source model on each LMentry task.
Table 21 and Table 23 aggregate the average and maximum performances for each model and task using only the original instruction templates. Similarly, Table 22 and Table 24 present the approximated average and maximum performances, computed with our heuristics, for each model and task using all paraphrased templates.
Table 25 contains the McNemar test p-values we used to assess the statistical significance of the differences in maximum performance between the original best prompt and the estimated optimal prompt.