Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Melanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane Suhr

Introduction

As the capabilities of LLMs have rapidly improved, their sensitivity to input prompt features has been used to optimize performance via prompt engineering (White et al., 2023). However, there has been little work in characterizing this sensitivity, especially to seemingly innocuous feature choices that preserve prompt meaning and intent. In this work, we analyze the sensitivity of widely used, open-source LLMs to a class of features that should not influence a prompt’s interpretation: formatting choices. We find that pre-trained LLMs are sensitive to these choices in unpredictable ways, with accuracy varying in up to 76 points for LLaMA-2-13B between equivalent formats, and ∼\sim10 accuracy points on average across 50+ tasks and several models. We also show that this variance is not eliminated by adding few-shot examples, increasing model size, or instruction tuning.

Designing prompt templates is a critical part of effectively using a pre-trained language model. This design process includes making choices about wording, choosing few-shot examples for in-context learning, and making decisions about seemingly trivial features like formatting. This process, and often even the resulting templates, is rarely reported or discussed in research papers, under the assumption that performance variance across these choices is insignificant compared to variance across data points or models. However, some anecdotal evidence points to formatting choices actually having a significant influence on model behavior (Aghajanyan, 2023). In some cases, researchers report a limited number of manually generated formats to show that scaling trends hold despite performance being significantly different (Schick et al., 2021). The assumption that formatting does not influence overall model performance may become problematic when improvements over existing approaches are attributed to the amount and source of training data, number of parameters, or model architecture, without also accounting for changes in prompt format. Ignoring variance across formats may also negatively affect user experience, e.g. if users inadvertently choose formats the LLM does not perform well on.

Our proposed tool, FormatSpread, enables a systematic analysis of these variances across a wide set of semantically equivalent prompt formats within a user-specified computational budget. We find that choices in formatting few-shot examples during in-context learning introduce spurious biases that may lead to significantly different conclusions in model performance. The sensitivity to formatting choices that we discover across widely-used, open-source models suggests that future research would benefit from reporting a performance spread over a sufficient sample of plausible formats, instead of simply reporting the formatting used and its performance, as is currently standard. Moreover, we argue that this reporting is crucial when comparing the performance of different models, as we show the influence of formatting choices only weakly correlates between models, thus making and fixing a formatting choice could introduce a significant confounding factor.

Fully exploring the space of prompt formats is intractable, as computation costs scale linearly with the number of formats considered. FormatSpread efficiently explores the space of prompt formats under a user-specified computational budget using Bayesian optimization. FormatSpread does not require access to the model weights, allowing its use on API-gated models: we find a spread up to 56 accuracy points with a median spread of 6.4 accuracy points with GPT3.5 across 320 formats and 53 tasks at a cost of under 10USD on average per task. Beyond facilitating evaluation, we also propose a suite of analyses to further characterize model sensitivity to formatting. Among other results, we show that the separability of continuous prompt embeddings correlates with the spread observed in task performance.

Overview

We evaluate LLM performance over the space of prompt formats that may plausibly be chosen by a non-adversarial user when designing a prompt for a target task, where the space of formats is defined by a grammar (§3.1). Our grammar’s definition naturally induces a definition of semantic equivalence among formats. We quantify model sensitivity in terms of performance range in a target task across the space of equivalent prompt formats to the original choice (§4.2). We cast the problem of searching across this space as a bandit problem, and propose FormatSpread (§3), which consists of a grammar (§3.1) and a procedure to estimate the minimum and maximum performance across a set of semantically equivalent formats given a pre-defined metric (§3.2). FormatSpread uses Bayesian optimization to identify the expected performance range with low additional computational cost (§4.5) all without requiring access to model weights, which enables use on API-gated LLMs. Furthermore, we perform in-depth analysis of this observed sensitivity, including by quantifying the contribution of individual feature choices to the final performance (§4.3) and measuring the identifiability of a format based solely on a model’s internal, continuous representation of any prompt via correlation with model performance (§4.4).

Measuring Sensitivity with FormatSpread

We construct a grammar that defines both the space of plausible prompt formats and semantic equivalence between formats. The grammar is manually constructed, as opposed to automatically induced from data, to guarantee a higher level of precision when defining the set of equivalent formats. Our grammar is directly tested by verifying that it can generate the formatting associated with 100+ Super-NaturalInstructions tasks (Wang et al., 2022).

Our grammar consists of fields that are composed to create a prompt format. For example, the format ‘Passage: || Answer: ’, has basic fields ‘Passage: ’, and ‘Answer: ’, denoted a1a_{1}, and a2a_{2}. Each basic field consists of a descriptor (e.g. ‘Passage’), a separator (e.g. ‘: ’), and a text placeholder to replace with each data point. We define basic fields as B1(d,s,f):=f(d)s<text>B_{1}(d,s,f):=f(d)s\text{{{{<text>}}}} using Backus-Naur notation, where dd is a descriptor string, s ⁣∈ ⁣S1s\!\in\!\mathcal{S}_{1} a separator, and f ⁣∈ ⁣Fcasingf\!\in\!\mathcal{F}_{\text{casing}} a function that alters dd while preserving meaning. Thus, in our example, a1 ⁣= ⁣B1(Passage,’ ⁣:  ’,id)a_{1}\!=\!B_{1}(\text{{{{Passage}}}},\text{{{{'\!:\ \ '}}}},id) and a2 ⁣= ⁣B1(Answer,’ ⁣:  ’,id)a_{2}\!=\!B_{1}(\text{{{{Answer}}}},\text{{{{'\!:\ \ '}}}},id), with idid the identity function. We define joining several fields as B2(n)(X1 ⁣,… ⁣,Xn, ⁣c):=X1cX2c…cXnB_{2}^{(n)}(X_{1}\!,\ldots\!,X_{n},\!c):=X_{1}cX_{2}c\ldots cX_{n}, with c ⁣∈ ⁣Cc\!\in\!\mathcal{C} being a space. Our example’s prompt format may be written as B2(2)(a1,a2,’ || ’)B_{2}^{(2)}(a_{1},a_{2},\text{{{{'\ ||\ '}}}}).

The grammar also supports enumeration, which is defined as joining several basic fields, each representing a different list item. For example, the enumeration ‘Option (A): , Option (B): , Option (C): ’ may be written as B2(3)(a1,a2,a3,’ || ’)B_{2}^{(3)}(a_{1},a_{2},a_{3},\text{{{{'\ ||\ '}}}}), where ai=B1(ei,’ ⁣:  ’,id)a_{i}=B_{1}(e_{i},\text{{{{'\!:\ \ '}}}},id). In our example, e1e_{1} represents ‘Option (A)’, and may in turn be written as the concatenation ei:=ds2fitem(i)e_{i}:=ds_{2}f_{\text{item}}(i) with d=‘Option’d=\text{{{{`Option'}}}}, s2=’ ’s_{2}=\text{{{{'\ '}}}} (single space), and fitem(1)=‘(A)’f_{\text{item}}(1)=\text{{{{`(A)'}}}}. Each fitemf_{\text{item}} transforms an item ii using a number format (e.g. letters or Roman numerals, denoted as Fitem2\mathcal{F}_{\text{item2}}) and an item wrapper (e.g. (A) or [A], denoted as Fitem1\mathcal{F}_{\text{item1}}).

In summary, we define valid prompt formats as those accepted by the following grammar:

Our grammar defines valid formats as finite compositions of B0,B0′,B1,B2,B3B_{0},B_{0}^{\prime},B_{1},B_{2},B_{3}. The sets S1,S2\mathcal{S}_{1},\mathcal{S}_{2}, C\mathcal{C}, Fcasing\mathcal{F}_{\text{casing}}, Fitem\mathcal{F}_{\text{item}} (two sets of separators, spaces, casing functions, and itemizing functions respectively) are pre-defined by the user. Throughout this work, we instantiate all sets with values typically observed in human-written prompt formats. We intentionally only modify the casing of descriptors (via Fcasing\mathcal{F}_{\text{casing}}) to guarantee semantic equivalence; one may also define a set of functions that paraphrases the descriptor, e.g., via synonym replacement. Appendix A.2 contains the full list of values we use for the constant sets, as well as a visualization of a prompt template generated from the grammar.

Two prompt formats p1p_{1}, p2p_{2} are equivalent if they represent the same rule application BiB_{i}, the descriptors (if any) are the same, and the sub-elements (if any) are equivalent. Appendix A.1 contains the formal definition of equivalence. The grammar’s strict definition allows us to assume that sets of equivalent formats share equivalent meanings. When measuring sensitivity (§3.2), we explore only the space of formats equivalent to a task’s original format.

We define restrictions to the combinations of spaces and separators to further ensure naturalness. For example, if B2(X1, ⁣…, ⁣Xn, ⁣c)B_{2}(X_{1},\!\ldots,\!X_{n},\!c) where cc does not contain a newline, then each XiX_{i}’s separators and any subcomponents’ separators should not contain a newline. This avoids unnatural formats like Input:\n Output:\n . We also allow for adding conditions that force constants (separators, spaces, etc.) in different applications of BiB_{i} to be equal. When measuring sensitivity to format perturbations, if two separators or spaces are equal in an original format, they are forced to jointly change to be considered equivalent. Appendix A.3 contains all contextual restrictions.

Given a valid format pp accepted by the grammar, the final prompt is constructed by concatenating with space cc an instruction string instinst, nn few-shot data points D1,…,DnD_{1},\ldots,D_{n} exemplifying the task, and a data point Dn+1D_{n+1} to be solved. All few-shot examples DiD_{i} are formatted using pp. Thus, the final prompt template is: inst c p(D1) c p(D2) c … c p(Dn) c p(Dn+1)inst\ c\ p(D_{1})\ c\ p(D_{2})\ c\ \ldots\ c\ p(D_{n})\ c\ p(D_{n+1}). Since Dn+1D_{n+1}’s output will be generated by the model, an empty string is added in place of the answer in the last field in the template. Prompt construction will modify instinst to match specific choices encoded in pp: concretely, if pp enumerates valid multiple-choice options as characters x1…xnx_{1}\ldots x_{n}, we ensure instinst refers to these choices as x1…xnx_{1}\ldots x_{n}.

2 Measuring Sensitivity

We measure how plausible choices in prompt formatting influence quantifiable metrics of generated outputs. Given a set of plausible formats {p1,…,pn}\{p_{1},\ldots,p_{n}\}, a dataset D\mathcal{D}, and a scalar metric mm, let the performance interval be [min⁡im(pi,D),max⁡im(pi,D)][\min_{i}m(p_{i},\mathcal{D}),\max_{i}m(p_{i},\mathcal{D})]. We define the performance spread or simply spread as max⁡im(pi,D)−min⁡im(pi,D)\max_{i}m(p_{i},\mathcal{D})-\min_{i}m(p_{i},\mathcal{D}). Higher spread indicates more sensitivity to variance within the space of plausible, semantically-equivalent formats. While our method is agnostic to the scalar metric mm used, and one could consider a number of metrics including text length, formality, or toxicity, throughout this work we focus our analysis on estimated task accuracy accacc. Due to ease in automatic evaluation, here we evaluate on classification tasks.

We assume a budget of EE total data point evaluations. We first search for the highest performing format with budget E/2E/2, and then for the lowest performing format with budget E/2E/2. Evaluations done for the first exploration are readily available for the second exploration, which yields a more informative prior for many formats. We consider two well-known regret minimization bandit algorithms: Thompson sampling (used in FormatSpread) and Upper Confidence Bound (UCB).

Thompson sampling allows for setting informative priors (αi,βi)(\alpha_{i},\beta_{i}) based on domain knowledge to accelerate runtime. Appendix A.4 details the exact priors we use. To our knowledge, we are the first to consider a Bayesian sampling method for prompt optimization.

UCB (Lai et al., 1985) computes an upper confidence bound to each arm’s performance, derived from Chernoff’s bound. The key difference with Thompson sampling is in how θi(t)\theta_{i}^{(t)} is defined. In UCB’s frequentist approach, θi(t)\theta_{i}^{(t)} is assigned the estimated accuracy plus the upper confidence bound: θi(t) ⁣← ⁣Si/Ni+clog(t)/Ni\theta_{i}^{(t)}\!\leftarrow\!S_{i}/N_{i}+c\sqrt{log(t)/N_{i}}. We use c=2c=2 following Pryzant et al. (2023), who find UCB with c=2c=2 to be most effective for prompt optimization.

Each prompt format is evaluated on E/nE/n points (with appropriate rounding).

Characterizing Prompt Format Variance with FormatSpread

We use a subset of 53 tasks from Super-NaturalInstructions (Wang et al., 2022) with diverse human-written formats and instructions, comprising 19 multiple-choice tasks and 34 classification tasks with {2,3,4}\{2,3,4\} basic fields. Appendix B.1 details the exact task selection procedure. To construct the final prompt template, we concatenate each task’s instruction and nn formatted few-shot examples using \n\n as spacing. While selection and ordering of few-shot examples is a component of prompt design influencing features of model output (Lu et al., 2022), our work focuses on prompt formatting. To remove this confounder, we fix the exact choice and ordering of examples for each task and for a given number of shots nn. Few-shot examples for each task are chosen randomly within each dataset and are not used for evaluation. We evaluate task data samples on an arbitrary order fixed across settings. Datasets are assumed to be of size 1,000 for fair evaluation across tasks.

We evaluate LLaMA-2-{7B,13B,70B} (Touvron et al., 2023), Falcon-7B and Falcon-7B-Instruct (Almazrouei et al., 2023), GPT-3.5-Turbo (Schulman et al., 2022), all autoregressive LMs.

We use two popular measures for computing accuracy: exact prefix matching and probability ranking. In exact prefix matching, we check if the output’s prefix matches the expected answer after normalization (casing, spacing, newlines). Ranking accuracy computes the rate that the expected answer is the highest-ranked valid option (in multiple choice and classification tasks) according to the model’s output distribution. Results are reported using ranking accuracy unless specified otherwise. Appendix B.2 shows additional analysis of exact prefix matching, with spreads even higher than those shown in Section 4.2, and including how formatting choice affects task degeneration (i.e., not answering any valid option).

For each evaluation task we randomly sample 10 plausible prompt formats and use FormatSpread to compute performance spread for each modeling and nn-shot choice (Figure 4). We find significant performance spread across tasks, with a median spread of 7.5 accuracy points across choices in the model and the number of few-shot examples. 20% of tasks consistently result in a spread of at least 15 accuracy points for all LLaMA-2 settings, and at least 9 points for all Falcon settings. We observe several tasks with performance spread over 70 accuracy points. Because this analysis uses only 10 randomly sampled formats, it represents a lower bound of the true spreads for each task. Furthermore, there exists significant performance spread regardless of increased model size (Figure 2(a) and Figure 12 for Llama-2-70B), instruction tuning (Figure 2(b)), or number of few-shot examples (Figure 2(c); Figure 2(a) and 2(b) plot 1- and 5-shot jointly). Appendix B.2 demonstrates similar results on a seletion of non-classification tasks.

Assuming model MM is better than M′M^{\prime} by at least dd accuracy using prompt pp, we compute how often M′M^{\prime} achieves at least dd higher accuracy than MM under a different format p′p^{\prime}. Figure 4 shows these trends are often reversed: LLaMA-2-13B and -70B reverse trend by at least d=d= 0.02 with probability 0.141; LLaMA-2-7B and Falcon-7B reverse trend by at least d=d= 0.02 with probability 0.140. Strikingly, often both experiments (first using pp, and then p′p^{\prime}) were statistically significant (p-value <0.05<0.05) on 1000 samplesWe use one-sided McNemar tests, also known as paired χ2\chi^{2} tests, since we evaluate models on the same set of samples. We test the significance of MM being better than M′M^{\prime} under pp, and MM being worse than M′M^{\prime} under p′p^{\prime}.: 76% and 47% respectively for the two model comparisons mentioned above. We find that formats yielding high performance for model MM may not yield high performance for M′M^{\prime}, implying that formats may not be inherently good or bad (Appendix B.2).

3 How do individual features contribute to performance?

We analyze how choices in particular constants (i.e. S1,S2\mathcal{S}_{1},\mathcal{S}_{2}, C\mathcal{C}, Fcasing\mathcal{F}_{\text{casing}}, Fitem\mathcal{F}_{\text{item}}) independently influence task performance across different formats. Figure 6 shows the distribution of accuracy for 500 sampled prompts conditioned on the choice of S1\mathcal{S}_{1} (the separator between a descriptor and the text placeholder) for one task in Super-NaturalInstructions. When comparing the individual influence of two feature choices, we measure both weak and strong notions of dissimilarity between distributions of accuracy across prompts conditioned on a chosen feature. We say two constant choices yield weakly different accuracy distributions if the values between the first quartile (Q1Q_{1}) and third quartile (Q3Q_{3}) do not intersect. This is equivalent to the boxes in a boxplot not overlapping. We say two constant choices yield strongly different accuracy distributions if the ranges [2.5Q1−1.5Q3,2.5Q3+1.5Q1][2.5Q_{1}-1.5Q_{3},2.5Q_{3}+1.5Q_{1}] do not overlap (adjusted to end in a data point). This is equivalent to two boxplots with their whiskers not overlapping. In Figure 6, ’ \n\t’ and ’: ’ (fourth and sixth) are only weakly different.

We compute accuracy for 500 random formats with 250 samples each on 31 tasks for 1-shot Llama-2-7B. Table 6 shows that choices in S2\mathcal{S}_{2}, Fitem1\mathcal{F}_{\text{item1}}, Fcasing\mathcal{F}_{\text{casing}} do not independently predict performance differences (weakly or strongly): although these features can have a large performance variance and thus should be explored with FormatSpread, they cannot be used to independently predict accuracy changes. Other constant sets have varying degrees of differences, with S1\mathcal{S}_{1} (separators) and Fitem2\mathcal{F}_{\text{item2}} (number format changes in enumerations) having the most individual impact. All tasks with strong dissimilarities are shown in Appendix B.4.

Table 1 shows a selection of tasks where changing a single constant on a format (e.g., casing in task322) results in large accuracy differences. Figure 8 shows that regardless of the scoring criterion used, a significant ratio of these atomic changes are associated with large accuracy changes. For example, 24% of atomic changes have an associated accuracy change of at least 5 points when using exact prefix matching as scoring criteria (11% when using probability ranking).

The space of prompt format accuracy is highly non-monotonic, which makes local search algorithms over the space less effective. Let (p1,p2,p3)(p_{1},p_{2},p_{3}) be a prompt format triple such that pi+1p_{i+1} is obtained by making an atomic change to pip_{i}. We argue that if the prompt format space is smooth, we should often see a triples’ accuracy to be strictly monotonic over ii. We choose 24 tasks (13 multiple choice, 11 non-multiple choice), sample 300 (p1,p2,p3)(p_{1},p_{2},p_{3}) triples for each, and the compute accuracy (using exact prefix matching) of each pip_{i} on 250 samples. 32.4 and 33.6% of triples were monotonic for multiple-choice and non-multiple-choice tasks respectively. Given that random shuffling within a triple will result in monotonicity 33.3% of the time, this suggests that local search mechanisms like simulated annealing may not be effective as they require a locally smooth search space.

4 Prompt formats are identifiable transformations of prompt embeddings

Prompt format choices represent a deterministic transformation of the input, even if its impact on the resulting performance is hard to predict. We represent prompt embeddings as the last hidden layer obtained when processing the whole input prompt (immediately before selecting the first token to generate). We demonstrate that format choice yields a highly identifiable transformation over this embedding, which suggests that formats can be seen as transformations of the output probability distribution.

For each task, and for both {1, 5}-shot settings, we collect prompt embeddings from LLaMA-2-7B corresponding to 10 randomly sampled valid formats for 1000 evaluation examples. We train an XGBoost (Chen & Guestrin, 2016) classifier that maps from the top nn principal components of a prompt embedding to the prompt format. We train with 800 vectors from each of the 10 formats (8000 vectors) and evaluate on the remaining 200. We find that although the original prompt embeddings are of size 4,096Equivalent to the dimension of hidden representations for LLaMA-2-7B., using just the top 100 principal components can result in a classifier with ≥\geq0.98 accuracy in format identification for all 31 tasks analyzed. Figure 8 shows the accuracy of format classification given a fixed number of principal components.Figure 20 in the Appendix visualizes examples of the top two principal components for ten prompt formats. We find that classifier accuracy given just the top two components correlates moderately with the spread of performance in the prompts they represent (0.4240.424, p=8.04⋅10−6p=8.04\cdot 10^{-6}; 0.5550.555 for the 5-shot setting; using exact prefix matching).

5 Fast exploration of the prompt formatting space: FormatSpread

In Section 4.2, we demonstrate that even when sampling just 10 formats from the space of plausible formats, we still observe significant performance spread on many tasks. However, this is only a lower bound of the spread a task may exhibit when increasing the number of formats: for example, about 17% of tasks are expected to increase their spread by at least 5 accuracy points when increasing from 10 to 20 sampled formats. Figure 10 quantifies the expected increase in spread when increasing the number of formats by evaluating 500 formats on 250 samples each and computing expected gains.

Figure 10 compares the efficiency of Thompson sampling, UCB, and naive sampling for estimating spread with respect to a budget EE (Section 3.2). To ensure accurate reports, we compute and show the true spread of the highest- and lowest-performing formats chosen by each method using all data. With a budget of 51,200 evaluations, Thompson sampling results in a spread within 1 accuracy point of the true spread, while naive sampling finds a spread within 4 points, and UCB within 11.

Finally, we use FormatSpread to measure sensitivity of several models where inference is expensive. With a budget of 40,000 evaluations and 320 prompt formats, we find that 1-shot LLaMA-2-70B–ran using 4-bit quantization (Dettmers et al., 2022)–yields a median spread of 0.171 (mean=0.221, std=0.200, using probability ranking across 53 tasks; 25% of tasks had a spread of 0.292 or higher, with a maximum spread of 0.876), and GPT-3.5 yields a median spread of 0.064 (mean=0.110, std=0.115, across 53 tasks using exact prefix matching given that we do not have access to the full logits; 25% of tasks had a spread of 0.148 or higher, with a maximum spread of 0.562), showing sensitivity to formatting is still present even on larger models. 5-shot LLaMA-2-70B still shows high spreads, with 25% of tasks having a spread of 0.310 and a maximum of 0.841. See spread visualization in Figure 24, and a list of best and worst formats found in Table LABEL:table:best_worst.

Related Work

The task of automatically finding the best-performing prompt for a given task without changing model parameters has recently gained attention, given the constantly improving yet somewhat unpredictable performance of LLMs. Prior work has often focused on discovering optimal prompts with gradient-based methods, which are effective, but often lead to disfluent or unnatural prompts (Shin et al., 2020), which can be mitigated with a Langevin dynamics-based method (Shi et al., 2022). Another approach is to learn, optimize, and insert continuous representations of prompts and tasks as input to models (Qin & Eisner, 2021; Lester et al., 2021; Ding et al., 2022; Ilharco et al., 2023). These methods also require access to the LLM’s parameters, thus cannot be applied to models behind an API. In contrast, FormatSpread does not assume access to any model internals. Prior gradient-free work has focused on edit-based enumeration over human-written prompts (Prasad et al., 2023), reinforcement learning (Deng et al., 2022), and by using LLMs themselves (Zhou et al., 2023; Gao et al., 2021). These works aim to achieve competitive task performance, even if the meaning of the prompt or instruction is modified. To our knowledge, we are the first to focus specifically on prompt formatting variance, a quintessential example of semantic equivalence.

Jailbreaking refers to the behavior of intentionally manipulating prompts to elicit inappropriate or sensitive responses, or otherwise reveal parts of the prompt that were intentionally not revealed. While the objective differs from our work, jailbreaking works (Wei et al., 2023; Zou et al., 2023) share the underlying technical question of finding the lowest-performing prompt. Our methods differ, since Wei et al. (2023) evaluate human-generated attacks to guide adversarial prompt design, and Zou et al. (2023) uses gradient-based search methods simultaneously across multiple models.

Some existing work has explored the influence of certain prompt design choices on model performance, for example the prompt’s language (Gonen et al., 2022) and the ordering of few-shot examples (Lu et al., 2022). Other work has focused on providing textual interpretations of continuous prompt representations (Khashabi et al., 2022). Beyond autoregressive LLMs, existing work has focused on performance variance in masked language models (Elazar et al., 2021; Jiang et al., 2020). Our work follows efforts in other domains that explore the influence of spurious features on research evaluations, e.g., in deep reinforcement learning (Islam et al., 2017; Henderson et al., 2018) and statistical machine translation (Clark et al., 2011).

Discussion

We introduce FormatSpread, an algorithm that estimates the performance spread across prompt formatting choices. We use FormatSpread to evaluate the spread of several widely-used open-source LLMs for classification tasks in few-shot learning settings. We find that spread is large regardless of model choice, even when increasing model size, number of few-shots, or when using instruction tuning. FormatSpread is designed to efficiently search the space of plausible prompt formats under a user-specified computational budget. For example, with a computational budget of exploring only 5% of the entire search space for task with 2,500 test examples and 320 plausible formats, we are able to estimate spread within 2 accuracy points of the true spread.

We also characterize the space of prompt formats, finding that it is largely non-monotonic and that few atomic features can be predictors of performance alone, although the separability of format embeddings is highly correlated with observed performance spread. These findings informed the design of our search procedure, where local search methods are not advantageous.

Our findings suggest that performance spread caused by arbitrary prompt formatting choices may influence conclusions made about model performance, especially when comparing models on benchmark tasks. Thus, we recommend that work evaluating LLMs with prompting-based methods would benefit from reporting a range of performance across plausible formats. However, we want to emphasize that single-format evaluation may still be sufficient for many use cases. For example, for researchers or practitioners who build systems on top of LLMs, choosing a single prompt format that works sufficiently well for use in this larger system is a valid methodological choice. However, we encourage future research to compute FormatSpread when comparing their systems to out-of-the-box models, to ensure fair baseline representation. Furthermore, FormatSpread can be used to identify lower-bound performance of a model or system. For example, when using a model for socially impactful tasks, such as stereotype classification in Figure 1, it is important to report the range of accuracy a non-adversarial user might encounter. Likewise, it is crucial to consider robustness to spurious features when claiming that models possess general abilities, such as theory of mind; and beneficial to report when e.g. exploring model biases. We leave it to future research to develop regularization procedures either during training or with an already-trained model to make models robust to diverse formatting choices.

Limitations

As defined by our grammar, all equivalent formats are semantically equivalent to human readers. However, some of them are more likely to be used by humans than others. Spaces and separators are inspired from naturally-occurring formats, but some values are more unusual, such as the spacing or the separator ::. Contextual restrictions enable disallowing undesired combinations of e.g. spaces and separators. However, formats may have multiple valid parses, and some may be more prone than others to unnatural character combinations. For example, let a data sample be ‘Passage: Lorem ipsum dolor sit amet. Answer: Yes’. Depending on if we consider the full stop . to be part of the passage or the format, we may parse it as B2(2)(B1(Passage,’ ⁣:  ’,id),B1(Answer,’ ⁣:  ’,id),’ ’)B_{2}^{(2)}(B_{1}(\text{{{{Passage}}}},\text{{{{'\!:\ \ '}}}},id),B_{1}(\text{{{{Answer}}}},\text{{{{'\!:\ \ '}}}},id),\text{{{{' '}}}}) or B2(2)(B1(Passage,’ ⁣:  ’,id),B1(Answer,’ ⁣:  ’,id),’. ’)B_{2}^{(2)}(B_{1}(\text{{{{Passage}}}},\text{{{{'\!:\ \ '}}}},id),B_{1}(\text{{{{Answer}}}},\text{{{{'\!:\ \ '}}}},id),\text{{{{'. '}}}}). In this work, we choose the former parsing throughout tasks to ensure full sentences. This sometimesLess than 2020% of cases, based on a manual inspection of 10 formats across 20 tasks. leads equivalent formats to have a less usual, yet trivially semantically equivalent resulting character combinations, e.g. B2(2)(B1(Passage,’ ⁣:  ’,id),B1(Answer,’ ⁣:  ’,id),’; ’)B_{2}^{(2)}(B_{1}(\text{{{{Passage}}}},\text{{{{'\!:\ \ '}}}},id),B_{1}(\text{{{{Answer}}}},\text{{{{'\!:\ \ '}}}},id),\text{{{{'; '}}}}). This last format would have the following string form on the example above: ‘Passage: Lorem ipsum dolor sit amet.; Answer: Yes’. We observe high performance spread both in these cases and beyond them. Contextual relations may also restrict these cases if desired by the end user.

Additionally, we focus our evaluation on tasks that have reasonably short input instructions and input field length (see task selection details in B.1). Future work may investigate on how input length affects final performance.

Acknowledgements

We thank Jillian Fisher, Sachin Kumar, Angela Zhou, and the Berkeley NLP group for valuable discussions. This work was conducted while A.S. was a Young Investigator at AI2. This material is based upon work partly funded by the DARPA CMO under Contract No. HR001120C0124, by DARPA MCS program through NIWC Pacific (N66001-19-2-4031), by NSF DMS-2134012, by NSF CAREER Grant No. IIS2142739, and an Alfred P. Sloan Foundation Fellowship. Any opinions, findings and conclusions or recommendations expressed in this material are those of the authors and do not necessarily state or reflect those of the United States Government or any agency thereof.

References

Appendix A Grammar Definition and Instantiation Details

Precisely, p1∼p2p_{1}\sim p_{2} if and only if at least one of the following hold: p1=p2=B0p_{1}=p_{2}=B_{0}; or pi=B0′(di,si)p_{i}=B_{0}^{\prime}(d_{i},s_{i}) with d1=d2d_{1}=d_{2}; or pi=B1(di,si,fi)p_{i}=B_{1}(d_{i},s_{i},f_{i}) with d1=d2d_{1}=d_{2}; or pi=B2(n)(X1,i,…,Xn,i,ci)p_{i}=B_{2}^{(n)}(X_{1,i},\ldots,X_{n,i},c_{i}) with Xj,1∼Xj,2 ∀1≤j≤nX_{j,1}\sim X_{j,2}\ \forall 1\leq j\leq n; or pi=B3(n)(di,j1,i,…,jn,i,s1,s2,c,f)p_{i}=B_{3}^{(n)}(d_{i},j_{1,i},\ldots,j_{n,i},s_{1},s_{2},c,f) where d1=d2d_{1}=d_{2} and jk,1=jk,2 ∀1≤k≤nj_{k,1}=j_{k,2}\ \forall 1\leq k\leq n. It is possible that generated formats equivalent in their string representation are not equivalent according to this equivalence relation.

Figure 11 shows a visualization of how a complex format is parsed using our defined grammar. A full prompt consists of an instruction, n few-shots and a data point to solve. For example, if the instruction was Given a sentence and two words that appear in it, answer which one of the two (A or B) appeared first in the sentence., a full prompt may look as follow. Note that we always use \n\n as space character between instruction and few-shots. The example below shows a 1-shot prompt. It is simply illustrative and does not correspond to any of the tasks considered.

Given a sentence and two words that appear in it, answer which one of the two (A or B) appeared first in the sentence. The quick brown fox jumps OPTIONS: CHOICE (A): fox ; CHOICE (B): brown ANSWER: B Over the lazy dog OPTIONS: CHOICE (A): lazy ; CHOICE (B): dog ANSWER:

FormatSpread forces all instantiations of a multiple choice variable to change jointly to maintain coherence, and this includes text in the instruction. Therefore, when changing the option items from A and B to I and II, the prompt will be generated as follows.

Given a sentence and two words that appear in it, answer which one of the two (I or II) appeared first in the sentence. The quick brown fox jumps OPTIONS: CHOICE (I): fox ; CHOICE (II): brown ANSWER: II Over the lazy dog OPTIONS: CHOICE (I): lazy ; CHOICE (II): dog ANSWER:

Enumerations are indexed from (i.e., “1, 2, 3” rather than “0, 1, 2”). ROMAN[x]\mathtt{ROMAN[x]} represents the Roman numerals written in regular ASCII characters. ‘0x215F’+x represent the series of Unicode characters for Roman numerals. ␣ denotes a spacing character for clarity.

A.3 Restrictions to Prompt Formats Spaces and Separators’ Combinations

We define several restrictions to ensure format naturalness. Users can additionally customize FormatSpread by defining their own rules and restrictions between values. Our rules are as follows:

If B2(X1, ⁣…, ⁣Xn, ⁣c)B_{2}(X_{1},\!\ldots,\!X_{n},\!c) where cc does not contain a newline, then each XiX_{i}’s separators and any subcomponents’ separators should not contain a newline.

Similar to the rule above, if B3(n)(d,j1,…,jn,s1,s2,c,f1,f2)B_{3}^{(n)}(d,j_{1},\ldots,j_{n},s_{1},s_{2},c,f_{1},f_{2}) such that some separator contains a newline (i.e. s1s_{1} contains a newline and/or s2s_{2} contains a newline) then the space cc must also contain a newline.

For B1(d,s,f):=f(d)s<text>B_{1}(d,s,f):=f(d)s\mathtt{<text>}, ss must not be the empty string (i.e., there has to be some separation between descriptor and text).

Having cc be an empty string space in B2(n)B_{2}^{(n)} is only allowed if the first n−1n-1 components are B1B_{1} fields with an empty . Similarly, the newline restrictions mentioned above only apply if the is not empty. This rarely happens in prompt formats, but there are formats such as Question: Options: A. B. where the Options: do not have a corresponding field.

A.4 Thompson Sampling Priors

For the first exploration (i.e., finding the best-performing prompt format), we set an informative prior Beta(α,β):=Beta(max⁡(β⋅x1−x,1.1),5)\text{Beta}(\alpha,\beta):=\text{Beta}\left(\max\left(\frac{\beta\cdot x}{1-x},1.1\right),5\right) for all arms pip_{i}, where xx is the original format’s accuracy. Our goal is to set an informative prior where the expected value of the prior distribution is the original format accuracy xx, since a priori it is the only information we have about performance. This restricts the parameters as follows:

Since β\beta will modulate how confident is the prior, and we want to avoid the model being overconfident, we fix β=5\beta=5. Because we want to have an informative prior Beta(α,β)\text{Beta}(\alpha,\beta) with a Gaussian-like PDF, we force α>1\alpha>1 and β>1\beta>1. In extreme cases, forcing α>1\alpha>1 might alter the expected value. The first exploration’s priors are thus exactly Beta(α,β)\text{Beta}(\alpha,\beta) with α=max⁡(β⋅x1−x,1.1)\alpha=\max\left(\frac{\beta\cdot x}{1-x},1.1\right) and β=5\beta=5 for all arms pip_{i}.

For the second exploration (i.e., finding the worst-performing prompt format), the model has access to the first explorations’ counters Si(E/B)S_{i}^{(E/B)} and Ni(E/B)N_{i}^{(E/B)}. Therefore, we set the second exploration’s priors to be Beta(α+Si(E/B),β+(Ni(E/B)−Si(E/B)))\text{Beta}\left(\alpha+S_{i}^{(E/B)},\beta+\left(N_{i}^{(E/B)}-S_{i}^{(E/B)}\right)\right).

Appendix B Additional Experiments’ Information and Plots

We use a number of heuristics to filter Super-NaturalInstructions tasks to our set of 53 evaluation tasks. Datasets should have at least 1000 samples to be considered. We also remove tasks whose instructions are too long (over 3,000 characters) and datasets with inputs longer than 2,000 characters, given that this makes performing inference at scale intractable. We also filter datasets whose valid outputs include more than 20 different strings, given that we focus on classification tasks.

We also removed tasks where we found a priori performance on the task was 0% accuracy using LLaMA-2-7B 1-shot. Some Super-NaturalInstructions tasks are derived from the same original dataset, but ask different questions. We did not include more than 4 tasks from the same original dataset.

Finally, we also searched for having socially impactful tasks. Those tasks were the only Super-NaturalInstructions tasks where we included a format if one was not provided by the dataset.

The selected tasks were the following 53: task050, task065, task069, task070, task114, task133, task155, task158, task161, task162, task163, task190, task213, task214, task220, task279, task280, task286, task296, task297, task316, task317, task319, task320, task322, task323, task325, task326, task327, task328, task335, task337, task385, task580, task607, task608, task609, task904, task905, task1186, task1283, task1284, task1297, task1347, task1387, task1419, task1420, task1421, task1423, task1502, task1612, task1678, task1724.

B.2 Additional Results for Section 4.2

Table 2 shows that if format p1p_{1} has lower performance than format p2p_{2} under model MM, there is <0.62<0.62 probability that this trend would hold under another model M′M^{\prime} (random chance is 0.50.5). This weak relative order preservation suggests that prompt format performance in a model may not be extrapolated to a different model, or in other words, that there are no inherently good or bad formats.

Here we show results with using exact prefix matching to compute accuracy. Often, failures in prefix matching are associated with degeneration, i.e., cases where the model does not answer any of the valid options, motivating the use of ranking accuracy. Degeneration makes models (specially smaller models) more unlikely to have high accuracy out of the box. As seen in Figure 8, prefix matching is linked to having higher changes when performing atomic changes. Moreover, exact prefix matching can lead to lower performance as generation is less constrained (see Figure 17).

Figure 14(c) shows spread remains regardless of model size increase, architecture change, or number of few-shot examples also when using exact prefix matching as accuracy metric. In line with the results shown for probability ranking in Section4.2, Figure 16 shows that the probability of reversing performance trends between two models just by changing prompt remains high when using exact prefix matching as metric. Strikingly, spread is significantly higher than in the probability ranking setting (see Figure 16), with median spread ranging from 12 to 28 accuracy points depending on the model used. This further motivates the need for running FormatSpread when benchmarking models with this accuracy metric. This increased spread may be partly due to degeneration, as we will detail next.

Sometimes when a model does not generate the correct answer with exact prefix matching, it also does not generate a valid response, i.e. it degenerates. We will now quantify this phenomenon using 53 SuperNaturalInstructions classification and multiple choice tasks.

Given a model, a task, and a format, let the centered mass be the ratio of examples where the model’s output matched with any valid option (regardless of correctness). Table 3 shows that the correlation between accuracy and centered mass is moderate or high depending on the model. This suggests that very often when a model does not return a valid answer, it does not return any valid answer at all. This is especially true for Falcon models, where we observe an almost perfect correlation between accuracy and centered mass. In conclusion, prompt format chosen often do not solely affect accuracy, but they also affect the frequency in which a model is actually able to perform a task. This will especially affect tasks for which there are no alternative metrics. Further research may focus specifically on targeting features that cause degeneration.

All experiments thus far focused solely on classification tasks. We will now focus on tasks that require generating (short) text, and cannot be framed as classification tasks. We selected 10 tasks from Instruction Induction (Honovich et al., 2023) that require generating a unique, valid string to be considered a correct response. Examples include identifying the second letter of a word, adding numbers, or answering a synonym to a given word. Instruction Induction tasks also show a wide range of difficulty, resulting in varied settings to be analyzed (see Figure 19(b)). Given that the collection does not contain human-generated formats, we applied a simple ‘Input: {}\n Output: {}’ format. Results for 1-shot and 5-shot settings show spread is still high across models and n-shot choices (see Figure 18).

B.3 PCA Examples

Section 4.4 systematically analyzes whether we can predict the prompt format that generated a given pre-softmax activation layer (i.e., prompt embeddings) by using solely its top-nn principal components. Figure 20 shows the top two principal components for two different tasks where all 10 formats considered are easily identifiable solely with a prompt embedding’s top two principal compoenents.

B.4 Notable Features

As discussed in Section 4.3, sometimes the choice of a constant may lead to significantly different accuracy ranges. Figures 21,22, and 23 show all strongly dissimilar choices of constants found on any given task, across 53 Super Natural-Instructions tasks, and on both accuracy metrics considered throughout the work. As can be appreciated, choices of constants do not consistently predict performance in isolation.

B.5 Thompson Sampling Results