GrIPS: Gradient-free, Edit-based Instruction Search for Prompting Large Language Models
Archiki Prasad, Peter Hase, Xiang Zhou, Mohit Bansal
Introduction
Recent advancements in prompting large language models (LMs) such as GPT-3 show that models can perform NLP tasks without any task-specific tuning Brown et al. (2020). Most of the work in this area focuses on few-shot learning, where models rely on textual prompts containing input-output example pairs (exemplar prompts). However, humans are often able to perform a new task when provided with a relevant set of instructions or a task description, not necessarily including any examples. In this direction, past works explore a new paradigm of instructional prompts where a prompt is tailored for a particular task by including natural language instructions Efrat and Levy (2020); Mishra et al. (2022a, b). Following Webson and Pavlick (2021), we characterize instructions as a natural language description of the task that includes what is required for a person to complete the task correctly. In general, whether an instruction is a sufficient description of a task depends on whom it is written for, i.e. people with less task expertise require more background information. Demonstrative examples of the task are not considered a part of the instructions.
For purposes of improving task performance via instructional prompts, Mishra et al. (2022b) provide a set of guidelines to manually rewrite raw instructions. Yet this kind of rewriting process requires substantial manual effort and subjective interpretation of the guidelines. In addition, an underlying assumption in Mishra et al. (2022b) is that instructions should be semantically coherent to humans. However, it is possible that the prompts that most improve model performance are semantically confusing to humans in some ways.
Past works attempt to automatically improve prompt quality for large LMs by means of prompt tuning Liu et al. (2021b). Existing prompt tuning methods use gradient-based approaches which have a few notable shortcomings. First, computing gradients with large LMs can be prohibitively computationally demanding. Second, this is entirely infeasible for models available only via APIs, because model gradients and weights are not standardly accessible. GPT-3 models can be finetuned on given data, but the model parameters and gradients remain unavailable (source). Third, output continuous representations may not directly map back onto tokens in the original vocabulary. Thus, we cannot verify whether models are responding to prompts reasonably Khashabi et al. (2021). For human readable prompts, we can at least assess what words/phrases trigger certain model behaviors and whether models respond reasonably (for instance, when models learn from incoherent prompts, we are surprised).
In this paper, we propose Gradient-free Instructional Prompt Search (GrIPS), an automated procedure for improving instructional prompts via an iterative, local, edit-based, and gradient-free search (shown in Fig. 1). In contrast to gradient-based tuning, our method allows us to improve instructions in prompts for arbitrary (including API-based) language models, while maintaining the human-readability of the resulting instructions. On eight classification tasks from the Natural-Instructions Mishra et al. (2022a), GrIPS improves the average accuracy of GPT-2 XL and InstructGPT (GPT-3) models by between and percentage points. We further show that when gradient-information is available, GrIPS is comparable if not outperforms parameter-efficient tuning methods Houlsby et al. (2019); Li and Liang (2021). Additionally, our searched instructions outperform manual rewritten instructions Mishra et al. (2022b) by percentage points on average for the InstructGPT curie engine. With the same data and computational budget, GrIPS outperforms search over in-context examples by about points for InstructGPT. Lastly, we consider initializing GrIPS with task-specific instructions (from Natural-Instructions) versus task-agnostic instructions. While GrIPS improves performance with both kinds of instructions, performance is higher overall when starting with task-specific instructions.
Contributions: In sum, our contributions include:
We propose GrIPS, an automated gradient-free search over instructional prompts that improves accuracy of GPT models by between and points on Natural-Instructions. We also show improvements for OPT, BLOOM, and FLAN-T5.
We show that GrIPS (a) outperforms manual rewriting Mishra et al. (2022b) and search over exemplar prompts, (b) is comparable to select gradient-based tuning methods, and (c) is effective for prompts containing both instructions and examples.
GrIPS can improve instructions when using as few as data points for a performance signal in scoring and when starting with either task-specific or task-agnostic instructions.
Related Work
Our work builds on recent work in prompting large language models, which Liu et al. (2021b) provide a comprehensive literature survey for. We focus on methods for improving model prompts here.
Few-shot learning for language models to perform NLP tasks is an active area of research Schick and Schütze (2021b); Le Scao and Rush (2021); Tam et al. (2021); Logan IV et al. (2021). Prompts in this line of work are mainly composed of a number of input-output examples Schick and Schütze (2021b); Le Scao and Rush (2021); Tam et al. (2021); Logan IV et al. (2021). Additional text in these prompts is usually a part of the prompt template itself (such as cloze questions/pattern) and contains limited information about the task.By prompt template, we are referring to the choice of cloze-question/pattern (typically a phrase or short sentence), verbalizer, or any structuring text around the training and test example(s). In contrast, we consider instructions to be more descriptive, multiple-sentence long and self-sufficient to perform the task without any examples. See illustrative examples of templates in Table 7 of Zhao et al. (2021). In contrast, our work focuses on instructional prompts as described below.
Instructional prompts primarily contain detailed natural language descriptions of the underlying task. Recent work focuses on utilizing instructions given to human annotators during data collection Efrat and Levy (2020); Mishra et al. (2022a). Mishra et al. (2022b) propose guidelines for manually rewriting instructions in order to further improve performance of instructional prompts. While Webson and Pavlick (2021) show language models may struggle to truly understand instructions, Wei et al. (2022); Sanh et al. (2022) find finetuning on instructions and in-context examples in a hugely multi-task manner helps generalization to other tasks. Lastly, Weller et al. (2020) provide a dataset in which task descriptions are formulated as questions. These questions are relatively short and domain-specific, whereas the instructions in Natural-Instructions Mishra et al. (2022a); Wang et al. (2022) are longer and correspond to more diverse tasks.
Instead of limiting prompts to natural language text, recent work explores training continuous vector tokens in prompts via gradient-based optimization Liu et al. (2021c); Lester et al. (2021); Li and Liang (2021); Qin and Eisner (2021). Sun et al. (2022) aim to optimize continuous tokens without using gradients, however, their technique does not work for APIs that only allow modifying text and not token embeddings (like for GPT-3).
Zhao et al. (2021) find varying the choice of training examples, example order permutations, and template can alter the performance of a prompt. Liu et al. (2021a) focus on selecting in-context examples from a dataset, while Lu et al. (2022); Kumar and Talukdar (2021) explore optimal ordering of examples. Others manually write effective prompt templates for NLP tasks Petroni et al. (2019); Brown et al. (2020); Schick and Schütze (2021a); Schick and Schütze (2021b); Schick and Schütze (2021c). In principle, all prompt search methods treat the prompt text as a parameter space to be optimized over Andreas et al. (2018). Jiang et al. (2020) and Gao et al. (2021) use automated paraphrasing of the prompt templates. Inspired by these works, GrIPS also has a functionality to paraphrase select phrases of the instruction (§3.2.2). Meanwhile, Shin et al. (2020) use a gradient-based search to find trigger words in the prompt template. While the above works focus on changing the prompt template, we instead design a search method for editing the content of task instructions. Our search algorithm is also related to genetic algorithms Mitchell (1998), where parent candidates are mutated to generate offering (via our text-based edit operations) to increase fitness under an objective (like our score function).
Methodology
In this section, we first describe and illustrate different prompt modes (§3.1). Then, in §3.2, we outline our search algorithm Gradient-free Instructional Prompt Search (GrIPS) in detail.
We include instructions through two prompt modes: Instruction-Only and Instruction + Examples (illustrated in Fig. 2). Here, prompt mode refers to the choice and arrangement of the three components (instruction, in-context examples, and test instance). These prompt modes are also used in Mishra et al. (2022a) (details in Appendix B). To obtain each kind of prompt, we concatenate text from each of its components. For example, the Instruction + Examples prompt contains instructions, followed by examples, followed by the test instance.
2 Gradient-free Instructional Prompt Search (GrIPS)
As illustrated in Fig. 1, the GrIPS algorithm starts with an initial base instruction, and then at each iteration, it generates new candidates by randomly selecting and applying phrase-level edit operations to each candidate. This results in a total of sampled operations in each iteration (phrase selection described below in §3.2.1 and edit operations in §3.2.2). These candidates are then scored based on the model performance on . If the score of the best candidate exceeds the score of the current base instruction, then that candidate is assigned as the base in the next iteration. Otherwise, the search continues with the same base instruction. The search stops when the score on does not improve for iterations or a maximum number of total iterations is reached.
While the above search is greedy, retaining only the best candidate in every iteration, we can alternatively retain the top- scoring candidates. Subsequent iterations, contain base candidates for which we perform search individually and the overall top- scoring candidates move to the next iteration until we reach the stopping criteria. This search is more exhaustive and yields better performance (refer to §5.3), however, it increases the number of model evaluations by -fold. We refer readers to Appendix C for full pseudo-code.
2.1 Splitting Instructions into Phrases
As each instruction is a collection of sentences, edit operations can be performed at the word, phrase, or sentence level. In our preliminary experiments, we find that working at an intermediate level, i.e. phrases, is most helpful. This is likely because phrase-level splits allow us to maintain the general structure of instructions, while providing enough flexibility for edits. In order to effectively split each sentence into phrases, we use a state-of-the-art CRF-based constituency parser Zhang et al. (2020a). Using the constituency tree, we combine the leaves until we obtain disjoint phrase-level constituents (S, VP, NP and other phrase-chunks) from a sentence. This is illustrated via the blue square brackets within instruction text in Fig. 1.
2.2 Edit Operations
Below, we describe edit operations used in GrIPS:
Delete (del). We remove all occurrences of the input phrase from the instruction. The deleted phrase is stored for subsequent use in the add operation.
Swap (swap). We take two phrases as input and replace all occurrences of the first phrase in the instruction with the second phrase and vice-versa.
Paraphrase (par). We replace all occurrences of the input phrase with a corresponding paraphrase generated using a publicly available PEGASUS-based Zhang et al. (2020b) paraphrase model from HuggingFace Wolf et al. (2020).Model available at: https://huggingface.co/tuner007/pegasus_paraphrase
Addition (add). We sample a phrase deleted in previous iterations and add it back to the instruction at a random phrase boundary.
These edit operations yield a broad space of possible instructions including simpler, less abstract instructions with fewer details. Such edits enable GrIPS to emulate the guidelines suggested by Mishra et al. (2022b) that also limit details and abstractions. Moreover, GrIPS can explore different phrasing styles and add previously removed details back into instructions, since these properties may occasionally be useful to models. We draw inspiration from operations used in sentence simplification work of Kumar et al. (2020). Empirically, the effectiveness of edits is shown in §5.1.
Experimental Setup
Natural-Instructions Mishra et al. (2022a); Wang et al. (2022) consists of a set of tasks, each comprised of task instructions, and labeled examples. Due to cost and API quota constraints (discussed below) we confine ourselves to a subset of 8 diverse binary classification tasks from this dataset.
Following Mishra et al. (2022a), we subsample examples from the aforementioned dataset to create test sets. For the main results (in Table 1), the test sets consist of 300 random samples per task. Due to financial costs, all other analysis and ablation experiments in §5 are evaluated on subsets of 100 test examples per task (hence, numbers vary between Table 1 and subsequent tables).
We use GPT models Radford et al. (2018, 2019); Brown et al. (2020) with 1B parameters, specifically GPT-2 XL (1.5B parameters), InstructGPT babbage, and curie.While we know that curie is larger than babbage, the exact model sizes for engines on OpenAI API are not officially available. The sizes of babbage and curie models are estimated as 1.3B and 6.7B parameters (source). Relative to standard GPT-3 models, InstructGPT models are specially designed to follow task instructions and therefore are a natural choice in our work Ouyang et al. (2022). In light of the cost constraints in running experiments (discussed below), we did not experiment with the davinci engine (largest model) that is known to exhibit stronger performance on several NLP tasks Brown et al. (2020).
A single run of GrIPS on a task requires model evaluations. We worked with a \600\20\text{-}25\text{ and }\125\text{-}175{\approx}. We note that after running GrIPS and obtaining the modified searched instruction, the cost of evaluation on the test set is significantly smaller, a total of \approx\150$ for all the results in this work.
We set number of edit operations per candidate , number of candidates per iteration , number of iterations , and patience . Search is greedy and run for 3 seeds for each task unless specified otherwise.
Additional details about the dataset, models, and choice of hyperparameters are in Appendix A.
Results and Discussion
In this section, we present the results of our experiments. First, we establish the effectiveness of GrIPS across models in §5.1. Then, we compare our search to other methods in §5.2 and §5.3 and provide additional analysis in subsequent sections.
Our main results are shown in Table 1. On average across tasks, GrIPS improves accuracy for GPT-2 XL, InstructGPT babbage and curie by , , and percentage points respectively that is statistically significant at the level.We perform two-sided hypothesis tests for these improvements by bootstrap with examples and random seeds resampled 100k times Efron and Tibshirani (1994). Accuracy for each method is averaged across test data, seeds, and tasks. Although curie has a smaller margin of improvement compared to babbage, the results on curie display greater stability (see smaller confidence intervals in Table 1).
Our results corroborate that larger InstructGPT models outperform smaller, non-InstructGPT counterparts Ouyang et al. (2022). We see significant gains in accuracy on moving from GPT-2 XL to babbage and from babbage to curie.
In Table 2, we evaluate several design choices in §3.2 on GPT-2 XL. First, we observe that removing the entropy term from the score function decreases accuracy by points. We find this term helps breaks ties between candidates with similar performance on in favor of less skewed-predictions and avoids local minima. Next, we re-run GrIPS with all but one edit operations and find that removing del, swap, par, and add operations drops accuracy by , , and points respectively, thus indicating that GrIPS benefits from all edit operations. Appendix C contains additional design ablations.
2 Comparing with Gradient-free Methods
Prior work in prompting often employs manual rewriting or searching good examples for -shot learning. Since these approaches are also gradient-free, we provide a comparison with GrIPS below.
Closest to our setting, Mishra et al. (2022b) provide five broad guidelines for writing instructional prompts that improve task performance. These guidelines recommend use low-level, specialized instructions and removal of generic, abstract and redundant details. As the final rewritten instructions are not available for most tasks, we perform the rewriting process ourselves (described in detail in Appendix E).
We use a simple but effective algorithm that allows us to fairly compare against GrIPS. At each step, we form a prompt by randomly sampling examples from the score set and then compute the model performance on the remaining points. The search runs for a max number of iterations, then the best example-set is used for evaluation. Note that will vary by task; we fit as many examples as we can in the space of 1024 tokens (between 8 and 28, for our tasks). We use the same score set for example search as GrIPS. Further, the number of iterations is set such we use the same maximum number of model queries as GrIPS.The financial cost of Examples-Only search is considerably higher than GrIPS. Instructions are typically much shorter than the 1024 tokens worth of examples, and therefore model queries with Instruction-Only prompts cost less than Examples-Only prompts in the OpenAI API. We note that relative to our example search, one could find a different example-set for each test instance Liu et al. (2021a), use a genetic algorithm Kumar and Talukdar (2021), or alternate search heuristics Lu et al. (2022).
First, Table 3 shows that our search outperforms manual rewriting for all models, by , and points for GPT-2 XL, InstructGPT babbage and curie, respectively. Next we observe that example search outperforms GrIPS for GPT-2 XL. However, when we use the InstructGPT models that have been designed to follow textual instructions better Ouyang et al. (2022), GrIPS outperforms the exemplar prompt search (by and points for babbage and curie respectively). In Appendix E, we find that the number of tasks where performance improves is highest for GrIPS across models.
3 Comparing with Gradient-based Methods
Our gradient-free design enables the use of GrIPS with larger API-based InstructGPT models. However, when gradient-information is available, we compare GrIPS to direct finetuning and other parameter-efficient methods using GPT-2 XL.
We explore three representative gradient-based approaches: direct finetuning, adapters Houlsby et al. (2019), and prefix-tuning Li and Liang (2021).These methods only use test input and not instructions. For the latter, we use prefix length and include a setting without MLP reparametrization. To ensure a fair comparison with GrIPS, for each task we perform an split of the score set into train and dev sets.
The comparison is presented in Table 4. Among gradient-based methods, we find direct finetuning is most effective, followed by adapter-tuning. Both approaches outperform GrIPS (greedy decoding) by and points respectively. However, exploring the search space more extensively using beam search improves performance of GrIPS by points, outperforming all methods without using any gradient information.Due to cost constrains, we do not use beam search with InstructGPT, although we expect it to improve performance. We also observe that GrIPS outperforms prefix-tuning by up to and points using greedy and beam search respectively. Since prefix-tuning upper bounds performance of AutoPrompt Shin et al. (2020); Li and Liang (2021), we expect GrIPS to outperform AutoPrompt as well. Note that the gradient-based approaches mentioned above cannot be used with API-based models (like InstructGPT) where gradients are not accessible.
4 Task Specific vs Agnostic Instructions
GrIPS is contingent on the instruction that we use to initialize the search. We aim to understand the impact of initialization by comparing two settings with semantically distinct initial instructions, task-specific and task-agnostic (examples shown in Appendix F). Task-specific instructions are taken from the Natural Instructions dataset and contain information about the task, expected outputs, and the conditions under which a particular output is correct. Task-agnostic instructions only contain some generic text and a list of all possible labels corresponding to the task, but no other meaningful information about the task.
In Table 5, we find that GrIPS is effective in both task-specific and task-agnostic settings with improvements up to and points, respectively. Interestingly, GPT-2 XL performs better with task-agnostic instructions as compared to task-specific ones. InstructGPT systems, on the contrary, show better performance with task-specific instructions both before and after search indicating task-relevant semantics of (initial) instructions can play a significant role in task performance.
5 GrIPS with other Open-Source Models
Similar to other instruction-based methods, GrIPS works best when models can follow declarative instructions and are responsive to changes to instructions (shown in Appendix D). While this may not be the case for standard pretrained large language models, we nevertheless show that GrIPS can be effectively used with other models such as GPT-J Komatsuzaki (2021), GPT-NeoX Black et al. (2022), OPT Zhang et al. (2022) and BLOOM Scao et al. (2022).
In Table 6, we observe that GrIPS can still improve performance of all the aforementioned models by nearly - points. Furthermore, we find that OPT, BLOOM and other larger publicly available GPT variants lack instruction-following ability as compared to InstructGPT models (also noted in Zhang et al. (2022)). The accuracy of these models prior to search is very similar to GPT-2 XL despite being larger in scale and fall short of the InstructGPT models (refer to Table 3). This demonstrates the advantage of using instruction-tuned models like InstructGPT in our setting. Finally, we use GrIPS on another publicly available instruction-tuned model named FLAN-T5 Chung et al. (2022) and find a point performance improvement. Here, we observe significantly higher average task accuracy even prior to search, which we attribute to the use of Natural-Instructions dataset in the instruction finetuning Chung et al. (2022), possibly exposing the model to the test instances as well as the task instructions.
6 GrIPS is Effective for Smaller Score Sets
While we use a score set of size by default, it would be preferable to use as little data as possible, all else equal. Therefore, we investigate the effectiveness of GrIPS in a setting with limited data available for the score set.
In Fig. 3, we experiment with a score set of size , or . We first observe that as the size of the score set decreases, the margin of improvement from the search declines as well ( point gain when versus point gain when ). This trend is expected because using fewer examples in the is equivalent to having a smaller train set, and thus we expect the model generalization to be worse. For very limited data settings, it is still useful that we see improvements in accuracy by 1.0 point using as few as data points. Our results also suggest that when more data is available, increasing can lead to further performance improvements.
7 Semantics of Searched Instructions
Table 7 (and Appendix G) contains some searched instructions by GrIPS. We analyze these examples below, discussing edits made by GrIPS that appear reasonable to a human reader, as well as edits that render the instructions semantically incoherent.
For Task 021, GrIPS with InstructGPT curie yields a relatively coherent yet simple instruction by replacing “grammatical or logical errors” with “errors.” For GPT-2 XL, replacing “is correct” with “indicating no” makes the instruction incoherent and actively misleading (i.e. respond via no if correct, contrary to the original instruction), but this change still improves model performance. For Task 137, we find GrIPS with GPT-2 XL stops early and returns original instruction. Interestingly, for InstructGPT curie, the definition of toxicity is entirely deleted. Finally, we see semantically incoherent edits occur for Task 195 with no information about possible labels (‘positive’ or ‘negative’). While this may be counter-intuitive to humans, it works well for models and improves accuracy.
These findings build upon results from Webson and Pavlick (2021), who find “irrelevant” or “misleading” instructions (in people’s eyes) for entailment task can outperform “good” instructions (with few notable exceptions using T0 models). Yet in §5.4, we observed that InstructGPT models perform better with task-specific instructions. Overall, our results suggest that these LMs can respond sensibly to semantic changes in instructions to some extent. As with the study of in-context learning mechanisms Xie et al. (2022); Razeghi et al. (2022); Min et al. (2022), how models utilize instructions remains largely unknown and merits further study.
8 Effectiveness of GrIPS on Instruction + Examples Prompts
Lastly, we show that GrIPS can also be applied to Instruction + Example prompts (refer to Fig. 2) that contain additional labeled examples before the test instance. Unlike in §5.2, we set the number of examples to across all tasks, as higher values of make the financial cost prohibitively large. In order to mitigate majority label bias in the prompts Zhao et al. (2021), we use equal number of examples from each label in the prompt. Since the choice of examples varies with the random seed, we use seeds instead of for these experiments.
Table 3 demonstrates that our search is effective in this setting across all models, improving accuracy by roughly points. For InstructGPT models, there is surprisingly little difference in performance between Instruction-Only and Instruction+Examples modes ( percentage points). For both babbage and curie, however, the prompts containing instructions outperform the Examples-Only prompts, by about points. Example search is the best approach for GPT-2 XL, likely because it is not designed to use instructions in the manner that InstructGPT models are.
Conclusion
We introduce GrIPS, an automatic search algorithm that edits task instructions to improve downstream task performance. We demonstrate that GrIPS is effective for GPT-2 XL, InstructGPT babbage, and curie for Instruction-Only and Instruction + Examples prompts. Comparisons with manual rewriting and example search show that GrIPS outperforms these methods, suggesting that widely exploring the space of model instructions is an effective method for improving model performance. Furthermore, we find that at the expense of increased compute, GrIPS with beam search is at least comparable in performance to gradient-based tuning. We show that our search is effective when starting with task-agnostic instructions and that it also works with as few as examples in the score set. Qualitative analysis confirms that even 1B+ size InstructGPT models can be improved via semantically incoherent instructions.
Acknowledgments
We thank the reviewers and the area chairs for their helpful comments and feedback. We thank OpenAI for providing academic access to their API. We also thank Derek Tam, Prateek Yadav, Yi-Lin Sung, Jaemin Cho, and Shiyue Zhang for their helpful comments. This work was supported by NSF-CAREER Award 1846185, DARPA Machine-Commonsense (MCS) Grant N66001-19-2-4031, ONR Grant N000141812871, and a Google PhD Fellowship. The views contained in this article are those of the authors and not of the funding agency.
Limitations
Our edit operations currently do not have the capability to add significantly new and pertinent information or sentences to the instruction, outside of what is available initially in the dataset. Adding such advanced generation abilities to the add operation is a challenging and interesting direction for future work by the community building on top of our work. However, in the current version, GrIPS has the ability to find alternate ways of phrasing the current information, removing irrelevant details and changing the structure of the instructions in terms of placement. Further, a framework like GrIPS may not be as effective for purely generation-based tasks due to lack of good metrics to replace the accuracy in the score function. Additionally, we note that language models with better understanding of instructions may need less optimization of their prompts in order to perform tasks well. Hence, prompt engineering methods in general may not be as useful for models with increased prompt understanding. Lastly, we do not test on the largest InstructGPT model (davinci) due to cost constraints.
Ethical Considerations
Instructions are a useful tool to convey extrinsic information to large language models and alter model outputs, e.g. by instructing models to generate less harmful content. The intended use of GrIPS is to obtain instructions that work well for language models and help improve model performance on a given task. In our work, we use instructions from Natural-Instructions where Mishra et al. (2022a); Wang et al. (2022) ensure quality control. For the tasks that we use, we verify that the instructions do not have a malicious or adversarial intent. Similar to methods prompting large language models, our proposed search can unfortunately be misused intentionally or unintentionally Weidinger et al. (2021) to elicit harmful, biased and problematic outputs for maliciously-designed or adversarial inputs and/or instructions. Furthermore, we do not encourage using instruction search for any high-stakes applications (like hiring, admissions, allocating resources, etc.). Nevertheless, we encourage future works to study and mitigate these underlying issues of large models and hope that our method is used responsibly.
References
Appendix
Appendix A Additional Experimental Details
In Table 8, we provide details about the 8 classification tasks from the Natural-Instructions dataset that are used in this work. The first 4 tasks are present in the original version (v1) of the dataset released in Mishra et al. (2022a). As shown in Table 8, the label distributions in these tasks examples are extremely skewed towards one label (). We chose the remaining 4 tasks from next release (v2), curated by Wang et al. (2022), such that (a) the label space and instructions are diverse in length, nature of the task, and label tokens; (b) the datasets are more balanced and less skewed towards one label; and (c) the dataset was stable on the github repository,Datset: https://github.com/allenai/natural-instructions. Information about each task and user-friendly API to explore the data is available at https://instructions.apps.allenai.org/ i.e. without any recent commits or modifications for at least 1 month. Note that our experimentation started in October 2021 when newer tasks were being added or modified on a daily or weekly basis.
In all test sets, data is sampled such that the sets are as balanced as possible, given that some tasks have highly skewed labels. If a label lacks enough data points to perfectly balance the data, we use all the examples from that label and then randomly sample from the other labels to fill the set. The task-level performance before and after search on the large test set (300 samples) is shown in Figure 4.
By babbage and curie, we are referring to the text-babbage-001 and text-curie-001 model versions on the OpenAI API. Following Zhao et al. (2021), classification tasks are performed by computing log-probabilities of the label tokens using the completion function of the OpenAI API. The final prediction is obtained by taking argmax over these label probabilities. Note that our setting is different from Mishra et al. (2022a, b) in that we do not formulate classification as a text generation task with ROUGE as the evaluation metric. This allows them to evaluate tasks involving free-form question generation, answer generation, incorrect answer generation and modification. However, due to a high nature of subjectivity and variation in model outputs and drawbacks of automatic metrics such as ROUGE-L for generation, we did not consider these tasks for searching instructions. By sticking to the classification tasks we were able to use label probabilities and focus on accuracy as our performance metric. We leave exploration of GrIPS for generative tasks for future work.
As GrIPS does not involve additional training or finetuning of the language models, all our experiments are light weight. Only GPT-2 XL requires GPU access which takes about minutes per task (only evaluation of prompts) and for all experiments combined uses little over GPU hours on an NVIDIA A100 40 GB GPU. Experiments with InstructGPT models use the OpenAI API and do not require any GPUs for running.
Due to financial constraints, hyperparameter tuning was conducted using line search using smaller (and cheaper) models like GPT-2 L and XL and on select tasks during preliminary experiments. We first considered the number of edit operations applied to each candidate in one iteration (), followed by a combination of number of candidates and number of iterations, i.e. . We increased patience as we reduced the number of candidates () in order to ensure that the search did not end prematurely. We observed that changing led to only marginal difference in performance and found to be most effective. We set in our experiments. We found that when using we explored several edited candidates for the same base instruction but ran the search for fewer iterations which turned out to be less effective. However, exploring too few candidates was also not effective as we often proceeded to the next iteration with sub-optimal edits. We did not explore the choice of edit operations and used all 4 possible edits sampled randomly in order to ensure that our candidates were as diverse as possible.
Appendix B Prompt Template vs Instructions
The terminology used in this paper differs slightly from Mishra et al. (2022a). The term ‘instructions’ in our work corresponds to their term ‘definition’. Additionally, to keep the prompt templates used in this work compatible with theirs, we still use the word ‘definition’ in the prompt template instead of ‘instruction’. This is also consistent with the schema in Natural-Instructions. Prompts in Fig. 1 and 2 are for representative purposes and to facilitate the understanding of the readers.
The above choices between ‘definition’ and ‘instruction’ is only one example of possible template-level changes. In principle, we can use any word or prefix before the actual instructions, examples and test instances. For example, for the prompt shown in Fig. 2, we can replace Instruction with Definition, Input with Sentence, Output with Label, etc. Each of these changes will result in a new prompt template. While these changes are subtle, empirically Zhao et al. (2021) show that models are sensitive to such changes. Since our objective is to explore better ways of leveraging instructions, we keep these template words unchanged in all our experiments so that the comparison of different searched instructions can be fair. Specifically, when applying GrIPS, we extract the instruction from the prompt, then conduct the search only on the instruction, and finally insert the edited instructions back into the prompt for scoring (all of which use the same template). Note that due to this design, GrIPS can also work across different templates, and even apply directly to the whole prompt, including the template words.
Appendix C Extensions and Variations of GrIPS
The full-pseudo code of GrIPS is shown in Algorithm 1 where we use greedy search. The beam search modification is described in Algorithm 2. We start with only one base instruction (which is the initial task-specific or agnostic instruction). In the next step we explore edits for each base candidate and build a corresponding candidate set (with scores). At the end of the iteration, we take the most promising or highest scoring path and proceed to the next iteration, effectively pruning the rest. When the search terminates, we find the best candidate from the filtered (remaining) set of candidates.
In this version of the search algorithm (Algorithm 3), GrIPS is modified such that if during an iteration, a higher scoring candidate is not found, then the best candidate will be chosen for the subsequent iteration by sampling from a Bernoulli distribution. The probability of success is given by:
Here, is the score of the highest scoring candidate, is the score of the base candidate, is the index of the iteration, , are hyperparameters. This formulation has been adapted from Pirlot (1996). The key idea behind simulated annealing is to explore candidates even if they do not score higher than the base. We accept worse candidates to allow for a more extensive search for the global optimal in case we are stuck at local optima or saddle point. The probability of exploration is and it is directly proportional to the difference in the scores. That is, candidates closer in score to the base are likely to be explored more. The parameter controls the overall degree of exploration and controls the decay in exploration as the iterations (index ) progress (i.e. move from exploration to exploitation). On comparing Simulated Annealing () with greedy search, we find that on average there is no statistically significant difference in performance. In fact, greedy search does slightly better with average performance of vs which is the average performance of simulated annealing search (on InstructGPT babbage). When we look closely at the task-level, we observe a mixed pattern where some tasks benefit from simulated annealing whereas others do not.
Fig. 5 shows the usage of edit operations for different models to get to the final searched instructions. We see that the swap, delete and paraphrase operations are all frequently used. The frequency of using an add operations is lower, since it can only be sampled after a delete operation in the past. Nonetheless, the add operation is used in search runs of roughly of the tasks. Next, we explore alternate choices of paraphrase and add operations. Instead of using a Pegasus-based paraphrase model, we replace it with another T5-based paraphrase modelModel available at: https://huggingface.co/prithivida/parrot_paraphraser_on_T5 and find the accuracy changes from to which is a minute difference. If the add operation is designed to add a random phrase from the initial instructions instead of phrases that are previously deleted, the average accuracy slightly reduces to (c.f. Table 2).
Appendix D Search Improvements Correlate with Model Sensitivity to Instructions
We observe that GrIPS works better on some tasks than others. Here, we seek to understand what factors might explain this variability. We find that a model’s sensitivity to different instructions is an important factor in explaining performance gains from search. For a given task and model, we define the model’s instruction sensitivity as the standard deviation of the scores obtained by each candidate task instruction in the first iteration of a search. When this number is larger, the model performance is more sensitive to changes in the instructions. Interestingly, in Table 9, we find that instruction sensitivity of a task correlates strongly (Pearson’s ) with the performance improvement margin for GPT-2 XL and InstructGPT babbage models (). However, for the curie engine the correlation is relatively weaker () and not significant at . Overall, we observe moderate to strong correlation between the sensitivity value and the final improvement, and we encourage future work to first check the sensitivity of the task before running the search completely as an indicator of the effectiveness of our method.
Appendix E Details on Gradient-free Methods
Mishra et al. (2022b) propose five broad suggestions to rewrite instructions described below:
Specialized-Reframing: replacing generic, redundant text and describe the low-level task
Pattern-Reframing: removing abstract details
Itemized-Reframing: split paragraphs into bulleted lists and rewriting negative sentences (phrases like do not X) as semantically equivalent positive instances (like do Y instead)
Decomposition-Reframing: break down tasks with multi-step reasoning into simpler tasks
Restraining-Reframing: re-emphasizing constraints on output (label space for classification)
In lieu of final rewritten instructions for our selected tasks, the rewriting process was done by the first three authors, after carefully studying the guidelines in the paper, in an iterative manner. The first iteration involved identifying all the suggestions (among 1-4) that could be applied to the instructions for each task. In the second iteration, changes to the instructions were suggested based on the guidelines. These changes were then reviewed by the other authors. Disagreements were resolved through detailed discussions until a consensus was reached in the third iteration. Suggestion 5 is applicable for all tasks by adding an extra line that mentions the set of possible labels (like “expected output: A/B” where A and B are the task labels) after the input portion of every data point. This was straightforward and did not require extensive discussions. The entire process was dedicated nearly hours of manual effort.
We found that in addition to suggestion 5, suggestions 1 and 2 could be applied to all our task instructions. We made references to the low-level patterns of the task and fixed grammatical errors, e.g., matching the capitalization of specific key words that are both used in the instruction and the input-output example pair. Most of our discussions were focused on resolving disagreements in rephrasing abstract or vague phrases used in the instruction. Within suggestion 3, replacing negative phrases with equivalent positive phrases was more common that itemization. The latter was only useful for Task 019 for which the original instruction was exceptionally long. We did not feel the need to decompose any task and use suggestion 4.
Unlike Mishra et al. (2022b), we find that including an extra sentence in the prompt to reiterate the label space (suggestion 5) indicated as Labels in Table 10) can hurt performance for InstructGPT models. The reverse is true for GPT-2 XL, where there is some performance gain. This might be because Mishra et al. (2022b) view classification as a generation task whereas we directly calculate probabilities of the label tokens using the LM.
E.2 Example Search
Fig. 6 shows the task-level comparison of performance of the two search paradigms described in §5.2. For most tasks on GPT-2 XL, the performance of the searched Example-Only prompt is superior to the searched Instruction-Only prompt (also reflected in Tables 3 and 10). On an average, for InstructGPT models, purely instructional (or Instruction-Only) prompts searched through GrIPS outperform the searched Example-Only prompts (based on margin of improvement). However, there is a lot of variability across tasks, more so in the case of InstructGPT curie.
Appendix F Task Agnostic Instructions
In Table 11, we compare task-specific and task-agnostic instructions. As mentioned in §5.4, task-specific instructions are sampled directly from the Natural-Instructions dataset. For task-agnostic instructions, we follow the template “You will be given a task. Read and understand the task carefully, and appropriately answer [list of labels].” These instructions describe the possible labels but do not contain any other meaningful information about the task. Given, that in §5.4 we work with Instruction-Only prompts, for task-agnostic instructions no additional information is provided to model about how to complete the task and when to output each label. The list of labels for each task is mentioned in Table 8. This means that tasks sharing the same label space correspond to the same task-agnostic instruction (shown in Table 11), even if the tasks are entirely different.
Appendix G Instructions after GrIPS
Tables 12, and 13 contain the original and searched instructions for the all the tasks not discussed in §5.7. Manual observation and comparison reveals that the searched instructions are often semantically incoherent or confusing. Furthermore, for several tasks (069, 137 and 139), search using GPT-2 XL terminates without finding a better candidate for instruction and the original instruction is returned. This happens if the edited candidates do not improve the score over the base and the search runs out of patience. We observe that of the searched instructions are shorter than the original, and of them contain some label information pertinent to the task.