Selective Annotation Makes Language Models Better Few-Shot Learners

Hongjin Su, Jungo Kasai, Chen Henry Wu, Weijia Shi, Tianlu Wang, Jiayi Xin, Rui Zhang, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, Tao Yu

Introduction

Much recent work builds approaches to natural language tasks on the impressive abilities of large language models (e.g., GPT-3; Brown et al., 2020). Large language models can perform downstream tasks by conditioning generation on a few task demonstrations, thereby avoiding the need for any parameter updates. This new, few-shot learning paradigm is called in-context learning and has become an attractive alternative to supervised finetuning (Liu et al., 2021). In this work, we study the implications of this remarkable capability of large language models for dataset creation and annotation. We extensively examine how to reduce the manual annotation cost while retaining high in-context learning performance.

Although in-context learning was originally proposed for few-shot learning, recent works show that retrieving prompts from a large set of annotated examples is necessary to achieve good performances (Liu et al., 2022; Rubin et al., 2022). In particular, they show that the performance substantially improves when similar examples (under some embedding function) are retrieved as in-context examples specifically for each test input (Liu et al., 2022). Each test sample only requires a few in-context examples in its prompt. Different test instances, however, require different in-context examples with their associated annotations, necessitating a large set of annotated examples.

Distinct from these recent efforts, we establish a two-step framework to better understand and improve the annotation efficiency (Fig. 1): the first step is selective annotation that picks a small number of instances to get annotated before test time, followed by prompt retrieval that retrieves in-context examples for each test instance from the annotated data. The total annotation budget is the number of examples selected and annotated in the first step. The second step is bounded by the number of examples that can fit as input to a language model. Based on this framework, we propose an unsupervised, graph-based selective annotation method, named vote-kk, that selects diverse and representative instances to be annotated.

Our extensive experiments over 10 datasets across diverse tasks (covering classification, commonsense reasoning, dialogue, and text/code generation; see Tab. 2) demonstrate that our graph-based selective annotation method, vote-kk (§2.1), substantially improves the in-context learning performance by balancing the diversity and representativeness of annotated samples. For instance, vote-kk, combined with similarity-based prompt retrieval (Liu et al., 2022; Rubin et al., 2022), achieves a 11.4% relative gain under a budget of 100 annotated examples and a 12.9% relative gain when only 18 examples are annotated; 18 samples can fit into language models’ input, meaning the prompt retrieval step is not needed. Moreover, the improvement is consistent across language models with varying sizes (2B-175B parameters) (§4.2). This finding is in contrast with finetuning, where we cannot see the effectiveness of selective annotation over random baseline, due to outliers (Karamcheti et al., 2021) or training instability (D’Arcy & Downey, 2022). We hypothesize that in-context learning with similarity-based prompt retrieval is more robust to small annotation sizes and outliers because only the most similar examples are retrieved for each test instance. Indeed, we observe that random prompt retrieval fails to benefit from selective annotation (§4.4), providing support for our hypothesis.

Besides performance comparisons within a fixed annotation budget, we show that selective annotation provides better few-shot performance with 5-100×\times less annotation cost for new natural language tasks. In-context learning with 18 examples selected by vote-kk achieves higher performance than 100 randomly selected examples on 6 out of the 10 tasks. It also outperforms strong finetuning methods by a large margin (Fig. 2) and requires 10-100×\times less annotations for similar performance (§4.1). We observe that in-context learning quickly (100 or 300 samples are annotated) converges to decent performance when vote-kk selective annotation is applied. These results suggest that large language models do not require large annotated datasets (e.g., 10K) due to their ability to adapt to new tasks through simple prompting.

Selective annotation also makes in-context learning much more stable. In real-world scenarios, even collecting unlabeled data is non-trivial and introduces randomness. We simulate such randomness in our experimental setting by subsampling the original unlabeled data multiple times. Our results suggest that vote-kk selective annotation largely reduces the variance of in-context learning even in this setting (Tab. 2). Further analysis shows larger improvements when there is a domain shift between training and test data (e.g., text from different Amazon users; Koh et al., 2021; §4.3). Finally, when compared with previous selective annotation methods designed for supervised training/finetuning, we demonstrate that vote-kk selective annotation consistently improves the performance (§4.5). As in-context learning has been applied to increasingly more natural language processing applications, we hope that our annotation-efficient framework will provide useful guidance for both researchers and practitioners.

Selective Annotation for In-Context Learning

In-context learning only requires a few annotated examples per test instance (few-shot learning), while avoiding expensive finetuning on the whole training data. It is, however, often assumed that all annotated training data are available for prompt retrieval (e.g., Liu et al., 2022; Rubin et al., 2022). Yet the implied total annotation costs are hardly discussed in previous work. We develop a better practice for few-shot learning with large language models by carefully studying the total annotation cost required for in-context learning. We also study how examples should be selected to annotate, in order to make in-context learning perform better for new tasks. We formulate a general framework (Fig. 1 left) that consists of two steps: selective annotation (§2.1) and prompt retrieval (§2.2).

The first step chooses examples to annotate before test time. This process thus determines the total annotation budget. This selective annotation process is largely ignored in the recent literature for in-context learning. We will demonstrate, however, that the annotation cost can be substantially reduced by choosing a small set of diverse, representative examples, while retaining the downstream performance (§3). Formally, given a set of unlabeled samples X={xi}i=1N\mathcal{X}=\{x_{i}\}_{i=1}^{N}, selective annotation aims at selecting a subset L⊂X\mathcal{L}\subset\mathcal{X} to be annotated, where ∣L∣=M|\mathcal{L}|=M is the annotation budget. We discuss our vote-kk selective annotation method and other selective annotation baselines below.

The goal of selective annotation for in-context learning is to select diverse and representative examples; representativeness will help many test instances to find similar demonstrations, while diversity increases the total coverage. We develop vote-kk, a graph-based method that promotes both diversity and representativeness. A detailed algorithm can be found in Appendix G. We first compute a vector representation for each unlabeled training instance using Sentence-BERT (Reimers & Gurevych, 2019) by averaging the resulting vectors over the text input words.https://huggingface.co/sentence-transformers/all-mpnet-base-v2. We then use the embedding vectors to create a directed graph G=(V,E)G=(V,E) where the vertices VV are the unlabeled instances X\mathcal{X} as defined above. For each vertex v∈Vv\in V, we create an edge to its kk nearest vertices in terms of the cosine similarity between the embeddings. Now let L\mathcal{L} and U\mathcal{U} denote the sets of already chosen (i.e., labeled) samples and remaining samples, respectively. Initially, L=∅\mathcal{L}=\emptyset. Every vertex u∈Uu\in\mathcal{U} is scored by a modified degree:

Random and Other Selective Annotation Methods

To quantify the effect of selective annotation, we also provide random and other baselines. For randomly-selected annotation, we conduct experiments three times and report the average score. We will show that these baselines substantially underperform the vote-kk method (§3.3), demonstrating the importance of the selective annotation step to reduce the total annotation cost.

2 Prompt Retrieval

Once we have a set of annotated examples L\mathcal{L} from selective annotation, we retrieve a few examples from the annotated set as in-context examples for each test instance. Following recent work (Liu et al., 2022), we will compute embeddings for all annotated samples using Sentence-BERT and find the most similar examples to each test instance in terms of cosine similarity.

Experiments

We conduct extensive experiments over 10 diverse datasets, spanning 9 distinct tasks, and show a better approach to few-shot learning than previously considered. In general, we find that the first step of selective annotation is particularly crucial to reduce the amount of required annotation.

We use 10 diverse NLP datasets across 9 tasks that are listed in Table 1. These datasets involve different task formulations, thereby allowing for extensive evaluations in varying scenarios. Some of those are included in the widely-used GLUE benchmark (Wang et al., 2019). Appendix A illustrates details of the 10 datasets with examples.

For each dataset, we use the standard train/dev./test split available from the Transformers library (Wolf et al., 2020). In the selective annotation step, we remove all labels in the training data. For the datasets that have test data available publicly, we use the the test data for evaluation (SST-5, XSUM, MWoZ, and DBpedia). For the others, we follow prior work (e.g., Jiang et al., 2020; Lan et al., 2020; Gao et al., 2021) and use the dev. data for evaluation.The one exception is GeoQuery, where we concatenated the dev. and test data to have reliable evaluations on larger data. We evaluate the methods by accuracy for all classification and multiple-choice selection datasets, joint accuracy (Budzianowski et al., 2018) for MWoZ, test suite accuracy (Zhong et al., 2020) for GeoQuery, exact matching (Rajpurkar et al., 2016) for NQ, and ROUGE-L (Lin, 2004) for XSum.

Given a set of unlabeled data, our vote-kk selective annotation algorithm is deterministic, without any randomness. However, we note that in real scenarios, even getting unlabeled samples is not trivial, and getting unlabeled samples can be a process with large variance. To simulate this real setting, we perform selective annotation from 3K instances that are randomly subsampled from the original training data for each task. For each experiment, we repeat this subsampling three times, and results are averaged over the three trials. We will find that vote-kk still substantially improves stability over alternative selective annotation methods.

2 In-Context Learning Models

We mainly perform experiments using GPT-J with 6B parameters (Wang & Komatsuzaki, 2021) due to our computational budget. The exceptions are the MWoZ, GeoQuery, and NQ datasets, where we use Codex-davinci-002 (Chen et al., 2021),The parameter size of Codex is not officially confirmed, but it is likely to be 175B. a variant of GPT-3 finetuned on code data from the web. Codex is particularly effective for structured prediction such as semantic parsing, and we found it is indeed effective on three datasets (MWoZ, GeoQuery, and NQ) in our preliminary experiments. We will explore the effectiveness of selective annotation on the largest publically available language models, OPT-175B (Zhang et al., 2022) for HellaSwag (Fig. 4) and Codex-davinci-002 for MWoZ, over varying annotation budgets. We will also explore other language models with different sizes for three representative tasks (HellaSwag, MWoZ, and SST-5) in §4.2: GPT-3 with 175B (Brown et al., 2020) and GPT-Neo with 2.7B parameters (Black et al., 2021). Our later experiments will show the same patterns among selective annotation methods over these different language models. For the classification and multiple-choice tasks, we compute the average log score for each choice and choose the maximum one. For generation tasks, we simply perform beam-search decoding.

See Appendix B for our in-context learning prompt templates for all 10 datasets. For every test instance, we feed as much retrieved samples as possible into the language model until the maximum token length is reached. On average, the number of samples NN fed into the language model is 13.4 across different experiments. The in-context examples are concatenated in the ascending order of the similarity so that more similar examples benefit from the recency bias (Lu et al., 2022).

3 Main Results

Seen in Table 2 are our results from all 10 diverse datasets with the annotation budgets of ∣L∣∈{18,100}|\mathcal{L}|\in\{18,100\}. 18 is chosen so that all annotated examples can be fit to the prompt for the language models without prompt retrieval. Over all datasets, vote-kk selective annotation outperforms the random baseline by a large margin (5.2% absolute gain and 11.4% relative gain on average) when the annotation budget is 100. Even when only 18 examples are annotated and fixed as the in-context examples for all testing instances (no prompt retrieval step), in-context learning with vote-kk still improves the randomly-selected annotation baseline (5.8% absolute gain and 12.9% relative gain on average). Particularly noteworthy is that in-context learning with 18 examples selected by vote-kk achieves higher performance than the one with 100 randomly selected examples on 6 out of 10 tasks. Moreover, vote-kk is a deterministic selective annotation method, conditioned on a set of unlabeled samples. Therefore, the variance of vote-kk comes solely from how the unlabeled samples are collected, largely improving the robustness of in-context learning. We therefore recommend that researchers and practitioners use selective annotation (e.g., our vote-kk method) to better benefit from the few-shot learning capability of large language models with stability. Our later experiments will also illustrate that vote-kk consistently outperforms alternative selective annotation methods (§4.5).

Analysis

Our extensive experiments demonstrated that selective annotation is important for the success of in-context learning. Here we conduct detailed analysis to provide further guidance for researchers and practitioners of few-shot in-context learning. We analyze selective annotation for in-context learning from a variety of perspectives: comparisons to finetuning methods (§4.1), varying language model sizes (§4.2), test data domain shifts (§4.3), prompt retrieval methods (§4.4), and alternative selective annotation methods (§4.5).

As discussed earlier, in-context learning is an alternative learning paradigm to conventional finetuning. Through the lens of our two-step framework, we observed that selective annotation and prompt retrieval are key to the success of in-context learning. A new question now arises: how does in-context learning compare with finetuning under limited annotation budgets? We empirically compare the two paradigms in this section.

We experiment with three representative tasks: MRPC (classification), HellaSwag (multiple-choice), and MWoZ (dialogue). Strong, state-of-the-art pretrained models are used for finetuning: large-sized RoBERTa (Liu et al., 2019) for MRPC and HellaSwag and DS2-T5 (Shin et al., 2022) for MWoZ. In-context learning uses GPT-J for MRPC, GPT-J and OPT 175B (Fig 4) for HellaSwag, and Codex-davinci-002 for MWoZ. Note that we do not aim to conduct head-to-head comparisons with exactly the same pretrained model; finetuning a large left-to-right language model (e.g., GPT-J and GPT-3) is computationally (and thus financially) infeasible in many cases. Here we examine the two paradigms from the practical perspective and benefit from the advantage of in-context learning, which requires no parameter updates of massive language models.

Fig. 2 compares the two paradigms across varying annotation sizes ({18,100,300,800}\{18,100,300,800\}). Over all three tasks, we observe that in-context learning with vote-kk selection outperforms the finetuning performance of state-of-the-art pretrained language models. Specifically, we find that to achieve similar performance to vote-kk with ∣L∣=|\mathcal{L}|= 18 or 100, finetuning requires 1000 annotated examples for HellaSwag and 800 for MWoZ (10-100×\times annotation cost). Note that the in-context learning performance usually converges when 100 or 300 examples are carefully selected and annotated, suggesting that a large annotated dataset is unnecessary for in-context learning to achieve strong performance. Interestingly, selective annotation helps in-context learning, but not finetuning. This result is consistent with recent work showing that many active learning algorithms perform similarly to random baseline, when pretrained language models are finetuned (Karamcheti et al., 2021; D’Arcy & Downey, 2022). They proposed that it might be due to outliers and the instability of finetuning on a limited number of annotated samples. We hypothesize that in-context learning with similarity-based prompt retrieval is more robust to outliers and small annotation sizes because only the most similar examples are retrieved for each test instance. We find two pieces of evidence for this hypothesis. First, § 4.4 shows that random (as opposed to similarity-based) prompt retrieval does not benefit from vote-kk selective annotation. Second, in Appendix E, we show that explicitly removing outliers also helps finetuning to benefit from vote-kk.

2 Language Models with Various Sizes

Fig. 3 shows performance with varying sizes of language models (GPT-Neo 2B, Black et al., 2021; GPT-J 6B, Wang & Komatsuzaki, 2021; GPT-3, Brown et al., 2020) on HellaSwag commonsense reasoning, SST-5 sentiment analysis, and MWoZ dialogue state tracking. In general, when a smaller model is used, the performance gap between random and vote-kk selection is larger. In the HellaSwag task, vote-kk outperforms randomly-selected annotation by 7.5% using GPT-Neo, but only 2.6% using GPT-3. Nonetheless, we see consistent performance gains from vote-kk selection over varying sizes.

3 Effects of Domain Shift

Recent work observed that when a large, pretrained language model is finetuned, the performance gain from active learning is limited (D’Arcy & Downey, 2022), but it can be larger if there is a domain shift between training and evaluation (Tamkin et al., 2022). We have demonstrated that selective annotation consistently improves in-context learning, but here we explore cases of domain shifts.

Following Tamkin et al. (2022), we use two natural language datasets from the WILDS benchmark (Koh et al., 2021): CivilComments (toxicity classification; Borkan et al., 2019) and Amazon (review classification; Ni et al., 2019). Each comes with both a random split and a domain split: the former splits data randomly and the latter is based on the domains (demographic identities for CivilComments and users for Amazon), simulating cases where a model is deployed in a new scenario unseen during annotations. Similar to §3.3, we conduct experiments with GPT-J under two settings: random/vote-kk selective annotation, followed by similarity-based prompt retrieval. Both selective annotation and prompt retrieval are conducted on the source domain.

Tab. 3 shows our results. We see that the gain from vote-kk is more pronounced under the domain splits: e.g., 9.9 vs. 5.5 accuracy point improvements on CivilComments. This suggests that selective annotation and prompt retrieval are particularly crucial when there is a domain shift in the evaluation data, as in many realistic scenarios (Koh et al., 2021; Longpre et al., 2022).

4 Random Prompt Retrieval

We have performed similarity-based prompt retrieval so far. Here we experiment with a random baseline for the prompt retrieval step to quantify the effect of prompt retrieval (Tab. 4). Interestingly, when random prompt retrieval is performed, vote-kk does not necessarily improve upon the randomly-selected annotation baseline: e.g., 62.5 vs. 63.2 on HellaSwag and 35.7 vs. 43.8 on MWoZ. This suggests that random prompt retrieval fails to benefit from diverse, representative 100 samples, selected by vote-kk selective annotation. Combining selective annotation and prompt retrieval is thus crucial for the success of in-context learning.

5 Alternative Selective Annotation Methods

Here we explore four additional methods for selective annotation:

Maximizing facility location (MFL; Lin & Bilmes, 2009) aims at optimizing the representativeness of the selected samples. Since this objective satisfies the submodular objective, maximization can be approximated via a greedy algorithm (see Appendix G.2).

Diversity focuses on maximizing the diversity of the embeddings for selected examples in the first step (Appendix G.3).

Least-confidence (Lewis & Gale, 1994) iteratively adds least-confident examples to the annotated set.

Fast vote-kk is a fast, efficient alternative to our vote-kk method (§2.1) that does not use confidence scores. It picks MM samples with the largest vote-kk scores. It avoids using the pretrained language model to compute a confidence score for every instance, resulting in a 10+ times speedup.

Notice that MFL, diversity, and least-confidence do not have hyperparameters other than the annotation budget. As shown in Tab. 5, vote-kk outperforms all the other methods. It is noteworthy, however, that fast vote-kk can achieve similar performance to vote-kk. Fast vote-kk is thus an attractive method for researchers and practitioners with a limited computational budget. Like vote-kk, MFL also optimizes representativeness and Diversity also optimizes diversity. In particular, MFL defines representativeness as a sum over distances from the selected examples to all other examples, and Diversity defines diversity as the distances between selected examples. Since they do not significantly outperform randomly-selected annotation, we conjecture that jointly optimize diversity and representativeness is needed for selective annotation. Moreover, the way vote-kk defines and diversity are also different from the baselines: vote-kk defines representativeness as the number of neighbors during similarity-based prompt retrieval, which is effectively tailored to in-context learning; vote-kk directly optimizes for the diversity of selected samples using the in-context learning model’s prediction confidence.

Related Work

In-context learning with large language models has recently received an increasing amount of interest, partly due to its flexibility and sample efficiency (Liu et al., 2021). Several recent works proposed methods to improve in-context learning in many aspects: e.g., meta-training (Chen et al., 2022; Min et al., 2022b), task instructions (Efrat & Levy, 2020; Mishra et al., 2022; Wei et al., 2021; Sanh et al., 2022), or task formulation (Holtzman et al., 2021; Zhao et al., 2021; Min et al., 2022a). In this paradigm, the choice of in-context (i.e., demonstration) examples has been shown crucial (Liu et al., 2022; Rubin et al., 2022; Lu et al., 2022), while recent work raised questions as to the degree to which correct labels are necessary (Min et al., 2022c). This work proposes an annotation-efficient in-context learning framework by focusing on the choice of examples and its implications on the annotation cost.

Active Learning

Active learning aims to enable machine learning models to achieve similar or greater performance with fewer labeled training instances (Cohn et al., 1994; Settles, 2009). Our selective annotation step for in-context learning shares the same goal of reducing the annotation cost. Most active learning methods involve iterative parameter updates (e.g., Wang et al., 2017; Kasai et al., 2019), which are computationally expensive for large language models used in in-context learning. Similar to our vote-kk algorithm, Lin & Bilmes (2009) used the facility location objective to optimize representativeness. We observed that this objective largely underperforms vote-kk for in-context learning, probably due to the fact the vote-kk (1) is effectively tailored to the prompt retrieval step of in-context learning and (2) directly optimizes the diversity of selected samples (see §4.5). More recently, the effectiveness of active learning has been questioned when large-scale pretrained models are finetuned for various tasks (Karamcheti et al., 2021; D’Arcy & Downey, 2022). Our experiments (§3) showed that selective annotation helps reduce the annotation cost of in-context learning, departing from the recent observations on finetuning with active learning. We hypothesize that it is because in-context learning with similarity-based prompt retrieval is more robust to outliers since each test instance only retrieves its most similar examples. This is supported by § 4.4, where random prompt retrieval does not benefit from selective annotation.

Conclusion

Much recent work illustrated the ability of large language models to adapt to new tasks simply from a few demonstration examples. We presented in-depth studies on the implications of this ability for dataset annotation through the lens of selective annotation and introduced an annotation-efficient practice. The best selective annotation method explored in this paper, our vote-kk method, selects diverse, representative examples to annotate. In terms of the task performance, vote-kk improves the performance on 10 diverse tasks by a large margin. Moreover, vote-kk selective annotation yields similar performance to state-of-the-art supervised finetuning with 10-100×\times less annotation cost. We further show that the effectiveness of vote-kk is consistent with different language model sizes and domain shifts between training and test data. We hope that our findings will help researchers and practitioners efficiently design new natural language tasks and beyond.

Acknowledgements

We thank Sewon Min, Pradeep Dasigi, Yanda Chen, Yushi Hu, Alisa Liu, and the ARK group at UW for their helpful feedback on this work.

References

Appendix A Datasets and Tasks

Appendix B Prompt Templates

B.2 MRPC

B.3 SST5

B.4 MultiWoz

B.5 GeoQuery

B.6 DBpedia

B.7 MNLI

B.8 RTE

B.9 Natural Question

B.10 XSUM

Appendix C Detailed Main Results

This section provides a detailed version of our main results in Table 2, where the maximum performance and minimum performances among the three trials are reported. Results are shown in Table 7 and Table 8.

Appendix D Evaluate HellaSwag on OPT-175B model

Here we show that vote-kk also improves model performance for OPT-175B

Appendix E Removing outliers for finetuning

Here we show that explicitly removing outliers also helps finetuning to benefit from vote-kk.

Appendix F Diversity and Representativeness of Selected Samples

We hypothesized that both representativeness and diversity are crucial for selective annotation (§2.1). Here we evaluate the diversity and representativeness of samples that are selected by different methods, using the methods from prior work on active learning (Margatina et al., 2021); their measures of diversity and representativeness use token overlap or embedding cosine similarities. As shown in Table 10, vote-kk improves both the diversity and the representativeness as compared to random selection.

Appendix G Details of Selective Annotation Methods

In this section, we provide details of selective annotation methods used in Section 4.5.

Algorithm 1 describes the vote-kk selective annotation method introduced in Section 2.1.

G.2 Greedy Algorithm for Maximizing Facility Location

Lin & Bilmes (2009) proposed to maximize the facility location objective to optimize representativeness of the selected samples. Since this objective satisfies the submodular property, they applied a greedy algorithm as an approximation. Algorithm 2 describes the selective annotation method adapted from this greedy algorithm.

G.3 Embedding Diversity

This method aims to find diverse samples to annotate using embedding vectors. We first compute a vector representation for each unlabeled training instance by Sentence-BERT (Reimers & Gurevych, 2019), which is a variant of BERT (Devlin et al., 2019), finetuned to detect paraphrases.https://huggingface.co/sentence-transformers/all-mpnet-base-v2. For instance, consider an example from SST-5 sentiment analysis in Table 6:A very well-made, funny and entertaining picture. We simply run Sentence-BERT on this text input and average the resulting vectors over the words to obtain a vector representation.