What Makes Good In-Context Examples for GPT-$3$?
Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, Weizhu Chen
Introduction
GPT- Brown et al. (2020) is a new breakthrough in NLP research. Previously, NLP models are pre-trained on large quantities of data and fine-tuned on a specific task and dataset. What sets GPT- apart from other pre-trained language models is its impressive “in-context” few-shot learning ability. Provided with a few in-context examples, GPT- is able to generalize to unseen cases without further fine-tuning. This opens up many new technological possibilities that are previously considered unique to human. For example, NLP systems can be developed to expand emails, extract entities from text, generate code based on natural language instructions with a few demonstration examples.
Despite its powerful and versatile in-context learning ability, GPT- has some practical challenges/ambiguities. The original paper Brown et al. (2020) utilizes task-relevant examples that are randomly sampled from the training set to construct the context. In practice, we observe that the performance of GPT- tends to fluctuate with different choices of in-context examples. As shown in Table 1, the variance of the empirical results with distinct in-context examples can be significant. The results are highly sensitive to the examples. Our work aims to carefully examine this issue to gain a deeper understanding on how to better select in-context examples to unleash GPT-’s few-shot capabilities and further improve its performance.
A brute-force approach would be to perform combinatorial search over the entire dataset. Unfortunately, this strategy is computationally expensive and thus impractical in many cases. To this end, we investigate the influences of employing different in-context examples on the empirical results. Interestingly, we found that the in-context examples that are closer to the test sample in the embedding space consistently give rise to stronger performance (relative to the farther ones). Inspired by this observation and the recent success of retrieval-augmented models Hashimoto et al. (2018), we propose to utilize nearest neighbors of a given test sample (among all the training instances available) as the corresponding in-context examples. The retrieved examples, along with the test sample, are provided to GPT- for the final prediction.
To verify the effectiveness of the proposed method, we evaluate it on several natural language understanding and generation tasks, including sentiment analysis, table-to-text generation and open-domain question answering. It is observed that the retrieval-based in-context examples unleash the few-shot capabilities of GPT- much more effectively than a random sampling baseline. Even with a smaller number of in-context examples, the proposed strategy empowers GPT- to achieve stronger performance. Moreover, we find that the specific sentence encoders employed for the retrieval procedure play a critical role. Thus, an extensive exploration regarding different pre-trained encoders is conducted, and it is shown that encoders fine-tuned on natural language matching tasks serve as more effective in-context examples selector on the QA task. Detailed analysis and case study further validate the effectiveness of proposed methods. In summary, our contributions in this paper are as follows:
i) to the best of our knowledge, we take a first step towards understanding the sensitivity of GPT-’s few-shot capabilities with respect to the selection of in-context examples;
ii) to alleviate the sensitivity issue, an additional retrieval module is introduced to find semantically-similar in-context examples of a test instance to construct its corresponding input, which greatly outperforms the baseline based on random sampled examples;
iii) fine-tuning the retrieval model on task-related dataset(s) leads to even stronger empirical results with GPT-;
iv) the performance of GPT- improves as the number of examples available for retrieval increases.
Method
The in-context learning scenario of GPT- can be regarded as a conditional text generation problem. Concretely, the probability of generating a target is conditioned on the context , which includes examples, and the source . Therefore, the prediction corresponding to the source can be expressed as:
where LM denotes the parameters of the language model, and is a context string. In GPT-, the is created by concatenating training instances along with their corresponding labels. As shown in the illustration of Figure 1, GPT- is asked to translate “mountain” to its German version based on the three examples given as part of the input.
For GPT-, this generation process is implemented through a giant transformer-based model architecture Vaswani et al. (2017); Brown et al. (2020). Given the large size of the GPT-3 model, it would be very computationally-involved to fine-tune it on task-specific samples. Thus, GPT- is typically leveraged in a in-context learning manner as described above. It has been shown that GPT- has powerful few-shot capabilities, where it can perform quite well with only a small number of demonstrations provided. Unfortunately, as shown in Table 1, the results of GPT tends to fluctuate significantly with different in-context examples chosen. Here we aim to alleviate this issue via judious in-context examples selection.
2 The Impact of In-Context Examples
Given the observation that the empirical results of GPT- are sensitive to the chosen in-context examples, we look at the role of in-context examples from an empirical perspective. Previous retrieve-and-edit literatures usually retrieve prototypes that are close to the test source in some embedding space. These examples and the test source often share semantic or lexical similarities. This hints on how we may select in-context examples for GPT-.
To this end, we examine the impact of the distance between in-context example and the test sample on GPT-’s performance. Concretely, a comparison is made on the the Natural Questions (NQ) dataset between two in-context example selection strategies. Given each test example, the first method utilizes the farthest training instances to construct the context provided to GPT-, while the second employs the closest neighbors. We use the CLS embeddings of a pre-trained RoBERTa-large model as the sentence representations to measure the proximity of two sentences (using the Euclidean distance).
For evaluation, test questions are randomly sampled and the average Exact Match (EM) scores with the two distinct strategies are reported in Table 2. It can be observed that the nearest neighbors, as the in-context examples, give rise to much better results relative to the farthest ones. Moreover, the pre-trained RoBERTa model serves as effective sentence embeddings for the retrieval procedure.
3 k𝑘kNN-augmented In-Context Example Selection
Based on the findings above, we propose KATEKATE: Knn-Augmented in-conText Example selection, a strategy to select good in-context examples for in-context learning. The process is visualized in Figure 2. Specifically, we first use a certain sentence encoder to convert sources in both the training set and test set to vector representations. For online prediction, we can convert the training set first and encode each test source on the fly. Then, for each test source , we retrieve its nearest neighbors from the training set (according to the distances in the sentence encoder’s embedding space). Given some pre-defined similarity measure such as the cosine similarity, the neighbors are ordered in such a way that when .
Afterwards, the sources are concatenated with their corresponding targets to form the context , which is further sent to GPT- along with the test input. The algorithm chart is presented in Algorithm 1. Note that different numbers of in-context examples can be employed here, and we conduct ablation study on its impact in a later section.
A core step for our context selection approach is mapping sentences into a latent semantic space, leaving a question as what sentence encoders we should choose. We compared among existing pre-trained text encoders and found them sufficient to retrieve semantically similar sentences. The sentence encoders can be divided into two categories.
The first category includes most generally pre-trained sentence encoders such as a pre-trained BERT, RoBERTa, or XLNet models. These models have been trained on large quantities of unsupervised tasks and achieved good performance on many natural language tasks. The corresponding embeddings contain rich semantic information from the original sentences.
The second category includes sentence encoders fine-tuned on specific tasks or datasets. For example, a sentence encoder trained on the STS benchmark dataset should be able to assess similarities among different questions better than a generally pre-trained sentence encoder. Reimers and Gurevych (2019, 2020) have shown that these fine-tuned encoders have achieved great performance on tasks such as sentence clustering, paraphrase mining, and information retrieval.
Experimental Setup
We apply the NN in-context selection method to the following three tasks: sentiment classification, table-to-text generation, and question answering (QA). Datasets and the common data split setups are shown in Table 3. In terms of the hyper-parameters in the GPT- API, we set the temperature to 0. We let GPT- keep generating tokens until there is a special token “\n”.
To retrieve semantically-similar training instances, we consider two types of sentence embeddings.
The original pre-trained RoBERTa-large model Liu et al. (2019), which is abbreviated as KATE;
The RoBERTa-large model fine-tuned on task-related datasets: i) fine-tuned on the SNLI and MultiNLI dataset (KATE); ii) first fine-tuned on the SNLI and MultiNLI dataset and then on the STS-B dataset (KATE).
Notably, all the sentence encoders share the same the architecture, where the only differences are the specific datasets used for fine-tuning. Euclidean distance is used for the KATE case, while cosine similarity is employed for KATE and KATE.
For sentiment classification, we select in-context examples under the transfer setting, where one dataset is treated as the training set and the evaluation is made on another dataset. This transfer setting is designed to simulate a real-world scenario where we would like to leverage an existing labeled dataset for a unlabeled one (of a similar task).
Specifically, we select in-context examples from the SST-2 training set Socher et al. (2013); Wang et al. (2018) and ask GPT- to make predictions on the IMDB test set Maas et al. (2011). To explore whether a sentence encoder fine-tuned on a similar task would benefit KATE’s performance, we also employ a pre-trained RoBERTa-large model fine-tuned on the SST-2 training set (dubbed as KATE). The performance is measured by the accuracy over the entire IMDB test set. The number of in-context examples is chosen to be since adding more examples does not further improve the performance.
Table-to-Text Generation
Given a Wikipedia table and a set of highlighted cells, this task focuses on producing human-readable texts as descriptions. ToTTo Parikh et al. (2020) is utilized for evaluation due to its popularity. We use BLEU Papineni et al. (2002) and PARENT Dhingra et al. (2019). metrics for evaluation. The ToTTo code base contains both evaluation and preprocessing scriptsThe ToTTo code base can be found at https://github.com/google-research/language/tree/master/language/totto. Due to the input length limit of GPT- (currently the token limit is 2048), we add an extra preprocessing step by deleting the closing angle brackets such as /cell and /table to save some space. The number of in-context examples is set as .
Question Answering
Given a factual question, the model is asked to generate the correct answer. Following prior studies, we use the Exact Match (EM) score to measure the performance of GPT- on open-domain QA tasks. The EM score is defined as the proportion of the number of predicted answers being exactly the same as (one of) the ground-truth answer(s). The matching is performed after string normalization, which includes article and punctuation removal. We conduct experiments on three open-domain QA benchmarks: Natural Questions (NQ) Kwiatkowski et al. (2019), Web Questions (WQ) Berant et al. (2013), and Trivia Question Answering (TriviaQA) Joshi et al. (2017). For this task, we pick the nearest 64 neighbors as the in-context examples for NQ and WQ and nearest 10 neighbors for TriviaQA (The retrieved 64 examples could not fit into 2048 token limit for TriviaQA. For fair comparison, we set the number of in-context examples to be 10 for TriviaQA for both the baseline and KATE method). The evaluation is done on the test sets of NQ and WQ and the dev set of TriviaQA.
2 Baseline Methods
For each test sentence, we randomly select in-context examples from the training set. We refer to this method as Random in the experimental results. To have a fair comparison with KATE, the number of in-context examples in this random baseline is the same as KATE to ensure fair comparison. On the test set, the random baseline is repeated for five times to obtain the average score and corresponding standard deviation.
k𝑘k-Nearest Neighbor
Additionally, to investigate whether the retrieval module is complementary to GPT-’s few-shot learning ability, we further consider a -nearest neighbor baseline. Specifically, for text generation tasks, the target associated with the first retrieved example is considered as the predicted target for the test sample. As to the sentiment analysis and QA tasks, the top retrieved examples are utilized, where the final prediction is determined by majority voting among the examples’ targets. If there is a tie case, we take the target of the example that is most similar to the test sentence as the prediction. To ensure fair comparison, we compare the baseline NN and KATE under the same embedding space of a pre-trained RoBERTa-large model. This baseline is abbreviated as NN.
Experimental Results
We first evaluate KATE on the sentiment analysis task. The results are shown in Table 4. It can be observed that KATE consistently produces better performance relative to the random selection baseline. Notably, there is no variance with the obtained results since the same set of retrieved in-context examples are employed. For the KATE method, when a pre-trained sentence encoder is fine-tuned on NLI or NLI+STS-B datasets, the performance slightly decreases. Since the objectives of the IMDB dataset and the NLI+STS-B datasets are different, this shows that fine-tuning on a dissimilar task can hurt KATE’s performance. Moreover, KATE performs worse than KATE because the sentence encoder has been further fine-tuned on the STS-B dataset. In contrast, KATE obtains the best accuracy, showing that fine-tuning on a similar task can benefit KATE’s performance. To verify that the gains are not merely from the retrieval step, we further compare KATE with the NN. It turns out that the performance of the NN method is similar to random guessing. This observation is consistent when one neighbor or three neighbors are retrieved. Notably, with the embeddings of the RoBERTa-large model fine-tuned on the SST-2 dataset, the accuracy of NN is 92.46, which is lower than that obtained with KATE. These results suggest that the GPT-3 model is critical to the final results, and the retrieval module is complementary to GPT-’s few-shot capabilities.
2 Table-to-text Generation
We utilize the ToTTo dataset to evaluate KATE on the table-to-text generation task. The results are shown in Table 5. The KATE method gives rise to considerable gains over the random baseline, according to both the BLEU and PARENT scores. On a finer scale, the evaluation can be done on the overlap subset and the nonoverlap subset. The overlap dev subset shares a significant number of header names with the training set, while the nonoverlap one does not share any header names. It can be observed that the KATE method improves the results on both the overlap and the nonoverlap subsets, meaning that the retrieval module is helpful for both situations where the test set follows the distribution of the training set and where the test set is out of distribution of the training set. Similar to sentiment analysis, there is a slight drop in performance from KATE to KATE and KATE. This is due to the difference between the objectives of the ToTTo dataset and NLI+STS-B datasets. The drop from KATE to KATE further validates the idea that fine-tuning on a dissimilar task can hurt KATE’s performance. For the NN baseline, it performs much worse than the random selection method and the KATE method, again suggesting that the retrieval process and GPT-3 work together collaboratively to achieve better results.
To understand how the retrieval mechanism helps GPT-’s predictions, we conduct a case study on the retrieved examples (see Table 6). By retrieving relevant examples from the training set, KATE provides useful detailed information within the table, e.g., the number of points, rebounds, and assists, to GPT- for more accurate description. On the other hand, the random selection method has the issue of hallucination, where the generated sequences contain information (i.e., “senior year” and “University of Texas”) not present in the table.
3 Questing Answering
We also evaluate KATE on the open-domain QA tasks, as shown in Table 7. For the QA tasks, we compare with some state-of-the-art methods such as RAG Lewis et al. (2020) and T5 Raffel et al. (2019). Both methods require fine-tuning on the specific datasets. The KATE method again improves GPT-’s few-shot prediction accuracies substantially across various benchmarks. It is worth noting that the fine-tuned transformer models serve as better sentence encoders for retrieval purpose (compared with the RoBERTa-large model without fine-tuning). KATE and KATE improve upon KATE because this time fine-tuning on NLI or STS-B datasets is helpful for retrieving semantically similar questions from the QA datasets. Moreover, on the NQ and TriviaQA datasets, further fine-tuning on the STS-B dataset improves KATE’s results. We also try reducing the number of in-context examples to be as small as five for both the random and KATE methods, where KATE outperforms the baseline as well. Therefore, the advantage of KATE over the random baseline holds for both small and large numbers of in-context examples. More details can be found in Section 5.1. We evaluate the other baseline NN by using the top- nearest neighbor. We also explore using 64 nearest neighbors (10 for TriviaQA) to determine the answer (by majority voting explained in Section 3.2). The EM score tends to be similar to retrieving the top- nearest neighbor. These NN baseline results again suggest that the retrieval module and GPT- work together to achieve better performance.
To investigate why the retrieval examples are helpful, we further present a case study. Concretely, the retrieved in-context examples from the NQ dataset are shown in Table 8. For the first and second cases, the random baseline provides wrong answers because GPT- is unable to recall the exact detail. However, the in-context examples selected by KATE contain the correct details, which facilitates GPT- to answer the questions. For the third test question, the random baseline leads GPT- to misinterpret the question as asking for a specific location. In contrast, KATE selects similar questions which ask for the origins of objects. Using these in-context examples, GPT- is able to interpret and answer the question correctly.
Analysis and Ablation Study
We first investigate the impact of the number of in-context examples on KATE’s performance. Concretely, on the NQ dataset, we choose the number of in-context examples to be , , , , and , and KATE is compared with the random baseline and KATE across different settings. As shown in the left plot of Figure 3, both KATE and the random baseline benefit from utilizing more in-context examples. However, KATE consistently outperforms the random selection method, even when the number of in-context examples is as few as . This result is interesting because in practice, employing less in-context leads to more efficient inference with GPT-.
2 Size of Training Set for Retrieval
We further examine how the size of the training set may influence the KATE method. On the NQ dataset, we create new subsets from the original training set, with sizes of 1k, 2k, 5k, 10k, 30k, and 70k, respectively. In-context examples are retrieved from these subsets instead of the original training set. The number of nearest neighbors is set to 64. We compare KATE with the random selection method and KATE, and the results are shown in the right plot of Figure 3. For KATE and KATE, as the size of the training set for retrieval increases, the EM scores also increase. In contrast, the result of the random sampling baseline does not change much. Intuitively, as the training size gets larger, it is more likely for KATE to retrieve relevant in-context examples to help GPT- answer a question correctly. As we have shown previously in Table 8, the retrieved in-context examples could provide critical detailed information to GPT-, thus helping GPT- to better answer the questions.
3 Order of In-context Examples
Moreover, we explore how the order of in-context examples may affect KATE’s results. As mentioned in Section 2.3, under the standard setting, the retrieved in-context examples are ordered such that whenever . Here, we randomly permute the order of in-context examples in the NQ dataset for the proposed KATE method, and conduct the experiments for different orders. Additionally, we explore the reverse order where whenever . The results are presented in Table 9. On this particular NQ dataset, the reverse order performs the best. One possible explanation is that since tokens next to each other have similar positional embeddings, putting the most similar sentences close to the test example may be helpful for GPT-3 to leverage the corresponding information. However, we also did the experiments on the WQ and TriviaQA and find that the default order performs slightly better than the reverse order. Hence, the choice of orders is data-dependent. Addtionally, it can be observed that the variation among the NQ results tends to be quite small (compared with the difference between the random baseline and KATE), indicating that the example order does not have a significant impact on KATE’s performance.
Related Work
NLP systems have made tremendous progress by pre-training models on unlabeled text. For text classification tasks, notable models include BERT Devlin et al. (2018), RoBERTa Liu et al. (2019), and XLNet Yang et al. (2019). For text generation tasks, notable models include BART Lewis et al. (2019), T5 Raffel et al. (2019), mT5 Xue et al. (2020), XLM Lample and Conneau (2019), GPT Radford et al. (2018), and GPT-2 Radford et al. (2019). These models encapsulate rich information to facilitate a wide range of downstream tasks ranging from natural language understanding to generation. These models can be adapted to many different tasks via fine-tuning. GPT- Brown et al. (2020), however, can be adapted to many downstream tasks without fine-tuning. Given just a few in-context examples, GPT- is able to quickly pick up patterns and produce answers analogously both in terms of the answer style and content. Thus, GPT- may be considered as a pattern recognizer to perform in-context learning. People have just started trying to understand GPT- from different perspectives. As mentioned in the introduction, Hendrycks et al. (2020) studies which categories of questions GPT- is more capable of answering. Our work focuses on how to choose good in-context examples.
Retrieval-based Text Generation
There is a long history of applying information retrieval in text generation Sumita and Hitoshi (1991). It is very related to the exemplar-based learning Jäkel et al. (2008); Ziyadi et al. (2020). The central idea is to treat retrieved samples as exemplars/prototypes and perform some editings on them. Some representative applications in the field of deep learning include machine translation Gu et al. (2018), sentiment transfer Li et al. (2018); Guu et al. (2018), QA Karpukhin et al. (2020); Mao et al. (2020), dialogue generation Yan et al. (2016); Cai et al. (2018); Song et al. (2016); Pandey et al. (2018); Weston et al. (2018); Wu et al. (2019), text summarization Cao et al. (2017); Peng et al. (2019), data-to-text generation Peng et al. (2019), and text-to-code generation Hashimoto et al. (2018). However, all these retrieve-and-edit frameworks require their decoders to be trained from scratch. This makes the editor network task- and data-specific. In contrast, GPT- in one perspective can be regarded naturally as a universal editor, adaptive to a wide range of tasks. Our work uniquely examines how to maximize the advantage of using GPT- without fine-tuning. For example, the more semantically similar context we provide to GPT-, the better results the model can generate. Other editors or generators do not have this ability.
Improve NLP Systems with k𝑘kNN
A recent line of work tries to incorporate nonparametric methods to improve a given model’s performance. These methods first access the test sample’s hidden representation and look for the nearest neighbors of this test sample in the database. Once the nearest neighbors are found, their labels are used to augment the model’s prediction. For example, the newly introduced NN-LM Khandelwal et al. (2019), NN-MT Khandelwal et al. (2020), and BERT-NN Kassner and Schütze (2020) generate the next token by retrieving the nearest neighbors from the datastore. Another related work is NN classification model Rajani et al. (2020), where they use NN as backoff when the confidence is low from the fine-tuned classification model. There are two key differences between our work and other approaches. First, other approaches modifies the model’s next token distribution using the nearest neighbors. However, we only changes the conditional text using the nearest neighbors. Second, other approaches can access the model’s parameters and embeddings which we do not have access to. Instead, we use some other independently pre-trained models to get the sentence embeddings to retrieve nearest neighbors.
Conclusion
This work presented a first step towards investigating the sensitivity of GPT- to in-context examples. To this end, we proposed KATE, a non-parametric selection approach that retrieves in-context examples according to their semantic similarity to the test samples. On several natural language understanding and generation tasks, the proposed method improves GPT-’s performance, over the random sampling baseline, by a significant margin. Moreover, we found that fine-tuning the sentence embeddings for retrieval on task-related datasets gave rise to further empirical gains. Detailed ablation studies were conducted to explore the robustness of KATE to different hyperprameters, such as the number of in-context examples, examples’ order, etc. We hope this work could provide insights for better understanding the behaviors of GPT- and represents a helpful step towards further improving its few-shot capabilities.