Pre-Training to Learn in Context
Yuxian Gu, Li Dong, Furu Wei, Minlie Huang
Introduction
Pre-trained language models (PLMs; Han et al., 2021; Qiu et al., 2020) have shown strong abilities of learning and performing unseen tasks conditioning on several task examples or instructions in its context, which is called in-context learning (ICL; Brown et al., 2020). Compared to conventional fine-tuning methods, ICL adapts PLMs to downstream tasks only through inference, without parameter updates, which is computationally cheaper in practice and is closer to general AI.
However, PLMs trained on massive corpora to predict the next word given previous words are not explicitly taught to learn in the context. This makes ICL a surprising emergent ability but also indicates that the ICL ability of PLMs is not fully exploited. Garg et al. (2022) has shown that by directly training to do ICL in a meta-learning paradigm, models show strong performance on learning simple function classes in the context. In practical NLP scenarios, previous works Min et al. (2022b); Chen et al. (2022b) also enhance the ICL performance by meta-fine-tuning PLMs on a large collection of downstream tasks and evaluating them on unseen tasks. However, the low diversity of human-annotated downstream tasks restricts the performance of the meta-tuned model. Direct training on downstream tasks also brings undesired bias on specific input formats, label spaces, or domains, which hurts the generalization of PLMs.
To enhance the ICL ability while maintaining generalization, we propose PICL (Pre-training for In-Context Learning), a framework that exploits the PLM’s ICL ability by pre-training models on data automatically constructed from the general plain-text corpus. Our framework is based on a simple observation that many paragraphs in the text documents contain “intrinsic tasks”. As shown in the left part of Figure 1, each paragraph in the document contains an intrinsic task. When doing language modeling on each paragraph, models implicitly perform the corresponding intrinsic tasks simultaneously. This shares a similar idea with the prompt-learning paradigm Liu et al. (2021), where downstream data examples from NLP tasks are transformed into text sequences, and the model learns to perform the original tasks when trained on the text sequences with language modeling. Different from the downstream data, text paragraphs contain more diverse intrinsic tasks and have little bias on input formats, label spaces, or domains because they are free-form texts from the large-scale general corpus. By gathering and concatenating paragraphs with the same intrinsic tasks (right part of Figure 1), we can construct a meta-training dataset to pre-train the model to perform the intrinsic task conditioning on paragraphs in the context, and thereby improve the ICL ability.
We adopt a retrieval-based approach to gather paragraphs sharing the same intrinsic tasks from a general corpus. We first train an encoder to represent each paragraph in a vector space where paragraphs with the same intrinsic task have close embeddings. The encoder is trained with contrastive learning Khosla et al. (2020) on a collection of downstream datasets by taking examples from the same tasks as positive pairs and those from different tasks as negative pairs. Then, treating any paragraph in the corpus as a query, we retrieve the paragraphs with close representations to the query, namely, sharing the same intrinsic task with the query. Finally, we concatenate the query and the retrieved paragraphs to get a pre-training instance. Note that although we use downstream datasets, the model is trained on instances constructed from the general corpus, which ensures its generalization.
We evaluate the ICL performance of the model pre-trained with PICL on seven widely-used text classification datasets and Super-NaturalInstructions Wang et al. (2022), a benchmark whose test split contains more than 100 tasks formulated into text generation. Empirical results show the effectiveness of PICL, enabling the model to reach or even outperform larger models with nearly 4x parameters. Besides, we find that the PICL-trained model is more generalizable on various tasks than previous meta-fine-tuning methods. We also conduct extensive experiments to analyze several key factors of PICL.
Method
We first present an overview of PICL and then describe the details in the following sections. As shown in the right part of Figure 1, we construct the pre-training instances from a corpus consisting of paragraphs split from full documents by “n”. For each paragraph in , we first use a retriever to find paragraphs sharing the same intrinsic task (Sentiment Analysis) with . Then the retrieved paragraphs are treated as demonstrations and concatenated with to form a pre-training instance: . Finally, we adopt a language modeling objective to pre-train the model on the constructed instances.
In this way, the pre-training stage can be regarded as a meta-training process, where the model learns to solve the intrinsic task in conditioning on its context . Since is a large-scale general corpus, it contains a variety of intrinsic tasks and little domain bias, which ensures the generalization of the pre-trained model.
The main component of the retriever is a task-semantics encoder that represents a text paragraph as a -dimensional vector in a space , where paragraphs with the same intrinsic tasks have similar representations. We define the similarity between two paragraphs and using the dot product of their representations: .
We use RoBERTa Liu et al. (2019) as the base model of . The output vector is computed by averaging the last-layer representation of each token in the input paragraph.
Retrieval
We approximate that paragraphs whose representations are close to each other in share the same intrinsic task. Therefore, for every paragraph in , searches for paragraphs with embeddings closest to :
We employ the FAISS library Johnson et al. (2019) for efficient searching.
Contrastive Training
We adopt contrastive learning Khosla et al. (2020); Karpukhin et al. (2020) to train the task-semantics encoder . As shown in Figure 2, we take two paragraphs with the same intrinsic task as positive pairs and those from different tasks as negative pairs. However, the annotation of a paragraph’s intrinsic task is usually unavailable. To this end, we use a collection of downstream NLP datasets from various tasks whose examples are converted into text sequences with human-written prompts to train . In this way, treating each text sequence as a paragraph, we can regard the corresponding downstream task as the intrinsic task annotation. We assume that the instances from all downstream tasks form a dataset . For each , we have a positive instance sharing the same task with and a set consisting of negative instances with different tasks than , the loss function takes the form:
Positive and Negative Instances
For each , we randomly sample a positive instance belonging to the same task with from . As shown in Figure 2, contains two kinds of negative instances: (1) Easy Negatives sampled from and belonging to different tasks than . (2) Hard Negatives sharing the same prompt with but containing mismatched tasks. For instance, in Figure 2, we apply the prompt from the sentiment task to the summarization task to create the hard negative instance . This prevents the model from hacking the contrastive objective using prompts like “Guess the sentiment” and learning a trivial pattern matching but forces the model to extract task semantics from the whole paragraph.
2 Data Construction
For each paragraph , we concatenate the retrieved paragraphs with to get a pre-training instance . To improve the quality of the constructed data, we derive an approach to filter out instances that are less informative to ICL. We consider the following score to measure the informativeness of an instance based on the perplexity difference of the paragraphs in the instance before and after they are concatenated as a sequence:
where is the length of a sequence and is the language modeling probability based on any uni-direct PLMs. Given a manually set threshold , we retain the instances that satisfy . This criterion leverages the original ICL ability of the PLM. If concatenating the paragraphs results in lower perplexity, they are more correlated and may be more informative for ICL. We finally construct a pre-training corpus containing instances .
3 Pre-Training
We pre-train the model with auto-regressive language modeling on . Unlike previous works Min et al. (2022b); Chen et al. (2022b), which only compute the language modeling loss on the label tokens, we compute the loss on the whole sequence. There are two reasons for this choice. First, the intrinsic tasks are already in the natural language format, and it is unnecessary to split the input and the label. Second, we argue that computing loss on the whole sequence ensures a large token number in a forward batch, which is critical to maintaining the basic in-weights ability Chan et al. (2022). Therefore, the loss function is:
where is the parameters of the model. In addition, we find that adding a language modeling loss on the original full documents before being split into paragraphs benefits the performance. Therefore, the final optimization objective is:
where we set in our main experiments.
Experimental Setup
We merge OpenWebText Gokaslan et al. (2019), WikiCorpus Foundation (2022), and BookCorpus Zhu et al. (2015) to construct the pre-training data, where full documents are split into paragraphs by “n”. The corpus consists of 80M paragraphs, totaling about 30GB. For each paragraph, we search for demonstrations and concatenate them until 1024 tokens, the maximum input length constraint of the language model we used. This ensures that the model sees various demonstration numbers during pre-training. We use GPT2-Large Radford et al. (2019) to compute in Equation 3 and set for filtering. More details of data processing and statistics are shown in Appendix A.
2 Baselines
We consider four baselines in our experiments:
VanillaICL directly prompts a PLM with the concatenation of training examples to do ICL.
ExtraLM further pre-trains the PLM on the original full documents before being split into paragraphs with the language modeling objective.
Self-Sup Chen et al. (2022a) designs four self-supervised pre-training objectives, including Next Sentence Generation, Masked Word Prediction, Last Phrase Prediction, and Classification, to enhance the ICL performance. We conduct the self-supervised pre-training on our merged corpus for a fair comparison.
MetaICL Min et al. (2022b) meta-trains the model on a large collection of downstream human-annotated datasets for learning to learn in context. The meta-training instances are constructed by concatenating several training examples in each dataset to a single text sequence. We replicate the method on the training set of our task-semantics encoder for a fair comparison.
3 Evaluation
We evaluate the model trained with PICL on two kinds of downstream tasks.
We consider seven widely-used text classification datasets, including SST-2 Socher et al. (2013), SST-5 Socher et al. (2013), Subj Pang and Lee (2004), MR Pang and Lee (2005), RTE Dagan et al. (2006), CB De Marneffe et al. (2019), and AGNews Zhang et al. (2015) to evaluate the few-shot ICL performance of the trained models (see Appendix B.1 for more details). Note that these tasks are not included in the training set of the task-semantics encoder. We randomly sample 4 or 8 demonstrations from the official training sets of each dataset. Effects of other demonstration numbers can be found in Section 4.3. We compute the average accuracy scores on at most 1000 samples from the validation split of each dataset across five random seeds for selecting demonstrations.
Instruction Following
To test the generalization of PICL, we also evaluate the trained model on a larger range of tasks with more free-form inputs, including both human instructions and few-shot examples. We use the test split of Super-NaturalInstructions Wang et al. (2022) as the benchmark and exclude the tasks that appear in the training set of the task-semantics encoder, resulting in 105 evaluation tasks (see Appendix B.2 for a full list of tasks). Each task is specified with a human-written instruction and two or three demonstrations. We follow Wang et al. (2022) to formulate all tasks to the text generation format and score the outputs with ROUGE-L Lin (2004).
4 Settings
We train the task-semantics encoder on 37 tasks (see Appendix C) using up to 10K examples per task. To enhance generalization, we apply multiple prompts from PromptSource Bach et al. (2022) to one example and use 320 prompts in all. We use the in-batch negative trick Chen et al. (2020) to compute the contrastive loss. We set the learning rate to , the batch size to 64, and construct 4 hard negatives for each instance. The encoder is trained from RoBERTa for 1 epoch.
Language Model
We test PICL based on the 770M GPT2-Large Radford et al. (2019) unless otherwise specified. Results on larger models can be found in Appendix E.1. To save computational resources, we train the model from its pre-trained checkpoints. We also test the VanillaICL performance of larger models, including GPT2-xLarge Radford et al. (2019) (1.5B) and GPT-Neo Black et al. (2021) (2.7B) for reference.
Pre-Training
We set the maximum learning rate to and use the “inverse square root” scheduler Vaswani et al. (2017) with 1000 steps warmup. The model sees 131K tokens in a step and is pre-trained for 100K steps. It takes less than a day to finish pre-training on 64 V100 32G GPUs.
Results
Table 1 shows the results of few-shot text classification, from which we have 3 observations.
First, among the baselines with 770M parameters, simply further training the model on our corpus with language modeling improves the performance (ExtraLM). This is likely due to the higher domain diversity of our corpus. MetaICL is helpful on most datasets, which verifies the effectiveness of meta-training for ICL. Self-Sup fails to bring benefits on most datasets against VanillaICL, probably because the constrained label space of the Classification training task (only contains “True” and “False”) brings bias to the model’s output. This emphasizes the importance of using training objectives with little bias.
Second, we observe that the PICL-trained model outperforms the baselines with the same model sizes by a large margin on most datasets across different shots, verifying the effectiveness of PICL. An exception is RTE, where MetaICL performs the best. We speculate the reason is that some training tasks of MetaICL share the same label space with RTE (“Yes”/“No”), such as paraphrase identification. Min et al. (2022c) has shown that the label space plays a vital role in ICL, which explains the good performance of MetaICL on RTE.
Thrid, comparing models across different sizes, we find that increasing the model parameters boosts the performance, but PICL enables the 770M model to beat a 2.7B counterpart. This indicates that the ICL ability can be enhanced not only through scaling up the parameters. Improving the structure of the pre-training data is also beneficial. In Appendix E.1, we can see that PICL is also effective when applied to a 1.5B model.
2 Instruction Following
The results on Super-NaturalInstructions are shown in Table 2. We can see that PICL achieves higher overall instruction following performance than the baselines, outperforming a larger model with about 4x parameters.
In Figure 3, we compare the per-task performance of PICL and MetaICL because they share the most similar setting where human-annotated downstream datasets are used. We observe that PICL outperforms MetaICL on about 3/4 of evaluation tasks, indicating that compared to fine-tuning directly on downstream tasks, pre-training on intrinsic tasks constructed from the general plain-text corpus brings better ICL ability and ensures higher generalization performance across a broad range of tasks (see Appendix E.2 for more details).
Most tasks where MetaICL beats PICL belong to text classification whose output spaces are “Yes/No” or “True/False”. This matches the second observation in Section 4.1, where MetaICL predicts “Yes/No” well because of training on tasks that share the same label spaces. On the other hand, PICL performs much better on text generation, or tasks whose output spaces share the same semantics with “Yes/No” but use label words not in the training tasks of MetaICL (e.g., “Correct/Wrong”). This indicates that direct training on downstream datasets causes overfitting to specific labels. There are also tasks where PICL performs similarly to MetaICL, such as reasoning and word analogy. We notice that the improvements of PICL and MetaICL on these tasks are also marginal against VanillaICL probably because these tasks rely more on the “in-weights learning” ability Chan et al. (2022), rather than in-context learning.
3 Analysis
We compare different approaches to retrieve paragraphs and test the final model performance. We try randomly selecting paragraphs (Random), retrieving using the non-parametric approach (BM25), encoding each paragraph with the original pre-trained encoder as it is (RoBERTa), or using the encoder for sentence similarity Reimers and Gurevych (2019) (SRoBERTa). We also study different numbers of hard negatives (0, 1, 4) and downstream tasks (7, 24, 37) to train the task-semantics encoder in PICL. From the results in Table 3, we can see that all retrieval methods except Random bring improvements against VanillaICL on both text classification and instruction following settings, indicating that improving the coherence of the paragraphs in the pre-training data benefits ICL. Using the task-semantics encoder in PICL achieves the best performance, showing the importance of retrieving paragraphs based on task semantics rather than word overlap or sentence meanings. Comparing different settings to train the task-semantics encoder, we observe that increasing the number of hard negatives and training tasks improves the final performance. This is in line with previous works Karpukhin et al. (2020); Chen et al. (2020); He et al. (2020) that more challenging hard negatives benefit contrastive learning.
Effect of Demonstration Numbers
Training with PICL brings two benefits: (1) PLMs learn a format where demonstrations from the same task are concatenated as the prefix, which is beneficial when the model is evaluated under the same number of demonstrations. (2) PLMs learn a better ability to infer and perform tasks from the context, even when the demonstration numbers in evaluation and pre-training do not match. To differentiate these effects, we conduct pre-training on instances containing only 4, 8, or 16 demonstrations and test the trained models under different text classification shots. Results in Figure 4 show that when pre-trained with different demonstration numbers, the models generalize well to unseen demonstration numbers in evaluation, achieving similar performance with the default setup where the model sees various demonstration numbers in pre-training (PICL-default). This indicates that the models learn more than the input formats in PICL.
Effect of Filtering
We try different threshold values for filtering and report the scores on text classification tasks in Figure 5(a), while controlling the sizes of the constructed pre-training data the same. We find that yields the best performance, which means we retain an instance if and only if the perplexity of individual paragraphs is higher than that of the concatenated sequence (Equation 3). For lower , the pre-training data contain too many uninformative instances for ICL. For larger , we speculate that the filtering process relies on the original GPT2-Large too much. Since we also pre-train based on GPT2-Large, the filtering process reduces the signals in the constructed data beyond the base model’s ability.
Effect of Full Documents
In Figure 5(b), we report the model performance on text classification tasks when using different choices of , which controls the proportion of the full-document data. We find that balancing the constructed and full-document data performs the best (). When is too large, the model is trained mostly on our constructed data and overfits its bias inevitably introduced by the task-semantics encoder in the data construction process. When is too small, our method degenerates into ExtraLM.
Effect of Data Amount
We study the size effect of the corpus used to construct the pre-training data in PICL and report the performance on text classification tasks in Figure 6(a). We conduct the data construction on 0.01%, 0.1%, 1%, and 10% of the original 80M paragraphs (100%) and pre-train models for at most 100K steps until the validation losses begin to increase. From the results, we conclude that when the corpus is small, pre-training with the constructed data hurts the performance because the search library is too small to find paragraphs sharing the same intrinsic tasks. Training on small data for multiple epochs also causes overfitting. When the corpus contains more than 80K paragraphs (0.1%), adding more data constantly improves the performance, which is consistent with the scaling law Kaplan et al. (2020).
Data Comparison
We compare the usefulness of different pre-training data to enhance the ICL ability. In addition to the final model performance, we borrow the thoughts for designing the filtering criterion in Section 2.2 to measure the usefulness of a pre-training instance by computing the perplexity using a reference large PLM: GPT-J Wang and Komatsuzaki (2021) with 6B parameters. Lower perplexity means the correlation within the instance is higher and is intuitively more helpful for enhancing the ICL ability. In Figure 6(b), we show the perplexity and the final model performance of 3 pre-training data: original full documents before being split into paragraphs (Full Doc), concatenation of randomly selected paragraphs (Rand ), and the concatenated same-intrinsic-task paragraphs gathered using the retrieval method in PICL before filtering (). We can see that the data constructed by retrieval has much lower perplexity and correspondingly yields higher accuracy scores, which verifies its usefulness. In Appendix F, we present several examples of the retrieved paragraphs and the corresponding intrinsic tasks.
Related Work
Recently, in-context learning (ICL), where models perform tasks simply conditioning on instructions or the concatenation of examples in the context Brown et al. (2020), has been found promising for using PLMs in various application scenarios. To this end, there emerge many works to improve the ICL performance by calibrating the model predictions Zhao et al. (2021); Han et al. (2022); Holtzman et al. (2021); Min et al. (2022a), selecting or reordering demonstrations Rubin et al. (2022); Liu et al. (2022); Lu et al. (2022), designing pre-training tasks Chen et al. (2022a), and breaking the context length limits Hao et al. (2022). However, the underlying mechanism of ICL is poorly understood Min et al. (2022c). Therefore, some works propose mathematical frameworks to reveal how ICL works Xie et al. (2021); Olsson et al. (2022); Elhage et al. (2021), or investigate the pre-training data to explain ICL’s good performance Chan et al. (2022); Shin et al. (2022).
Multi-Task Fine-tuning for Cross-Task Generalization
Fine-tuning PLMs on a large collection of downstream tasks enables generalization to unseen tasks under zero-shot Wei et al. (2022); Sanh et al. (2022); Ouyang et al. (2022); Chung et al. (2022) and few-shot Min et al. (2022b); Chen et al. (2022b); Mishra et al. (2022); Garg et al. (2022) scenarios. However, the performance of multi-task fine-tuning is largely restricted by the diversity of the annotated training tasks Gu et al. (2022b), which requires massive human efforts to scale up. In addition, direct training on downstream tasks easily brings undesired bias. In this work, we propose to meta-train the model with the intrinsic tasks automatically collected from the large-scale general corpus, which is easier to scale up and introduces little bias.
Pre-training Data Programming
The conventional pre-training paradigm trains the model on plain-text corpora with the language modeling objective Radford et al. (2018, 2019); Brown et al. (2020). Recently works have found that carefully designed pre-training instances can further boost specific abilities like prompt adaption Gu et al. (2022a), reasoning Razeghi et al. (2022), or sentence representation Levine et al. (2021). Our work studies constructing pre-training instances to improve the PLM’s ICL ability while still maintaining its generalization on various NLP tasks.
Conclusion
This paper presents PICL, a framework that exploits the in-context learning ability of PLMs by pre-training models on concatenations of text paragraphs sharing the same “intrinsic tasks” gathered from the large-scale general corpus. In PICL, models learn to perform various intrinsic tasks conditioning on their context while preserving their generalization due to the little bias of the pre-training data. Extensive experiments show that PICL improves the ICL performance on various datasets against several baselines, enabling a 770 M model to outperform a larger model with about 4x parameters while maintaining good generalization across a wide range of tasks. For future work, we would like to consider adding human instructions to our pre-training framework to enhance more abilities of PLMs like zero-shot instruction following.
Limitations
One limitation of our paper is that the exact distribution of the intrinsic tasks in the original corpus and the constructed data is still unknown. Knowing the distribution can offer a better interpretation of the effectiveness of PICL, even of the strong performance of large language models. Besides, although we can find many constructed instances that share obvious intrinsic tasks (see Appendix F), there still exist some instances where the intrinsic tasks are hard to identify. How to better evaluate the contribution of these instances to the ICL ability or designing better filtering approaches to select more informative data for ICL is worth studying.
Our task-semantics encoder inevitably contains some bias because it is trained on downstream datasets, although we have tried to ensure a large number and diversity of the dataset collection. However, the final language model is pre-trained on the general corpus, and we add the full document loss, which eliminates the bias to some extent.
Regarding computing power, we acknowledge that our framework takes relatively large training resources in the retrieval and pre-training process. Therefore, we did not conduct experiments based on extra-large language models.
Acknowledgements
This work was supported by the NSFC projects (Key project with No. 61936010 ). This work was also supported by the Guoqiang Institute of Tsinghua University, with Grant No. 2020GQG0005.
References
Appendix A Details of the Pre-training Corpus
This section presents details of the data processing of the pre-training corpus and its statistics.
Our pre-training corpus is a merge of OpenWebText Gokaslan et al. (2019), WikiCorpus Foundation (2022), and BookCorpus Zhu et al. (2015), downloaded from the HuggingFace datasets repository1. We first split each document in the corpus into paragraphs with “n”. To avoid training with too short paragraphs, we concatenate a paragraph with previous paragraphs if the token number after concatenation is lower than 128. We also exclude paragraphs longer than 500 tokens because they are not likely to fit into an instance with more than 1 paragraph. The filtering process in Section 2.3 drops about 24% instances. The licenses of all corpora allow for scientific research.
Statistics
We plot the distribution of the mean paragraph length per instance in Figure 7(a) and the distribution of the paragraph number per instance in Figure 7(b). The average paragraph length is 150.0, and the average paragraph number in an instance is 11.7. We can see that the model sees various demonstration numbers in PICL pre-training.
Appendix B Details of the Evaluation Data
The details of each text classification dataset and the corresponding prompt in evaluation are listed in Table 6. All datasets are downloaded from the HuggingFace datasets repository1. We simplify the evaluation prompts as much as possible to reduce the effect of prompt engineering. Following previous works Brown et al. (2020); Sanh et al. (2022), the model is evaluated by the ranking score strategy, where we compare the perplexity of each classification label under the model and choose the label with the lowest perplexity. The licenses of all datasets allow for scientific research.
B.2 Instruction Following
The original test split of the benchmark Super-NaturalInstructions Wang et al. (2022) contains 119 tasks. We exclude tasks that appear in the training tasks of the task-semantics encoder or whose input length is too long to fit in the context of our model. Our final evaluation includes 105 tasks. A full list of the tasks is shown in Table 8. We use the same template to combine few-shot examples with task instructions as Wang et al. (2022). The license of this benchmark is Apache License 2.0.
Appendix C Details of the Downstream Training Data
The downstream datasets we use to train the task-semantics encoder are a merge of the training data used in Sanh et al. (2022) and the HRLR setting in Min et al. (2022b). All datasets are downloaded from the HuggingFace datasets repository https://huggingface.co/datasets/ and all prompts come from the PromptSource library Bach et al. (2022)https://github.com/bigscience-workshop/promptsource. We exclude datasets from the sentiment classification task, the topic classification task, and the natural language inference task because they are included in our text classification evaluation. We finally get a collection of 37 datasets, as listed in Table 5. The licenses of all datasets allow for scientific research.
Appendix D More Experimental Details
All model checkpoints we used come from the HuggingFace models repositoryhttps://huggingface.co/models. The searching interval of each hyper-parameter is listed in Table 4.
Appendix E More Results
We test PICL based on the GPT2-xLarge Radford et al. (2019) with 1.5B parameters. From the results in Figure 7, we can see that PICL is also applicable to larger models, outperforming the baselines based on the same-sized model on most datasets.
E.2 Instruction Following
We present the performance comparison between PICL and MetaICL per evaluation task in Figure 8. PICL outperforms MetaICL on 77 / 105 tasks, indicating that PICL ensures the better generalization of the trained model. The name of each task is also listed in Figure 8. We can see that the top three tasks where MetaICL performs the best are:
which are all “Yes/No” classification tasks. The top three tasks where PICL performs the best are:
which are text generation, text generation, and “Correct/Wrong” classification tasks respectively. The four tasks where PICL and MetaICL have the same scores are:
bard_analogical_reasoning_trash_or_treasure,
which belong to commonsense reasoning and word analogy tasks.
Appendix F Case Studies
In Table 9 and 10, we present several cases of the retrieved paragraphs and the corresponding intrinsic tasks. We can see that there exists a large range of intrinsic tasks in the constructed data and many of them do not appear in the training data of the task-semantics encoder, which shows the generalization of the encoder.