CrossFit: A Few-shot Learning Challenge for Cross-task Generalization in NLP

Qinyuan Ye, Bill Yuchen Lin, Xiang Ren

Introduction

Pre-trained language models fine-tuned with abundant task-specific data have become the predominant recipe for state-of-the-art results in NLP. However, these approaches are heavily dependent on large-scale labeled datasets that are expensive to create, and the resulting models still generalize poorly to out-of-distribution inputs created with small, harmless perturbations Ribeiro et al. (2020). In retrospect, researchers have advocated for building more human-like, general linguistic intelligence that can “reuse previously acquired knowledge about a language and adapt to a new task quickly” Yogatama et al. (2019); Linzen (2020).

Existing work has approached this problem via better few-shot fine-tuning, by re-formulating target tasks into cloze questions that resembles the pre-training objective Schick and Schütze (2020a, b), generating prompts and using demonstrations Gao et al. (2020). Such progress primarily focus on improving instance-level generalization, i.e., how to better generalize from few labeled instances to make predictions about new instances, within the scope of one individual task. From a broader perspective, human-like learning ability also benefits from task-level generalization, or cross-task generalization, i.e., how to learn a new task efficiently given experiences of learning previous tasks.

Such ability has been widely studied in computer vision and robotics community Yu et al. (2020); Triantafillou et al. (2020), but is relatively under-explored in NLP. Pruksachatkun et al. (2020) and Vu et al. (2020) study transferability between one intermediate task and a given target task, while it’s possible to further improve performance with multiple intermediate tasks. Han et al. (2018) and Bansal et al. (2020a) focus on cross-task generalization within the scope of classification tasks, whereas humans can generalize across different task formats (classification, multiple choice, generation, etc.), goals (question answering, fact checking, etc.) and domains (biomedical, social media, etc.).

Towards developing general linguistic intelligence, we present CrossFit, a few-shot learning challenge to acquire, evaluate and analyze cross-task generalization in a realistic setting, with standardized training pipeline, data access and evaluation protocol. The CrossFit challenge requires a model to first learn from a set of seen tasks in an upstream learning stage, and then perform few-shot learning on a set of unseen tasks, as illustrated in Fig. 1. In accompany, we introduce the NLP Few-shot Gym, a repository of 160 few-shot NLP tasks gathered from open-access resources, covering a wide range of capabilities and goals, and formulated into a unified text-to-text format. To analyze the capability and limitation of existing approaches to the CrossFit challenge, we design eight specific seen/unseen task partitions.

With the CrossFit Challenge and the NLP Few-shot Gym, we aim to investigate the following research questions:

Q1. Can we teach cross-task generalization ability to pre-trained models with existing methods?

Q2. During upstream learning, is it better to be “well-rounded” (learning from diverse tasks) or be “specialized and targeted” (learning from tasks in the same category with unseen tasks)?

Q3. Does it help if we have more labelled data for seen tasks during upstream learning?

To address the above questions, we empirically analyze the performance of multi-task learning and three meta-learning algorithms (MAML Finn et al. (2017), first-order MAML and Reptile Nichol et al. (2018)). We observe that these approaches can indeed lead to better few-shot performance on unseen tasks. Interestingly, simple multi-task learning outperforms existing meta-learning methods in many cases, encouraging future research on identifying the reasons and developing improved meta-learning methods. For Q2, we observe that performance of individual unseen tasks varies with different selection of seen tasks, calling for more thorough investigation of the relationship between task similarity and transferability. As for Q3, we find that enlarging the size of upstream data does not necessitate better cross-task generalization abilities. We envision cross-task generalization to be an integral component towards general linguistic intelligence, and we hope CrossFit serves as a useful testbed for driving related progress.

Related Work

Few-shot learning refers to teaching models a new task with a small number of annotated examples. Large-scale pre-trained language models (e.g., BERT Devlin et al. (2019)) have demonstrated great ability to learn new tasks efficiently via fine-tuning Zhang et al. (2021). Schick and Schütze (2020a, b) proposed pattern-exploiting training (PET), which formulates text classification and NLI tasks into cloze questions (or “prompts”) that resemble masked language modeling. PET can be further improved by generating prompts automatically and incorporating demonstrations into the input Gao et al. (2020); and by densifying the supervision signal with label conditioning Tam et al. (2021). While successful, in these approaches the downstream tasks are learned in isolation. Our work aims to boost few-shot learning ability on unseen tasks via acquiring cross-task generalization ability from diverse seen tasks.

Meta-learning in NLP.

Recent works have explored meta-learning methods for relation classification Han et al. (2018); Gao et al. (2019), general text classification Dou et al. (2019); Bansal et al. (2020a, b), low-resource machine translation Gu et al. (2018), cross-lingual NLI/QA Nooralahzadeh et al. (2020). In general, these works apply meta-learning algorithms to a set of sub-tasks; however the sub-tasks are either synthetic (e.g., classifying a new set of five relations is a new sub-task) or drawn from a rather narrow distribution (e.g., QA in one language is a sub-task). In our work, we explore a more realistic setting – learning from a set of NLP tasks with diverse goals: classification, question answering, conditional generation, etc. This setting is attracting attention in NLP community rapidly and is also explored in very recent work Zhong et al. (2021); Mishra et al. (2021); Bragg et al. (2021); Wei et al. (2021).

Unifying NLP Task Formats.

Researchers have explored unifying the formats of different tasks, in order to better enable knowledge transfer, e.g., DecaNLP McCann et al. (2018), UFO-Entail Yin et al. (2020) and EFL Wang et al. (2021). Following T5 Raffel et al. (2020), we adopt a unified text-to-text format that subsumes all text-based tasks of interest. Related to our work, UnifiedQA Khashabi et al. (2020) examines the feasibility of training a general cross-format QA model with multi-task learning. Our work extends from these ideas, and we significantly enlarge the task repository to 160 to broaden the coverage, in hopes to build a general-purpose few-shot learner.

The CrossFit Challenge

In this section, we present the CrossFit Challenge, a problem setting for acquiring and evaluating cross-task generalization. Ideally, a strong CrossFit system can capture cross-task generalization ability from a set of seen tasks and thus adapts to new unseen tasks efficiently.

The meaning of “task” is overloaded: “tasks” can be categorized at different granularity (e.g., text classification vs. QA, yes/no QA vs. machine reading comprehension), and from different aspects (e.g., domain, label space). Herein we take a general formulation by defining a “task” with its training and testing examples. We define a task TT as a tuple of (Dtrain,Ddev,Dtest)(\mathcal{D}_{train},\mathcal{D}_{dev},\mathcal{D}_{test}). Each set D\mathcal{D} is a set of annotated examples {(xi,yi)}\{(x_{i},y_{i})\} in text-to-text format. In few-shot setting, the size of Dtrain\mathcal{D}_{train} and Ddev\mathcal{D}_{dev} are required to be small (e.g., 16 example per class for classification tasks).

Existing work mostly focuses on improving instance-level generalization for individual task by using task-specific templates. Performance on individual tasks is used as the measure of success. For the CrossFit Challenge, we aim to acquire cross-task generalization and build better general-purpose few-shot learners, which calls for a different problem setting with distinct training procedure and evaluation protocol.

2 Problem Setting

To acquire and evaluate cross-task generalization, we first gather a large repository of few-shot tasks T\mathcal{T}, and partition them into three non-overlapping sets Ttrain\mathcal{T}_{train}, Tdev\mathcal{T}_{dev}, Ttest\mathcal{T}_{test}. In hopes to examine the capability and limitation of an approach in different settings, and to answer our research questions, we design multiple task partitions with different focuses. Details of the repository and partitions, or as we name them, the NLP Few-shot Gym, are deferred to §4.

Learning Stages.

A CrossFit method may learn from Ttrain\mathcal{T}_{train} and perform necessary tuning with Tdev\mathcal{T}_{dev} in the upstream learning stage; it is then evaluated with few-shot tasks in Ttest\mathcal{T}_{test}:

Upstream learning stage. Here, the algorithm has access to the Dtrain\mathcal{D}_{train} and Ddev\mathcal{D}_{dev} for each training task in Ttrain\mathcal{T}_{train}, while Dtest\mathcal{D}_{test} is unavailable. The algorithm also has access to all data in Tdev\mathcal{T}_{dev}, but for validation purpose only (i.e., it is not allowed to use Tdev\mathcal{T}_{dev} to update model weights).

Few-shot learning stage. In this stage, Ttest\mathcal{T}_{test} became available. Models resulting from the upstream learning stage are required to learn from Dtrain\mathcal{D}_{train} via a particular few-shot learning method (e.g., direct fine-tuning). The final few-shot learning performance is evaluated on Dtest\mathcal{D}_{test}. For clarification, the performance on the Ddev\mathcal{D}_{dev} of a task in Tdev\mathcal{T}_{dev} or Ttest\mathcal{T}_{test} will be used for tuning hyper-parameters during fine-tuning. The overall performance on Tdev\mathcal{T}_{dev} is used for tuning tuning hyper-parameters during upstream learning.

Evaluation Metric.

Evaluating the performance of a model on a diverse collection of NLP tasks is inherently challenging, as different tasks use different metrics. It is thus not reasonable to simply aggregate performance of classification tasks (e.g., accuracy, F1) and generation tasks (e.g., ROUGE, BLEU) by taking the average.

To address this problem, we first narrow down to a collection of 7 evaluation metrics: classification F1, accuracy, QA F1, exact match (EM), Rogue-L, Matthew correlation, and Pearson correlation, which cover all tasks in our experiments. Then, we define Average Relative Gain (ARG), a metric that computes relative performance changes before and after the upstream learning stage for each test task, and finally take the average across all test tasks.

For example, suppose we have Ttest={TA,TB}\mathcal{T}_{test}=\{T_{A},T_{B}\}. If an upstream learning algorithm helps improve the few-shot learning performance from 50%50\% F1 score to 70%70\% on task TAT_{A} (i.e., a 40%40\% relative improvement), and from 40%40\% accuracy to 30%30\% on task TBT_{B} (i.e., −25%-25\% relative improvement), the final ARG on Ttest\mathcal{T}_{test} would be computed as 40%+(−25%)2=7.5%\frac{40\%+(-25\%)}{2}=7.5\%.

The ARG metric reflects the overall performance gain on all tasks in Ttest\mathcal{T}_{test}, no matter what specific metrics each task uses. We use ARG for a high-level comparison, and we still analyze the performance for each task (e.g., absolute performance metrics, performance growth with “more shots”, sensitivity to different selection of Ttrain\mathcal{T}_{train}) in our in-depth analysis.

NLP Few-shot Gym

Towards learning to generalize across tasks in CrossFit challenge, we need a resource that contains sufficient number of tasks, covering a wide range of NLP applications, and presented in a unified text-to-text format. Herein, we introduce the NLP Few-shot Gym, a repository of 160 few-shot tasks gathered from existing open-access datasets.

We choose to use Huggingface Datasetshttps://huggingface.co/datasets. It is an extensible library that provides access to 626 open-access NLP datasets (as of Feb 25th, 2021) with a unified, open-source API. Lhoest et al. (2021) as the pool of our candidate tasks. We filter these datasets on a case-by-case basis, mainly using the following criteria: (1) We focus on English monolingual datasets. (2) We exclude datasets that require information retrieval, as they require a separate retriever. (3) We exclude sequence labeling tasks (e.g., dependency parsing, NER), which are highly dependent on tokenization, and are hard to evaluate in text-to-text format. (4) We exclude datasets dealing with extremely long documents (e.g., a scientific paper) as input, as most pre-trained models cannot process such long input sequences. We finalize our selection with 160 datasets which are detailed in Appendix A.

2 A Unified Text-to-Text Format

We follow Raffel et al. (2020) and convert all of our datasets into a unified text-to-text format. For example, the task of natural language inference (originally a sentence-pair classification problem) becomes: premise: hypothesis: , and the target sequence is either the word entailment, contradiction or neutral. As for machine reading comprehension tasks, the input format is question: context: and the target sequence is the correct answer span. We also reference the format for QA tasks from UnifiedQA Khashabi et al. (2020).

3 Formulating Few-shot Tasks

We mainly follow the practice in Gao et al. (2020) for few-shot sampling. For classification and regression tasks, we include 16 training examples per class in DtrainD_{train}. For other types of tasks, we include 32 examples in DtrainD_{train}. In conformity with real-world situations where labeled data are scarce, we assume a development set DdevD_{dev} which shares the same size with DtrainD_{train}.

We sample Dtrain\mathcal{D}_{train} and Ddev\mathcal{D}_{dev} splits from each dataset’s original train set with 5 different random seeds. This helps us reduce variance during few-shot evaluation, and also enlarges the number of few-shot tasks used for learning. Consequently, the “effective size” of our NLP Few-shot Gym is 160×5=800160\times 5=800, while we use the number 160160 throughout the paper to avoid possible confusion.

We use the original development set for each dataset as Dtest\mathcal{D}_{test}, or withhold 20%20\% of the dataset when the official development split is not available. The held-out test examples are sampled once before sampling Dtrain\mathcal{D}_{train} and Ddev\mathcal{D}_{dev}.

4 Task Ontology and Partitions

As mentioned in §3.2, a CrossFit method is expected to first acquire cross-task generalization on a set of Ttrain\mathcal{T}_{train} and evaluate such ability on Ttest\mathcal{T}_{test}. To comprehensively analyze to what extent a trained model can generalize, and how its behavior differs in different scenarios, we need to build different partitions of (Ttrain,Tdev,Ttest)(\mathcal{T}_{train},\mathcal{T}_{dev},\mathcal{T}_{test}).

Towards this goal, we first manually classify the 160 tasks and form a task ontology with categories and sub-categories, as shown in Fig. 2. The first-level categories include classification, question answering, conditional generation, and others.We later discuss the limitation of this design in §6-Q2 Further, we design eight different partitions of (Ttrain,Tdev,Ttest)(\mathcal{T}_{train},\mathcal{T}_{dev},\mathcal{T}_{test}). We illustrate four partitions in Fig. 3 and provide more details in Table 1.

Our Partition 1 randomly split all 160 few-shot tasks into the three sets, where ∣Ttrain∣=120|\mathcal{T}_{train}|=120 and ∣Tdev∣=∣Ttest∣=20|\mathcal{T}_{dev}|=|\mathcal{T}_{test}|=20. The design of Partition 1 mimics the real-world language learning environment where the goal is to build a general-purpose few-shot learner, and a set of diverse tasks (Ttrain\mathcal{T}_{train}) are used to train the learner. Our Partition 2.1-2.3 withhold 10 classification tasks for development and 10 more for testing. The Ttrain\mathcal{T}_{train} is controlled to have either 100% classification tasks, 100% non-classification tasks, or half-and-half. These three partitions help us to understand the influence brought by different task distribution in Ttrain\mathcal{T}_{train}. The remaining four partitions still focus on crossing task boundaries, but in a finer granularity: seen and unseen tasks are in the same category, but not the same sub-category. For example, Partition 3.1 has 57 non-NLI classification tasks as Ttrain\mathcal{T}_{train}, and 8 NLI tasks as Ttest\mathcal{T}_{test}. These partitions help us to understand whether cross-task generalization in this finer granularity is easier for model to acquire.

Methods to CrossFit

We mainly use BART-Base Lewis et al. (2020) as the text-to-text transformer for our analysis in the CrossFit setup. We leave confirmatory experiments with T5-v1.1-Base and BART-Large model in Appendix C.

This serves as the basic baseline method for the CrossFit challenge, which does not make use of Ttrain\mathcal{T}_{train} or Tdev\mathcal{T}_{dev}, or go through the upstream learning stage. For each task T∈TtestT\in\mathcal{T}_{test}, we directly fine-tune the text-to-text model with its Dtrain\mathcal{D}_{train}, tune the hyper-parameters with Ddev\mathcal{D}_{dev}, and assess its performance with the test set Dtest\mathcal{D}_{test}. We use the performance of direct fine-tuning as the base for computing ARG scores of other CrossFit approaches. We expect a model trained with upstream learning would capture cross-task generalization ability and thus have better ARG scores.

Multi-task Learning (MTL).

A straight-forward yet effective method is to combine the dataBoth Dtrain\mathcal{D}_{train} and Ddev\mathcal{D}_{dev} are used, as Ddev\mathcal{D}_{dev} is used for gradient updates in meta-learning algorithm. We do so to make sure that the data access for the two methods is fair. in the training tasks to learn a multi-task model, before fine-tuning it on each test task. Specifically, we gather source-target examples for all tasks in Ttrain\mathcal{T}_{train} and fine-tune the text-to-text model with these examples. Then we use the resulting checkpoint as initialization and perform the same procedure in “direct fine-tuning” for each test task in Ttest\mathcal{T}_{test}. The performance gain over the direct fine-tuning is used for computing its overall ARG score.

Model-Agnostic Meta-learning (MAML).

Cross-task generalization ability, closely aligns with the concept of learning to learn. Hence, we use MAML Finn et al. (2017), a representative meta-learning approach during upstream learning. The core concept of MAML is to learn a set of initialization weight, from which the model adapts fast to a new task within few gradient updates. In MAML training, we iterate through tasks in Ttrain\mathcal{T}_{train} to update the model. For each train task (Dtrain,Ddev)(\mathcal{D}_{train},\mathcal{D}_{dev}), we first sample a support batch Bsupport\mathcal{B}_{support} from Dtrain\mathcal{D}_{train} and a query batch Bquery\mathcal{B}_{query} from Ddev\mathcal{D}_{dev}. We use fθf_{\theta} to denote the text-to-text model with parameters θ\theta. Using Bsupport\mathcal{B}_{support}, we first compute the updated parameters θ′\theta^{\prime} with gradient descent (i.e., the inner loop). Due to the large size of pre-trained text-to-text models, we use one gradient update in the inner loop, i.e., θ′=θ−α∇θL(fθ,Bsupport).\theta^{\prime}=\theta-\alpha\nabla_{\theta}\mathcal{L}(f_{\theta},\mathcal{B}_{support}). Then we apply the updated text-to-text model fθ′f_{\theta^{\prime}} to Bquery\mathcal{B}_{query}, and do one step of meta-optimization (i.e., the outer loop), with θ←θ−β∇θL(fθ′,Bquery)\theta\leftarrow\theta-\beta\nabla_{\theta}\mathcal{L}(f_{\theta^{\prime}},\mathcal{B}_{query}).

First-order MAML.

First-order MAML Finn et al. (2017) avoids second-order optimization and improves training stability using the first-order approximation by differentiating with respect to the fast weights θ′\theta^{\prime} instead of the original parameters θ\theta for the gradient ∇θL(fθ′,Bquery)\nabla_{\theta}\mathcal{L}(f_{\theta^{\prime}},\mathcal{B}_{query}), i.e., θ←θ−β∇θ′L(fθ′,Bquery).\theta\leftarrow\theta-\beta\nabla_{\theta^{\prime}}\mathcal{L}(f_{\theta^{\prime}},\mathcal{B}_{query}).

Reptile.

Reptile Nichol et al. (2018) is another memory-efficient, first-order meta-learning algorithm that first makes multiple gradient updates in the inner loop, then directly uses θ′−θ\theta^{\prime}-\theta to approximate ∇θL(fθ′,Bquery)\nabla_{\theta}\mathcal{L}(f_{\theta^{\prime}},\mathcal{B}_{query}), i.e., θ←θ+β(θ′−θ)\theta\leftarrow\theta+\beta(\theta^{\prime}-\theta).

Empirical Analysis

In this section we look to interpret the results and answer our research questions. We summarize the ARG scores in Table 1 and plot the performance of each test task (for each partition) in Fig. 4-5.

[couleur= msftBlack!05, epBord= 1, arrondi=0.1, logo=\bclampe,marge= 2, ombre=true, blur, couleurBord=msftBlack!10, tailleOndu=3, sousTitre =Q1. Can we teach pre-trained LMs to generalize across tasks with existing methods? ]

From Table 1, we observe that, on average, the tested upstream learning methods indeed improve cross-task generalization: their ARG scores are positive, meaning that they are better than direct fine-tuning (ARG=0%). Further, by aggregating results from all upstream learning methods and task partitions, we find that the performance on 51.47% test tasks are significantly improved (>5%>5\% relative improvement compared to direct fine-tuning); 35.93% tasks are relatively unaffected (between ±5%\pm 5\%); and 12.60% tasks suffer from worse performance (<−5%<-5\%).

Correlated Performance Gains.

The performance gain obtained with different upstream learning methods are correlated with each other – i.e., tasks that benefit from multi-task learning is likely to also benefit from meta-learning. For the Random partition, the Spearman Correlation between the relative improvement brought by MTL and MAML is 0.660.66, with pp value equals to 0.00150.0015. This suggests that different upstream learning methods, while taking different optimization objectives, capture similar inductive bias from Ttrain\mathcal{T}_{train}.

MTL is a strong baseline.

Surprisingly, the most straight-forward multi-task learning method is hard to beat. This could be counter-intuitive, as meta-learning methods are specifically designed for rapid generalization to unseen tasks, sharing the same goal with our CrossFit challenge. We think there are three possible reasons: (1) Due to memory constraints, we limit the number of inner-loop updates to be one, which may be insufficient. Also, meta-learning methods are highly sensitive to hyper-parameters and even random seeds Antoniou et al. (2019), which we do not tune exhaustively for practical reasons. (2) Text-to-text transformers have much more complex architectures, while most meta-learning methods are typically applied to small feed-forward/convolutional networks. (3) The CrossFit challenge has a highly diverse set upstream tasks, which may introduce under-explored difficulties. That being said, we believe it is important to identify the true cause, and to develop improved meta-learning methods for the CrossFit challenge as future work.

Forgetting Pre-Trained Knowledge.

A few test tasks have negative performance gain after upstream learning, including Glue-COLA (measuring linguistic acceptability) and Domain Crawl (separating domain names into tokens) in the Random Partition setting. For Glue-COLA, similar observations are reported by Pruksachatkun et al. (2020) in an intermediate-task transfer learning setting, where the authors conjecture catastrophic forgetting of the masked language modeling (MLM) tasks may be the cause. BART uses denoising pre-training objective, a variant of MLM. Intuitively, Domain Crawl is also one of the most similar tasks to denoising in all test tasks, which further supports this hypothesis. We thus conjecture that for test tasks that resemble pre-training objectives, upstream learning could hurt performance due to the catastrophic forgetting phenomena.

Understanding negative transfer Wu et al. (2020) and selecting source tasks to avoid negative transfer Vu et al. (2020) are also growing research topics. In this work we refrain from further investigation; however we believe combating negative transfer and thus improving CrossFit performance is a promising future direction.

[couleur= msftBlack!05, epBord= 1, arrondi=0.1, logo=\bclampe,marge= 2, ombre=true, blur, couleurBord=msftBlack!10, tailleOndu=3, sousTitre =Q2. Well-rounded or specialized? Which is a better strategy of upstream learning?]

“Learning to be well-rounded vs. learning to be specialized” is a common dilemma that human learners struggles with. For the CrossFit challenge, the former refers to learning from a set of diverse tasks in upstream learning; the latter refers to learning from a set of tasks closer to target few-shot tasks. To study this research question, we want to find out which option works better in upstream learning. Put differently, we aim to analyze the influence of upstream task selection for a fixed set of the downstream tasks.

Setup.

We first conduct controlled experiments with Partition 2.1-2.3, where Ttest\mathcal{T}_{test} is a fixed set of classification tasks, and Ttrain\mathcal{T}_{train} varies. In Partition 2.1, all tasks in Ttrain\mathcal{T}_{train} are classification tasks (i.e., “specialized and targeted”); in Partition 2.2, half of the tasks are classification tasks (i.e., “well-rounded”); in Partition 2.3, all tasks are non-classification tasks (i.e., “specialized in an opposite direction”, for a controlled experiment).

Analysis and Discussion.

It is surprising at first that non-classification tasks and classification tasks are equivalently helpful in terms of ARG scores (see Fig. 5). On a second thought, this observation is encouraging as it demonstrates that acquiring cross-task generalization is feasible and promising, even when Ttrain\mathcal{T}_{train} and Ttest\mathcal{T}_{test} are drastically different. It also suggests that our categorization of tasks (§4.4) may not align with how models learn transferable skills: selecting Ttrain\mathcal{T}_{train} tasks that have the same format and goal as the test task may not lead to optimal transfer.

In retrospect, we acknowledge that our design of ontology and partitions based on task format and goal is flawed. This is merely one aspect of “task similarity”. However, understanding the complex relationship between tasks is another challenging and under-explored problem. We consider our ontology as a starting point, rather than a fixed final one. We use the current ontology to guide our experiment and analysis, and we hope future analysis could help build a more informative ontology.

Case Studies.

We further look at cases where a test task appear in Ttest\mathcal{T}_{test} of multiple partitions. For example, AI2_ARC and Race-High are in the Ttest\mathcal{T}_{test} of both Random partition and Held-out-MCQA partition. We present the results in Table 2. In general, the performance of these tasks varies when different Ttrain\mathcal{T}_{train} sets are used. However, we have not found consistent patterns of what type of Ttrain\mathcal{T}_{train} lead to better performance for a specific test task.

[couleur= msftBlack!05, epBord= 1, arrondi=0.1, logo=\bclampe,marge= 2, ombre=true, blur, couleurBord=msftBlack!10, tailleOndu=3, sousTitre =Q3. Does it help if we have more labelled data for upstream tasks?]

As described in §4.3, we limit our upstream tasks to be also few-shot: classification tasks have 16 examples per class, and non-classification tasks have 32 examples. This decision is empirically determined following prior works Schick and Schütze (2020a, b); Gao et al. (2020) and makes our extensive analysis practical and efficient. It is possible that using more data for each upstream task can significantly improve cross-task generalization. To investigate this, we conduct a set of controlled experiments where the number of examples in upstream tasks are changed to $$ times of the original size. We use the Held-out-Para Partition and multi-task learning for the experiments, and present the result in Fig. 6. Surprisingly, we find that the effect from using more upstream data is inconsistent on different target tasks. The overall ARG for all sizes are close: even 8x larger upstream data leads to only 4% improvement in ARG. We conclude that enlarging the size of data during upstream learning does not necessitate better cross-task generalization ability. This also justifies our decision to keep upstream tasks few-shot.

[couleur= msftBlack!05, epBord= 1, arrondi=0.1, logo=\bclampe,marge= 2, ombre=true, blur, couleurBord=msftBlack!10, tailleOndu=3, sousTitre =Q4-Q6. Additional Analysis] Due to space limit, we summarize our other findings below and defer the details to Appendix C.

Few-Shot →→\rightarrow More-Shot (Q4).

In practice, users may continue to collect data over time. We wonder if cross-task generalization ability is still helpful for medium/high-resource target tasks. We find that the performance gain from upstream learning is still evident when 1024 shots are available. The performance gap diminishes with millions of training examples.

Using Different Base Models (Q5).

We extend our analysis on BART-base (139M) to larger pre-trained text-to-text Transformers: BART-Large (406M) and T5-v1.1-Base (248M). Generally, the performance grows with models sizes with only few exceptions, which suggests that upstream learning methods we use are model-agnostic, and can be applied to larger models to further improve few-shot performance.

Integration with PET Training (Q6).

Pattern-exploiting training (PET) Schick and Schütze (2020a, b) was originally proposed for classification tasks and encoder language models. We test a few variants of PET training with BART-Base and try applying PET training after upstream learning. In general we observe deteriorated performance compared to direct fine-tuning. We hypothesize that PET methods are not directly applicable to encoder-decoder language models used in our study.

Conclusion and Future Work

In this paper, we study the problem of building better few-shot learners via acquiring cross-task generalization ability from diverse NLP tasks. Towards our goal, we introduce the CrossFit Challenge, an task setup that standardizes the training pipeline, data access and evaluation protocol. We also present the NLP Few-shot Gym, a repository of 160 diverse few-shot NLP tasks, to support CrossFit learning in different scenarios. We empirically demonstrated that cross-task generalization can be acquired via multi-task learning and meta-learning; confirmed that the selection of seen tasks would influence the few-shot performance on unseen tasks.

We have highlighted several unexpected or undesired observations in our analysis, for which we invite future work in understanding and combating related issues. In addition, we envision the CrossFit Challenge and the NLP Few-shot Gym to serve as the testbed for many interesting “meta-problems”, such as (1) learning to generate prompt for diverse task formats and further improve learning efficiency Shin et al. (2020); Gao et al. (2020); (2) learning to select appropriate source tasks to learn from during upstream learning Zamir et al. (2018); Standley et al. (2020), potentially with task2vec methods Achille et al. (2019); Vu et al. (2020); (3) applying task augmentation strategies to prevent over-fitting Murty et al. (2021); (4) learning to accumulate knowledge and avoid catastrophic forgetting in an continual learning setup Jin et al. (2021); (5) decomposing complex tasks into atomic tasks and exploring cross-task generalization through the lens of compositionality Andreas et al. (2016); Khot et al. (2021).

Acknowledgments

We thank authors and crowd-workers of all datasets used in our study. We thank huggingface datasets team for making datasets more accessible. We thank anonymous reviewers and members of USC INK Lab for their valuable feedback. This work is supported in part by the Office of the Director of National Intelligence (ODNI), Intelligence Advanced Research Projects Activity (IARPA), via Contract No. 2019-19051600007; the DARPA MCS program under Contract No. N660011924033; the Defense Advanced Research Projects Agency with award W911NF-19-20271; NSF IIS 2048211.

References

Appendix A Selected Tasks in NLP Few-shot Gym

Appendix B Details about Task Partition

B.2 Partition 2.1. 45cls

B.3 Partition 2.2. 23cls+22non-cls

B.4 Partition 2.3. 45non-cls

B.5 Partition 3.1. Held-out-NLI

B.6 Partition 3.2. Held-out-Para

B.7 Partition 4.1. Held-out-MRC

B.8 Partition 4.2. Held-out-MCQA

B.9 Partition 5. Held-out-GLUE

To examine whether combining our methods with template-based training Schick and Schütze (2020a, b); Gao et al. (2020) results in even better few-shot performance, we add another partition that uses all non-GLUE classification tasks as Ttrain\mathcal{T}_{train}, and all GLUE tasks as Ttest\mathcal{T}_{test}.

Appendix C Additional Results and Analysis

[couleur= msftBlack!05, epBord= 1, arrondi=0.1, logo=\bclampe,marge= 2, ombre=true, blur, couleurBord=msftBlack!10, tailleOndu=3, sousTitre =Q4. Does the improved cross-task generalization ability go beyond few-shot settings?]

In real-world applications, annotated data usually grow for a few-shot task over time. Is upstream learning still helpful when a target task has more shots? To study this question, we study CommonsenseQA (in Held-out-Multiple-Choice Partition), ROPES (in Held-out-MRC Partition), and MNLI (in Held-out-NLI Partition) as target tasks in medium and high-resource scenarios. We take their corresponding checkpoints after upstream learning and conduct experiments in medium and high-resource scenarios. That is, we randomly sample {32,64,…,4096}\{32,64,\dots,4096\} examples from the three datasets, and use them as Dtrain\mathcal{D}_{train}. Then, we sample a Ddev\mathcal{D}_{dev} with the same size as Dtrain\mathcal{D}_{train}, or has the size of 1024 if ∣Dtrain∣>1024|\mathcal{D}_{train}|>1024. We also try fine-tuning with the full dataset.We do five random samples of 1024 examples as Ddev\mathcal{D}_{dev} and use the remaining examples in the original train set as Dtrain\mathcal{D}_{train}. We use the original dev set for testing. The performance of these settings is shown in Fig. 7.

From Fig. 7, we see that the benefits brought by upstream learning methods extend into medium resource cases with up to 2048 training examples. For CommonsenseQA, checkpoints from upstream learning outperform direct fine-tuning significantly, even with the full dataset. This finding encourages the use of upstream learning before task-specific fine-tuning when the target task has limited annotation. On the other hand, for resource-rich tasks (e.g., MNLI), the improvement brought by upstream learning diminishes. This aligns with the findings of (Wang et al., 2020) who discuss the benefits of pre-training on resource-rich tasks.

[couleur= msftBlack!05, epBord= 1, arrondi=0.1, logo=\bclampe,marge= 2, ombre=true, blur, couleurBord=msftBlack!10, tailleOndu=3, sousTitre =Q5. Can we further improve few-shot performance by using different/larger pre-trained models?]

We have been mainly using BART-Base (139M parameters) as the main network, while it is possible to further push the limits of few-shot learning by using scaling up to larger models or using different model architectures. Previous work has shown that scaling up model size leads to better performance Raffel et al. (2020); Brown et al. (2020). Moreover, since meta-learning algorithms are naturally unstable, it is important to verify whether they function as expected with larger models. In Q5, we experiment with T5-v1.1-Base (248M)T5-Base was trained on a mixture of downstream tasks during its pre-training; such practice strays from the purpose of our study. Therefore, we use T5-v1.1-Base model, which is trained with the C4 Corpus only. and BART-Large (406M) model with Held-out-Para Partition to verify these assumptions. We only consider first-order methods, as second-order optimization with these larger models is impossible with our available computation.

Our results are plotted in Fig. 8. In Fig. 8(a) we compare the few-shot performance of direct fine-tuning on these three pre-trained models. On average, few-shot performance grows with models size, with a few exceptions such as QQP+T5-v1.1-Base and MRPC+Bart-Large. In Fig. 8(b-c) we plot the effect brought by upstream learning method for larger models. Except for FoMAML+T5-v1.1-BaseWe observe instability in training loss during FoMAML training for T5-v1.1-Base., upstream learning methods consistently improves few-shot performance on Ttest\mathcal{T}_{test}, which verifies that upstream learning methods we use are model-agnostic, and can be applied to larger models to further improve few-shot performance.

[couleur= msftBlack!05, epBord= 1, arrondi=0.1, logo=\bclampe,marge= 2, ombre=true, blur, couleurBord=msftBlack!10, tailleOndu=3, sousTitre =Q6. Can we use pattern-exploiting training to replace direct fine-tuning to achieve even better performance?] Pattern-exploiting training (PET) is a novel method that formulate a target task into cloze-style questions Schick and Schütze (2020a, b); Gao et al. (2020). This approach narrows the gap between the masked language modeling objective during pre-training and downstream task fine-tuning, and therefore leads to more efficient transfer. PET is demonstrated to be effective with encoder models (e.g., RoBERTa), however, whether it is applicable to text-to-text models with auto-regressive decoders is underexplored to the best of our knowledge. In Q6, we study whether applying PET-style methods to text-to-text models is feasible, and whether combining the two methods further pushes the few-shot performance.

To align with the experiment settings in Schick and Schütze (2020a, b); Gao et al. (2020), we introduce a new task partition “Held-out-GLUE”, which uses non-GLUE classification tasks as Ttrain\mathcal{T}_{train}, and GLUE tasks as Ttest\mathcal{T}_{test}. We use the top 3 patterns in Gao et al. (2020) for each GLUE task, and use the ensemble of the three models to produce the final prediction.

Since pattern-exploiting training is originally designed for encoder models (e.g., BERT/RoBERTa), we first tried two of its variants that adapts it to our auto-regressive transformer models. The first variant generates complete sentence, e.g., generate “The movie is great. A wonderful piece” from “The movie is great. A piece” for sentiment classification. The second variant generates only the word “wonderful”, from “The movie is great. A piece”. Though the first variant is more similar to the denoising pre-training objective of BART, we find the second variant to have better performance.

We then launch pattern-exploiting training using variant two with the original BART-Base models. We observe negative performance on average (leftmost blue bar in Fig. 9). Performance is improved with CoLA and MRPC, but not with the remaining GLUE tasks. We further launch experiments with/without pattern-exploiting training, with our upstream learning checkpoints. Still pattern-exploiting training leads to deteriorated performance on average.

We stop further investigation since this is out of the scope of our study. Still we believe it is important to identify the reasons and develop pattern-exploiting methods for auto-regressive models.

Appendix D Reproducibility

All our experiments are implemented with Huggingface Transformershttps://github.com/huggingface/transformers Wolf et al. (2020). For higher-order optimization in the meta-learning approach optimization, we use higher libraryhttps://github.com/facebookresearch/higher. Our code has been uploaded in supplementary materials, and is also open-sourced at https://github.com/INK-USC/CrossFit.

Hyper-parameters.

We mainly follow the practice in Gao et al. (2020). During few-shot fine-tuning, we select the learning rate from {1e−5,2e−5,5e−5}\{1e-5,2e-5,5e-5\}, and the batch size from {2,4,8}\{2,4,8\}, based on DdevD_{dev} performance. We set the total number of updates to be 1000, number of warmup updates to be 100. We evaluate the model on DdevD_{dev} every 100 steps.

Infrastructure and Runtime.

Upstream learning are done with one single Quadro RTX 8000 (48GB). Upstream learning jobs finishes within 3 hours on average. Fine-tuning experiments are all done with one single GPU, with either NVIDIA Quadro GP100, NVIDIA Quadro RTX 8000, NVIDIA Quadro RTX 6000, NVIDIA GeForce RTX 1080 Ti, or NVIDIA GeForce RTX 2080 Ti, based on availability. Fine-tuning on one few-shot task (with hyperparmeter tuning for all 5 random samples) takes approximately 4 hours on average.

Number of Parameters.

BART-Base model contains 139 million parameters. T5-v1.1-Base model contains 246 million parameters. BART-Large model contains 406 million parameters.