FLEX: Unifying Evaluation for Few-Shot NLP

Jonathan Bragg, Arman Cohan, Kyle Lo, Iz Beltagy

Introduction

Few-shot learning, the challenge of learning from a small number of examples, is critical for developing efficient, robust NLP techniques . In recent years, separate threads of few-shot NLP research have pursued goals like generalization to new classes [e.g., 5, 25], adaptation to new domains and tasks [e.g., 3, 4, 21], and direct application of pretrained language models (LMs) [e.g., 10, 55, 56, 24]. Unfortunately, despite the shared goal of advancing few-shot NLP techniques, the community does not know which techniques work best or even if they perform better than simple baselines. Evaluation suites across these research threads are disjoint, lack challenging-yet-realistic testing setups (e.g., class imbalance, variable training set sizes, etc.), and do not employ careful experimental design to ensure accurate and precise evaluation estimates and minimal computational burden. Prior work in few-shot learning outside of NLP serves as a stark warning of the consequences of improper measurement: Dhillon et al. showed that techniques from several years of prior work did not make clear progress due to large overlapping accuracy distributions and, moreover, do not outperform a simple, carefully-tuned baseline.

Need for systematic benchmark design As such, a high-quality benchmark is urgently needed to enable rigorous comparison of techniques across disjoint, highly-active threads of few-shot NLP research. But what should such an evaluation suite look like? Some best practices for evaluation of few-shot methods have been introduced in the computer vision (CV) literature and should be applied to NLP. However, unifying few-shot NLP work introduces new challenges. For example, the benchmark needs to test all types of transfer studied in separate research threads to measure progress on new techniques that make gains in each of these important generalization settings (§2). Also, given the importance of zero-shot learning and learning from task descriptions , the benchmark needs to include zero-shot episodes and textual labels to enable measuring progress for models that do not use conventional supervised training, including methods that leverage the latent knowledge in pretrained LMs . Further, the benchmark must accommodate new, computationally-expensive approaches, without overly reducing the number of evaluation episodes at the expense of statistical accuracy .

Need for a robust few-shot model Recent prompt-based models have shown strong results in few-shot learning. These models leverage the power of (often large) pretrained language models and adapt the format of downstream tasks to the underlying pretraining objective (e.g., Masked Language Modeling). This way, given the right natural language prompt (and sometimes verbalizers and additional demonstrative examples), the model can quickly fine-tune on the downstream task . However, adapting task formats to the underlying (masked) language modeling objectives is not straightforward; such models have been shown to be sensitive to varying choices of the prompt/demonstrations, training settings, hyperparameters, and learning algorithms , often requiring large held out sets and/or complex methods to overcomes such challenges. Can models eschew complex prompt engineering by unifying pretraining and downstream task formats?

In this paper, we tackle these key issues by introducing FLEX—Few-shot Language Evaluation across (X) many transfer types—and contributing the following:

FLEX Principles (§3), a set of requirements and best practices for few-shot NLP evaluation that enables unified, rigorous, valid, and cost-sensitive measurements.

Sample Size Design: In support of valid, cost-sensitive measurement, we introduce a novel approach to few-shot sample size design (§5) that optimizes for a benchmark’s statistical accuracy and precision while keeping computational costs accessible to a broad range of researchers.

FLEX benchmark (§4), an implementation of the FLEX Principles. It tests across four few-shot transfer settings,Prior work evaluated at most two settings. and includes a public leaderboard for few-shot NLP that covers 20 datasets across diverse NLP tasks (e.g., NLI, relation classification, entity typing). Table 1 summarizes key differences between FLEX and other few-shot NLP evaluation suites.

UniFew (§6), a prompt-based model for few-shot learning in NLP. While most existing methods leverage pre-trained LMs for few-shot learning, LM pre-training tasks do not closely match natural downstream task formats, requiring complex methods (e.g., extensive prompt-engineering, use of verbalizers, episodic hyperparameter tuning, custom learning algorithms) to make these models work in few-shot setting. Instead, the key idea of our model, UniFew, is to close the gap between pre-training and fine-tuning formats by posing tasks as multiple-choice QA and using an underlying model that is pre-trained on a similar natural QA task format. This eliminates the need for complexities of adapting downstream tasks to the LM objectives, while resulting in competitive performance with both recent few-shot and meta-learning methods.

To aid similar efforts, our release of FLEX includes a toolkit for benchmark creation and few-shot NLP model development, which we used to create the FLEX benchmark and train UniFew.

Background and Related Work

We first provide background and notation for few-shot learning and evaluation, then discuss related work in NLP and outside NLP that motivated us to create the FLEX Principles and benchmark.

Broadly, modern approaches to few-shot learning are evaluated in a three-phase procedure . In the first phase, a general-purpose pretrained model is obtained. In the subsequent “meta-training” phase,Meta-training may include a “meta-validation” component, for validating generalization. techniques aim to adapt the model to be well-suited for few-shot generalization. Finally, a “meta-testing” phase evaluates the adapted model in new few-shot prediction settings.

Let D\mathcal{D} be a dataset of (x,y)(x,y) examples with full label set YD\mathcal{Y}_{\mathcal{D}}. From it, we construct three sets of episodes, corresponding to meta-training, meta-validation, and meta-testing and denoted by Etrain\mathcal{E}_{\textrm{train}}, Eval\mathcal{E}_{\textrm{val}}, and Etest\mathcal{E}_{\textrm{test}}, respectively. Each episode in each of these sets is a few-shot problem with its own test set and other attributes. Formally, each episode EE is a tuple (DtrainE,DvalE,DtestE,YDE)(\mathcal{D}_{\textrm{train}}^{E},\mathcal{D}_{\textrm{val}}^{E},\mathcal{D}_{\textrm{test}}^{E},\mathcal{Y}^{E}_{\mathcal{D}}), where YDE\mathcal{Y}^{E}_{\mathcal{D}} is a sampled subset of labels in YD\mathcal{Y}_{\mathcal{D}} and Dtrain∣val∣testE\mathcal{D}^{E}_{\textrm{train}|\textrm{val}|\textrm{test}} are disjoint sets of examples from D\mathcal{D} with labels in YDE\mathcal{Y}^{E}_{\mathcal{D}}.In the few-shot literature, DtrainE\mathcal{D}_{\textrm{train}}^{E} and DtestE\mathcal{D}_{\textrm{test}}^{E} are also called the support and query sets, and ∣YDE∣|\mathcal{Y}^{E}_{\mathcal{D}}| the way. For each episode, the model’s objective is to correctly predict labels for examples DtestE\mathcal{D}_{\textrm{test}}^{E}. To accomplish this, models make use of labeled examples in DtrainE\mathcal{D}_{\textrm{train}}^{E}, which is typically configured such that each label ii in YDE\mathcal{Y}^{E}_{\mathcal{D}} has KiEK_{i}^{E} provided examples; KiEK_{i}^{E} is known as the shot, and the setting when a class has no examples in DtrainE\mathcal{D}_{\textrm{train}}^{E} (i.e., KiE=0K_{i}^{E}=0) is called zero-shot.

Few-shot evaluation in NLP

Research in few-shot NLP has proceeded in several parallel threads, each focused on a different type of transfer ability . Each thread has separate evaluation practices, and the vast majority of few-shot NLP research has limited evaluation to a single transfer type (see Table 1). Here, we describe these types of transfer and their evaluation practices.

Following the CV literature , one thread of few-shot NLP focuses on class transfer, the problem of generalizing from a supervised set of classes at meta-train time to a different set of classes from the same dataset at meta-test time. Evaluation typically involves splitting classes YD\mathcal{Y}_{\mathcal{D}} into YtrainD\mathcal{Y}^{\mathcal{D}}_{\textrm{train}}, YvalD\mathcal{Y}^{\mathcal{D}}_{\textrm{val}} and YtestD\mathcal{Y}^{\mathcal{D}}_{\textrm{test}} disjoint subsets. Class transfer has been studied on many text classification tasks , including relation classification , intent classification , inter alia. In contrast, domain transfer keeps the same classes between meta-training and meta-testing but changes the textual domain (e.g., generalizing from MNLI to science-focused SciTail ). Evaluation then requires identifying pairs of datasets with the same classes YD\mathcal{Y}_{\mathcal{D}}, where one dataset’s episodes are assigned to Etrain\mathcal{E}_{\textrm{train}} and the other’s to Etest\mathcal{E}_{\textrm{test}}. Domain transfer has also been studied on many tasks , including dialogue intent detection & slot tagging , sentiment classification , NLI , and machine translation .

Researchers have also begun to study task transfer, the problem of generalizing from a set of tasks at meta-train time to unseen tasks at meta-test time. Evaluation requires tasks (e.g., NLI) appearing in Etest\mathcal{E}_{\textrm{test}} not to appear in Etrain\mathcal{E}_{\textrm{train}} or Eval\mathcal{E}_{\textrm{val}}. Prior work has used GLUE tasks for meta-training before meta-testing on tasks such as entity typing , while other work instead used GLUE for meta-testing . Very recent work has studied task transfer over a large set of datasets . A limited amount of work evaluates both domain and task transfer . An important emerging line of work (not noted by Yin ) is pretraining transfer, the problem of whether pretrained language models can perform well at meta-test time without any meta-training. Evaluation in this setting requires Etrain,Eval=∅\mathcal{E}_{\textrm{train}},\mathcal{E}_{\textrm{val}}=\emptyset. Prior work has shown that pretrained language models are capable of surprising performance on many few-shot tasks, even without fine-tuning . More recent work, mainly focusing on text classification, has reported further gains with cloze-style formats , prompt engineering , or calibration . FLEX is designed to exercise all four of these transfer types from previous work.

Few-shot evaluation outside NLP

The few-shot learning literature has largely focused on image classification, with the introduction of increasingly complex meta-learning algorithms [e.g., 68, 61, 54, 39, 23]. However, more recent work has shown that simple fine-tuning baselines are in fact competitive, and attribute this delayed discovery to problematic evaluation methodology . FLEX adopts recommended methodology , and we introduce an analogous baseline (UniFew) to provide a strong measurement foundation for few-shot NLP.

FLEX Principles for Few-Shot NLP Evaluation

We now enumerate key desiderata for a few-shot NLP benchmark capable of solving the urgent problems with few-shot NLP evaluation, including separate evaluations for each transfer type and failure to incorporate best measurement practices from other domains (§2).

Diversity of transfer types To make NLP models broadly useful, few-shot NLP techniques must be capable of class, domain, and task transfer. Moreover, techniques should make use of the relevant supervision provided during meta-training to increase performance compared to the pretraining transfer setting. The benchmark should measure all four transfer settings to ensure that the community develops techniques that improve on strong pretraining transfer baselines, and enable comparison across these currently separate threads of research.

Variable number of shots and classes To better simulate a variety of real-world scenarios, the benchmark should include a variety of training set sizes and numbers of classes . Testing robustness to these factors is crucial; few-shot techniques are often sensitive to changes in these factors , yet all prior few-shot NLP evaluations we are aware of used a fixed number of training shots and classes, known in advance during meta-training.

Unbalanced training sets The benchmark should also include unbalanced training sets with different training shots per class, another realistic setting adopted by CV benchmarks . Class imbalance has also been observed to degrade performance , yet prior few-shot NLP evaluations do not include this setting either.

Textual labels While numerical label values are often used in classification tasks, descriptive textual labels are also present for many tasks. Making these textual labels available for use by few-shot techniques enables the development of techniques that can leverage the class name, like in-context learning , template generation , and meta-learning . Textual labels are crucial in particular for zero-shot evaluation.

Zero-shot evaluation We believe zero-shot evaluation is integral to the goals of few-shot evaluation. Similar to the motivation for measuring pretraining transfer, zero-shot evaluation is an important use case and also provides a strong baseline for some tasks. In the absence of training examples, textual class labels or richer task descriptions must be provided. Some recent few-shot NLP work [e.g., 10, 24] evaluated with zero training shots, but most [e.g., 5, 3, 75] did not.

No extra meta-testing data We believe the benchmark should not provide validation data (DvalE=∅,∀E∈Etest\mathcal{D}_{\textrm{val}}^{E}=\emptyset,\forall E\in\mathcal{E}_{\textrm{test}}) or unlabeled data for meta-testing tasks, since few-shot learning seeks to enable high performance in environments where collecting additional data is costly.Unlabeled data collection can be costly too, e.g. due to manual filtering . Variation in these dimensions in prior NLP work makes comparison of results extremely difficult because it is often under-reported and gives unfair advantage to approaches that leverage such data . For example, per-episode hyperparameter tuning on extra data has been shown to greatly inflate evaluation scores . A few researchers follow our suggested approach, but others have used many different settings, from validation sets of various sizes to no validation set but a large set of unlabeled examples .

Principled sample size design Promising few-shot techniques can incur significant computational cost per episode, e.g., due to fine-tuning model parameters , searching for prompts , inter alia. To alleviate these costs, related works often evaluate with a limited number of episodes, which precludes statistically accurate or precise performance estimates. We believe the benchmark’s test sample size should be optimized to enable proper performance evaluation for such techniques, while ensuring the computational burden is inclusive toward researchers without large compute resources.

Proper reporting of CIs, SDs, and individual results The benchmark should report confidence intervals (CIs) of performance estimates and follow recent guidelines to report standard deviations (SDs) for understanding variability. Moreover, we newly advocate for controlling for the same sampled few-shot episodes across all methods and reporting individual episode results, so that researchers can run higher-powered paired statistical tests when comparing results , crucial when the benchmark has been optimized for low evaluation budgets.

FLEX Benchmark

The FLEX benchmark is a unifying, rigorous evaluation suite for few-shot learning in NLP, which implements the desiderata outlined in the previous section. In this section, we describe detailed design decisions and our accompanying few-shot NLP toolkit (§4.4), which we are releasing to facilitate easily adding NLP datasets and advanced sampling options to future benchmarks. We also describe the FLEX leaderboard (§4.5).

Following GLUE and other prior work , we focus on tasks formatted as classification. Despite recent advances, NLP state-of-the-art models remain significantly worse than human performance on many text classification tasks, particularly in the few-shot setting. Automatic scoring of classification tasks is also more reliable than text generation tasks.

We selected datasets across three recent few-shot NLP evaluation suites, which separately studied class transfer , domain and task transfer , and pretraining transfer . Our benchmark includes a broad mix of tasks (NLI, question classification, entity typing, relation classification, and sentiment analysis) and formats (document, sentence, sentence pair). More complete dataset and license details are available in the following subsection and Appendix A.

2 Meta-Evaluation Protocols

As discussed earlier, FLEX evaluates four different types of transfer: Class, Domain, Task, and Pretraining Transfer. To support all types, we report results to the FLEX benchmark both without meta-training (pretraining-only) and with meta-training. This reporting scheme evaluates the performance of the basic pretrained model and the benefit (or lack thereof) of meta-training. A similar reporting scheme was proposed by Triantafillou et al. for CV.

Pretraining-Only In this setting, the pretrained model is directly meta-tested on our benchmark without any additional training. This is the Pretraining Transfer setting, and it is the most difficult, but given the recent success of pretrained models in NLP for few-shot learning , we believe that comparison to models without any meta-training is important for NLP tasks.

Meta-Trained In this setting, the model is meta-trained then meta-tested on our benchmark. We carefully selected and split datasets across meta-train/validation/test in order to enable testing of Class, Domain, and Task transfer with a single meta-training phase (to reduce computational burden). Datasets involved in each transfer setting (detailed split information in Table 4 in Appendix A):

Class Transfer: FewRel , HuffPost , Amazon , 20News , and Reuters take part in meta-training and meta-testing but with different classes.

Domain Transfer: MR , CR , SNLI , and SciTail are only in the meta-testing phase, but the corresponding sentiment and NLI datasets exist in the meta-training phase (MNLI , QNLI , and SST-2 ).

Task Transfer: Subj , TREC , and CoNLL are also for meta-testing only, and they represent tasks that the model does not encounter during meta-training.

Instead of per-episode hyperparameter tuning, we provide meta-validation episodes Eval\mathcal{E}_{\textrm{val}} for learning (during meta-training) global hyperparameters that work across all episodes. Specifically, the meta-validation dataset splits (see Table 4) consist of CoLa for task transfer, WNLI for domain transfer, and the validation splits used by Bao et al. for all class transfer datasets. Following , we also include meta-training datasets MRPC , RTE , and QQP .

3 Episode Sampling

We describe how our benchmark samples meta-testing episodes Etest\mathcal{E}_{\textrm{test}}. For meta-training, we allow users to sample from Etrain\mathcal{E}_{\textrm{train}}, Eval\mathcal{E}_{\textrm{val}} in any way, or directly use the underlying dataset splits.

4 Extensible Toolkit for Benchmark Creation and Model Training & Evaluation

Alongside the FLEX benchmark, we release an extensible, highly-configurable Python toolkit, which we used to generate the benchmark, and train and evaluate our models. Unlike existing meta-learning frameworks (e.g., Torchmeta , learn2learn ), our framework makes available a wide range of community-contributed NLP datasets and utilities via HuggingFace Datasets .Apache License 2.0. Full license details for all software dependencies available in Appendix F. Our code also provides advanced sampling utilities (e.g., for class imbalance), ensures reproducibility by checksumming generated episodes, and reports all recommended statistics.

5 Public Leaderboard

We provide public leaderboards for each of the meta-evaluation protocols: Pretraining-Onlyhttps://leaderboard.allenai.org/flex/ and Meta-Trained.https://leaderboard.allenai.org/flex_meta/ Submissions take the form of a text label predictions file, which is produced by our toolkit. Results are reported with confidence intervals, standard deviations, and individual predictions on request. See Appendix G for a screenshot of the results interface.

Sample Size Design: Balancing Statistical Measurement & Compute Cost

We demonstrate a principled approach to determining the optimal sample size configuration in our few-shot benchmark. A proper benchmark should produce performance estimates that are accurate, close to the true value, and precise, low variance. A large (test) sample size can achieve this, yet must be considered alongside computational cost so that a broad community of researchers with differing amounts of compute resources can participate. This decision is further complicated in the few-shot setting, where sample size refers to both the number of test episodes ∣Etest∣|\mathcal{E}_{\textrm{test}}| and the number of test examples ∣DtestE∣|\mathcal{D}_{\textrm{test}}^{E}| per episode E∈EtestE\in\mathcal{E}_{\textrm{test}}. For practicality, we consider ∣Dtest∣‾\overline{|\mathcal{D}_{\textrm{test}}|}, the mean ∣DtestE∣|\mathcal{D}_{\textrm{test}}^{E}| across all episodes, rather than every ∣DtestE∣|\mathcal{D}_{\textrm{test}}^{E}|. It remains unknown how one should best distribute test examples between ∣Etest∣|\mathcal{E}_{\textrm{test}}| and ∣Dtest∣‾\overline{|\mathcal{D}_{\textrm{test}}|}: More episodes each with fewer examples, or fewer episodes each with many examples? Prior work has been inconsistent in this regard. For example, Gao et al. used ∣Etest∣=5|\mathcal{E}_{\textrm{test}}|=5 and large ∣Dtest∣‾\overline{|\mathcal{D}_{\textrm{test}}|}, while Bao et al. used ∣Etest∣=1000|\mathcal{E}_{\textrm{test}}|=1000 and much smaller ∣Dtest∣‾\overline{|\mathcal{D}_{\textrm{test}}|}.

Inspired by simulation techniques for informing statistically-powered experimental design , we study how different configurations of ∣Etest∣|\mathcal{E}_{\textrm{test}}| and ∣Dtest∣‾\overline{|\mathcal{D}_{\textrm{test}}|} across different compute budgets CC impact the accuracy and precision of our estimated CIs, specifically with respect to coverage probability and width. First, we estimate per-episode and per-test-example costs of our few-shot model (§6) to obtain valid (C,∣Etest∣,∣Dtest∣‾)(C,|\mathcal{E}_{\textrm{test}}|,\overline{|\mathcal{D}_{\textrm{test}}|}) configurations s.t. the full benchmark completes within given CC (GPU-hours).Costs estimated using a Quadro RTX-8000 GPU with 48Gb memory. For few-shot settings, model was trained with 300 steps. Per-episode and per-test-example costs were approx. 95–98 and 0.7–0.11 GPU-sec, respectively. Using a model with high per-episode cost for this analysis allows us to define a lower-bound sample size requirement; we can always test inexpensive or zero-shot models on more ∣Etest∣|\mathcal{E}_{\textrm{test}}| or Dtest‾\overline{\mathcal{D}_{\textrm{test}}} within budget. Then, for each (C,∣Etest∣,∣Dtest∣‾)(C,|\mathcal{E}_{\textrm{test}}|,\overline{|\mathcal{D}_{\textrm{test}}|}), we perform 1000 simulation runs, in which each run samples predictions under a true model accuracy μacc\mu_{acc} and computes a single 95% CI, its width, and whether it correctly covers μacc\mu_{acc}. Averaging over simulation runs gives us estimates for the coverage probability and width of our benchmark’s CI for a single (C,∣Etest∣,∣Dtest∣‾)(C,|\mathcal{E}_{\textrm{test}}|,\overline{|\mathcal{D}_{\textrm{test}}|}). We repeat this whole procedure for different μacc∈{0.3,0.35,…,0.95}\mu_{acc}\in\{0.3,0.35,\dots,0.95\} to cover a wide range of possible model performances observed across many datasets (see Table 3).

Figure 1 shows CI coverage probability and width for many (C,∣Etest∣,∣Dtest∣‾)(C,|\mathcal{E}_{\textrm{test}}|,\overline{|\mathcal{D}_{\textrm{test}}|}) configurations. First, we find in Figure 1(a) that sufficiently-many test episodes (i.e., ∣Etest∣>60|\mathcal{E}_{\textrm{test}}|>60) is needed to guarantee coverage probability of our CIs is within one percentage point of the target 95%, a trend that holds regardless of compute budget. Small ∣Etest∣|\mathcal{E}_{\textrm{test}}| also corresponds to large CI widths across all considered budgets in Figure 1(b). This suggests that the choices of ∣Etest∣=1,5,10|\mathcal{E}_{\textrm{test}}|=1,5,10 in prior work can mean inaccurate and wide CIs, while choices of ∣Etest∣=1000|\mathcal{E}_{\textrm{test}}|=1000 can be prohibitively costly for methods with high training cost.

Next, Figure 1(b) reveals (i) diminishing returns in CI width (decrease in yy-axis) as compute increases, and (ii) existence of an optimal balance between ∣Etest∣|\mathcal{E}_{\textrm{test}}| and ∣Dtest∣‾\overline{|\mathcal{D}_{\textrm{test}}|} for each budget. Restricting our consideration to budgets with optima satisfying sufficient coverage probability (∣Etest∣>60|\mathcal{E}_{\textrm{test}}|>60), the minimum viable budget is 36 GPU-hours. Then, assessing the marginal benefit of each 12 GPU-hour budget increase in terms of marginal reduction in CI width between optima, we arrive at our FLEX configuration of ∣Etest∣=90|\mathcal{E}_{\textrm{test}}|=90 and ∣Dtest∣‾≈470\overline{|\mathcal{D}_{\textrm{test}}|}\approx 470 under a budget of C=48C=48 GPU-hours.Consider budget increases 36→4836\to 48, 48→6048\to 60, 60→7260\to 72 and 72→8072\to 80. The first reduces CI width by 13%. Further increases reduce CI width by an additional 9%, 7%, and 5%, respectively. We choose C=48C=48 based on these diminishing returns. Further details are in Appendix B.

UniFew: A Few-Shot Learning Model by Unifying Pre-training and Downstream Task Formats

Despite their encouraging results, existing works on few-shot learning in NLP are based on either customized and often complex meta-learning algorithms , heavy manual/automated engineering of textual descriptions or prompts , ordering of training examples , extensive hyperparameter tuning on held-out sets , or custom learning algorithms . We present UniFew, a strong few-shot learning model across all transfer settings and datasets tested, that eschews the need for incorporating the above-mentioned complexities and challenges.

UniFew is a prompt-based model , a class of models that tailor the input/output format of their data to match the format used during pretraining. While this technique allows them to perform a task without the need for additional classification layers, prompt-based models are typically sensitive to the choice of the prompts, which can require extensive search, trial-and-error, and even additional models to get right . To avoid this issue while still leveraging the strong capabilities of pretrained models, UniFew (1) converts examples into multiple-choice question-answer (QA) format, and (2) uses UnifiedQA , a T5 model further pretrained on a large collection of QA pairs.UnifiedQA and T5 both use Apache License 2.0. We use publicly-released large-size model weights.,None of the supervised datasets in the pretraining of UnifiedQA or T5 are in FLEX.

Compared to other prompt-based models, UniFew has two main strengths. First, the prompt design problem is much simpler because UnifiedQA questions had well-defined formats. For example, we only need four general prompt templates which cover all 20 datasets in the FLEX benchmark, while prior works have needed specialized prompts for each dataset. Second, UnifiedQA’s multiple-choice format ensures the model outputs a valid class label, without the need for learned or manually-defined mappings or verbalizers required for other prompt-based methods .In rare cases, especially for zero-shot, UnifiedQA may generate an invalid answer (e.g., “Yes, Yes, No” instead of “Yes”). We use simple heuristics to normalize the answer in such cases. In concurrent work, Zhong et al. also show the benefit of performing meta-tuning on a variety of datasets; while their task setup as Q/A is similar to UniFew, they focus exclusively on binary zero-shot classification tasks and, unlike UniFew, do not handle multi-class or few-shot problems.

We experiment with UniFew both without and with meta-training on the FLEX benchmark’s meta-training data, following the FLEX protocol (§4.2). We call the meta-trained variant UniFewmeta. We use simple prompts in the format of question followed by choices followed by the answer (according to the UnifiedQA original format). The exact prompts used are provided in Appendix C.

Training details For meta-training and meta-validation of UniFew, we sampled Etrain\mathcal{E}_{\textrm{train}} and Eval\mathcal{E}_{\textrm{val}} with 5-class, 5-training-shot sampling with the same number of shots per class.Users of FLEX can specify the sampling configuration of Etrain\mathcal{E}_{\textrm{train}} and Eval\mathcal{E}_{\textrm{val}} as desired. We trained the model for total number of 30K steps, using a linear learning rate scheduler with peak rate of 3e−53e{-}5, 200 warmup steps, and batch size of 4; we selected the best checkpoint based on Eval\mathcal{E}_{\textrm{val}} performance. At meta-test time, for each episode, we trained the model on the episode’s training examples (if they exist) and predicted the outputs on test examples. For training at meta-test time, we used constant learning rate of 3e−53e{-}5 and batch size of 4, and trained the model for 400 steps.For comparison with we trained the model for 600 steps. We used NVidia RTX8000 GPUs, which take about 7 GPU-hours for meta-training and 48 GPU-hours for meta-testing. For meta-testing we split the episodes among 8 GPUs to speed up evaluations.

Experiments

To demonstrate the efficacy of UniFew, we evaluate it against state-of-the-art approaches for few-shot and meta-learning in NLP: LM-BFF , a language model prompt-based fine-tuning method, as well as Distributional Signatures (DS) and H-SMLMT , two state-of-the-art meta-learning techniques. Refer to Appendix D for details on these methods.

We compare to these methods using the datasets in the FLEX benchmark to establish the quality of our model. Since we constructed our benchmark from disjoint subsets of datasets evaluated in each of these prior works (§4.1), we compare each method with its corresponding subset of datasets. Each of these prior works evaluates their methods using different experimental setups (classes, number of episodes, shots) than our benchmark and was not designed to handle FLEX’s challenging episode characteristics like class imbalance. To enable fair comparison, we test UniFew on the exact data splits released by the authors when available (H-SMLMT and LM-BFF). For DS, we sample (balanced) episodes using our framework after matching their test settings (number of shots and classes, class splits, etc.) and reproduce their reported results to within 1% absolute difference using their model code; we use these episodes for our experiments. The results in Table 2 show that UniFewmeta outperforms both H-SMLMT and DS meta-learning approaches by relatively large margins, while achieving competitive results compared with LM-BFF. Note that UniFew’s strong results are without meta-learning approaches, extensive prompt-engineering, or per-episode hyperparameter search.

Evaluating UniFew on the FLEX benchmark

Having established UniFew as a strong model comparable to recent, state-of-the art techniques, we present its results on the final version of our benchmark (with class imbalance, etc.). From Table 3, we observe three findings. First, pretraining is an effective technique for infusing an NLP model with the ability to perform few-shot generalization even without any meta-training, as UniFew is able to score Δfew=+12.8\Delta_{\text{few}}=+12.8 higher when provided few rather than zero examples. Second, by comparing UniFewmeta and UniFew, we see that meta-training has a substantial impact on zero-shot performance (Δmeta=+14.5\Delta_{\text{meta}}=+14.5), but its benefit, while still substantial, is less in the few-shot setting (Δmeta=+8.6\Delta_{\text{meta}}=+8.6). Third, while meta-training adds roughly the same benefit to zero and few-shot performance for both domain and task transfer settings, meta-training disproportionately benefits zero-shot class transfer (Δmeta=+16.2\Delta_{\text{meta}}=+16.2) over few-shot class transfer (Δmeta=+4.3\Delta_{\text{meta}}=+4.3). Such observations are made possible through unified evaluation and comparison across different transfer types. The full FLEX benchmark results broken down by individual datasets are in Appendix E.

Limitations and Future Work

While the initial FLEX benchmark is focused on classification tasks, we aim to use our benchmark creation toolkit (§4.4) to incorporate additional task formats like span selection or text generation. Furthermore, the benchmark currently only supports English language tasks; to study language transfer, we aim to incorporate new datasets using our toolkit. Adding diverse datasets has its own challenges; while we’ve selected datasets for our benchmark based on prior work adoption and have attempted to verify their licensing for research use, we were unable to find license details for some datasets (Appendix A). We believe it is crucial to continually evolve the suite of datasets to remain challenging for the best models and to tackle real-world challenges .

In addition, Sample Size Design (§5) simulations currently rely on our own available training estimates. We plan to gather a more representative sample from community leaderboard submissions.

Our public leaderboard could benefit from extended support for detailed comparisons between submissions based on properties of techniques. For example, approaches may vary in terms of model characteristics (e.g., number of parameters), data and supervision used during pretraining, amount of compute, etc. We encourage reporting all these factors to enable the community to analyze and make progress on important sub-spaces in the overall few-shot technique design space.

Finally, we believe the benefits of improving few-shot NLP techniques outweigh potential risks, but we acknowledge potential harms associated with language models . Few-shot models learn a task from a few examples but rely heavily on knowledge encoded in the pretrained model. Thus, few-shot models are more likely to inherit the biases of the pretrained models, compared to more fully supervised models; as the community focuses more on few-shot learning, it is more important than ever for future pretrained models to be careful about biases in the underlying pretraining corpora.

Conclusion

In this work, we unify and bring rigor to few-shot NLP evaluation. We formulate the FLEX Principles, a set of requirements and best practices that enables unified, rigorous, valid, and cost-sensitive measurement. We advance the principles with new Sample Size Design methodology for optimizing statistical accuracy and precision while keeping costs low. The FLEX benchmark is our instantiation of the FLEX Principles; it employs Sample Size Design and includes four few-shot transfer settings, zero-shot evaluation, and a public leaderboard with diverse NLP tasks. We present UniFew, a prompt-based model that aligns pretraining and downstream task formats, achieving results competitive with recent few-shot methods despite using trivial prompt engineering. Finally, we release an extensible, open-source toolkit (used to train UniFew and generate the FLEX benchmark) to support future benchmark creation and few-shot NLP model training.

Acknowledgments and Disclosure of Funding

We would like to thank Chandra Bhagavatula, Matt Gardner, Matt Peters, Doug Downey, Dan Weld, and the four anonymous reviewers for helpful comments, suggestions and feedback. We would also like to acknowledge the large community effort involved in the creation of the datasets and open-source tools we utilize.

References

Appendix A Datasets

Table 4 summarizes the tasks and datasets used for meta-training and meta-testing. To enable automated benchmark construction and maximize access, we restrict datasets to those that are freely available for automated download.We exclude RCV1 (used by ) and MPQA (used by ), since they require agreeing to license terms through web forms at download time. We include the GLUE tasks used by Bansal et al. for meta-trainingWe follow Bansal et al. and use the matched+mismatched version of MNLI and exclude WNLI and STS-B from meta-training due to the small training size and regression task format, respectively and thus exclude GLUE tasks used by Gao et al. from meta-testing. Although Bansal et al. additionally use SNLI for meta-training, we reserve it for meta-testing for comparison to Gao et al. and because NLI is already represented in the meta-training datasets.

Textual labels and licenses for datasets

We made CoNLL labels more descriptive from their original PER,ORG,LOC,MISC. For TREC, we used the more readable labels from the manual template in . For readability, Amazon labels are shown without underscores and Amazon and HuffPost capitalization has been removed. License information is shown in parentheses.

MR (license unavailablehttps://www.cs.cornell.edu/people/pabo/movie-review-data/rt-polaritydata.README.1.0.txt): Test: negative, positive

CR (license unavailablehttps://www.cs.uic.edu/~liub/FBS/CustomerReviewData.zip): Test: negative, positive

Subj (license unavailablehttps://www.cs.cornell.edu/people/pabo/movie-review-data/subjdata.README.1.0.txt): Test: objective, subjective

TREC (license unavailablehttps://cogcomp.seas.upenn.edu/Data/QA/QC/): Test: description, entity, expression, human, location, number

FewRel (MIT Licensehttps://huggingface.co/datasets/few_rel): Train: applies to jurisdiction, architect, child, competition class, constellation, contains administrative territorial entity, country, country of citizenship, country of origin, crosses, father, field of work, followed by, follows, genre, has part, head of government, headquarters location, heritage designation, instance of, instrument, league, licensed to broadcast to, located in or next to body of water, located in the administrative territorial entity, located on terrain feature, location, location of formation, manufacturer, member of, member of political party, military branch, military rank, mother, mountain range, mouth of the watercourse, movement, notable work, occupant, occupation, operating system, operator, owned by, part of, participant, participant of, participating team, place served by transport hub, position held, position played on team / speciality, record label, religion, residence, said to be the same as, sibling, sport, sports season of league or competition, spouse, subsidiary, successful candidate, taxon rank, tributary, voice type, winner, work location; Val: developer, director, original network, performer, publisher; Test: after a work by, characters, composer, distributor, language of work or name, main subject, nominated for, original language of film or TV show, platform, screenwriter

HuffPost (CC0: Public Domainhttps://www.kaggle.com/rmisra/news-category-dataset): Train: arts, arts & culture, black voices, comedy, culture & arts, fifty, food & drink, good news, green, impact, latino voices, media, money, parenting, religion, sports, style, the worldpost, travel, women; Val: crime, queer voices, science, weird news, worldpost; Test: business, college, divorce, education, entertainment, environment, healthy living, home & living, parents, politics, style & beauty, taste, tech, weddings, wellness, world new

CoNLL (license for research by Reutershttps://www.clips.uantwerpen.be/conll2003/ner/): Test: location, organization, other, person

SNLI (Creative Commons Attribution-ShareAlike 4.0 International Licensehttps://huggingface.co/datasets/snli): Test: contradiction, entailment, neutral

SciTail (license unavailablehttps://allenai.org/data/scitail): Test: entailment, neutral

Amazon (license unavailablehttp://jmcauley.ucsd.edu/data/amazon/): Train: automotive, baby, beauty, cell phones and accessories, grocery and gourmet food, health and personal care, home and kitchen, patio lawn and garden, pet supplies, sports and outdoor; Val: apps for android, cds and vinyl, digital music, toys and games, video games; Test: amazon instant video, books, clothing shoes and jewelry, electronics, kindle store, movies and tv, musical instruments, office products, tools and home improvement

20News (license unavailablehttps://huggingface.co/datasets/newsgroup): Train: rec.autos, rec.motorcycles, rec.sport.baseball, rec.sport.hockey, sci.crypt, sci.electronics, sci.med, sci.space; Val: comp.graphics, comp.os.ms-windows.misc, comp.sys.ibm.pc.hardware, comp.sys.mac.hardware, comp.windows.x; Test: alt.atheism, misc.forsale, soc.religion.christian, talk.politics.guns, talk.politics.mideast, talk.politics.misc, talk.religion.misc

Reuters (license for research by Reutershttps://kdd.ics.uci.edu/databases/reuters21578/README.txt): Train: acq, alum, bop, cocoa, coffee, copper, cotton, cpi, crude, earn, gnp, gold, grain, interest, ipi; Val: iron-steel, jobs, livestock, money-fx, money-supply; Test: nat-gas, orange, reserves, retail, rubber, ship, sugar, tin, trade, veg-oil, wpi

CoLa (released under fair usehttps://nyu-mll.github.io/CoLA/): Val: acceptable, unacceptable

MNLI (multiple licenseshttps://www.aclweb.org/anthology/N18-1101.pdf): Train/Val: contradiction, entailment, neutral

MRPC (license unavailablehttps://www.microsoft.com/en-us/download/details.aspx?id=52398): Train/Val: equivalent, not_equivalent

QNLI (CC BY-SA 4.0https://rajpurkar.github.io/SQuAD-explorer/): Train/Val: entailment, not_entailment

QQP (non-commercial usehttps://www.kaggle.com/quora/question-pairs-dataset): Train/Val: duplicate, not_duplicate

RTE (license unavailablehttps://gluebenchmark.com/): Train/Val: entailment, not_entailment

SST-2 (license unavailablehttps://nlp.stanford.edu/sentiment/): Train/Val: negative, positive

WNLI (CC BY 4.0https://cs.nyu.edu/~davise/papers/WinogradSchemas/WS.html): Val: entailment, not_entailment

Appendix B Sample Size Simulations

We describe how we performed the simulations described in §5.

The cost of meta-testing on FLEX for a given dataset is the sum of the cost of both few-shot and zero-shot evaluations:

where Cfew|zeroEC_{\text{few|zero}}^{E} is (average) time spent per-episode during model setup and training, Cfew|zeroIC_{\text{few|zero}}^{I} is (average) time spent per-episode per-test-instance on evaluation. We estimate these quantities on a single Titan RTX-8000 GPU with 48Gb memory by conducting meta-testing runs with the UniFew model (300 steps in few-shot setting) across all datasets in FLEX with arbitrary choices for ∣DtestE∣|\mathcal{D}_{\textrm{test}}^{E}|. These tended to be around 95–98 sec for CfewEC_{\text{few}}^{E}, 1–3 sec for CzeroEC_{\text{zero}}^{E}, and 0.7–0.11 sec for Cfew|zeroIC_{\text{few|zero}}^{I}. From this, we derived possible (C,∣Etest∣,∣Dtest∣‾)(C,|\mathcal{E}_{\textrm{test}}|,\overline{|\mathcal{D}_{\textrm{test}}|}) configurations by solving for ∣Dtest∣‾\overline{|\mathcal{D}_{\textrm{test}}|} over grids of C=24,36,…,84C=24,36,\dots,84 and ∣Etest∣=5,15,30,45,…,150|\mathcal{E}_{\textrm{test}}|=5,15,30,45,\dots,150.

Simulating confidence intervals

We describe a single simulation run by a given (C,∣Etest∣,∣Dtest∣‾)(C,|\mathcal{E}_{\textrm{test}}|,\overline{|\mathcal{D}_{\textrm{test}}|}). First, we need to generate ∣Dtest∣‾\overline{|\mathcal{D}_{\textrm{test}}|} model predictions for every episode E∈EtestE\in\mathcal{E}_{\textrm{test}}. To do this, we assume each episode has a latent episode-specific model accuracy μacc(1),…,μacc(∣Etest∣)\mu_{acc}^{(1)},\dots,\mu_{acc}^{(|\mathcal{E}_{\textrm{test}}|)}, where each μacc(⋅)\mu_{acc}^{(\cdot)} is drawn from a Normal distribution with mean μacc\mu_{acc} and variance σacc2\sigma^{2}_{acc}. Here, μacc\mu_{acc} represents the unknown overall model accuracy that is our target of estimation, and σacc2\sigma^{2}_{acc} represents inherent variability in task difficulty across episodes (e.g., due to different number of classes or imbalance). In our simulations, we set σacc=0.05\sigma_{acc}=0.05. For each episode EE, we generate prediction outcomes (i.e. correct or incorrect) from a Bernoulli with success probability μaccE\mu_{acc}^{E}. This allows us to compute episode-specific accuracy estimates μ^acc(1),…,μ^acc(∣Etest∣)\hat{\mu}_{acc}^{(1)},\dots,\hat{\mu}_{acc}^{(|\mathcal{E}_{\textrm{test}}|)} and finally compute the mean, standard deviation, and (bootstrap) CI across these episodes. In doing so, a single simulation run represents a possible submission outcome to FLEX for a given model, and we can obtain the resulting CI’s width and verify whether it contains the true model accuracy μacc\mu_{acc}.

Appendix C Prompts

We use the following prompts for FLEX benchmark tasks based on the input type:

The format of question, followed by the document followed by answer choices, as well as the use of the special delimiter of \\n is according to UnifiedQA’s original pretraining. We follow ’s format of NLI for sentence pair tasks and T5 for relation classification.

Appendix D Baseline Models

This section briefly describes the baselines we use for comparison.

LM-BFF is a language model prompt-based fine-tuning method with extensive automated and manual approaches for prompt generation. It also uses a strategy for dynamically and selectively incorporating demonstrations into each context which is an extension to GPT-3’s in-context learning technique .

Distributional Signatures (DS) A meta-learning method designed for class transfer. DS uses lexical “distributional signatures,” characteristics of the underlying word distributions to transfer attention patterns across tasks within a meta-learning framework.

SMLMT A self-supervised approach for domain and task transfer. SMLMT creates the target task distribution from a large set of unlabeled sentences used within a meta-learning framework for optimal transfer. We compare with the strongest model variant in this paper, Hybrid-SMLMT which is trained on both self-supervised and supervised tasks.