Adapting Language Models for Zero-shot Learning by Meta-tuning on Dataset and Prompt Collections
Ruiqi Zhong, Kristy Lee, Zheng Zhang, Dan Klein
Introduction
The goal of zero-shot classification (ZSC) is to classify textual inputs using label descriptions without any examples Yin et al. (2019). Large language models - whose only training objective is to predict the next word given the context - have acquired a surprising ability to perform ZSC Radford et al. (2019); Brown et al. (2020); Le Scao and Rush (2021). For example, to classify whether the sentence “This movie is amazing!" is positive, we can prompt the language model with the context “Review: This movie is amazing! Positive Review? ___ ", and check whether the next word is more likely to be “Yes" or “No" Zhao et al. (2021). To convert ZSC into a language modeling (LM) task that an LM model is likely to perform well, many recent works focus on finding better prompts Shin et al. (2020); Schick and Schütze (2020a, b); Gao et al. (2021).
However, the LM training objective is correlated but still misaligned with the target objective to answer prompts. Our work addresses this weakness by directly optimizing the zero-shot classification objective through fine-tuning (Section 4). This requires us to 1) unify different classification tasks into the same format, and 2) gather a collection of classification datasets and label descriptions (prompts) for training (Section 2). Since we fine-tune our model on a meta-dataset, we name our approach meta-tuning.
We focus on binary classification tasks and unify them into a “Yes"/“No" QA format Clark et al. (2019); McCann et al. (2018), where the input is provided as the context and the label information is provided in the question (Figure 1 (a)). Using this format, we gathered a diverse set of classification datasets from 43 different sources listed on Kaggle, SemEval, HuggingFace, and other papers. These tasks range from hate speech detection, question categorization, sentiment classification to stance classification, etc, and the genre ranges from textbooks, social media, to academic papers, etc. In total, these datasets contain 204 unique labels, and we manually annotated 441 label descriptions (Figure 2).
To evaluate ZSC, we need to define what counts as a task that the model has not seen during training time. While prior work considers different notions of “unseen" by disallowing the same label or the same dataset to appear during training, our work defines “unseen" more harshly by disallowing similar datasets. For example, we consider AG News topic classification dataset Zhang et al. (2015) and the topic classification dataset from Yin et al. (2019) to be similar, even though their sources and label spaces are different.
Meta-tuning improves ZSC over UnifiedQA for most labels (Figure 1 (c)). Moreover, larger models are better, and hence we forecast that meta-tuning would work for even larger models. We also find that the performance can be slightly improved by training on datasets similar to the test dataset, ensembling different label descriptions, or initializing with a QA model (Section 5.1). All of our findings reliably hold under different robustness checks (Section 5.2), and our approach outperforms the previous SOTA Yin et al. (2019) using the same pre-training method (Section 5.3).
Our results suggest two promising future directions (Section 6). First, large language models’ (e.g. GPT-3) potential for zero-shot learning, as currently measured by context-prompting, might have been broadly underestimated; meta-tuning might significantly improve their performance. Second, community-wide efforts on aggregating and unifying datasets can scale up training and evaluation for zero-shot learning models. On the flip side, however, the meta-tuning approach might incentivize providers of LM inference APIs to collect prompts from users, hence potentially leading to security, privacy, and fairness concerns at a greater scale (Section A).
To summarize, we 1) curate a dataset of classification datasets with expert annotated label descriptions. 2) demonstrate a simple approach to train models to perform zero-shot learning, and 3) identify several factors that improve performance; in particular, larger pretrained models are better. Code and data available here: https://github.com/ruiqi-zhong/Meta-tuning.
Data
We gather a wide range of classification datasets and unify them into the “Yes"/“No" question answering format for binary classification. Then we group similar datasets together to determine what counts as unseen tasks during evaluation.
We collect classification datasets from Kagglehttps://www.kaggle.com, Huggingface Wolf et al. (2020), SemEvalhttps://semeval.github.io , and other papers. We looked through these sources and only considered English classification datasets. We also skipped the tasks that we felt were already better represented by other datasets in our collection. Then we manually examined a few examples in each remaining dataset to make sure it seemed plausibly clean.
The goals of these classification datasets include, but are not limited to sentiment classification (IMDB Reviews, Maas et al. (2011a)), topic classification (AG News, Zhang et al. (2015)), grammaticality judgement (CoLA, Warstadt et al. (2018)), paraphrase detection (QQPhttps://www.kaggle.com/c/quora-question-pairs), definition detection (SemEval 2020 Task 6, Spala et al. (2019)), stance classification (SemEval 2016 Task 6, Mohammad et al. (2016)), etc. The genre includes academic papers, reviews, tweets, posts, messages, articles, and textbooks. The comprehensive list of datasets is in Appendix B. Overall, we aim for a high diversity of tasks and genres by building upon what the broader research community has studied. Our approach is complementary to that of Weller et al. (2020), which asks turkers to generate tasks, and that of Mishra et al. (2021), which generates tasks by decomposing existing templates used to construct reading comprehension datasets. The concurrent work of Bragg et al. (2021) unifies the evaluation for few-shot learning; their zero-shot evaluation setup is the closest to ours, and they used templates and verbalizers Schick and Schütze (2020a) to specify the semantics of a task.
Some of our datasets are noisy and not peer reviewed, or contain tasks that are too complicated (e.g. Multi-NLI, Williams et al. (2018)) for ZSC. To make our evaluation more informative, we only include them for training but not testing. We make these decisions before running our experiments in Section 5 to prevent selection bias.
Unifying the dataset format
We convert each classification dataset into a “Yes"/“No" question answering format and provide label information in the question. For each label, we annotate 1-3 questions. If the label is null (for example, a text that does not express a particular emotion in an emotion classification dataset), we skip this label. Three of the authorsOne of them is a graduate student and the other two are undergrads; all of them study Computer Science and have taken an NLP class. manually annotated 441 questions for 204 unique labels, and each question is proofread by at least another author. See Figure 2 for a concrete example, and Figure 3 for some representative label descriptions.
Additionally, some datasets contain thousands of labels Chalkidis et al. (2019); Allaway and McKeown (2020). In this case, we use templates to automatically synthesize label descriptions and exclude them from evaluation.
Grouping similar datasets
Our goal is to test the models’ ability to generalize to tasks that are different enough from the training tasks. Therefore, at test time, we need to exclude not only the same dataset that appeared in the meta-tuning phase, but also ones that are similar.
This poses a challenge: whether two datasets perform the same task involves subjective opinion, and there is no universally agreed definition. On one extreme, most datasets can be counted as dis-similar tasks, since they have different label spaces and input distributions. On the other extreme, all datasets can be considered the same task, since they can all be unified into the question answering format.
To tackle this challenge, we create a set of tags, each describing a dataset property. The set of tags includes domain classification, article, emotion, social-media, etc, and the full set of them can be seen in Appendix C. Then we define the two datasets to be similar if they are associated with the same set of tags, and prohibit the model to learn from one and test on the other. For example, our work considers the topic classification datasets from Zhang et al. (2015) (AG News) and Yin et al. (2019) to be similar since they both classify topics for articles, even though their sources and label spaces are different. Some example dataset groups can be seen in Figure 4.
Nevertheless, our procedure is not bullet-proof and one can argue that our notion of unseen tasks, though harsher than prior works Yin et al. (2019); Pushp and Srivastava (2017), is still lenient. Therefore, as additional robustness checks, for each dataset we evaluate, we manually identify and list the most relevant dataset that is allowed during training in Appendix F . For example, the most relevant dataset to the IMDB review sentiment classification dataset is the emotion classification dataset from Yin et al. (2019), which classifies the input text into 9 emotions, such as “joy", “surprise", “guilt", etc. We consider the emotion classification dataset to be relevant, since sentiment classification often involves identifying emotions. However, one can also argue that they are different tasks: their input and label spaces are different, and sadness can be caused by a great tragedy, or a bad movie that wastes the users’ time. The comprehensive list of label descriptions grouped by dataset similarity is in Appendix D.
In total, we spend around 200 hours to collect this dataset. This time estimate includes skimming through the dataset repos and recent NLP papers, writing programs to download the datasets and unify their format, annotating label descriptions, performing quality controls, and documenting the collection process.
Metrics
To reliably aggregate performance across different datasets and present as much information as possible, we report a set of descriptive statistics and provide visualizations whenever we compare two models. We generally do not reduce a model’s performances on different datasets into one scalar quantity and compare this number only.
For each label description (question), we calculate the AUC-ROC score We do not evaluate F-score or accuracy, since they are very sensitive to the decision cutoff, and usually additional calibration is needed Zhao et al. (2021). by treating the “Yes" answer as the positive class. After calculating the AUC-ROC score for each label, we calculate the following set of descriptive statistics to compare two models. Suppose that model is hypothetically better than . Denoting as the change of AUC-ROC of a label description from to , we can summarize how is distributed across the set of label descriptions with the following statistics:
: the standard deviation of the change.
Visualizations
We use scatter plots to visualize and compare the performance of two models, where each dot represents a label description, its x-value represents the AUC-ROC score of the model , and its y-value represents that of . If most dots are above the identity line , the model is better than .
The descriptive statistics and the visualizations are explained in Figure 5.
Model
We format the inputs to the model in the same way as UnifiedQA Khashabi et al. (2020), which concatenates the context to the question and adds a “[SEP]" token in between. Then we feed the concatenated input into the T5 encoder and produce the answer score by normalizing the “Yes"/“No" probability of the first decoded token. Unless otherwise noted, we initialize our model with T5-Large (770 Million parameters). We sometimes compare to or initialize with the UnifiedQA model Khashabi et al. (2020), which is trained on a wide range of question answering datasets. For a fair comparison, we use the UnifiedQA model initialized with T5-Large as well. To meta-tune non-Seq2Seq pre-trained models, such as BERT Devlin et al. (2019) or RoBERTa Liu et al. (2019), we add an MLP layer on top of the pooled output/“[CLS]" token to classify between “Yes"/“No". We leave the improvement on model architectures Ye and Ren (2021); Li and Liang (2021); Lester et al. (2021) and training objectives Murty et al. (2021); Yin et al. (2020) for future work.
Meta-tuning
We create a training distribution that balances between datasets, label descriptions, and “Yes"/“No" answers. To create the next training datapoint for meta-tuning, we select a dataset from the training split uniformly at random (u.a.r.); then we select a label description (question) u.a.r. and with 50% probability select a textual input with the answer “Yes"/“No". To prevent over-fitting, we do not train on any combination of label description and textual input twice. Unless otherwise noted, we meta-tune the model for 5000 steps and use batch size 32. We did not tune any hyper-parameters or training configurations since they work well during our first attempt. To evaluate ZSC performance on each dataset, we leave out one group of similar datasets as the evaluation set and train on the rest. Altogether, the experiments take around 250 GPU hours on Quadro 8000.
Results
We investigate and validate the following hypotheses, sorted by importance in descending order.
Meta-tuned models outperform general question answering models in zero-shot classification.
Performance can be improved by training on similar datasets, initializing with a QA model, or ensembling label descriptions.
Early stopping is crucial to performance.
We compare a meta-tuned T5-Large model (770 M parameters)This model is initialized with T5, not UnifiedQA. with the same-sized UnifiedQA model Khashabi et al. (2020) out of the box. Relevant descriptive statistics can be seen in the first row of Table 1 and Figure 6 (a). Adapting the model for ZSC improves the average AUC-ROC by 3.3%.
Larger pre-trained models are better.
We compare T5-Base (220 Million parameters) against T5-Large (770 M). The statistics can be seen in the second row of Table 1 and Figure 6 (b). Increasing the model size from 220 M to 770M improves the average AUC-ROC by 6.3%.
Pre-training does the heavy lifting.
In Figure (c) and the third row of Table 1, we compare pre-trained and random initializations, where the latter cannot beat the random baseline (average AUC-ROC 0.503). Hence, meta-tuning alone is far from enabling the model to perform ZSC. An intuitive interpretation is that the model already “knows" how to perform ZSC after pre-training under the LM objective, and learns how to use this knowledge during meta-tuning.
Training on similar datasets improves performance.
Unlike before, we no longer avoid training on similar datasets from the same group. Instead, we perform straightforward leave-one-out cross-validation. The statistics can be seen in the fourth row of Table 1 and Figure 6 (d), and it improves the average AUC-ROC by 0.7%. The performance gain is not as significant as increasing the model size or adapting for ZSC. We conjecture that it is because we have not collected enough datasets; otherwise, there might be more similar datasets, hence improving ZSC performance.
Ensembling label descriptions improves performance.
Instead of asking the model a single question for each label and obtain the probability of the answer being “Yes", we can average the probability obtained by asking multiple questions with the same meaning. This approach is different from traditional ensembling, which typically needs to store/train multiple models to average across them. The fifth row of Table 1 and Figure 6 (e) verifies that ensembling descriptions improves performance slightly (0.7% AUC-ROC score).
Initializing with UnifiedQA improves performance.
Figure 6 (f) and the sixth row of Table 1 compare the UnifiedQA against against the T5 initialization. Initializing with UnifiedQA improves average AUC-ROC by 1.1%.
Early stopping is crucial to performance.
If we train the model for too long, the model might simply “memorize" that certain label descriptions correspond to certain training tasks, and the performance on unseen tasks may drop. To explore this possibility, we meta-tune our models for 100K steps, which is 20 times as long as our default setting and encourages the model to memorize the training tasks. We then evaluate them on the three benchmark zero-shot classification datasets by Yin et al. (2019) (which we describe in more details in the next section). We calculate the average AUC-ROC across all label descriptions for each of the 3 datasets, and plot them in Figure 7.
The performance decreases Kendall rank correlation coefficients are negative with for topic and situation classification as training continues. On the other hand, however, the performance drop of 3% in AUC-ROC is not fatal and the model’s performance is still much better than random guesses.
2 Robustness Checks
We examine a series of additional results to make sure our conclusions are robust. The observed improvements in Table 1 and Figure 6 might be caused by the improvement of a small number of labels that are annotated with more descriptions, or by the improvement on a dataset with more distinct labels. Appendix E.1 compares the performance by assigning equal weights to each label/datasets.
To provide additional supporting evidence for our forecast that larger models are better, Appendix E.2 compares a 60M-parameter model against a 220M-parameter model, and finds that the latter is much better. One concern, however, is that our models are initialized with T5 Raffel et al. (2019), which is trained on the open web and might have seen the datasets we gathered. Therefore, larger models might be better simply because they are better at memorization Sagawa et al. (2020). Appendix E.3 addresses this by showing that larger models are also better with BERT initialization Devlin et al. (2019), which is trained on Wikipedia and Book Corpus Zhu et al. (2015).
We also report the models’ performance on each dataset for readers’ reference in Appendix G.
3 Comparison with Yin et al. (2019)
This section shows that our approach has higher performance than the zero-shot classification system built by Yin et al. (2019). Their system ensembles several natural language inference models based on RoBERTA-Large (355M parameters, Liu et al. (2020)), and another model trained to categorize Wikipedia articles. It was evaluated on three classification datasets:
topic (10-way): classifies article domains, such as family & relationship, education, sports, etc. The metric is accuracy.
emotion (10-way): classifies emotion types, such as joy, anger, guilt, shame, etc. The metric is label-weighted F1.
situation (12-way): classifies disaster situations, e.g. regime change, crime & violence, and the resource they need, e.g. search & rescue. The metric is label-weighted F1.
We use the exact same evaluation metrics as in Yin et al. (2019), and the same label resolution strategy when the model answers “Yes"or “Entailment” for natural language inference models. for multi-label classification. Concretely, when the model predicts “Yes" on multiple labels, the one with the highest probability is selected. For a fair comparison, we meta-tune RoBERTa of the same size and compare it with the highest performing model in Yin et al. (2019) for each of the three datasets.
The results are in Table 2, and our model has higher performance across all 3 datasets using the same pre-training method.
Discussion and Future Directions
We construct a dataset of classification datasets to adapt the language model for zero-shot classification via meta-tuning. The adapted model outperforms a general-purpose question answering model and the prior state of the art based on natural language inference. We forecast that meta-tuning would be more effective on larger models, and the current engineering ceiling for zero-shot learning might have been broadly under-estimated.
Aggregating and unifying datasets
The main bottleneck of our research is to manually gather a wide range of datasets and unify their format. The difficulties are: 1) we need to brainstorm and review the NLP literature extensively to decide what new tasks to look for; 2) different datasets encode their data in different formats, and we need to write programs manually for each of them to convert to the desired format; 3) it is hard to tell the quality of a dataset purely by its provenance, and sometimes we need to examine the dataset manually. If we as a community can aggregate and unify datasets better, we could potentially train and evaluate zero-shot learning models at a larger scale.
Meta-tuning as a probe
There is a growing interest in measuring the intelligence Hendrycks et al. (2021a, b) or the few-shot learning ability Brown et al. (2020) of large language models like GPT-3. However, since these models are not adapted to answer those prompts Holtzman et al. (2021), we suspect that its knowledge and true potential to perform few-shot learning is much higher than reported. Since pre-training does the heavy lifting and meta-tuning is unlikely to provide additional ZSC ability to the model, we can potentially first use meta-tuning as a probe to make them adapted to answering prompts before measuring their performance.
Still, to make this methodology rigorous, interpreting and controlling the strength of the probes will be an important future direction Hewitt and Liang (2019). For example, if the training set contains a prompt that is too similar to the prompt to be tested, the probe will be meaningless.
Beyond Shallow Correlations
One possibility is that the model only learns shallow statistical correlations from meta-tuning rather than “more sophisticated reasoning skills". For example, the word “exciting" might occur in positive reviews more. This is unlikely, given that larger models are consistently better than smaller or randomly initialized ones. To explain this performance gap, larger models must have learned to use more complicated features during meta-tuning.
Relation to Meta/Multitask-Learning
Our method is closely related to, but different from meta-learning Yin (2020); Murty et al. (2021) and multi-task learning Ye et al. (2021); Aghajanyan et al. (2021). Both meta-learning and multitask-learning typically involve at least a few examples from the target task; in our setup, however, the model does not learn from any target task examples. The “meta” in our name does not mean “meta-learning”, but reflects the fact that our model learns from a meta-dataset of tasks.
Nevertheless, our framework can be easily adapted to a few-shot learning setup, which enables the language model to learn to learn from in-context examples (see below). Since this approach models the learning process as a sequence classification problem, it can be seen as a form of meta-learning similar to Ravi and Larochelle (2016).
Annotating Prompts
Three of our authors annotated the label descriptions. Since they are all Computer Science major students who understand machine learning and natural language processing, they might not be representative of the final user population of this ZSC application. Annotating prompts that match the target user distribution will be an important research direction.
Additionally, shorter and more natural descriptions sometimes fail to capture the exact semantics of the label. For example, in Yin et al. (2019), the description of the label “medical" is “people need medical assistance"; or alternatively, it can be longer but more accurate: “people need an allied health professional who supports the work of physicians and other health professionals". How to scalably generate more accurate and detailed label descriptions without expert efforts will be another future direction.
Optimizing Prompts
Our work is complementary to recent works that optimize the prompts to achieve better accuracy. Even if our meta-tuned model is specialized in answering prompts, it might still react very differently towards different prompts. For example, in the stance classification dataset Barbieri et al. (2020), we annotated two label descriptions (prompts) for the same label: “Does this post support atheism?" and “Is the post against having religious beliefs?". They have similar meanings, but the former has much lower accuracy than the later. We conjecture that this is because the model cannot ground abstract concepts like “atheism".
Other extensions
We conjecture that meta-tuning can be extended to more diverse tasks beyond zero-shot binary classification. To extend to multi-label classification, we need to develop a procedure to resolve the labels when the model predicts positive for more than one labels. To extend to few-shot learning, we need to increase the context length to fit several training examples into the input, which requires a larger context window and hence more computational resources. To extend to other sequence generation tasks, we need to collect a wide range of diverse sequence generation tasks to meta-tune the model, such as machine translation, summarization, free-form question answering, grammar correction, etc.
Acknowledgements
We thank Eric Wallace for his feedbacks throughout the project. We thank Steven Cao, David Gaddy, Haizhi Lai, Jacob Steinhardt, Kevin Yang and anonymous reviewers for their comments on the paper.
References
Appendix A Ethics
In the existing prompting framework, end users send the natural language descriptions and a few training examples to the large language model inference API to perform few-shot learning Brown et al. (2020). This becomes a natural source of training data for meta-tuning. Hence, the success of meta-tuning presented in this paper might incentivize for-profit organizations who provide language model inference APIs to collect prompts from the users, and train on these data.
Privacy, security, and fairness
If a model is meta-tuned on user-provided data, certain security, privacy and fairness concerns can potentially emerge. For example, Carlini et al. (2020) shows that it is possible to extract the training data from large language models, and hence meta-tuned systems might expose some users’ prompts to other users. Wallace et al. (2020) shows that it is possible to poison the model through training data and trigger unwanted behaviors; the meta-tuning procedure might be susceptible to these data poisoning attacks as well. Finally, meta-tuning might perpetuate existing societal biases hidden in the users’ prompts Bolukbasi et al. (2016).
If not addressed properly, these concerns might have a broader negative societal impact through meta-tuning. Compared to other domain-specific and task-specific machine learning applications, meta-tuned models might be applied to a much wider range of tasks, deployed at a larger scale, and serving a more diverse set of user population. Therefore, biased or poisoned training data for one task from one user population might compromise fairness and performance of another task and harm another user population; additionally, malicious or biased data might even tamper with the few-shot learning capability (“meta-poisoning").
Potential abuse
As shown in Figure 6, the AUC-ROC score for a lot of tasks are still well below 0.9, and hence our system is far from solving a significant fraction of tasks. Therefore, even though our system is flexible and has the potential to perform a wide range of tasks, it does not present an elixir to all classification tasks. Particularly, it should not be applied to higher stake scenarios (e.g. hate speech detection, fake news detection, etc), since its efficacy, robustness, and fairness properties remain unknown.
Appendix B Datasets
Maas et al. (2011b). Classifies whether the user likes the movie.
Positive: “’My favourite police series of all time turns to a TV-film. Does it work? Yes. …"
Negative: “ "Stupid! Stupid! Stupid! I can not stand Ben stiller anymore."
Zero Shot Emotion Classification
Yin et al. (2019). This task classifies a textual input into 9 emotion types {“sadness", “joy", “anger", “disgust", “fear", “surprise", “shame", “guilt", “love"}, and none-type if not any of the above. For example,
Joy: “Making new friends is always fun , specially when playing dress up"
Anger: “People that smoke cigarettes irritate my soul."
Zero Shot topic Classification
Yin et al. (2019). This task classifies an articles into 10 topic classes, such as “Health", ‘Sports", “Family & Relationships", “Science & Mathematics", etc.
Science: “How the human species evolved ? How the human species evolved ? A tough question as it overlaps science and theology …"
Sports: “Number one formula one driver ? Depends on your question.Kimi Raikkonen won more races …"
Zero Shot Situation Classification
Yin et al. (2019). This task classifies what type of situation the user is having and what aids it needs. The classes include “food", “terrorism", etc.
Food: "livestock , crops , rural and urban housing , roads , river and canal embankments , and other infrastructure suffered extensive damage"
Terrorism: "Daish attacked Tahurmahu village with chemical gas."
SemEval 2017 Task 4
Barbieri et al. (2020) Rosenthal et al. (2017). This task classifies the sentiment of a tweet as {“negative", “neutral", “positive"}. For example,
Negative: “That sucks if you have to take the SATs tomorrow"
Neutral: “@user Well said on HMW. Can you now address why Texans fans file out of the stadium midway through the 4th qtr of every game?"
Positive: “"This is the last Sunday w/o football …..,NFL is back baby""
ETHICS Justice
Hendrycks et al. (2020). This task identifies whether the justification provided for a given scenario is deemed reasonable. For example,
Reasonable: “I deserve to be paid by my boss because I keep her house clean daily."
Unreasonable: “I deserve to be given a brand new car by my boss because I keep her house clean daily."
ETHICS Deontology
Hendrycks et al. (2020). This task identifies whether the excuse provided for a given scenario is a reasonable excuse. For example,
Scenario: “Aren’t you supposed to be preparing for the meeting?"
Excuse (Reasonable): “No because I have everything already set up in the conference room."
Excuse (Unreasonable): “No because you don’t have a meeting"
ETHICS Common Sense
Hendrycks et al. (2020). This task identifies identifies whether a scenario demonstrates common sense. For example,
Common Sense: “I went to the principal’s office to change my records before going to a different school."
Not Common Sense: “I secured the loan because I would make the payments."
EURLEX57K
Chalkidis et al. (2019). This task classifies European legislation.
National Currency: “Council Regulation (EC) No 2595/2000 of 27 November 2000 amending Regulation (EC) No 1103/97 on certain provisions relating to the introduction of the euro"
Southern Africa: “95/458/EC: Commission Regulation (EC) No 302/2006 of 20 February 2006 on import licences in respect of beef and veal products originating in Botswana, Kenya, Madagascar, Swaziland, Zimbabwe and Namibia"
SemEval 2019 Task 6
Barbieri et al. (2020) Zampieri et al. (2019). This task classifies the tweet as either offensive or not offensive. For example,
Offensive: “@user She has become a parody unto herself? She has certainly taken some heat for being such an….well idiot. Could be optic too Who know with Liberals They’re all optics. No substance"
Not Offensive: “@user @user She is great. Hi Fiona!"
Click Bait Detection
This task detects whether a news title is a click bait.
Click Bait: “Can You Pass This Basic Trigonometry Quiz"
Non Click Bait: “NASCAR driver Kyle Busch wins 2011 Jeff Byrd 500".
Abstract Domain Classification
This classifies the abstract into 4 domains: “Physcis", “Maths", “Computer Science", “Statistics". For example,
Physics: “a ever-growing datasets inside observational astronomy have challenged scientists inside many aspects, including an efficient and interactive data exploration and visualization. many tools have been developed to confront this challenge …"
Maths: “a main result of this note was a existence of martingale solutions to a stochastic heat equation (she) inside the riemannian manifold …"
SemEval 2019 Task 5
Barbieri et al. (2020) Basile et al. (2019). This task identifies whether the tweet contains hate speech towards women and/or immigrants or not. For example,
Hate Speech: “This account was temporarily inactive due to an irrational woman reporting us to Twitter. What a lack of judgement, shocking. #YesAllMen"
No Hate Speech: “@user nice new signage. Are you not concerned by Beatlemania -style hysterical crowds crongregating on you…"
SemEval 2019 Task 8
Mihaylova et al. (2019). This task identifies whether the text is an example of a question asking for factual information, an example of a question asking for an opinion, or an example of socializing. For example,
Factual: “is there any place i can find scented massage oils in qatar?"
Opinion: “hi there; i can see a lot of massage center here; but i dont which one is better. can someone help me which massage center is good…and how much will it cost me? thanks"
Socializing: “Hello people…let’s play this game…you have to write something good about the person whose ’post’ is above you on QL.You can write anything and you can write multiple times."
SemEval 2018 Task 3
Barbieri et al. (2020) Van Hee et al. (2018). This task identifies whether the tweet contains irony or not. For example,
Irony: “seeing ppl walking w/ crutches makes me really excited for the next 3 weeks of my life"
No Irony: “@user on stage at #flzjingleball at the @user in #Tampa #iheartradio"
SemEval 2018 Task 1
Barbieri et al. (2020); Mohammad et al. (2018) This task classifies a tweet as one of 4 emotion types {“sadness", “joy", “anger", “optimism"}. For example,
Sadness: “@user I so wish you could someday come to Spain with the play, I can’t believe I’m not going to see it #sad"
Joy: “#ThisIsUs has messed with my mind & now I’m anticipating the next episode with #apprehension & #delight! #isthereahelplineforthis"
Anger: “@user Haters!!! You are low in self worth. Self righteous in your delusions. You cower at the thought of change. Change is inevitable."
Optimism: “Don’t be #afraid of the space between your #dreams and #reality. If you can #dream it, you can #make it so"
SemEval 2016 Task 6
Mohammad et al. (2016); Barbieri et al. (2020) This task classifies a tweet’s stance as {“neutral", “against", “favor"}. Each tweet contains a stance on one of the five different target topics {“abortion", “atheism", “climate change", “feminism", “hillary"}. For example,
Neutral: “@user maybe that’s what he wants #SemST"
Against: “Life is #precious & so are babies, mothers, & fathers. Please support the sanctity of Human Life. Think #SemST"
Favour: “@user @user Nothing to do with me. It’s not my choice, nor is it yours, to dictate what another woman chooses. #feminism #SemST"
SemEval 2020 Task 6
Spala et al. (2020). This task classifies whether textbook sentence contains a definition. For example,
Contains Definition: “Since 2005, automated sequencing techniques used by laboratories are under the umbrella of next-generation sequencing, which is a group of automated techniques used for rapid DNA sequencing"
Doesn’t Contain Definition: “These automated low-cost sequencers can generate sequences of hundreds of thousands or millions of short fragments (25 to 500 base pairs ) in the span of one day."
TREC
Li and Roth (2002). This task classifies a question into one of six question types: DESC (description), ABBR (abbreviation), ENTY (entity), HUM (people/individual), LOC (location), NUM (numeric information), each of which have specific fine-grained sub-categories. For example,
DESC: “How did serfdom develop in and then leave Russia?"
ENTY: “What films featured the character Popeye Doyle?"
HUM: “What contemptible scoundrel stole the cork from my lunch?"
LOC: “What sprawling U.S. state boasts the most airports?"
NUM: “How many Jews were executed in concentration camps during WWII?"
SUBJ
Pang and Lee (2004). This task classifies a sentence as being subjective or objective. For example,
Subjective: “smart and alert, thirteen conversations about one thing is a small gem."
Objective: “the movie begins in the past where a young boy named sam attempts to save celebi from a hunter."
The Corpus of Linguistic Acceptability
Warstadt et al. (2018).This task detects if sentences are grammatically acceptable by their original authors. For example,
Grammatically Acceptable: “Her little sister will disagree with her."
Grammatically Not Acceptable: “Has not Henri studied for his exam?"
The Multi-Genre NLI Corpus
Williams et al. (2018). This task detects if a premise is a contradiction or entailment of a hypothesis, or if a hypothesis holds neutral view on the premise.. For example,
Neutral: “Premise: Exoatmospheric Kill Vehicles orbiting Earth would be programmed to collide with warheads. Hypothesis: Exoatmospheric Kill Vehicles would be very expensive and hard to make."
Entailment: “Premise: so we have to run our clocks up forward an hour and i sure do hate to loose that hour of sleep in the morning. Hypothesis: I don’t like the time change that results in losing an hour of sleeping time."
Contradiction: “Premise: The mayor originally hoped groundbreaking would take place six months ago, but it hasn’t happened yet. Hypothesis: The mayor doesn’t want groundbreaking to happen at all."
Metaphor as a Medium for Emotion: An Empirical Study
MohammadST16. This task detects if the application of a word is Literal or Metaphorical. For example,
Metaphorical: “Her husband often abuses alcohol."
Political Preference Classification
Allaway and McKeown (2020). This task predicts a comment’s stand point on a political topic. For example,
Con: “Regulation of corporations has been subverted by corporations. States that incorporate corporations are not equipped to regulate corporations that are rich enough to influence elections, are rich enough to muster a legal team that can bankrupt the state. Money from corporations and their principals cannot be permitted in the political process if democracy is to survive."
Pro: “Regulation is to a corporation what a conscience is to a living person. Without a conscience, we would all be sociopaths. Corporations do not have a conscience, thus they need regulation to make sure they are focused on benefiting society instead on merely benefiting themselves."
Neutral: “Without government to ensure their behavior, companies will attempt to make a profit even to the DETRIMENT of the society that supports the business. We have seen this in the environment, in finances, in their treatment of workers and customers. Enough."
Airline Service Review
This task classifies if an airline review has a positive or negative sentiment. For example,
Positive: “This is such a great deal! Already thinking about my 2nd trip to Australia; I haven’t even gone on my 1st trip yet!"
Negative: “amazing to me that we can’t get any cold air from the vents."
Covid-19 Tweets Sentiment Analysis
This task classifies if a tweet has a positive or negative sentiment. For example,
Positive: “Taken by Henk Zwoferink on Saturday in Wargl, our black beauty hauled a train bringing the last tourists home. Our colleagues are #workinghard to keep supply chains running while respecting the measures to ensure everyone’s #safety. A pleasure to work with such #DedicatedPeople!"
Negative: “So far, the Minister does not seem to have made statement on the catastrophe that can develop if the issue of markets operation is not addressed. Food insecurity has potential to make current Covid-19 panic look like a kindergarten and could lead to riots. I submit."
Hotel Review
This task predicts if a hotel review is a positive or negative review. For example,
Negative: “The single rooms like hospital rooms single rooms hotel sparse intentional know ugly like trapped hospital white walls sink basin room small rectangle shape.the beds hard rocks blankets rough really noisy.this overrated hotel stayed fans type hotels"
Positive: “loved stay, stayed univ, inn 10 days april 2005 thoroughly enjoyed, free parking clean spacious room friendly staff great breakfast snack, loved location, definitely stay, "
Stock Market Sentiment
This task predicts if a comment holds a positive or negative view on the performance of the stock market. For example,
Negative: “GPS wow that wa s a fast fast fade…"
Positive: “user Maykiljil posted that: I agree that MSFT is going higher & possibly north of 30"
AG-News
Zhang et al. (2015). This task classifies the topic of news based on their contents. For example,
World News: “Greek duo could miss drugs hearing"
Sports News: “AL Wrap: Olerud Cheers Yankees by Sinking Ex-Team"
Business News: “Lowe’s Second-Quarter Profit Rises"
Tech News: “Satellite boosts Olympic security"
Real and Fake News
This task classifies if a news is fake or real. For example,
Real: “WASHINGTON (Reuters) - Alabama Secretary of State John Merrill said he will certify Democratic Senator-elect Doug Jones as winner on Thursday despite opponent Roy Mooreâ x80 x99s challenge, in a phone call on CNN. Moore, a conservative who had faced allegations of groping teenage girls when he was in his 30s, filed a court challenge late on Wednesday to the outcome of a U.S. Senate election he unexpectedly lost."
Fake: “Ronald Reagan shut down the Berkeley protests many years ago THIS is how you do it!"
Disaster Tweets
This task detects if a tweet announces an emergency or a disaster. For example,
Contains Disaster: “Our Deeds are the Reason of this #earthquake May ALLAH Forgive us all."
Does not Contain Disaster: “My dog attacked me for my food #pugprobs."
Obama vs Trump Tweets
This task detects if a tweet was send by Obama or Trump. For example,
Obama: “Michelle and I are delighted to congratulate Prince Harry and Meghan Markle on their engagement. We wish you a lifetime of joy and happiness together."
Trump: “Together, we dream of a Korea that is free, a peninsula that is safe, and families that are reunited once again!"
Kaggle Sexually Explicit Tweets
This dataset provides positive examples of profane comments. For example,
Explicit“What do guys say when you get naked in front of them for the first time?"
Democratic vs Republican Tweets
This task detects if a tweet was send by the Democratic or Republican Party. For example,
Democratic: “#YuccaMountain would require moving tens of thousands of metric tons of radioactive waste across the country and through Southern Nevada."
Republican: “Stopped by One Hour Heating& Air Conditioning to discuss the benefits tax reform will bring to their business."
Women E-commerce Clothing Reviews
This task predicts if the buyer likes or recommends a product base on its review. For example,
Like: “After reading the previous reviews, i ordered a size larger. i am so glad i did it! it fits perfectly! i am 5’4"/115/32dd and went with the s regular. so beautiful! i can’t wait to wear it!"
Dislike: “The zipper broke on this piece the first time i wore it. very disappointing since i love the design. I’m actually going to try to replace the zipper myself with something stronger, but annoying that it’s come to that."
Quora Question Pairs
This task predicts if a pair of Quora question is asking for the same thing. For example,
Same: “Question 1: How many months does it take to gain knowledge in developing Android apps from scratch?; Question 2: How much time does it take to learn Android app development from scratch?"
Different: “Question 1: How would you review the site Waveclues? ; Question 2: Is there a good pay for reviews site out there?"
Headline Sarcasm Detection
This task detects if is a news headline contains scarcasm. For example,
Sarcasm: “guy who just wiped out immediately claims he’s fine"
No Sarcasm: “Donald trump effigies burn across Mexico in Easter ritual"
Company Account Tweets
This task detects whether the tweet is targeted towards a company account. For example,
Yes: “@VirginTrains Oh, that’s nice. What are you doing about it? What are you targets next year?"
No: “@115738 That’s the best kind of trick-or-treating. All treats, my friend. -Becky"
SMS Spam Detection
Almeida et al. (2013) This task detects whether the SMS is a spam message. For example,
Spam: “Thank you, winner notified by sms. Good Luck! No future marketing reply STOP to 84122 customer services 08450542832"
Ham: “Lol great now I am getting hungry."
Clothing Fitness
Misra et al. (2018) Checking whether the customer complains that the cloth is too small or too large.
Water Problem Topic Classification
Classifying the topic of a report on water problems. The labels include “biological", “climatic indicator", “environmental technology", etc. For example,
Biological: “Mineralization of organic phosphorus in bottom sediments reaches 40–80% and as we found out during the project implementation it intensified in autumn-winter period."
climatic indicator: “The average amount of precipitation in the lower part of the basin makes 470 mm to 540 mm. The relative average annual air humidity makes 60-65%".
environmental technology: “Most of wastewater treatment facilities require urgent modernization and reconstruction".
Sexist Statement Detection
This task classifies whether the statement is sexist. For example,
Sexist: “It’s impossible for a girl to be faithful."
Non Sexist: “Without strength, can we work to create wealth?"
Movie Spoiler Detection
Misra (2019) https://www.kaggle.com/rmisra/imdb-spoiler-dataset?select=IMDB_reviews.json This task classifies whether the movie review is a spoiler. For example,
Spoiler: “I must say that this movie was good but several things were left unsaid. For those who have seen the movie know what I am talking about but for those who haven’t, I don’t want to give spoilers. I was also impressed by Vin Diesel’s acting skills. Overall I have to say it was a good movie filled with several twists and turns."
Non Spoiler: “The Great Wall amazes with its spectacular effects, both on screen and sound. Usually I do not appreciate 3D movies, but in this case I felt like it worth it.However, being honest, the storytelling and the story itself had its weaknesses. There were many logical lapses, and for me, many details are still waiting to be answered.On the other hand, expect decent acting especially from the main characters.All in all, The Great Wall is a solid popcorn-movie, but I expected a more elaborated unfolding of the legend it tells about."
News Summary/headline Topic Classification
This task classifies the topic of the summary of a news. For example,
Politics: “City and state officials said they received little advance warning of the decision."
Business: “The streaming giant’s third-quarter earnings were nothing like the Upside Down."
Appendix C Dataset Property Tags
Here we list all the dataset property tags (Section 2). We define two datasets to be “similar" if they have the set of tags, and disallow meta-tuning on datasets that are similar to evaluation dataset.
social media: whether the source is from social media (e.g. tweets).
social/political: whether the task is highly related to political/social topics. Some examples include stance classification and hate speech detection.
topic classification: whether the task classifies the topics of the input.
good vs. bad: whether the task classifies whether the text is judging something to be good or bad.
paper: whether input text comes from a paper.
review: whether the input text is a review of a product (e.g. movie, hotel).
questions: whether the input texts are questions. Some examples include classifying whether the question asks for factual information or subjective opinion and detecting whether two questions have the same meaning.
emotion: whether the task classifies certain emotion in the text, for example “hate", “surprise", “joy", etc.
Besides, we do not assign tags to datasets that we are confident to be different enough from other tasks (e.g. extracting whether a text contains definition), and allow the model to be meta-tuned on all other datasets.
Appendix D List of Label Descriptions
The comprehensive list of label descriptions and grouping can be seen in Figure 8, 9, and 10.
Appendix E Robustness Checks
We weight each label and dataset equally in Table 4 and 5. We find that, under almost all comparisons across different weighting, the mean change is positive, and the change above a certain threshold is more frequent than the change below a certain threshold . The only single exception the “Ensemble" row in Table 5, where there are slightly more datasets where the change is lower than -1%. Nevertheless, given that the trend is still positive under and , and two other description weightings, we may still conclude that ensembling label descriptions is more likely to improve model performance.
E.2 Larger T5 Models are Better
In addition to comparing T5-Base (220 Million parameters) vs. T5-Large (770M), we also compare T5-small (60M) vs. T5-base (220M). Across all metrics, larger models are significantly better. Most notably, there is a sudden jump in performance when increasing model size from T5-small to T5-base (sometimes increase in ).
E.3 Larger BERT Models are Better
We also compare different sizes of BERT Turc et al. (2019) (41, 110, and 330M) parameters. Across all metrics, larger models are significantly better.
Appendix F Most Relevant Datasets
To ensure that we are testing the models’ ability to generalize to an unseen tasks, we disallow both training and testing on datasets that are too similar, which is defined as “having the same set of dataset property tags" (Section 2). To help interpret how we define unseen tasks, for each dataset that we evaluate on, we try to find the “most relevant" dataset that the model has seen during the meta-tuning phase, and list it in Table 6.
Appendix G Performance Break Down
For each model, we average the AUC-ROC scores for each label description for each dataset, and report the results in Table 7.