Are Larger Pretrained Language Models Uniformly Better? Comparing Performance at the Instance Level
Ruiqi Zhong, Dhruba Ghosh, Dan Klein, Jacob Steinhardt
Introduction
Historically, large deep learning models Peters et al. (2018); Devlin et al. (2019); Lewis et al. (2020); Raffel et al. (2019) have improved the state of the art on a wide range of tasks and leaderboards Schwartz et al. (2014); Rajpurkar et al. (2016); Wang et al. (2018), and empirical scaling laws predict that larger models will continue to increase performance Kaplan et al. (2020). However, little is understood about such improvement at the instance (datapoint) level. Are larger models uniformly better? In other words, are larger pretrained models better at every instance, or are they better at some instances, but worse at others?
Prior works hint at differing answers. Hendrycks et al. (2020) and Desai and Durrett (2020) find that larger pretrained models consistently improve out-of-distribution performance, which implies that they might be uniformly better at a finer level. Henighan et al. (2020) claim that larger pretrained image models have lower downstream classification loss for the majority of instances, and they predict this trend to be true for other data modalities (e.g. text). On the other hand, Sagawa et al. (2020) find that larger non-pretrained models perform worse on rare subgroups; if this result generalizes to pretrained language models, larger models will not be uniformly better. Despite all the indirect evidence, it is still inconclusive how many instances larger pretrained models perform worse on.
A naïve solution is to finetune a larger model, compare it to a smaller one, and find instances where the larger model is worse. However, this approach is flawed, since model predictions are noisy at the instance level. On MNLI in-domain development set, even the same architecture with different finetuning seeds leads to different predictions on 8% of the instances. This is due to under-specification D’Amour et al. (2020), where there are multiple different solutions that can minimize the training loss. Since the accuracy improvement from our BERT-baseThis is not the original release by Devlin et al. (2019); we pretrained models ourselves. to BERT-large is 2%, most signals across different model sizes will be dominated by noise due to random seeds.
To account for the noise in pretraining and finetuning, we define instance accuracy as “how often a model correctly predicts an instance” (Figure 1 left) in expectation across pretraining and finetuning seeds. We estimate this quantity by pretraining 10 models with different seeds, finetuning 5 times for each pretrained models (Figure 1 middle), and averaging across them.
However, this estimate is still inexact, and we might falsely observe smaller models to be better at some instances by chance. Hence, we propose a random baseline to estimate the fraction of false discoveries (Section 3, Figure 1 right) and formally upper-bound the false discoveries in Section 4. Our method provides a better upper bound than the classical Benjamini-Hochberg procedure with Fisher’s exact test.
Using the 50 models for each size and our improved statistical tool, we find that, on the MNLI in-domain development set, the accuracy “decays” from BERT-large to BERT-mini on at least 4% of the instances, which is significant given that the improvement in overall accuracy is 10%. These decaying instances contain more controversial or wrong labels, but also correct ones (Section 4.2). Therefore, larger pretrained language models are not uniformly better.
We make other interesting discoveries at the instance level. Section 5 finds that instance-level accuracy has momentum: improvement from mini to medium correlates with improvement from medium to large . Additionally, Section 6 attributes variance of model predictions to pretraining and finetuning random seeds, and finds that finetuning seeds cause more variance for larger models. Our findings suggest that instance-level predictions provide a rich source of information; we therefore recommend that researchers supplement model weights with model predictions. In this spirit, we release all the pretrained models, model predictions, and code here: https://github.com/ruiqi-zhong/acl2021-instance-level.
Data, Models, and Predictions
To investigate model behavior, we considered different sizes of the BERT architecture and finetuned them on Quora Question Pairs (QQPhttps://www.quora.com/q/quoradata/First-Quora-Dataset-Release-Question-Pairs), Multi-Genre Natural Language Inference (MNLI; Williams et al. (2020)), and the Stanford Sentiment Treebank (SST-2; Socher et al. (2013)). To account for pretraining and finetuning noise, we averaged over multiple random initializations and training data order, and thus needed to pretrain our own models rather than downloading off the internet. Following Turc et al. (2019) we trained 5 architectures of increasing size: mini (L4/H256, 4 Layers with hidden dimension 256), small (L4/H512), medium (L8/H512), base (L12/H768), and large (L24/H1024). For each architecture we pre-trained models with 10 different random seeds and fine-tuned each of them 5 times (50 total) on each task; see Figure 1 middle. Since pretraining is computationally expensive, we reduced the context size during pretraining from 512 to 128 and compensated by increasing training steps from 1M to 2M. Appendix A includes more details about pretraining and finetuning and their computational cost, and Appendix B verifies that our cost-saving changes do not affect accuracy qualitatively.
Unless otherwise noted, we present results on the MNLI in-domain development set in the main paper.
Comparing Instance Accuracy
To find the instances where larger models are worse, a naïve approach is to finetune a larger pretrained model, compare it to a smaller one, and find instances where the larger is incorrect but the smaller is correct. Under this approach, BERT-large is worse than BERT-base on 4.5% of the instances and better on 7%, giving an overall accuracy improvement of 2.5%.
However, this result is misleading: even if we compare two BERT-base model with different finetuning seeds, their predictions differ on 8% of the instances, while their accuracies differ only by 0.1%; Table 1 reports this baseline randomness across model sizes. Changing the pretraining seed also changes around additional predictions beyond finetuning.
Table 1 also reports the standard deviation of overall accuracy, which is about 40 times smaller. Such stability starkly contrasts with the noisiness at the instance level, which poses a unique challenge.
Recall that our goal is to find instances where larger models are less accurate, which we refer to as decaying instances. We therefore study the instance difference between two model sizes and , defined as
Quantifying the Decaying Instances
We have the following theorem, which we formally state and prove in Appendix D:
(Informal) If all the random seeds are independent, then for all thresholds ,
Suppose we observe and , where there are different random seeds for each model size We assumed even number of random seeds since we will mix half of the models from each size to compute the random baseline. Then
For the random baseline estimator, we have
We compute this lower bound for other pairs of model sizes in Table 2, and the full results across other tasks and model size pairs are in Appendix C. In all of these settings we find a non-zero fraction of decaying instances, and larger model size differences usually lead to more decaying instances.
Unfortunately, applying Theorem 1 as above is not fully rigorous, since some finetuning runs share the same pretraining seeds and hence are dependent.Although we anticipate such dependencies do not cause a substantial difference, as discussed in Appendix D.1. To obtain a statistically rigorous lower bound, we slightly modify our target of interest. Instead of examining individual finetuning runs, we ensemble our model across 5 different finetuning runs for each pretraining seed; these predictions are essentially the same as individual finetuning runs, except that the finetuning randomness is averaged out. Hence we obtain 10 independent sets of model predictions with different random seeds, which allows us to apply Theorem 1.
1 Fisher’s Test + Benjamini-Hochberg
Here is a more classical approach to lower-bound the decaying fraction. For each instance, we compute a significance level under the null hypothesis that the larger model is better, using Fisher’s exact test. We sort the significance levels ascendingly, and call the percentile . Then we pick a false discovery rate (say, 25%), find the largest s.t. , and estimate the decaying fraction to be at least . This calculation is known as the Benjamini-Hochberg procedure Benjamini and Hochberg (1995).
To compare our method with this classical approach, we estimate the lower bound of the decaying fraction for different pairs of model sizes with different numbers of pretrained models available. To make sure our choice of the false discovery rate does not bias against the classical approach, we adaptively choose to maximize its performance. Appendix F includes the full results and Table 4 is a representative subset.
We find that our approach is more powerful, particularly when the true decaying fraction is likely to be small and only a few models are available, which is usually the regime of interest. For example, across all pairs of model sizes, our approach only needs 2 random seeds (i.e. pretrained models) to provide a non-zero lower bound on the decaying fraction, while the classical approach sometimes fails to do this even with 10 seeds. Intuitively, when fewer seeds are available, the smallest possible significance level for each instance is larger than the decaying fraction, hence hurting the classical approach.
2 Understanding the Decaying Instances
We next manually examine the decaying instances to see whether we can find any interpretable patterns. One hypothesis is that all the decaying fractions are in fact mislabeled, and hence larger models are not in fact worse on any instances.
For each task of MNLI, QQP, and SST-2, the first author annotated 100 instances (decay + control group) (Table 5). We present all the annotated decaying instances in Appendix J.
However, we cannot find an interpretable pattern for these correctly labeled decaying instances by simply eyeballing. We discuss future directions to discover interpretable categories in Section 7.
Correlation of Instance Difference
We next investigate whether there is a momentum of instance accuracy increase: for example, if the instance accuracy improves from to , is it more likely to improve from to ?
Therefore, we bucket instances by their estimated medium accuracy into intervals of size 0.1, and we find the correlation to be positive within each bucket (Table 6, row 2). This fixes the problem with the naïve approach by getting rid of the negative correlation, which could have misled us to believe that improvements by larger models are uncorrelated.
We additionally find that the correlations between improvements become stronger when model size differences are smaller. Table 6 row 1 reports results for another model size triplet with smaller size difference, i.e. (, , ) = (small, medium, base), and the correlation is larger for all buckets. Results for more tasks and size triplets are in Appendix G and the same conclusions hold qualitatively.
Variance at the Instance Level
Section 3 found that the overall accuracy has relatively low variance, but model predictions are noisy. This section formally analyzes variance at the instance level. For each instance, we decompose its loss into three components: Bias2, variance due to pretraining randomness, and variance due to finetuning randomness. Formally, we consider the loss:
where is a random variable 0/1 indicating whether the prediction is correct or incorrect, with respect to randomness in pretraining and finetuning. Therefore, by bias-variance decomposition and total variance decomposition, we have
where, by using and as pretraining and finetuning random seeds:
capturing “how wrong is the average prediction”, variance due to pretraining, and variance due to finetuning seeds, respectively.
We find that 1) larger models have larger finetuning variance, 2) large has smaller pretraining variance than base ; however, the ordering between other sizes varies across tasks and losses, and 3) finetuning variance is 28 times as large as pretraining variance, and the ratio is bigger for larger models.
Discussion and Future Directions
To investigate model behaviors at the instance level, we produced massive amounts of model predictions in Section 2 and treated them as raw data. To extract insights from them, we developed better metrics and statistical tools, including a new method to control the false discoveries, an unbiased estimator for the decomposed variances, and metrics that compute variance and correlation of improvements conditioned on instance accuracy. We find that larger pretrained models are indeed worse on a non-trivial fraction of instances and have higher variance due to finetuning seeds; additionally, instance accuracy improvements from mini to medium correlate with improvements from medium to large .
Overall, we treated model prediction data as the central object and built analysis tools around them to obtain a finer understanding of model performance. We therefore refer to this paradigm as “instance level understanding as data mining”. We discuss three key factors for this paradigm to thrive: 1) scalability and the cost of obtaining prediction data, 2) other information to collect for each instance, and 3) better statistical tools. We analyze each of these aspects below.
Data mining is more powerful with more data. How easy is it to obtain more model predictions? In our paper, the main bottleneck is pretraining. However, once the pretrained models are released, individual researchers can download them and only need to repeat the cheaper finetuning procedure.
Furthermore, model prediction data are under-shared: while many recent research papers share their code or even model weights to help reproduce the results, it is not yet a standard practice to share all the model predictions. Since many researches follow almost the same recipe of pretraining and finetuning McCoy et al. (2020); Desai and Durrett (2020); Dodge et al. (2020), much computation can be saved if model predictions are shared. On the other hand, as the state of the art model size is increasing at a staggering speede.g. BERT Devlin et al. (2019) has 340M parameters, while Switch-Transformer has over 1 trillion parameters Fedus et al. (2021). , most researchers will not be able to run inference on a single instance. The trend that models are becoming larger and more similar necessitate more prediction sharing.
Data mining is more powerful with more types of information. One way to add information to each instance is to assign “meta-labels”. In the HANS McCoy et al. (2019) dataset, the authors tag each instance with a heuristic For example, “the label [entailment] is likely if the premise and the hypothesis have significant lexical overlap”. that holds for the training distribution but fails on this instance. Naik et al. (2018a) and Ribeiro et al. (2020) associate each instance with a particular stress test type or subgroup, for example, whether the instance requires the model to reason numerically or handle negations. Nie et al. (2020) collects multiple human responses to estimate human disagreement for each instance. This meta-information can potentially help us identify interpretable patterns for the disagreeing instances where one model is better than the other. On the flip side, identifying disagreeing instances between two models can also help us generate hypothesis and decide what subgroup information to annotate.
We can also add performance information on other tasks to each instance. For example, Pruksachatkun et al. (2020) studied the correlation between syntactic probing accuracy Hewitt and Liang (2019) and downstream task performance. Turc et al. (2019) and Kaplan et al. (2020) studied the correlation between language modelling loss and the downstream task performance. However, they did not analyze correlations at the instance level. We may investigate whether their results hold on the instance level: if an instance is easier to tag by a probe or easier to predict by a larger language model, is the accuracy likely to be higher?
Data mining is more powerful with better statistical tools. Initially we used the Benjamini-Hochberg procedure with Fisher’s exact test, which required us to pretrain 10 models to formally verify that the decaying instances exist. However, we later realized that 2 is in fact enough by using our approach introduced in Section 4. We could have saved 80% of the computation for pretraining if this approach was known before we started.
Future work can explore more complicated metrics and settings. We compared at most 3 different model sizes at a time, and higher order comparisons require novel metrics. We studied two sources of randomness, pretraining and finetuning, but other sources of variation can be interesting as well, e.g. differences in pretraining corpus, different model checkpoints, etc. To deal with more sophisticated metrics, handle different sources and hierarchies of randomness, and reach conclusions that are robust to noises at the instance level, researchers need to develop new inference procedures.
To conclude, for better instance level understanding, we need to produce and share more prediction data, annotate more diverse linguistic properties, and develop better statistical tools to infer under noises. We hope our work can inform researchers about the core challenges underlying instance level understanding and inspire future work.
Acknowledgement
We thank Steven Cao, Cathy Chen, Frances Ding, David Gaddy, Colin Li, and Alex Wei for giving comments on the initial paper draft. We would also like to thank the Google Cloud TPU team for their hardware support.
References
Appendix A Pretraining and Finetuning Details
Here we explain how to obtain the model predictions, which are analyzed in later sections. To obtain these predictions under the “pretraining and finetuning” framework Devlin et al. (2019), we need to decide a model size, perform pretraining, finetune on a training set with a choice of hyper-parameters, and test the model on an evaluation set. We discuss each bolded aspects below.
Similar to Turc et al. (2019), we experimented with the following five model sizes, listed in increasing order: mini (L4/H256) 4 Layers with hidden dimension 256 , small (L4/H512), medium (L8/H512), base (L12/H768), and large (L24/H1024).
We used the pretraining code from Devlin et al. (2019) and the pre-training corpus from Li et al. (2020). Compared to the original BERT release, we used context size 128 instead of 512, since computation cost grows quadratically with respect to context size; we also pretrained for 2M steps instead of 1M.
We consider 3 datasets: Quora Question Pairs (QQP) https://www.quora.com/q/quoradata/First-Quora-Dataset-Release-Question-Pairs, Multi-Genre Natural Language Inference (MNLI; Williams et al. (2020)), and the Stanford Sentiment Treebank (SST-2; Socher et al. (2013)). For QQP we used the official training split. For MNLI we used 350K out of 400K instances from the original training split, and added the remaining 50K to the evaluation set, since the original in-domain development set only contains 10K examples. For SST-2, we mix the training and development set of the original split, split the instances into 5 folds, train on four of them, and evaluate on the remaining fold.
As in Turc et al. (2019), we finetune 4 epochs for each dataset. For each task and model size, we tune hyperparameters in the following way: we first randomly split our new training set into 80% and 20%; then we finetune on the 80% split with all 9 combination of batch size and learning rate [1e-4, 5e-5, 3e-5], and choose the combination that leads to the best average accuracy on the remaining 20%.
After finetuning our pretrained models, we evaluate them on a range of in-domain, out-of-domain, or challenging datasets to obtain model predictions. Models trained on MNLI are also evaluated on Stanford Natural Language Inference (SNLI; Bowman et al. (2015)), Heuristic Analysis for NLI Systems (HANS; McCoy et al. (2019)), and stress test evaluations (STRESS; Naik et al. (2018b)). Models trained on QQP are also evaluated on Twitter Paraphrase Database (TwitterPPDB; Lan et al. (2017)).
Since pretraining introduces randomness, for each model size , we pretrain 10 times with different random seed ; since finetuning also introduces noise, for each pretrained model we pretrain 5 times with different random seed ; besides, we also evaluate the model at the checkpoints after epochs, where .
Pretraining 10 models for all 5 model sizes altogether takes around 3840 hours on TPU v3 with 8 cores. Finetuning all of them 5 times for all three tasks in our paper requires around 1200 hours.
Appendix B Compare Our Models to the Original
Since we decreased the pre-training context length to save computation, these models are not exactly the same as the original BERT release by Devlin et al. (2019) and Turc et al. (2019). We need to benchmark our model against theirs to ensure that the performance of our model is still reasonable and the qualitative trend still holds. For each each size and task, we finetune the original model 5 times and calculate the average of overall accuracy.
The comparison can be seen in Table 8. We find that our model does not substantially differ from the original ones on QQP and SST-2. On MNLI, the performance of our BERT-base and BERT-large is 23% below the original release, but the qualitative trend that larger models have better accuracy still holds robustly.
Appendix C More Instance Difference Results
Additionally, for each pair of model sizes and , we estimate “how much instances are getting better/worse accuracy?” by taking the maximum difference between the red curve and the blue curve. We report these results for MNLI, SST-2, and QQP in Table 9. We find that larger model size gaps lead to larger decaying fraction, but also larger improving fraction as well.
Appendix D Proof of Theorem 1
Our goal is to show that if all the random seeds are independent,
More concretely, suppose each instance is indexed by , the set of all instances is , and the random seed is ; then is a random dimensional vector, where if the model of size correctly predicts instance under the random seed . We are comparing model size and , where is larger; to keep notation uncluttered, we omit these indexes whenever possible.
Suppose we observe and , where there are different random seeds for each model size We assumed even number of random seeds since we will mix half of the models from each size to compute the random baseline. Then
For the random baseline estimator, we have
To reiterate, the definition of the true decay rate is
By re-arranging terms and linearity of expectation, Equation 22 is equivalent to the following
Hence, we can declare victory if we can prove that for all ,
which will be proved in Lemma 1.
D.1 Independent Seed Assumption
Suppose we observe 2 independent pretraining seeds for each size and infinite number of finetuning seeds for each pretraining seed, and let us consider the threshold -0.8. Then
The key idea behind this counter-example is that even if the larger model has better average, the distribution of average finetuning accuracy for different pretraining seeds might not stochastically dominate the one with lower average because of outliers. Hence, a priori, this is unlikely to happen in practice, since pretraining variance is generally small, and we have multiple pretraining seeds to average out the outliers. Nevertheless, future work is needed to make a more rigorous argument.
Appendix E Upward Bias of Adaptive Thresholds
In section 3 we picked the best threshold that can maximize the lowerbound, which can incur a slight upward bias. Here we estimate that the bias is at most 10% relative to the unbiased lowerbound with a bootstrapping method.
We report all results in Table 10. In general, we find that the upward bias is negligible, which is at most around .
Appendix F Comparison with Significance Testing
We also experimented with the classical approach that calculates the significance-level for each instance and then use the Benjamini-Hochberg procedure to lowerbound the decaying fraction. To make sure that we are comparing with this approach fairly, we lend it additional power by picking the false discovery rate that can maximize the true discovery counts. We report the decaying fraction on MNLI in-domain development set found by this classical method and compare it with our method for different model size differences in Table 11; we also simulate situations when we have fewer models.
In general, we find that our method always provide a tighter (higher) lowerbound than the classical method, and 2 models are sufficient to verify the existence (i.e. lowerbound 0) of the decaying fraction; in contrast, the classical method sometimes fails to do this even with 10 models, e.g., when comparing base to large .
Appendix G More Results on Momentum
For nearly all buckets, the improvements are positively correlated.
When model size gap becomes larger (e.g. has the largest model size differences), the correlation decreases.
Appendix H Loss Decomposition and Estimation
In this section, under the bias-variance decomposition and total variance decomposition framework, we decompose loss into four components: bias, variance brought by pretraining randomness, by finetuning randomness, and across different checkpoints throughout training. We formally define the quantities we want to estimate in Appendix H.1, present an unbiased estimator for these quantities in Appendix H.2, and show that our method can be generalized to arbitrary number of source of randomness in Appendix H.3.
Specifically, the main paper focused on scenarios with 2 sources of randomness: pretraining and finetuning. We discuss the case with 3 sources of randomness in the appendix, rather than 2 as in the main paper, because it is easier to understand the general estimation strategy in the case of 3.
The expected squared loss of model on instance can then be written as
and that by fluctuations across checkpoints
H.2 Unbiased Estimation
We first describe the data on which we apply our estimator. Suppose we pretrain models with different random seeds, for each of the pretrained models we finetune with different random seeds, and we evaluate at different checkpoints. Then [] := , we observe and , where each observed and are i.i.d. distributed. Our goal is to estimate from the four quantities described in the previous section.
Therefore, , we have
As before, by linearity of expectation, we can declare victory if we can develop an unbiased estimator for the following quantity and then average across :
which verbally means ”variance across different finetuning seeds of the mean of over different checkpoints , conditioned on the pretraining seed .”
A naive solution is to take first take the mean of for each , i.e.
We introduce the following general theorem to correct this bias.
In this estimator, the first term “pretends” that are perfect estimator for the population mean and calculate the variance, while the second term corrects for the fact that the empirical mean estimation is not perfect. Notice the theorem only requires that and are unbiased, and is agnostic to the actual computation procedure by these estimators.
We define the population mean of to be , i.e.
and the population mean of across randomness in to be , i.e.
We look at the first term of the estimator in equation 50:
There are 5 summands within , and we look at them one by one:
Putting these five terms together, we continue calculating Equation 55:
Then from Equation 50, we can tell that is unbiased.
Now we come back to the topic of developing an unbiased estimator for as defined in Equation 46. To utilize Theorem 2, we need two components:
An unbiased estimator for the variance of , i.e.
Therefore, to develop an unbiased estimator for , it suffices to have an unbiased estimate of . We define
An unbiased estimator for the variance of , i.e. .
However, we cannot straightforwardly estimate as before, since samples are no longer independent. We need to use Equation 57 to develop an unbiased estimator (the LHS is exactly what we want!), i.e.
It is easy to see that the following is an unbiased estimator for the loss .
By linearity of expectation and loss decomposition in Equation 34,
H.3 Generalization
We can generalize this estimation strategy to decompose variance into arbitrary number of randomness. In general, we want to estimate some quantity of the following form
from the data that has an hierarchical tree structure of randomness.
For the goal of developing an unbiased estimator, we can get rid of the outer expectation easily by linearity of expectation: simply estimate the Variance conditioned on and average them together, as discussed in Section H.2.1.
an unbiased estimator for the variance of . If , we can directly compute the sample variance of the as our estimate (e.g. in Equation 63). Otherwise, we use Equation 57 to decompose the desired quantities into two, and estimate them recursively by applying Theorem 2 and Equation 57.
For readability we wrote the proof with the assumption that, in the tree of randomness, the number of branches for each node at the same depth is the same. However, our proof does not make use of this assumption and can be applied to a general tree structure of randomness as long as the the number of children is larger or equal to 2 for each non-terminal node.
Appendix I Variance Conditioned on Bias
We also experimented with the squared loss:
where is the probability the assigned to the correct label for instance . We plot the same curve in Figure 10 and observe the same trend.
Appendix J Example Decaying Instances
We show below all the annotated instances from this decaying fraction and their categories for MNLI (Section J.1), QQP , and SST-2(Section J.3).
MNLI is the abbreviation of Multi-Genre Natural Language Inference (Williams et al. (2020)). In this task, given a premise and a hypothesis, the model needs to classify whether the premise entails/contradicts the hypothesis, or otherwise. The instances can be seen below. Premise : and that you’re very much right but the jury may or may not see it that way so you get a little anticipate you know anxious there and go well you know Hypothesis : Jury’s operate without the benefit of an education in law. Label : Neutral Category : Correct Premise : In fiscal year 2000, it reported estimated improper Medicare Fee-for-Service payments of $11. Hypothesis : The payments were improper. Label : Entailment Category : Fine Premise : is that what you ended up going into Hypothesis : So that must be what you chose to do? Label : Entailment Category : Correct Premise : INTEREST RATE - The price charged per unit of money borrowed per year, or other unit of time, usually expressed as a percentage. Hypothesis : Interest rate is defined as the total amount of money borrowed. Label : Entailment Category : Wrong Premise : The analyses comply with the informational requirements of the sections including the classes of small entities subject to the rule and alternatives considered to reduce the burden on the small entities. Hypothesis : The rules place a high burden on the activities of small entities. Label : Contradiction Category : Correct Premise : Isn’t a woman’s body her most personal property? Hypothesis : Women’s bodies belong to themselves, they should decide what to do with it. Label : Neutral Category : Unsure Premise : The Standard , published a few days before Deng’s death, covers similar territory. Hypothesis : The Washington Post covers similar territory. Label : Neutral Category : Correct Premise : Shoot only the ones that face us, Jon had told Adrin. Hypothesis : Jon told Adrin and the others to only shoot the ones that face us. Label : Entailment Category : Wrong Premise : But if you take it seriously, the anti-abortion position is definitive by definition. Hypothesis : If you decide to be serious about supporting anti-abortion, it’s a very run of the mill belief to hold. Label : Neutral Category : Unsure Premise : yeah well that’s the other thing you know they talk about women leaving the home and going out to work well still taking care of the children is a very important job and and someone’s got to do it and be able to do it right and Hypothesis : It is not acceptable for anybody to refuse work in order to take care of children. Label : Contradiction Category : Correct Premise : The researchers found expected stresses like the loss of a check in the mail and the illness of loved ones. Hypothesis : The stresses affected people much diffferently than the researchers expected. Label : Contradiction Category : Correct Premise : so you know it’s something we we have tried to help but yeah Hypothesis : We did what we could to help. Label : Entailment Category : Correct Premise : Czarek was welcomed enthusiastically, even though the poultry brotherhood was paying a lot of sudden attention to the newcomers - a strong group of young and talented managers from an egzemo-exotic chicken farm in Fodder Band nearby Podunkowice. Hypothesis : Czarek was welcomed into the group by the farmers. Label : Entailment Category : Correct Premise : ’I don’t suppose you could forget I ever said that?’ Hypothesis : I hope that you can remember that forever. Label : Contradiction Category : Wrong Premise : Oh, my friend, have I not said to you all along that I have no proofs. Hypothesis : I told you from the start that I had no evidence. Label : Entailment Category : Correct Premise : I should put it this way. Hypothesis : I should phrase it differently. Label : Entailment Category : Correct Premise : An organization’s activities, core processes, and resources must be aligned to support its mission and help it achieve its goals. Hypothesis : An organization is successful if its activities, resources, and goals align. Label : Entailment Category : Fine Premise : A more unusual dish is azure, a kind of sweet porridge made with cereals, nuts, and fruit sprinkled with rosewater. Hypothesis : Azure is a common and delicious food made with cereals, nuts and fruit. Label : Entailment Category : Wrong Premise : once you have something and it’s like i was watching this program on TV yesterday in nineteen seventy six NASA came up with Three D graphics right Hypothesis : I was watching a program about gardening. Label : Contradiction Category : Correct Premise : , First-Class Mail used by households to pay their bills) and the household bill mail (i.e. Hypothesis : Second-Class Mail used by households to pay their bills Label : Contradiction Category : Unsure Premise : Rightly or wrongly, America is seen as globalization’s prime mover and head cheerleader and will be blamed for its excesses until we start paying official attention to them. Hypothesis : America’s role in the globalization movement is important whether we agree with it or not. Label : Entailment Category : Correct Premise : After being diagnosed with cancer, Carrey’s Kaufman decides to do a show at Carnegie Hall. Hypothesis : Carrey’s Kaufman is only diagnosed with cancer after doing a show at Carnegie Hall. Label : Contradiction Category : Correct Premise : Several pro-life Dems are mounting serious campaigns at the state level, often against pro-choice Republicans. Hypothesis : Serious campaigns are being run by a few pro-life Democrats. Label : Entailment Category : Correct Premise : On the northwestern Alpine frontier, a new state had appeared on the scene, destined to lead the movement to a united Italy. Hypothesis : The unite Italy movement was waiting for a leader. Label : Neutral Category : Fine Premise : well we bought this with credit too well we found it with a clearance uh down in Memphis i guess and uh Hypothesis : We bought non-sale items in Memphis on credit. Label : Contradiction Category : Correct Premise : He slowed. Hypothesis : He stopped moving so quickly. Label : Entailment Category : Correct Premise : As legal scholar Randall Kennedy wrote in his book Race, Crime, and the Law , Even if race is only one of several factors behind a decision, tolerating it at all means tolerating it as potentially the decisive factor. Hypothesis : Race is one of several factors in some judicial decisions Label : Entailment Category : Correct Premise : Although all four categories of emissions are down substantially, they only achieve 50-75% of the proposed cap by 2007 (shown as the dotted horizontal line in each of the above figures). Hypothesis : All of the emission categories experienced a downturn except for one. Label : Contradiction Category : Correct Premise : He sat up, trying to free himself. Hypothesis : He was trying to take a nap. Label : Contradiction Category : Correct Premise : Impossible. Hypothesis : Cannot be done. Label : Entailment Category : Correct Premise : But, as the last problem I’ll outline suggests, neither of the previous two objections matters. Hypothesis : I will not continue to outline any more problems. Label : Entailment Category : Correct Premise : As the Tokugawa shoguns had feared, this opening of the floodgates of Western culture after such prolonged isolation had a traumatic effect on Japanese society. Hypothesis : The Tokugawa shoguns had feared that, because they understood the Japanese society very well. Label : Neutral Category : Fine Premise : In the ancestral environment a man would be likely to have more offspring if he got his pick of the most fertile-seeming women. Hypothesis : Only a man who stayed with one female spread his genes most efficiently. Label : Contradiction Category : Fine Premise : Tommy was suddenly galvanized into life. Hypothesis : Tommy had been downcast for days. Label : Neutral Category : Correct Premise : Improved products and services Initiate actions and manage risks to develop new products and services within or outside the organization. Hypothesis : Managed risks lead to new products Label : Entailment Category : Fine Premise : Coast Guard rules establishing bridgeopening schedules). Hypothesis : The Coast Guard is in charge of opening bridges. Label : Entailment Category : Correct Premise : The anthropologist Napoleon Chagnon has shown that Yanomamo men who have killed other men have more wives and more offspring than average guys. Hypothesis : Yanomamo men who kill other men have better chances at getting more wives. Label : Entailment Category : Fine Premise : The Varanasi Hindu University has an Art Museum with a superb collection of 16th-century Mughal miniatures, considered superior to the national collection in Delhi. Hypothesis : The Varanasi Hindu University has an art museum on its campus which may be superior objectively to the national collection in Delhi. Label : Entailment Category : Correct Premise : Part of the reason for the difference in pieces per possible delivery may be due to the fact that five percent of possible residential deliveries are businesses, and it is thought, but not known, that a lesser percentage of possible deliveries on rural routes are businesses. Hypothesis : We all know that the reason for a lesser percentage of possible deliveries on rural routes being businesses, is because of the fact that people prefer living in cities rather than rural areas. Label : Neutral Category : Correct Premise : right oh they’ve really done uh good job of keeping everybody informed of what’s going on sometimes i’ve wondered if it wasn’t almost more than we needed to know Hypothesis : I don’t think I have shared enough information with everyone. Label : Contradiction Category : Correct Premise : To reach any of the three Carbet falls, you must continue walking after the roads come to an end for 20 minutes, 30 minutes, or two hours respectively. Hypothesis : There are three routes to the three Carbet falls, each a different length and all continue after the road seemingly ends. Label : Entailment Category : Correct Premise : But when the cushion is spent in a year or two, or when the next recession arrives, the disintermediating voters will find themselves playing the roles of budget analysts and tax wonks. Hypothesis : The cushion will likely be spent in under two years. Label : Entailment Category : Correct Premise : But, Slate protests, it was [Gates’] byline that appeared on the cover. Hypothesis : Slate was one hundred percent positive it was Gates’ byline on the cover. Label : Neutral Category : Correct Premise : But it’s for us to get busy and do something.” Hypothesis : ”We don’t do much, so maybe this would be good for us to bond and be together for the first time in a while.”. Label : Neutral Category : Fine Premise : Pearl Jam detractors still can’t stand singer Eddie They say he’s unbearably self-important and limits the group’s appeal by refusing to sell out and make videos. Hypothesis : A lot of people consider Eddie to be a bad singer. Label : Neutral Category : Correct Premise : it’s the very same type of paint and everything Hypothesis : It’s the same paint formula, it’s great! Label : Entailment Category : Fine Premise : Exhibit 3 presents total national emissions of NOx and SO2 from all sectors, including power. Hypothesis : In Exhibit 3 there are the total regional emissions od NOx and SO2 from all sectors. Label : Entailment Category : Correct Premise : uh-huh and is it true i mean is it um Hypothesis : It’s true. Label : Entailment Category : Wrong Premise : When a GAGAS attestation engagement is the basis for an auditor’s subsequent report under the AICPA standards, it would be advantageous to users of the subsequent report for the auditor’s report to include the information on compliance with laws and regulations and internal control that is required by GAGAS but not required by AICPA standards. Hypothesis : The report is required by GAGAS but not AICPA. Label : Entailment Category : Correct Premise : i’m on i’m in the Plano school system and living in Richardson and there is a real dichotomy in terms of educational and economic background of the kids that are going to be attending this school Hypothesis : The Plano school system only has children with poor intelligence. Label : Contradiction Category : Correct
J.2 QQP
QQP is the abbreviation of Quora Question Pairshttps://www.quora.com/q/quoradata/First-Quora-Dataset-Release-Question-Pairs. Given two questions, the model needs to tell whether they have the same meaning (i.e. Paraphrase/Non-paraphrase). Question 1 : Which universities for MS in CS should I apply to? Question 2 : Which universities should I apply to for an MS in CS? Label : Paraphrase Category : Correct Question 1 : What should I do to make life worth living? Question 2 : What makes life worth living? Label : Paraphrase Category : Fine Question 1 : Why did Quora remove my question? Question 2 : Why does Quora remove questions? Label : Paraphrase Category : Correct Question 1 : How do I get thousands of followers on Instagram? Question 2 : How can I get free 10k real Instagram followers fast? Label : Paraphrase Category : Fine Question 1 : What is the basic knowledge of computer science engineers? Question 2 : What is basic syallbus of computer science engineering? Label : Non-paraphrase Category : Fine Question 1 : How many mosquito bites does it take to kill a human being? Question 2 : How many times can a single mosquito bite a human within 8 hours? Label : Non-paraphrase Category : Correct Question 1 : How does it feel to become attractive from unattractive? Question 2 : What does it feel like to go from physically unattractive to physically attractive? Label : Paraphrase Category : Correct Question 1 : Who is answering the questions asked on Quora? Question 2 : Who can answer the questions asked on Quora? Label : Paraphrase Category : Correct Question 1 : What machine learning theory do I need to know in order to be a successful machine learning practitioner? Question 2 : What do I need to know to learn machine learning? Label : Paraphrase Category : Wrong Question 1 : If you could go back in time and change one thing, what would it be and why? Question 2 : If you could go back in time and do one thing, what would it be? Label : Paraphrase Category : Correct Question 1 : Will there be a civil war if Trump doesn’t become president? Question 2 : Will there be a second civil war if Trump becomes president? Label : Paraphrase Category : Correct Question 1 : Do Quora contributors get paid? Question 2 : How do contributors get paid by Quora? Label : Paraphrase Category : Correct Question 1 : Did India meet Abdul Kalam’s 2020 vision so far? Question 2 : How far do you think India has reached on President APJ Kalam’s vision in the book India 2020? Label : Non-paraphrase Category : Correct Question 1 : How do I stop my dog from whining after getting spayed? Question 2 : How do I stop my dog from whining? Label : Paraphrase Category : Wrong Question 1 : What difference are exactly between Euclidean space and non Euclidean space? Question 2 : What is the difference between Euclidean and non-Euclidean? Label : Non-paraphrase Category : Wrong Question 1 : Why doesn’t Hillary Clinton win the White House if she won the popular vote? Question 2 : How did Hillary Clinton win the popular vote but Donald Trump win the election? Label : Paraphrase Category : Correct Question 1 : How is public breastfeeding seen where you live? Question 2 : How is breastfeeding in public seen in your country? Label : Paraphrase Category : Correct Question 1 : What are some ways to change your Netflix password? Question 2 : How do you change your Netflix password and email? Label : Paraphrase Category : Fine Question 1 : What do you think, is your best answer on Quora? Question 2 : What is your best answer on Quora? Label : Paraphrase Category : Correct Question 1 : How can I travel to Mexico without a passport? Question 2 : Can I travel to Mexico without a passport? Label : Paraphrase Category : Correct Question 1 : How do modern Congolese people view Mobutu in retrospect? Question 2 : How do Congolese currently view Mobutu Sese Seko? Label : Non-paraphrase Category : Correct Question 1 : How is Tanmay Bhat losing weight? Question 2 : Tanmay Bhat: How did you manage to reduce your fat? Label : Non-paraphrase Category : Unsure Question 1 : Is Xiaomi a brand to trust (comparing it with brands like Samsung and HTC)? What is better: Xiaomi MI3 or HTC Desire 816? Question 2 : Is xiaomi a trusted brand? Label : Non-paraphrase Category : Correct Question 1 : Why did Buddhism spread in East Asia and not in its native land India? Question 2 : How was Buddhism spread in Asia? Label : Non-paraphrase Category : Correct Question 1 : Can I become a multi billionaire betting on horses? Question 2 : How much money can I make betting on horses? A month? Can I make 20,000 a month? Label : Paraphrase Category : Fine Question 1 : What is a diet for gaining weight? Question 2 : What is a way to gain weight? Label : Non-paraphrase Category : Correct Question 1 : How do I use Instagram on my computer? Question 2 : How can I get Instagram on my computer? Label : Paraphrase Category : Fine Question 1 : What is the legal basis of a ”you break it, you buy it” policy? Question 2 : Is a ”you break it you buy it” policy actually legal? Label : Paraphrase Category : Correct Question 1 : How I should fix my computer while it is showing no boot device found? Question 2 : How do I fix the ”Boot device not found” problem? Label : Paraphrase Category : Correct Question 1 : What innovative name can I use for an interior designing firm? Question 2 : What can i name my interior designing firm? Label : Paraphrase Category : Correct Question 1 : What would it realistically cost to go to Tomorrowland? Question 2 : How much is a ticket to Tomorrowland? Label : Non-paraphrase Category : Fine Question 1 : Is there a gender pay gap? If so why? Question 2 : Is the gender pay gap a myth? Label : Paraphrase Category : Correct Question 1 : How can I get rid of a canker sore on the bottom of my tongue? Question 2 : How can I get rid or a canker sore on the tip of my tongue? Label : Paraphrase Category : Correct Question 1 : How can I sleep better and early in night? Question 2 : How can I sleep better at night? Label : Paraphrase Category : Fine Question 1 : Why did DC Change Captain Marvel’s name? Question 2 : Why did DC have to change Captain Marvel’s name but Marvel didn’t have to change Scarecrow’s name? Label : Paraphrase Category : Fine Question 1 : Should be there any difference between IIT and non IIT students in terms of placement package from a company if both of them are equally talented? Question 2 : Should be there any difference between IIT and non IIT students in terms of placement package from a company if both of them are equally capable? Label : Paraphrase Category : Correct Question 1 : What happened to The Joker after The end of The Dark Knight? Question 2 : What happens to the Joker at the end of The Dark Knight (2008 movie)? Label : Non-paraphrase Category : Wrong Question 1 : I love my wife more then anything. Why do I fantasize about her with other men? Question 2 : Why do I fantasize about other men having sex with my wife? Label : Paraphrase Category : Fine Question 1 : What is your opinion on the new MacBook Pro Touch Bar? Question 2 : What do you think about the OLED touch bar on the new MacBook Pro? Label : Paraphrase Category : Correct Question 1 : How do I get rid of my negative alter ego? Question 2 : How do you get rid of your negative alter ego? Label : Paraphrase Category : Correct Question 1 : How can I get wifi driver for my hp laptop with windows 7 os? Question 2 : How can I get wifi driver for my laptop with windows 7 os? Label : Paraphrase Category : Fine Question 1 : What’s your attitude towards life? Question 2 : What should be your attitude towards life? Label : Paraphrase Category : Wrong Question 1 : What books would I like if I loved A Song of Ice and Fire? Question 2 : Are there books which are similar to A Song of Ice and Fire? Label : Paraphrase Category : Fine Question 1 : Why do Muslims think they will conquer the whole world? Question 2 : Do you think Muslims will take over the world? Label : Non-paraphrase Category : Correct Question 1 : Is dark matter a sea of massive dark photons that ripple when galaxy clusters collide and wave in a double slit experiment? Question 2 : Does a superfluid dark matter which ripples when Galaxy clusters collide and waves in a double slit experiment relate GR and QM? Label : Paraphrase Category : Correct Question 1 : What is Batman like? Question 2 : What is Batman’s personality like? Label : Non-paraphrase Category : Correct
J.3 SST-2
SST-2 is the abbreviation of Stanford Sentiment Treebank Socher et al. (2013). In this task, the model needs to recognize whether the phrases or sentences reflect positive or negative sentiments. Input : predictability is the only winner Label : Negative Category : Correct Input : abandon their scripts and go where the moment takes them Label : Negative Category : Correct Input : chases for an hour and then Label : Positive Category : Unsure Input : provide much more insight than the inside column of a torn book jacket Label : Negative Category : Unsure Input : a children ’s party clown Label : Negative Category : Fine Input : perhaps even the slc high command found writer-director mitch davis ’s wall of kitsch hard going . Label : Negative Category : Correct Input : own placid way Label : Negative Category : Correct Input : get on a board and , uh , shred , Label : Negative Category : Correct Input : asks what truth can be discerned from non-firsthand experience , and specifically questions cinema ’s capability for recording truth . Label : Positive Category : Correct Input : puts the dutiful efforts of more disciplined grade-grubbers Label : Positive Category : Correct Input : filter out the complexity Label : Positive Category : Correct Input : told what actually happened as if it were the third ending of clue Label : Negative Category : Correct Input : is more in love with strangeness than excellence . Label : Positive Category : Wrong Input : i found myself howling more than cringing Label : Positive Category : Correct Input : goldbacher draws on an elegant visual sense and a talent for easy , seductive pacing … but she and writing partner laurence coriat do n’t manage an equally assured narrative coinage Label : Positive Category : Unsure Input : for a thirteen-year-old ’s book report Label : Negative Category : Correct Input : a problem hollywood too long has ignored Label : Negative Category : Correct Input : twisted sense Label : Negative Category : Correct Input : a stab at soccer hooliganism Label : Negative Category : Correct Input : sinuously plotted Label : Negative Category : Correct Input : shiner can certainly go the distance , but is n’t world championship material Label : Positive Category : Correct Input : holding equilibrium up Label : Negative Category : Wrong Input : i am highly amused by the idea that we have come to a point in society where it has been deemed important enough to make a film in which someone has to be hired to portray richard dawson . Label : Positive Category : Wrong Input : waters Label : Negative Category : Wrong Input : what might have emerged as hilarious lunacy in the hands of woody allen or Label : Positive Category : Correct Input : of those airy cinematic bon bons whose aims – and by extension , accomplishments – seem deceptively slight on the surface Label : Positive Category : Correct Input : do n’t blame eddie murphy but Label : Negative Category : Correct Input : melodramatic paranormal romance Label : Negative Category : Correct Input : could possibly be more contemptuous of the single female population Label : Negative Category : Correct Input : cremaster 3 ” should come with the warning “ for serious film buffs only ! ” Label : Negative Category : Correct Input : softheaded metaphysical claptrap Label : Negative Category : Correct Input : owed to benigni Label : Negative Category : Unsure Input : to be a suspenseful horror movie or a weepy melodrama Label : Positive Category : Correct Input : genuinely unnerving . Label : Positive Category : Correct Input : gaping enough to pilot an entire olympic swim team through Label : Negative Category : Correct Input : this is popcorn movie fun with equal doses of action , cheese , ham and cheek ( as well as a serious debt to the road warrior ) , but it feels like unrealized potential Label : Positive Category : Fine Input : feeling like it was worth your seven bucks , even though it does turn out to be a bit of a cheat in the end Label : Negative Category : Correct Input : pull it back Label : Negative Category : Correct Input : , this is more appetizing than a side dish of asparagus . Label : Negative Category : Correct Input : crime drama Label : Negative Category : Unsure Input : like most movies about the pitfalls of bad behavior Label : Negative Category : Fine Input : befallen every other carmen before her Label : Positive Category : Unsure Input : appeal to those without much interest in the elizabethans ( as well as rank frustration from those in the know about rubbo ’s dumbed-down tactics ) Label : Negative Category : Unsure Input : about existential suffering Label : Negative Category : Fine Input : , if uneven , Label : Negative Category : Unsure Input : succumbs to sensationalism Label : Positive Category : Wrong Input : that turns me into that annoying specimen of humanity that i usually dread encountering the most Label : Negative Category : Fine Input : at least a minimal appreciation Label : Positive Category : Unsure Input : underlines even the dullest tangents Label : Negative Category : Correct Input : heard before Label : Positive Category : Unsure Input : i like my christmas movies with more elves and snow and less pimps and ho ’s . Label : Negative Category : Unsure Input : can aspire but none can equal Label : Negative Category : Unsure Input : fathom Label : Negative Category : Unsure Input : attempt to bring cohesion to pamela ’s emotional roller coaster life Label : Negative Category : Unsure Input : movie version Label : Positive Category : Wrong Input : of spontaneity in its execution and a dearth of real poignancy Label : Positive Category : Correct