Dataset Cartography: Mapping and Diagnosing Datasets with Training Dynamics

Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, Yejin Choi

Introduction

The creation of large labeled datasets has fueled the advance of AI Russakovsky et al. (2015); Antol et al. (2015) and NLP in particular Bowman et al. (2015); Rajpurkar et al. (2016). The common belief is that the more abundant the labeled data, the higher the likelihood of learning diverse phenomena, which in turn leads to models that generalize well. In practice, however, out-of-distribution (OOD) generalization remains a challenge Yogatama et al. (2019); Linzen (2020); and, while recent large pretrained language models help, they fail to close this gap Hendrycks et al. (2020). This urges a closer look at datasets, where not all examples might contribute equally towards learning Vodrahalli et al. (2018). However, the scale of data can make this assessment challenging. How can we automatically characterize data instances with respect to their role in achieving good performance in- and out-of- distribution? Answering this question may take us a step closer to bridging the gap between dataset collection and broader task objectives.

Drawing analogies from cartography, we propose to find coordinates for instances within the broader trends of a dataset. We introduce data maps: a model-based tool for contextualizing examples in a dataset. We construct coordinates for data maps by leveraging training dynamics—the behavior of a model as training progresses. We consider the mean and standard deviation of the gold label probabilities, predicted for each example across training epochs; these are referred to as confidence\mathbf{confidence} and variability\mathbf{variability}, respectively (§2).

Fig. 1 shows the data map for the SNLI dataset Bowman et al. (2015) constructed using the RoBERTa-large model Liu et al. (2019). The map reveals three distinct regions in the dataset: a region with instances whose true class probabilities fluctuate frequently during training (high variability\mathbf{variability}), and are hence ambiguous for the model; a region with easy-to-learn instances that the model predicts correctly and consistently (high confidence\mathbf{confidence}, low variability\mathbf{variability}); and a region with hard-to-learn instances with low confidence\mathbf{confidence}, low variability\mathbf{variability}, many of which we find are mislabeled during annotation .All terms are defined with respect to the model. Similar regions are observed across three other datasets: MultiNLI Williams et al. (2018), WinoGrande Sakaguchi et al. (2020) and SQuAD Rajpurkar et al. (2016), with respect to respective RoBERTa-large classifiers.

We further investigate the above regions by training models exclusively on examples from each region (§3). Training on ambiguous instances promotes generalization to OOD test sets, with little or no effect on in-distribution (ID) performance.We define out-of-distribution (OOD) test sets as those which are collected independently of the original dataset, and ID test sets as those which are sampled from it. Our data maps also reveal that datasets contain a majority of easy-to-learn instances, which are not as critical for ID or OOD performance, but without any such instances, training could fail to converge (§4). In §5, we show that hard-to-learn instances frequently correspond to labeling errors. Lastly, we discuss connections between our measures and uncertainty measures (§6).

Our findings indicate that data maps could serve as effective tools to diagnose large datasets, at the reasonable cost of training a model on them. Locating different regions within the data might pave the way for constructing higher quality datasets., and ultimately models that generalize better. Our code and higher resolution visualizations are publicly available.https://github.com/allenai/cartography

Mapping Datasets with Training Dynamics

Our goal is to construct Data Maps for datasets to help visualize a dataset with respect to a model, as well as understand the contributions of different groups of instances towards that model’s learning. Intuitively, instances that a model always predicts correctly are different from those it almost never does, or those on which it vacillates. For building such maps, each instance in the dataset must be contextualized in the larger set. We consider one contextualization approach, based on statistics arising from the behavior of the training procedure across time, or the “training dynamics”. We formally define our notations (§2.1) and describe our data maps (§2.2).

Consider a training dataset of size NN, D={(x,y∗)i}i=1N\mathcal{D}=\{(\boldsymbol{x},y^{*})_{i}\}_{i=1}^{N} where the iith instance consists of the observation, xi\boldsymbol{x}_{i} and its true label under the task, yi∗y^{*}_{i}. Our method assumes a particular model (family) whose parameters are selected to minimize empirical risk using a particular algorithm.In this paper, the model is RoBERTa Liu et al. (2019), currently established as a strong performer across many tasks. We assume the model defines a probability distribution over labels given an observation. We assume a stochastic gradient-based optimization procedure is used, with training instances randomly ordered at each epoch, across EE epochs.

The training dynamics of instance ii are defined as statistics calculated across the EE epochs. The values of these measures then serve as coordinates in our map. The first measure aims to capture how confidently the learner assigns the true label to the observation, based on its probability distribution. We define confidence\mathbf{confidence} as the mean model probability of the true label (yi∗y^{*}_{i}) across epochs:

where pθ(e)p_{\boldsymbol{\theta}^{(e)}} denotes the model’s probability with parameters θ(e)\boldsymbol{\theta}^{(e)} at the end of the ethe^{\text{th}} epoch.Note that μ^i\hat{\mu}_{i} is with respect to the true label yi∗y^{*}_{i}, not the probability assigned to the model’s highest-scoring label (as used in active learning, for example). In some cases we also consider a coarser, and perhaps more intuitive statistic, the fraction of times the model correctly labels xi\boldsymbol{x}_{i} across epochs, named correctness\mathbf{correctness}; this score only has 1+E1+E possible values. Intuitively, a high-confidence\mathbf{confidence} instance is “easier” for the given learner.

Lastly, we also consider variability\mathbf{variability}, which measures the spread of pθ(e)(yi∗∣xi)p_{\boldsymbol{\theta}^{(e)}}(y^{*}_{i}\mid\boldsymbol{x}_{i}) across epochs, using the standard deviation:

Note that variability\mathbf{variability} also depends on the gold label, yi∗y^{*}_{i}. A given instance to which the model assigns the same label consistently (whether accurately or not) will have low variability\mathbf{variability}; one which the model is indecisive about across training, will have high variability\mathbf{variability}.

Finally, we observe that confidence\mathbf{confidence} and variability\mathbf{variability} are fairly stable across different parameter initializations.The average Pearson correlation coefficient between five random seeds’ resulting training runs is 0.75 or higher (for both measures, on WinoGrande). Training dynamics can be computed at different granularities, such as steps vs. epochs; see App. A.1.

2 Data Maps

We construct data maps for four large datasets: WinoGrande Sakaguchi et al. (2020)—a cloze-style task for commonsense reasoning, two NLI datasets (SNLI; Bowman et al., 2015; and MultiNLI; Williams et al., 2018), and QNLI, which is a sentence-level question answering task derived from SQuAD Rajpurkar et al. (2016). All data maps are built with models based on RoBERTa-large architectures. Details on the model and datasets can be found in App. §A.2 and §A.3.

Fig. 1 presents the data map for the SNLI dataset. As is evident, the data follows a bell-shaped curve with respect to confidence\mathbf{confidence} and variability\mathbf{variability}; correctness\mathbf{correctness} further determines discrete regions therein. The vast majority of instances belong to the high confidence\mathbf{confidence} and low variability\mathbf{variability} region of the map (Fig. 1, top-left). The model consistently predicts such instances correctly with high confidence; thus, we refer to them as easy-to-learn (for the model). A second, smaller group is formed by instances with low variability\mathbf{variability} and low confidence\mathbf{confidence} (Fig. 1, bottom-left corner). Since such instances are seldom predicted correctly during training, we refer to them as hard-to-learn (for the model). The third notable group contains ambiguous examples, or those with high variability\mathbf{variability} (Fig. 1, right-hand side); the model tends to be indecisive about these instances, such that they may or may not correspond to high confidence\mathbf{confidence} or correctness\mathbf{correctness}. We refer to such instances as ambiguous (to the model).

Fig. 2 shows the data map for WinoGrande, which exhibits high structural similarity to the SNLI data map (Fig. 1). The most remarkable difference between the maps is in the density of the hard-to-learn region, which is much lower for WinoGrande, as is evident from the histograms below. One explanation for this might be that WinoGrande labels were rigorously validated post annotation. App. §C includes data maps for all four datasets, with respect to RoBERTa-large, in greater relief.

Different model architectures trained on a given dataset could be effectively compared using data maps, as an alternative to standard quantitative evaluation methods. App. §C includes data maps for WinoGrande (Fig. 9(b)) and SNLI (Fig. 10 and Fig. 11) based on other (somewhat weaker) architectures. While data maps based on similar architectures have similar appearance, the regions to which a given instance belongs might vary. Data maps for weaker architectures still display similar regions, but the regions are not as distinct as those in RoBERTa based data maps.

Tab. 1 shows examples from WinoGrande belonging to the different regions defined above. easy-to-learn examples are straightforward for the model, as well as for humans. In contrast, most hard-to-learn and some ambiguous examples could be challenging for humans (see green highlights in Tab. 1), which might explain why the model shows lower confidence\mathbf{confidence} on them. These categories could be harder for models either because of labeling errors (blue highlights) or simply because the model is indecisive about the correct label. See App. §A.4 for similar examples from SNLI.

The next four sections include a diagnosis of the different data regions defined above. The effect of training models on each region on both in- and out-of-distribution performance is studied in §3. The effect of selecting decreasing amounts of data is discussed in §4. We investigate the presence of mislabeled instances in the hard-to-learn regions of the data maps in §5. Lastly, we demonstrate connections between training dynamics measures and measures of uncertainty in §6.

Data Selection using Data Maps

Data maps reveal distinct regions in datasets; it is natural to wonder what roles do instances from different regions play in learning and generalization. We answer this empirically by training models exclusively on instances selected from distinct regions, followed by standard in-distribution (ID), as well as out-of-distribution (OOD) evaluation.

Our strategy is straightforward—we train the model from scratch on a subset of the training data selected by ranking instances based on the different training dynamics measures.Hyperparameters are also tuned from scratch (App. §A.3). We hypothesize that ambiguous and hard-to-learn regions could be the most informative for learning, since these examples are the most challenging for the model Shrivastava et al. (2016). We compare these two settings to RoBERTa-large models trained on data subsets, selected using several other methods. All subsets considered contain 33% of the training data (to control for the effect of train data size on performance).

The two most natural baselines are those where all data is used (100% train), and where a 33% random sample is used (random). Our second set of baselines considers subsets which are the most easy-to-learn for the model (high-confidence\mathbf{confidence}), and those that the model is most decisive about (low-variability\mathbf{variability}), which comprises a mixture of easy-to-learn and hard-to-learn examples. We also consider baselines based on high- and low-correctness\mathbf{correctness}. Finally, we also compare with our implementation of the following methods for data selection from prior work (discussed in §7): forgetting Toneva et al. (2018), AFLite LeBras et al. (2020), AL-uncertainty Joshi et al. (2009), and AL-greedyK Sener and Savarese (2018).

Results

We test our selections on the same datasets from the previous section—WinoGrande, SNLI, MultiNLI and QNLI. We report ID validation performance, and OOD performance on test sets either created independently of the dataset (NLI Diagnostics Wang et al. (2019) for SNLI and MultiNLI, and WSC Levesque et al. (2011) for WinoGrande), or specifically to be adversarial to the dataset (Adversarial SQuAD Jia and Liang (2017) for QNLI); see App. §A.2 for details.

Tab. 2 shows our results on WinoGrande.The official test set for WinoGrande has been filtered with AFLite, making ID evaluation more challenging than OOD. However, we apply all our selection methods (including the AFLite selection) on WinoGrande’s unfiltered training data. Training on the most ambiguous data results in the best OOD performance, exceeding that of 100% train, even with just a third of the data. A similar effect is seen with hard-to-learn, as well as its coarse-grained counterpart, low-correctness\mathbf{correctness}. In each of the three cases, ID performance is also higher than all other 33% baselines, though we observe some degradation compared to the full training set; this is expected as with larger amounts of data models tend to fit the dataset distribution rather than the task Torralba and Efros (2011). The only selection methods that underperform the random baseline are forgetting, and the ones where we select data the model is highly confident and decisive about (high-confidence\mathbf{confidence}, high-correctness\mathbf{correctness}, and low-variability\mathbf{variability}). The latter pattern, as well as our overall results, highlight the important role played by examples which are challenging for the model, i.e., ambiguous and hard-to-learn examples.

Given that our selection methods outperform baselines from prior work, we only report random and 100% train selection baselines on the remaining datasets, where we see similar trends. Tab. 3 shows results for SNLI and MultiNLI, where the random selection baseline is already within 1% of the 100% train baseline.The ID performance of all models exceeds human accuracy (88%) for SNLI. However, the difference in ID and OOD performance in SNLI is quite high, showing that there is still room for improvement in the NLI task. Selecting 33% of the most ambiguous examples achieves even better ID performance, within 0.2% of the 100% train baseline, while exceeding OOD performance substantially on each of the linguistic categories in the NLI Diagnostics test set.While MultiNLI-mismatched is technically out-of-domain, performance is close to matched (ID). While hard-to-learn does not perform as well as ambiguous on most cases, it still matches or outperforms the 100% train baseline on OOD test sets. Tab. 4 shows a similar trend for QNLI, where we gain over 2% performance on the OOD Adversarial SQuAD test set, with minimal loss in ID accuracy.

Overall, regions revealed by data maps provide ways to substantially improve OOD performance across datasets. Regional selections of data not only improve model generalization, but also do so using substantially less data, providing a method to potentially speed up training. We note, however, that discovering such examples requires computing training dynamics, which involves training a model on the full dataset. Future directions involve building more efficient data maps, to better fulfill the training speedup potential.

Role of Easy-to-Learn Instances

Data maps uncover ambiguous regions, small subsets from which lead to improved OOD performance, with minimal degradation of ID performance (§3). We next investigate how performance is affected as we vary the size of the ambiguous subsets. We retrain our model with subsets containing the top 50%, 33%, 25%, 17%, 10%, 5% and 1% ambiguous instances of WinoGrande (Fig. 3, left and center). Large ambiguous subsets (25% or more) result in high ID and OOD performance. Surprisingly however, for smaller ambiguous subsets (17% or less), the model performs at chance level, despite random restarts.This is common for large models trained on small datasets Devlin et al. (2019); Phang et al. (2018); Dodge et al. (2020). In contrast, a baseline that randomly selects subsets of similar sizes is able to learn (while naturally performing worse as data decreases, eventually failing at 1%). This indicates that ambiguous instances alone might be insufficient for learning.

Given that the model barely struggles with easy-to-learn instances (by definition), we next replace some ambiguous examples with easy-to-learn examples in the 17% most ambiguous subset. Interestingly, replacing just a tenth of the ambiguous data with easy-to-learn instances, the model not only successfully learns, but also outperforms the random selection baseline’s ID performance (Fig. 3 right). This indicates that for successful optimization, it is important to include easier-to-learn instances. However, with too many replacements, performance starts decreasing again; this trend was seen in the previous section with the high-confidence\mathbf{confidence} baseline (Tab. 2). OOD performance shows a similar trend, but matches or is worse than the baseline. Selection of the optimal balance of easy-to-learn and ambiguous examples in low data regimes is an open problem; we defer this exploration to future work.

Detecting Mislabeled Examples

Crowdsourced datasets are often subject to noise attributed to incorrect labeled annotations Sheng et al. (2008); Krishna et al. (2016); Ekambaram et al. (2017), which may lead to training models that are not representative of the task at hand. Recent studies have shown that over-parameterized neural networks can fit the incorrect labels blindly Zhang et al. (2017), which might hurt their generalization ability Hu et al. (2020). For large datasets, identifying mislabeled examples can be prohibitively expensive. Our data maps provide a semi-automated method to identify such mislabeled instances, without significantly more effort than simply training a model on the data. We hypothesize that hard-to-learn examples—those with low confidence\mathbf{confidence}—might be mislabeled, as has also been suggested in prior work Manning (2011); Toneva et al. (2018).

To verify this hypothesis, we design an experimental setting where we artificially inject noise in the training data, by flipping the labels of 1% of the training data for WinoGrande. Motivated by our qualitative analysis (Tab 1), we select the candidates for flipping from the easy-to-learn region—this minimizes the risk of selecting already mislabeled examples. We retrain RoBERTa with the partly noised data, and recompute confidence\mathbf{confidence} and variability\mathbf{variability} of all instances. Fig. 4 shows the training dynamics measures, before and after re-training. Flipped instances move to the lower confidence\mathbf{confidence} regions after retraining, with some movement towards higher variability\mathbf{variability}. This indicates that perhaps the hard-to-learn region (low confidence\mathbf{confidence}) of the map contains other mislabeled instances. We next explore a simple method to automatically detect such instances.

We train a linear model to classify examples as mislabeled (noise) or not, using a single feature: the confidence\mathbf{confidence} score from the retrained RoBERTa model on WinoGrande. This model is trained using a balanced training set for this task by sampling equal numbers of noisy (label-flipped) and clean examples from the original train set. This simple classifier is quite effective—a sanity check evaluation on a similarly constructed test set yields 100% F1.A similar experiment with only variability\mathbf{variability} scores as features resulted in a much poorer classifier—70% F1.

Next, we run the trained noise classifier on the entire original training set, with features extracted from the original training dynamics measures (computed without added noise). We first observe that despite training on a balanced dataset, our classifier predicts only a few examples as mislabeled—only 31 WinoGrande instances (out a total of 40K). A similar experiment on SNLI results in 15K noisy examples (out of 500K). These results are encouraging and follow our intuitions that most instances in data are indeed labeled “correctly”. Indeed, WinoGrande contains a lower portion of noisy examples, as indicated by our data maps (Fig. 2).

We further investigated these trends via a human evaluation on the output of the classifier. We created an evaluation set by randomly selecting 50 instances from each predicted class as per our classifier. Two of the authors re-annotated these 100 instances (without access to the original or predicted labels); some instances were annotated as too ambiguous for the authors. After discussions to resolve their differences, both annotators agreed on 96% of the instances in each dataset. Using our annotations as a new gold standard, we found that for WinoGrande, 67% of the instances predicted as noisy by the linear classifier are indeed either mislabeled or ambiguous, compared to only 13% of the ones predicted as correctly labeled. Similar patterns are observed for SNLI (76% vs. 4%).

Our results demonstrate the potential of using data maps as a tool to “clean-up” datasets, by identifying mislabeled or ambiguous instances.In preliminary experiments, retraining WinoGrande after removal of noise did not yield a large difference in performance, given the relatively small amount of noise. Notably, our results were obtained using a simple method; this encourages exploration of methods that might lead to more accurate noise-detectors.

Training Dynamics as Uncertainty Measures

We introduced data maps, and used training dynamics measures as coordinates for data points in §2. We now take a closer look at these measures, and find intuitive connections with measures of uncertainty. When a model fails to predict the correct label, the error may come from ambiguity inherent to the example (intrinsic uncertainty), but it may also come from the model’s limitations (often referred to as model uncertainty).These are also sometimes called the aleatoric and epistemic uncertainty, respectively Gal (2016). To understand how examples contribute to a dataset, it is important to separate these two sources of error.

We start by studying the relationship between intrinsic uncertainty and our training dynamics measures. Human agreement can serve as a proxy for intrinsic uncertainty. We estimate human agreement using the multiple human annotations available in SNLI’s development set.Only SNLI dev. and test set have multiple annotations on all instances. We obtain training dynamics with RoBERTa-large run on train and dev. combined. For each annotator, we compute whether they agree with the majority label from the other four, breaking ties randomly and then averaging over annotators.Normally, this provides the minimum-variance unbiased estimate, though SNLI’s development set throws away examples without a majority, which introduces some bias.,Note that the model only has seen the majority vote, while we take into account all annotator labels to quantify agreement.

Fig. 5 visualizes the relation between our training dynamics measures (confidence\mathbf{confidence} and variability\mathbf{variability}) and human agreement, averaged over the examples. We observe a strong relationship between human agreement and confidence\mathbf{confidence}: high confidence\mathbf{confidence} indicates high agreement between annotators, while low confidence\mathbf{confidence} often indicates disagreement on the example. In contrast, once confidence\mathbf{confidence} is known, variability\mathbf{variability} does not provide much information about the agreement.

The connection between our second measure, variability\mathbf{variability}, and model uncertainty is more straightforward: variability\mathbf{variability}, by definition, captures exactly the uncertainty of the model. See App. B.1 for an additional discussion (with empirical justifications) on connections between training dynamics measures and dropout-based Srivastava et al. (2014), first-principles uncertainty estimates.

These relations are further supported by previous work, which showed that deep ensembles provide well-calibrated uncertainty estimates Lakshminarayanan et al. (2017); Gustafsson et al. (2019); Snoek et al. (2019). Generally, such approaches ensemble models trained from scratch; while ensembles of training checkpoints lose some diversity Fort et al. (2019), they offer a cheaper alternative capturing some of the benefits Chen et al. (2017a). Future work will involve investigation of such alternatives for building data maps.

Related Work

Our work builds data maps using training dynamics measures for scoring data instances. Loss landscapes Xing et al. (2018) are similar to training dynamics, but also consider variables from the stochastic optimization algorithm. Toneva et al. (2018) also use training dynamics to find train examples which are frequently “forgotten”, i.e., misclassified during a later epoch of training, despite being classified correctly earlier; our correctness\mathbf{correctness} metric provides similar discrete scores, and results in models with better performance. Variants of such approaches address catastrophic forgetting, and are useful for analyzing data instances Pan et al. (2020); Krymolowski (2002).

Prior work has proposed other criteria to score instances. AFLite LeBras et al. (2020) is an adversarial filtering algorithm which ranks instances based on their “predictability”, i.e. the ability of simple linear classifiers to predict them correctly. While AFLite, among others Li and Vasconcelos (2019); Gururangan et al. (2018), advocate removing “easy” instances from the dataset, our work shows that easy-to-learn instances can be useful. Similar intuitions have guided other work such as curriculum learning Bengio et al. (2009) and self-paced learning Kumar et al. (2010); Lee and Grauman (2011) where all examples are prioritized based on their “difficulty”.

Other approaches have used training loss Han et al. (2018); Arazo et al. (2019); Shen and Sanghavi (2019), confidence Hovy et al. (2013), and meta-learning Ren et al. (2018), to differentiate instances within datasets. Perhaps our measures are the closest to those from Chang et al. (2017); they propose prediction variance and threshold closeness—which correspond to variability\mathbf{variability} and confidence\mathbf{confidence}, respectively.They also consider confidence intervals; our preliminary experiments, with and without, yielded similar results. However, they use these measures to reweight all instances, similar to sampling effective batches in online learning Loshchilov and Hutter (2016). Our work, instead, does a hard selection for the purpose of studying different groups within data.

Our methods are also reminiscent of active learning methods (Settles, 2009; Peris and Casacuberta, 2018; P.V.S and Meyer, 2019), such as uncertainty sampling Lewis and Gale (1994) which selects (unlabeled) data points, which a model trained on a small labeled subset, has least confidence in, or predicts as farthest (in vector space, based on cosine similarity) (Sener and Savarese, 2018; Wolf, 2011). Our approach uses labeled data for selection, similar to core-set selection approaches Wei et al. (2013). Active learning approaches could be used in conjunction with data maps to create better datasets, similar to approaches proposed in Mishra et al. (2020). For instance, creating datasets with more ambiguous examples (with respect to a given model) could make it beneficial for OOD generalization.

Data error detection also involves instance scoring. Influence functions Koh and Liang (2017), forgetting events Toneva et al. (2018), cross validation Chen et al. (2019), Shapely values Ghorbani and Zou (2019), and the area-under-margin metric Pleiss et al. (2020) have all been used to identify mislabeled examples. Some approaches avoid hard examples altogether Bottou et al. (2005); Northcutt et al. (2017) to reduce fit to noisy data. Our use of training dynamics to locate mislabeled examples involves minimal additional effort beyond training a model on the dataset.

Conclusion

We presented data maps: an automatic method to visualize and diagnose large datasets using training dynamics. Our data maps for four different datasets reveal similar terrains in each dataset: groups of ambiguous instances useful for high performance, easy-to-learn instances which aid optimization, and hard-to-learn instances which often correspond to data errors. While our maps are based on RoBERTa-large, the methods to build them are model-agnostic (App. §C.1). Our work shows the effectiveness of simple training dynamics measures based on mean and standard deviation; exploration of more sophisticated measures to build data maps is an exciting future direction. Data maps not only help diagnose and make better use of existing datasets, but also hold potential for guiding the construction of new datasets. Moreover, data maps could facilitate comparison of different model architectures trained on a given dataset, resulting in alternative evaluation methodologies. Our implementation is publicly available to facilitate such efforts.https://github.com/allenai/cartography

Acknowledgements

This research was supported in part by DARPA under the MCS program through NIWC Pacific (N66001-19-2-4031) grant, and by the Allen Distinguished Investigator Award. We thank the anonymous reviewers, and our colleagues from AI2 and UWNLP, especially Ana Marasović, and Suchin Gururangan, for their helpful feedback.

References

Appendix A Supplemental Material

Both confidence\mathbf{confidence} and variability\mathbf{variability} are computed across epochs, but could alternatively be computed over other granularities, e.g. over every few steps of optimization. This might enable more efficient computation of the same. However, care must be taken to ignore the first few steps till optimization stabilizes. In our experiments, we considered all epochs including the first to compute the training dynamics, since the first epoch contains multiple steps of optimization for large training sets.

Moreover, it is possible to stop training early, or before the training converges for computing training dynamics. This early burn-out scheme results in confidence\mathbf{confidence} and variability\mathbf{variability} measures which correlate well with confidence\mathbf{confidence} and variability\mathbf{variability} (see Fig. 6). For our experiments, we use later burn-outs corresponding to model convergence.

A.2 Datasets

This appendix provides further details on datasets. We perform our experimental evaluation on four large datasets, each with at least 10K instances. Sizes of the different datasets are reported in Tab. 5. Instances in each of the original datasets are labeled by crowdworkers, whereas the OOD test sets are either manually or semi-automatically created. The performance in each case is reported as accuracy.

This dataset contains a large scale crowd-sourced collection of Winograd schema challenge (WSC Levesque et al., 2011) style questions. Commonsense reasoning is required to select an entity from a pair of entities to complete a sentence. Following Sakaguchi et al. (2020), we use the multiple choice architecture based on RoBERTa Liu et al. (2019). For OOD evaluation, we use the validation set from the original WSC as provided under the SuperGLUE benchmark Wang et al. (2019). We used a rule-based method to convert WSC validation and training data to the cloze-style format followed in WinoGrande, removing all the repetitions included in the training data. Figure 2 shows the data map for WinoGrande.

SNLI and MultiNLI

The task of natural language inference involves prediction of the relationship between a premise and hypothesis sentence pair. The label determines whether the hypothesis entails, contradicts or is neutral to the premise. We experiment with the Stanford natural language inference (SNLI) dataset Bowman et al. (2015) and its multi-genre counterpart, MultiNLI Williams et al. (2018). For MultiNLI, we use the version released under the GLUE benchmark Wang et al. (2018). Several challenge sets have been proposed to evaluate models OOD. As an OOD test set, we consider NLI diagnostics Wang et al. (2018) which contains a set of hand-crafted examples designed to demonstrate NLI model performance on several fine-grained semantic categories, such as lexical semantics, logical reasoning, predicate argument structure and commonsense knowledge. In addition, we also report performance on the OOD mismatched MultiNLI validation set. Figure 8(a) shows the data map for MultiNLI.

QNLI

Rajpurkar et al. (2016) proposed the SQuAD dataset containing question and document pairs, where the answer to the question is a span in the document. The QNLI dataset, provided as part of the GLUE benchmark Wang et al. (2018) reformulates this as a sentence-level binary classification task. Here, the goal is to determine if a candidate sentence from the document contains the answer to the given question. As an OOD test set, we consider the Adversarial SQuAD challenge set Jia and Liang (2017) where distractor sentences are added to the document to confound the model. We automatically convert this to the QNLI format. Figure 9(a) shows the data map for QNLI.

A.3 Experimental Settings

For each of our classifiers, we minimize cross entropy with the Adam optimizer Kingma and Ba (2014) following the AdamW learning rate schedule from PyTorchpytorch.org. Each experiment is run with 3 random seeds and a learning rateLearning rate is chosen using a log-uniform sampling strategy from the range (5e-6, 2e-5). chosen using the AllenTune package Dodge et al. (2019). Initializations greatly affect performance, as noted in Dodge et al. (2020). WinoGrande and SNLI RoBERTa-large models are trained for 6 epochs, and MultiNLI and QNLI are trained for 5 epochs each. Each experiment is performed on a single Quadro RTX 8000 GPU. Based on the available GPU memory, our experiments on all datasets use a batch size of 96, except for WinoGrande, where a batch size of 64 is used. Our implementation uses the Huggingface Transformers library Wolf et al. (2019). For the active learning baselines, we train a acquisition model using RoBERTa-large on a randomly sampled 1% subset of the full training set.

A.4 SNLI Qualitative Analysis

Qualitative samples from different regions of the SNLI data map are provided in Tab. 6.

Appendix B Additional Results

Results on the SNLI validation set are provided in Tab. 7.

To empirically test the hypothesis that confidence\mathbf{confidence} and variability\mathbf{variability} from the training dynamics respectively quantify intrinsic and model uncertainty, we compare confidence\mathbf{confidence} and variability\mathbf{variability} against an established method of capturing intrinsic and model uncertainty from the literature based on dropout Srivastava et al. (2014). Dropout can be seen as variational Bayesian inference Gal and Ghahramani (2016), with predictions from different dropout masks corresponding to predictions sampled from the posterior. Thus, confidence\mathbf{confidence} and variability\mathbf{variability} computed from sampled dropout predictions measure the average and standard deviation of the gold label’s probability under the posterior—quantifying the intrinsic and model uncertainty in a principled way.

We computed confidence\mathbf{confidence} and variability\mathbf{variability} from both training dynamics and dropout on WinoGrande’s development set.To compute training dynamics, we trained a model on the combined training and development sets for WinoGrande. In contrast, the dropout model was trained only on WinoGrande’s training set then run on development, to avoid over-fitting and provide higher quality uncertainty estimates. Figure 7 visualizes a regression analysis of the relationship between confidence\mathbf{confidence} and variability\mathbf{variability} from training dynamics and dropout. confidence\mathbf{confidence} from training dynamics and dropout correlate between 0.450 and 0.452 for Pearson’s rr at 95% confidence. Likewise, variability\mathbf{variability} from training dynamics and dropout share a Pearson’s rr from 0.390 to 0.393 at 95% confidence. Thus, the training dynamics empirically demonstrate a positive, predictive relationship with these first-principles estimates of the intrinsic and model uncertainty. Compared to dropout, however, training dynamics have the pragmatic advantage that all information required to calculate them is already available from training, without additional work or computation.

Appendix C Additional Data Maps

All the data maps have been provided in Fig. 8 and Fig. 9.

While training dynamics are inherently model dependent, data maps can be built for any model, and might reveal similar structures. Since models can be of varying capacities with respect to a task or dataset, instances might receive different co-ordinates on data maps built based on different models. For instance, BERT is known to be worse at reasoning than RoBERTa Sakaguchi et al. (2020); Talmor et al. (2019), and RoBERTa being a larger model is likely very sample efficient Kaplan et al. (2020). However, the overall structure of data maps based on different models remains the same; Fig. 9(b) shows the data map built for WinoGrande using a BERT-large classifier.

Four different architectures for the SNLI dataset are compared in Fig. 10 and Fig. 11.