Active Learning Helps Pretrained Models Learn the Intended Task

Alex Tamkin, Dat Nguyen, Salil Deshpande, Jesse Mu, Noah Goodman

Introduction

Modern pretrained models can be adapted to new tasks with remarkably little data, enabling downstream applications for tasks with only tens or hundreds of examples (Brown et al., 2020; Radford et al., 2021). However these low-data applications magnify a fundamental problem in machine learning—namely, that datasets are often incomplete proxies for the desired behavior. For example, a sentiment classifier may behave unpredictably during the holidays if its training data lacks examples of toy reviews. Likewise, an object classifier may struggle to classify objects in atypical environments (e.g. camels in the Arctic) if the training data does not make it adequately clear which features are salient for the task. While these dataset flaws can be solved by labeling better examples, identifying such examples can be challenging, especially as the nature of these flaws may not be known in advance.

One way to describe these kinds of challenges is through the lens of task ambiguity: the failure of the training data to fully specify the user’s intended behavior for all possible inputs. We consider whether pretrained models can automatically resolve their own task ambiguity through active learning (AL). In principle, AL allows models to resolve task ambiguity by identifying examples whose labels would be informative; for example, in Figure 1, the provided training data only contains red squares and blue circles, leaving the intended behavior for blue squares ambiguous. Asking for labels of blue squares resolves this ambiguity. This kind of interaction could clarify the intended behavior for different kinds of examples without the expectation that model developers must anticipate all potential gaps in the model’s abilities.

However, AL has often seen limited success in practice. In traditional settings with smaller, non-pretrained models, several problems limit the effectiveness of AL, including label noise, examples that are challenging for models to learn, and a lack of generalizability of AL heuristics (Lowell et al., 2019; Karamcheti et al., 2021). But pretrained models possess several strengths which may enable them to overcome these challenges: they extract high-level semantic features, such as shape and color (Olah et al., 2018; Goh et al., 2021), which can be used to identify informative examples, and they produce more calibrated uncertainty estimates which can be used for selecting ambiguous inputs (Hendrycks et al., 2019; Desai and Durrett, 2020).

We consider the use of active learning (AL) on a range of real-world image and text datasets where task ambiguity arises. We compare several AL acquisition functions against a random-sampling baseline, and compare the difference in performance with and without the use of pretrained models. Our contributions demonstrate that:

Pretrained models trained with AL can select examples that resolve task ambiguity in the finetuning data

The resulting accuracy gains can be quite large in practice: up to 5×5\times reduction in labeled data points for the same performance, or +11% absolute gain for the same labeling budget

This ability to actively learn is an emergent property of pretraining: AL has a neutral or even harmful effect on non-pretrained models

Method

We study the pool-based active learning (AL) setup common in the literature (Settles, 2009), where we have a (possibly pretrained) model M\mathcal{M}, a small seed set of training data S={(xi,yi)}\mathcal{S}=\{(x_{i},y_{i})\}, and a larger pool of unlabeled data P⁡={xi}\operatorname{\mathcal{P}}=\{x_{i}\}. The AL procedure proceeds as follows: first, finetune M\mathcal{M} on S\mathcal{S} until convergence; then, select points xix_{i} from P⁡\operatorname{\mathcal{P}} that are deemed most informative according to an acquisition function a(x;M)a(x;\mathcal{M}), obtaining the corresponding labels yiy_{i}, until some budget kk of data points is exhausted. This newly labeled batch B={(xi,yi)}\mathcal{B}=\{(x_{i},y_{i})\} is then added to the existing data S\mathcal{S}, and the original model is retrained on S\mathcal{S} for the next acquisition step. This process is repeated for a fixed number of acquisition steps.

For our acquisition function, we adopt the classic uncertainty sampling approach to AL, in particular the least confidence heuristic, where we acquire points for which our model is least confident in its predicted label (Settles, 2009). Specifically, treating the outputs of the model M\mathcal{M} as a probability distributionWhile in general there is no guarantee that this probability distribution will be well-calibrated, recent work has found that pretraining improves model calibration across a variety of settings, including on out-of-domain data (Desai and Durrett, 2020; Hendrycks et al., 2019). over possible labels p(y∣x;M)p(y\mid x;\mathcal{M}), we define the acquisition function to be

This least confidence heuristic has shown to be simple and effective in a variety of settings (Settles, 2009; Hendrycks and Gimpel, 2017; Mussmann et al., 2020), and we similarly find good results here (we refer to least confidence sampling as “uncertainty sampling” except in Section 5.2 where we explore other uncertainty-based acquisition functions including entropy and margin sampling).

On top of this standard AL pipeline, we propose the following change to improve the practical applicability of pretrained models in data-scarce settings:

The AL cycle begins by finetuning the pretrained model on the seed set. If the size of the seed set is large enough, the seed set may be partitioned into a training set and a validation set, and early stopping may be performed on the validation set. However, in few-shot settings, labeling costs may be high, and the seed set may be too small to meaningfully partition.

Instead of an arbitrary fixed number of finetuning steps, we propose an alternative method to terminate finetuning in the absence of a validation set. Specifically, we find that a simple but effective heuristic is to stop finetuning when the training loss decreases to 0.1% of the original training loss at the start of finetuning. In our experiments, this heuristic performs as well as early stopping on an actual validation set (see Appendix D for more details). We also leverage standard learning rates and other hyperparameters recommended by model developers (see Appendix B).

By using a standardized recipe across tasks and removing the need for a separate validation set, our AL pipeline is more robust to the real-world difficulties of deploying AL where use of a validation set is impractical (Lowell et al., 2019; Perez et al., 2021), although further work is needed to capture the full extent of this recipe’s generalizability.

Datasets

We consider a variety of datasets where task ambiguity manifests through a scarcity of particular kinds of examples. We consider two such kinds of examples: those defined by combinations of causal and spurious features (typical vs atypical backgrounds) as well as those defined by unseen attributes that shift during deployment (product categories and camera trap locations). These datasets provide an empirical testbed for the ability of pretrained models to choose disambiguating examples using active learning (AL).

Spurious correlations arise when multiple features are predictive of the label in a training dataset, yet it is ambiguous which ones are causally linked to the task label (Geirhos et al., 2020). We consider two such datasets, and see whether AL can choose the disambiguating examples where the spurious features are not copresent with the causal features:

The Waterbirds dataset (Sagawa et al., 2019) consists of photographs of landbirds or waterbirds digitally edited onto land or water backgrounds. The task is to classify whether the bird is a landbird or a waterbird. In the train set, 77% of the pictures feature landbirds and 23% waterbirds. 95% of both landbirds and waterbirds appear on land and water backgrounds, respectively. In the validation and test sets, this percentage is decreased to 50%, instead. Thus, the image background is a spurious feature the model may come to rely on when making the prediction.

Treeperson

As the Waterbirds dataset was synthetically generated, we also consider a dataset where we perform classification over real, unedited images with spuriously correlated objects. We use the object annotations in Visual Genome (Krishna et al., 2016) to create a new dataset of 8,638 images called Treeperson, for which the task is to predict whether a person is in a given image. While 50% of the images contain a person in this dataset, each image also contains either a tree or a building, and the presence of these objects is spuriously correlated with the presence of people. At train time, 90% of training images with people contain a building, while 90% of training images without people contain a tree. Thus, a model may be incentivized to form representations that classify according to the presence of trees and buildings, rather than the presence of the actual causal variable of interest (people). These values are changed to 50% at test time, removing this correlation to evaluate how well the model learned the actual task of interest. For more details on this dataset, see Appendix C.

2 Measuring robustness to distribution shift

Distribution shifts occur when algorithms are evaluated on different data distributions than the ones they were trained on. Examples include changing the location or time of day that photos were taken, or changing the topic or author of a particular textual source. These shifts can reduce performance, and we consider whether AL can help choose diverse, informative examples that clarify how the model should behave over a range of natural distribution shifts.

This dataset considers the task of species classification from a database of photos taken from wildlife camera traps (Beery et al., 2020; Koh et al., 2021). The dataset is unbalanced, with most images containing no animal, and the distribution of camera locations and species changes between the in domain (ID) and out-of-domain (OOD) subsets.

Amazon-WILDS

This dataset considers the task of predicting the number of stars associated with the text of a given Amazon review (Ni et al., 2019; Koh et al., 2021). The reviewers are different in the training set versus the test set, and the task is to perform as well as possible on this set of new reviewers. In addition to the number of stars, we also consider model performance stratified by different product types, which highlights minority subgroups whose categorization is not visible to the model.

Experimental Setup

For computer vision datasets, we finetune BiT (Kolesnikov et al., 2020), a recently-proposed family of vision models which have achieved state-of-the-art performance on several vision tasks. We primarily consider the BiT-M-R50x1 model, pretrained on ImageNet-21k (Deng et al., 2009). To explore the effectiveness of larger architectures and pretraining sources, in Section 6.2 we also consider performance achieved by the same-size BiT-S-R50x1, trained on ImageNet-1k, and the deeper BiT-M-R101x1 model, also trained on ImageNet-21k. These models have been shown to have emergent few-shot learning abilities, where strong classifiers for new tasks can be obtained by simply finetuning on tens or hundreds of examples with typical gradient descent techniques (rather than meta-learning techniques, for example).

Text

For the text dataset (Amazon), we use RoBERTa-Large (Liu et al., 2019), another pretrained model with similar properties as BiT, and a representative of the BERT (Devlin et al., 2019) family of models which together have obtained state-of-the-art scores on modern NLP benchmarks (Wang et al., 2018).

Other details, including hyperparameters and seed set/acquisition sizes are deferred to Appendix B.

Random acquisition baseline

As a running baseline, we compare to the same model finetuned with a random acquisition function (equivalent to not doing AL). That is, a(x;M)=rand(0,1)a(x;\mathcal{M})=\text{rand}(0,1), so we simply sample a random batch of data from the pool at each acquisition step.

Comparison with non-pretrained models

To examine whether effective AL is a result of the pretraining process, we also compare to the performance observed when applying AL to a randomly initialized, instead of pretrained, BiT-M-R50x1.

Results

For a general measure of success, in Figure 2 we plot the accuracy of AL versus random sampling on the validation datasets as a function of the number of samples acquired during training.

Waterbirds is evaluated on a balanced dataset where the foreground and background are not correlated. In this setting, uncertainty sampling achieves a +11% improvement in average validation accuracy over random sampling (Figure 2(a)). This comes primarily from a +25% average increase across the landbird-on-water and waterbird-on-land images (i.e. those without the spurious correlation; Figure 2(b)). Uncertainty sampling required 5x fewer labels than random sampling to achieve random sampling’s final accuracy.

Amazon

In the Amazon dataset, we also see gains from AL in the OOD setting, including +1% on average across reviewers, and +2.5% on the worst 10th percentile (Figure 2(g)). This result suggests that our AL recipe may be of use outside of BiT or computer vision settings more broadly. While the final difference between uncertainty and random sampling is not large, it is statistically significant. Uncertainty sampling required 1.3x fewer labels than random sampling to achieve random sampling’s final accuracy.

iWildCam

With the iWildCam dataset, uncertainty sampling achieved a +9% improvement upon random sampling. Uncertainty sampling also required 1.8x fewer labels than random sampling to achieve random sampling’s final accuracy (Figure 2(e)).

Treeperson

In the Treeperson dataset, uncertainty sampling is +2% improved over random sampling by the end of training (Figure 2(c)). Uncertainty sampling required 1.6x fewer labels than random sampling to achieve random sampling’s final accuracy.

2 Additional AL methods

We consider two additional AL methods in addition to least confidence sampling: 1) entropy sampling, which chooses the example that maximizes the entropy of the model’s predictive distribution, and 2) margin sampling, which chooses the example with the smallest difference between the first and second most probable classes (Scheffer et al., 2001; Settles, 2009). We run experiments with all methods on the 182-class iWildCam dataset.Note that least confidence, entropy, and margin sampling are identical in the case of binary classification. All methods significantly outperform random sampling (Figure 3).

Analysis

Overall, we attribute improved performance to pretrained models’ ability to identify and preferentially sample examples that resolve task ambiguity in the data. For example, for Waterbirds and Treeperson, we see the model select examples with atypical background combinations, as one might intuitively hope. Similarly, for Amazon and iWildCam we see the model upsample rare types of examples, even when they are not explicitly marked in the data.

Figure 4(a) depicts the rate at which uncertainty sampling acquires examples of each subgroup compared to the expected rate at which random sampling would acquire examples from those same subgroup. Examples where the bird and background are mismatched are heavily oversampled. We emphasize that these minority examples are not simply members of the minority class (waterbirds). Instead, the model identifies and preferentially upsamples disambiguating examples where the spurious feature (background) and the causal feature (bird type) disagree.

Treeperson

For Treeperson we see the same pattern as in Waterbirds: the model identifies and upsamples examples where only one of the spurious or causal features is present, despite the spurious feature being latent (Figure 4(b)).

Amazon

We also see similar behavior in the Amazon dataset, indicating our method’s applicability to multiple modalities and pretrained models. Not only does the model upsample lower star ratings, which are less common, it is also able to upsample rarer product categories—a latent attribute (Figure 5).

2 Pretraining is the key ingredient in our experiments

What drives the success of AL in our experiments? We hypothesize that better AL is an emergent property of the pretraining process, and evaluate this hypothesis by comparing pretrained models to their corresponding non-pretrained counterparts. We also examine the effect of model scale on success at AL: if a model has been trained on more data (and has presumably learned to extract more semantically-relevant features), does this enable more effective AL?

Concretely, we conduct AL experiments with 3 pretrained models: BiT-S-R50x1, BiT-M-R50x1, BiT-M-R101x1, and their corresponding non-pretrained versions. We find that for our experiments with pretrained models, uncertainty acquisition outperformed random acquisition (Figures 6(a), LABEL:, 6(b) and 6(c)). Importantly, AL on the non-pretrained models provided no or even negative benefit, even after exploring a range of different hyperparameter configurations. These controlled experiments provide strong evidence that pretraining is indeed crucial in our setting. That said, we do not claim that pretraining is the only way to enable good AL in settings of task ambiguity; other methods of addressing the so-called “cold start” problem (e.g. Gao et al. (2020); Yuan et al. (2020)) may also prove fruitful or complementary.

In one case, we also see a demonstrable effect of scale on the efficacy of the AL process: the BiT-S-R50x1 model, which was pretrained on a smaller dataset than the BiT-M models (ImageNet-1k vs ImageNet-21k) fails to outperform random sampling on iWildCam, in contrast to the two other models pretrained on more data. This suggests a potential “scaling trend” for AL, where gains from AL may continue to grow as pretrained models are trained for longer on more data. However, we did not see a difference between BiT-M-50x1 and BiT-M-101x1, which were trained on the same dataset but have different numbers of parameters—suggesting that dataset size may be a more crucial variable than parameter count in some settings.

Impact of pretraining on acquisition patterns

As an additional cross-check, we also observe that pretrained models acquire disambiguating subgroups much more efficiently than their non-pretrained counterparts. See Appendix G for additional figures and results.

3 Pretraining yields a better feature space for AL

While pretraining clearly improves the AL process, the mechanisms behind this improvement remain unclear. Given the strong theoretical results AL enjoys in the linear setting (Balcan et al., 2007; Balcan and Long, 2013; Mussmann and Liang, 2018), we hypothesize that pretraining may aid AL by linearizing the features salient for task ambiguity. This hypothesis is further inspired by recent studies finding that a wide range of features are linearly separable in the feature spaces of large pretrained models (Chen et al., 2020).

To quantify this, we train linear classifiers on the second to last layer of the BiT models. The classifiers are trained to predict each image’s bird type and background type (4 classes total, rebalanced to comprise 25% of the data). As shown in Figure 7, these classes are indeed far more linearly separable in pretrained models, providing evidence for this hypothesis.

We present additional investigations of t-SNE plots for pretrained and non-pretrained models in Appendix H, which demonstrates increased separation of latent classes for pretrained models, as well as how the acquired examples are closer to the class boundaries.

4 How does the degree of task ambiguity impact AL?

Finally, we measure the impact of the strength of task ambiguity on AL. To do so, we construct variants of the Waterbirds dataset where the percentage of mismatched examples range from 95% to 50% but the marginal class probabilities remain fixed.We construct these datasets using the code at https://github.com/kohpangwei/group_DRO. We then proceed with AL and report results in Figure 8. We observe a clear trend where the gains from AL gradually increase as the percentage of mismatched examples increases to 95%. We find similarly clear trends in the upsampling ratio of mismatched backgrounds, shown in Figure 9. These results provides further evidence that task ambiguity is the key driving factor behind the success of AL in this setting.

5 Failure Cases

We also encounter some failure cases when training on datasets far from the distribution of the pretrained model. We perform preliminary experiments on Camelyon17-WILDS (Bándi et al., 2019; Koh et al., 2021), which considers tumor identification from tissue patches, and FMoW-WILDS (Christie et al., 2018; Koh et al., 2021), which considers land-use classification from satellite images. AL performs comparably or worse than random sampling on these datasets, even when using a pretrained BiT model, suggesting that specialized models may be necessary to see gains in domains very different from ImageNet-21k.However, we note that Camelyon17-WILDS is also known to exhibit high variance across seeds (see https://wilds.stanford.edu/get_started/) which may also be a contributing factor.

Related work

Several works address ambiguity or poor specification in machine learning problems. Taylor et al. (2020) describe the problem of “inductive ambiguity identification,” and describe AL as a promising potential solution that has failed to see practical success. D’Amour et al. (2020) describe the problem of underspecification, where high variance, instability, and poor model performance result from training overparameterized models that are underconstrained by their training datasets. Geirhos et al. (2020) describes how task ambiguity can arise when both desirable and undesirable features are predictive of the training labels, a problem which several works seek to better characterize and address (Nagarajan et al., 2021; Sagawa et al., 2019; Srivastava et al., 2020; Sagawa et al., 2020). Finally, Finn et al. (2018) address task ambiguity in few-shot settings via a probabilistic meta-learning algorithm, and perform an AL experiment in a 1D regression setting. We build on these works by demonstrating that simple uncertainty sampling with pretrained models can be an effective approach to the task ambiguity problem across a wide variety of high-dimensional classification settings—including when the sources of task ambiguity are not known.

Uncertainty and distribution shift

In the face of these challenges, several works have tried to quantify how much pretrained models know about problems or their own uncertainty about them. Rajpurkar et al. (2018) propose a question answering dataset with unanswerable questions, where a model must abstain rather than proceeding with an answer. Pretraining can also improve the calibration of model uncertainty (Hendrycks et al., 2019) and pretrained features can be used for out-of-distribution detection (Reiss et al., 2021; Wu and Goodman, 2020)—observations that align with our findings that uncertainty sampling can identify minority subgroups in datasets. A related stream of work seeks to identify high-confidence examples that are predicted incorrectly (Attenberg et al., 2011; Lakkaraju et al., 2017); by contrast, our focus is on improving model behavior across the full range of examples. Finally, our observation that upsampling latent minority groups results in better performance aligns well with Sagawa et al. (2020); Idrissi et al. (2021); Liu et al. (2021), which explore various upweighting or upsampling strategies. Importantly, however, our approach does not require these groups to be known in advance.

AL and example selection

Active learning (AL) (Lewis and Catlett, 1994; Settles and Craven, 2008; Settles, 2009; Houlsby et al., 2011) is a well-studied field that investigates how machine learning algorithms might automatically select helpful additional data points to maximize their performance. Such strategies are especially helpful in imbalanced settings (Ertekin et al., 2007; Mussmann et al., 2020) and have been fruitfully applied to deep models (Gal et al., 2017; Beluch et al., 2018), including pretrained models (Yuan et al., 2020; Margatina et al., 2021; Shelmanov et al., 2021). Past work has also considered AL for few-shot learning (Woodward and Finn, 2017). We extend these works by considering AL for resolving task ambiguity, showing that pretrained models successfully choose examples based on their high-level semantics, such as atypical backgrounds or rare latent attributes. Also in contrast to prior work, we investigate the role of pretraining itself by performing equivalent experiments with non-pretrained models, and providing a potential mechanism for the difference.

Pretrained models and their emergent properties

Our work contributes to a broader literature on how pretraining enables new kinds of model capabilities (Bommasani et al., 2021; Tamkin et al., 2021a), especially those holding across multiple domains (Tamkin et al., 2021b). For example, (Brown et al., 2020) identify the phenomenon of in-context learning, where tasks can be specified for models through a language modeling prompt, while Caron et al. (2021) discover that a self-supervised vision model implicitly learns high-quality segmentation maps visible through attention scores. Kaplan et al. (2020); Henighan et al. (2020) conduct scaling laws experiments which chart how capabilities emerge with scale. We identify a new model capability that is significantly improved by pretraining: the capacity to actively learn and resolve task ambiguity in high-dimensional settings.

Discussion and Limitations

We show that pretrained models can preemptively resolve task ambiguity through active learning (AL), without requiring humans to anticipate these possible failure modes in advance. We find that AL helps across a variety of settings where data is spuriously correlated, undergoes domain shift, or contains unlabeled subgroups. These behaviors emerge most clearly as a result of large-scale pretraining, suggesting that AL may be an underappreciated tool for increasing the reliability of systems in real-world settings.

Of course, AL is no cure-all for resolving task ambiguity. First, it requires a human in the loop, which increases the time required to train a model compared to random sampling. Second, it requires the labeling method to be relatively free of noise—this may be acceptable if annotators are domain experts or are well-trained, but may also increase the cost per acquired example. Third, it is limited by the range of examples present in the unlabeled dataset—a model cannot elicit labels for examples that do not exist.

Finally, we note the opportunity for an exciting array of future work, including broader investigation of these methods across domains such as medical, scientific, or industrial settings, as well as better understanding how pretraining shapes AL as models continue to scale.

Acknowledgments

We would like to thank Shyamal Buch, Shreya Shankar, Megha Srivastava, and Ethan Perez for helpful comments. AT is supported by an Open Phil AI Fellowship.

References

Appendix A Code release

Code and training scripts will be released soon.

Appendix B Additional experimental details

Here we provide some additional experimental details.

All BiT runs use the same default settings specified in the BiT paper (Kolesnikov et al., 2020) and the BiT GitHub repo: https://github.com/google-research/big_transfer

Settings that are constant across all BiT runs:

Learning Rate - Base learning rate 0.003. Linear warm up to this rate, then staircase decay. Exact schedule depends on dataset size, but for our few shot setting, this means: (1) linear warm up in the first 100 steps to 0.003, then (2) decay 10 fold every 100 steps; (3) after 500 steps, stop and move on to the next acquisition round.

Data Augmentation - Random cropping and flipping. See our repo/utils/datasets/load for details. Also available in BiT repo.

Batch Size - 32 when training, split into gradient accumulation microbatches of size 8.

Early stopping condition - When the training loss reaches below 0.001 times the original training loss.

Settings that varied between image datasets:

Size of Initial Seed Set - 40 for Waterbirds, 40 for Treeperson, 182 for iWildCam

Size of Training Pool - Entire training set for Waterbirds, Treeperson. A random subset of 12000 examples (re-drawn for each acquisition) from the entire training set of iWildCam.

Number of Acquired Examples - 320 for Waterbirds, 320 for Treeperson, 1456 for iWildCam.

Number of Examples Acquired Each Acquisition - 5 for Waterbirds, 5 for Treeperson, 20 for iWildCam.

Settings used for Amazon + RoBERTa-Large runs:

Optimizer - AdamW with default hyperparameters (β1=0.99,β2=0.999\beta_{1}=0.99,\beta_{2}=0.999, weight decay = 0.1).

Batch Size - 2 (due to memory considerations).

Early Stopping Condition - When the training loss reaches below 0.001 times the original training loss.

Size of Training Pool - A random subset of 2000 examples (re-drawn for each acquisition) from the entire training set.

Number of Examples Acquired Each Acquisition - 2.

Appendix C Treeperson dataset

The Treeperson dataset is composed of images from Visual Genome (Krishna et al., 2016) with different compositions of detected objects.

The validation set contains 498 examples of each subclass.

The following annotated objects were used to form the different subclasses:

Person: person, people, man, men, woman, women

The training set and validation set were drawn randomly from qualifying images in Visual Genome’s training set and validation set, respectively.

Appendix D Early stopping condition

At the outset of this work, we explored if we could identify a heuristic for stopping training when there was no validation set present. We compared how the BiT model would perform if it stopped the finetuning step based off of when the validation accuracy plateaued versus when the training loss decayed to be 0.1% of its original value.

We ran a smaller experiment than the standard Waterbirds parameters we described in Appendix B. Namely, our seed set was of size 32, we acquired 64 examples on top of that, and we acquired one example at a time. We also defined the validation accuracy as having plateaued the fifth time it did not increase.

We ran twelve paired experiments where for twelve different randomized seed sets, we performed the Waterbirds experiment four times - with random sampling and stopping when validation accuracy plateaued, with random sampling and stopping when training loss decayed, with uncertainty sampling and stopping when validation accuracy plateaued, and with uncertainty sampling and stopping when training loss decayed.

We found that these experiments achieved:

Random sampling + Stop from validation accuracy: average accuracy = 69.11%, average duration = 3.5 hours

Random sampling + Stop from training loss: average accuracy = 69.93%, average duration = 38 minutes

Uncertainty sampling + Stop from validation accuracy: average accuracy = 83.64%, average duration = 4.5 hours

Uncertainty sampling + Stop from training loss: average accuracy = 85.62%, average duration = 74 minutes

Thus, we concluded that the by stopping our finetuning step just by waiting for the training loss to decay to 0.001 of its original values, we could achieve comparable if not better accuracies, spend a fraction of the time, and remove the need for a labeled validation set.

We did not consider other values besides 0.001 when conducting this search, though one could consider optimize this number in a fair way on a separate set of labeled datasets (in general one should not optimize AL pipelines using the validation sets they are being evaluated on).

Appendix E How the degree of task ambiguity impacts acquisitions

In Figure 9, we also plot the degree of upsampling that occurs for mismatched examples for different degrees of task ambiguity. As disambiguating examples become more scarce, uncertainty sampling increasingly upsamples them relative to their prevalence in the unlabeled training set.

Appendix F Vision transformer model

We broaden our coverage of computer vision models to include vision transformers (Dosovitskiy et al., 2021), the other major architecture family currently in use. We train ViT-16/B,https://huggingface.co/google/vit-base-patch16-224 on Treeperson. (Figure 10). This vision transformer was pretrained on ImageNet-21k for fewer epochs then BiT (9 vs 70), which is reflected by its lower oversampling of minority classes and comparatively smaller gains in accuracy compared to BiT (Figure 11).

Appendix G Effect of model scaling and pretraining on acquisition

For the Waterbirds model scaling experiment, we track the examples each model acquires. These are presented in Figure 12. The acquisition patterns of all the pretrained models look fairly similar—they upsample both minority (landbird/water-background, waterbird/land-background) subclasses. However, the non-pretrained model is unable to capture that distinction, and is only able to upsample images with a water background, resulting in worse performance. A summary of final class acquisition ratios is available in Figure 13.

Appendix H Visualization Of Image Embeddings Of Pretrained Bit-M

We visualize the second-to-last layer embeddings of BiT (without any finetuning) using t-SNE (van der Maaten and Hinton, 2008).

As seen in Figure 14(b), when performing t-SNE on the Waterbirds image embeddings, we find that the t-SNE of the pretrained model has much more structure than that of the non-pretrained model. In the non-pretrained model, the image embeddings of the landbird/land background, landbird/water background, and waterbird/water background classes are distributed uniformly about the center of the t-SNE projection. However, we observe much more structure in the pretrained model’s t-SNE. In particular, the most noticible difference is that the landbird/water background and waterbird/land background classes can be found near the other landbird and waterbird images, respectively. This shows that even without any finetuning, the pretrained model already has learned features to help process the type of bird that appears in the image.

In the t-SNE plots, the black stars represent the embeddings of the first 10 images that were acquired by each model. For the pretrained model, the acquired examples very clearly fall along the waterbird-landbird boundary in the projected feature space. However, in the non-pretrained model, the acquired examples are distributed randomly. This suggests that the features acquired during pretraining enables them to select examples that fall near the decision boundary of a new unseen task.