On Robustness and Transferability of Convolutional Neural Networks

Josip Djolonga, Jessica Yung, Michael Tschannen, Rob Romijnders, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Matthias Minderer, Alexander D'Amour, Dan Moldovan, Sylvain Gelly, Neil Houlsby, Xiaohua Zhai, Mario Lucic

Introduction

Deep convolutional networks have attained impressive results across a plethora of visual classification benchmarks where the training and testing distributions match. In the real world, however, the conditions in which the models are deployed can often differ significantly from the conditions in which the model was trained. It is thus imperative to understand the impact dataset shifts have on the performance of these models. This problem has gained a lot of traction and several systematic investigations have shown unexpectedly high sensitivity of image classifiers to various dimensions, including photometric perturbations , natural perturbations obtained from video data , as well as model-specific adversarial perturbations .

The problem of dataset shift, or out-of-distribution (OOD) generalization, is closely related to a learning paradigm known as transfer learning [56, §13]. In transfer learning we are interested in constructing models that can improve their performance on some target task by leveraging data from different related problems. In contrast, under dataset shift one assumes that there are two environments, namely training and testing , with the constraint that the model cannot be adapted using data from the target environment. As a consequence, the two environments typically have to be more similar and their differences more structured than in the transfer setting (c.f. Section 2).

In the context of transfer learning, detailed scaling laws characterizing the interplay between the in-distribution and transfer performance as a function of pre-training data set size, model size, architectural choices such as normalization, and transfer strategy have been established recently . Model and dataset scale were identified as key factors for transfer performance. The similarities between transfer learning and OOD generalization suggests that these axes are also relevant for OOD generalization and raises the question of what the corresponding scaling laws are. While some axes have been partially explored by prior work , the big picture is largely unknown. Even more importantly, is in-distribution performance enough to characterize OOD performance, or can transfer performance give a more fine-grained characterization of OOD performance of a population of models than in-distribution performance? To the best of our knowledge, this question has not been systematically explored before in the literature.

Contributions We systematically investigate the interplay between the in-distribution accuracy of image classification models on the training distribution, their generalization to OOD data (without adaptation), and their transfer learning performance with adaptation in the low-data regime (see Fig. 1 for an illustration). Specifically:

We present the first meta-analysis of existing OOD metrics and transfer learning benchmarks across a wide variety of models, ranging from self-supervised to fully supervised models with up to 900M parameters. We show that increasing the model and data scale disproportionately improves transfer and OOD performance, while only marginally improving the performance on the ImageNet validation set.

Focusing on OOD robustness, we analyze the effects of the training set size, model scale, and the training regime and testing resolution, and find that the effect of scale overshadows all other dimensions.

We introduce a novel dataset for fine-grained OOD analysis to quantify the robustness to object size, object location, and object orientation (rotation angle). We believe that this is a first systematic study to show that the models become less sensitive (and hence more robust) to each of these factors of variation as the dataset size and model size increase.

Background

Dataset shift types ImageNet-v2 is a recollected version of the ImageNet validation set . The authors attempted to replicate the data collection process, but found that all models drop significantly in accuracy. Recent work attributes this drop to statistical bias in the data collection . ImageNet-C and ImageNet-P are obtained by corrupting the ImageNet validation set with classical corruptions, such as blur, different types of noise and compression, and further cropping the images to 224×224224\times 224. These datasets define a total of 15 noise, blur, weather, and digital corruption types, each appearing at 5 severity levels or intensities. ObjectNet presents a new test set of images collected directly using crowd-sourcing. ObjectNet is particular as the objects are captured at unusual poses in cluttered, natural scenes, which can severely degrade recognition performance. Given this clutter, and arguably better suitability as a detection than recognition task , Y∣XY|X might be hard to define and the dataset goes beyond a covariate shift. In contrast, the ImageNet-A dataset consists of real-world, unmodified, and naturally occurring examples that are misclassified by ResNet models. Hence in addition to the covariate shift due to the data source, this dataset is not model-agnostic and exhibits a strong selection bias .

Attempting to focus on naturally occurring invariances, annotated two video datasets: ImageNet-Vid-Robust and YouTube-BB-Robust, derived from the ImageNet-Vid and YouTube-BB datasets respectively. In the authors propose the pm-kk metric—given an anchor frame and up to kk neighboring frames, a prediction is marked as correct only if the classifier correctly classifies all 2k+12k+1 frames around and including the anchor. We present the details of each dataset in Appendix A.

Recently, a suite of datasets has been collected to benchmark modern image classification transfer techniques . The Visual Task Adaptation Benchmark (VTAB) defines 19 datasets with 1000 labeled samples each, categorized into three groups: natural (most similar to ImageNet) consists of standard natural classification tasks (e.g., CIFAR); specialized contains medical and satellite images; and structured (least similar to ImageNet) consists mostly of synthetic tasks that require understanding of the geometric layout of scenes. We compute an overall transfer score as the mean across all 19 datasets, as well as scores for each subgroup of tasks. We provide details for all of the tasks in Appendix A.

A meta-analysis of robustness and transferability metrics

While many robustness metrics have been proposed to capture different sources of brittleness, it is not well understood how these metrics relate to each other. We investigate the practical question of how useful the various metrics are in guiding design choices. Further, we empirically analyze the relationship between robustness and transferability metrics, which is lacking in the literature, despite their close relationship. To analyze these questions, we evaluated 39 different models over 23 robustness metrics and the 19 transfer tasks.

Metrics For robustness, we measure the model accuracy on the ImageNet, ImageNet-v2 (the matched frequency variant) and ObjectNet datasets. We also consider video datasets, ImageNet-Vid and YouTube-BB; we use both the accuracy metric and the pm-1010 metric (suffix -W). On ImageNet-C we report the AlexNet-accuracy-weighted accuracy over all corruption times (called mean corruption error in ). To evaluate the transferability of the models, we use the VTAB-1K benchmark introduced in Section 2. We evaluate average transfer performance across all 19 datasets, with 1000 examples each, as well as per-group performance. To transfer a model we performed a sweep over two learning rates and schedules. We report the median testing accuracy over three fine-tuning runs with parameters selected using a 800-200 example train-validation split.

Models To perform this meta-analysis we consider several model families.We evaluate ResNet-50 and six EfficientNet (B0 through B5) models including variants using AutoAugment and AdvProp , which have been trained on ImageNet. We include self-supervised SimCLR (variants: linear classifier on fixed representation (lin), fine-tuned on 10% (ft-10), and 100% (ft-100) of the ImageNet data), and self-supervised-semi-supervised (S4L) models that have been fine-tuned to 10% and 100% of the ImageNet data. We also consider a set of models that use other data sources. Specifically, three NoisyStudent variants which use ImageNet and unlabelled data from the JFT dataset, BiT (BigTransfer) models that have been first trained on ImageNet, ImageNet-21k, or JFT and then transferred to ImageNet by fine-tuning, and the Video-Induced Visual Invariance (VIVI) model , which uses ImageNet and unlabelled videos from the YT8M dataset . Finally, we consider the BigBiGAN model which has been first trained as a class-conditional generative model and then fine-tuned as an ImageNet classifier. All details can be found in Appendix E.

How informative are robustness metrics for discriminating between models? The goal of a metric is to discriminate between different models and thus guide design choices. We therefore quantify the usefulness of each metric in terms of how much it improves the discriminability between the various models beyond the information provided by ImageNet accuracy. Specifically, we train logistic regression classifiers to discriminate between the 12 model groups outlined above. We compared the performance of a classifier using only ImageNet accuracy as input feature, to a classifier using ImageNet and up to two of the other metrics, see Fig. 4 and Appendix A. We found that most of the tested metrics provide little increase in model discriminability over ImageNet accuracy. We further, similarly to , found that all metrics are highly rank-correlated with each other, which we present in Appendix A. Of course, these results are conditioned on the size and composition of our dataset, and may differ for a different set of models. However, based on our collection of 39 models in 12 groups, the most informative metrics are those based on different datasets and/or video, rather than ImageNet-derived datasets.

How related are OOD robustness and transfer metrics? Next, we turn to transfer learning. It has been observed that better ImageNet models transfer better . Since robustness metrics correlated strongly with ImageNet accuracy, we might expect a similar relationship. To get an overall view, we compute the mean of all robustness metrics, and compare it to transfer performance. Figure 2 (center) shows this average robustness plotted against transfer performance, while Figure 2 (left) shows transfer versus ImageNet accuracy. Indeed, we observe a large correlation coefficient ρ=0.73\rho=0.73 between robustness and transfer metrics; however, the correlation is not stronger than between transfer and ImageNet. Further, we compute the correlation of the residual robustness score (mean robustness minus ImageNet accuracy) against transfer score, and find only a weak relationship of ρ=0.12\rho=0.12. This indicates that robustness metrics, on aggregate, do not provide additional signal that predicts model transferability beyond that of the base ImageNet performance. We do, however, see some interesting differences in the relative performances of different model groups. Certain model groups, while attaining reasonable ImageNet/robustness scores, transfer less well to VTAB. Therefore, there are factors unrelated to robust inference that do influence transferability. One example is batch normalization which is outperformed by group normalization with weight standardization in transfer . Next, we break down the correlation by robustness metrics and transfer datasets in Fig. 2 (right). We see that each metric correlates similarly with the task groups. However, for the groups that require more distant transfer (Specialized, Structured), no metric predicts transferability well. Perhaps surprisingly, raw ImageNet accuracy is the best predictor of transfer to structured tasks, indicating that robustness metrics do not relate to challenging transfer tasks, at least not more than raw ImageNet accuracy.

Summary Metrics based on ImageNet have very little additional discriminative power over ImageNet accuracy, while those not based on ImageNet have more, but their additional discriminative power is still low—popular robustness metrics provide marginal complementary information. Transferability is also related to ImageNet accuracy, and hence robustness. We observe that while there is correlation, transfer highlights failures that are somewhat independent of robustness. Further, no particular robustness metric appears to correlate better with any particular group of transfer tasks than ImageNet does. Inspired by these results, we next investigate strategies known to be effective for ImageNet and transfer learning on the OOD robustness benchmarks.

Scaling laws for OOD performance

Increasing the scale of pre-training data, model architecture, and training steps have recently led to diminishing improvements in terms of ImageNet accuracy. By contrast, it has been recently established that scaling along these axes can lead to substantial improvements in transfer learning performance . In the context of robustness, this type of scaling has been explored less. While there are some results hinting that scale can improve robustness , no principled study decoupling the different scale axes has been performed. Given the strong correlation between transfer performance and robustness, this motivates the systematic investigation of the effects of the pre-training data size, model architecture size, training steps, and input resolution. While paramount to the out-of-distribution performance, as we find, these pretraining design choices have not yet received a great deal of attention from the community.

Setup We consider the standard ImageNet training setup as a baseline, and scale up the training accordingly. To study the impact of dataset size, we consider the ImageNet-21k and JFT datasets for the experiments, as pre-training on either of them has shown great performance in transfer learning . We scale from the ImageNet training set size (1.281.28M images) to the ImageNet-21k training set size (13M images, about 1010 times larger than ImageNet). To explore the effect of the model size, we use a ResNet-50 as well as the deeper and 3×3\timeswider ResNet-101x3 model. We further investigate the impact of the training schedule as larger datasets are known to benefit from longer training for transfer learning . To disentangle the impact of dataset size and training schedules, we train the models for every pair of dataset size and schedule.

We fine-tune the trained models to ImageNet using the BiT HyperRule , and assess their OOD generalization performance in the next section. Throughout, we report the reduction in classification error relative to the model which was trained on the smallest number of examples and for the fewest iterations, and which hence achieves the lowest accuracy. Other details are presented in Appendix B.

Pre-training dataset size impact The results for the ResNet-101x3 model are presented in Fig. 3. When pre-trained on ImageNet-21k, the OOD classification error significantly decreases with increasing pre-training dataset size and duration: We observe relative error reductions of 2020-30%30\% when going from 112k steps on 1M data points to 1.12M steps on 13M data points. The reductions are least pronounced for YouTube-BB(-W). Note that training for 1.121.12M steps leads to a lower accuracy than training for only 457k steps unless the full ImageNet-21k dataset is used. For models trained on JFT we observe a similar behavior except that training for 1.12M steps often leads to a higher accuracy than training for 457k steps even when only 1M or 5M data points are used (c.f. Appendix B). These results suggest that, if the models have enough capacity, increasing the amount of pre-training data, without any additional changes, leads to substantial gains in all datasets simultaneously which is in line with recent results in transfer learning .

Model size impact Figure 3 shows the relative reduction in classification error when using a ResNet-101x3 instead of a ResNet-50 as a function of the number of training steps and the dataset size. It can be seen that increasing the model size can lead to substantial reductions of 55–20%20\%. For a fixed training duration, using more data always helps. However, on ImageNet-21k, training too long can lead to increases in the classification error when the model size is increased, unless the full ImageNet-21k is used. This is likely due to overfitting. This effect is much less pronounced when JFT is used for training. JFT results are presented in Appendix B. Again, reductions in classification error are least pronounced for YouTube-BB/YouTube-BB-W.

Testing resolution and OOD robustness During training, images are typically cropped randomly, with many crop sizes and aspect ratios, to prevent overfitting. In contrast, during testing, the images are usually rescaled such that the shorter side has a pre-specified length, and a fixed-size center crop is taken and then fed to the classifier. This leads to a mismatch in object sizes between training and testing. Increasing the resolution at which images are tested leads to an improvement in accuracy across different architectures . Furthermore, additional benefits can be obtained by applying FixRes — fine-tuning the network on the training set with the test-time preprocessing (i.e. omitting random cropping with aspect ratio changes), and at a higher resolution. We explore the effect of this discrepancy on the robustness of different architectures. As some of the robustness datasets were collected differently from ImageNet, discrepancies in the cropping are likely. We investigate both adjusting test-time resolution and applying FixRes. For FixRes, we use a simple setup with a single schedule and learning rate for all models (except using a 10×10\times smaller learning rate for the BiT models), and without heavy color augmentation as in or label smoothing as in . We did not extensively tune hyperparameters, but chose a setup that works reasonably well across architectures and training datasets. Note that changing the resolution can be seen as scaling the computational resources available to the model, as both training and inference costs grow with the resolution.

Following the protocol of the FixRes paper , we evaluate each model for all resolutions in {64,128,224,288,320,384,512,768}\{64,128,224,288,320,384,512,768\} to illustrate the potential of adapting the testing resolution (in practice we do not have access to an OOD validation set so we cannot select the optimal solution in advance). For conciseness, we show the accuracy for ImageNet-A and ObjectNet at the testing resolution proposed by the authors of the respective architecture along with the highest accuracy across testing resolutions (Figure 5). The results for other datasets and resolutions are deferred to Appendix C.

We start by discussing observations that apply to most models, excluding the BiT models which will be discussed below. While FixRes only leads to marginal benefits on ImageNet, it can lead to substantial improvements on the robustness metrics. Choosing the optimal testing resolution leads to a significant increase in accuracy on ImageNet-A and ObjectNet in most cases, and applying FixRes often leads to additional substantial gains. For ObjectNet, fine-tuning with testing preprocessing (i.e. fine-tuning with central cropping instead of random cropping as used during training) can help even without increasing resolution.

Increasing the resolution and/or applying FixRes often slightly helps on ImageNet-V2. For ImageNet-C, the optimal testing resolution often corresponds to the resolution used for training, and applying FixRes rarely changes this picture. This is not surprising as the ImageNet-C images are cropped to 224 pixels by default, and increasing the resolution does not add any new information to the image. For the video-derived robustness datasets ImageNet-Vid-Robust and YouTube-BB-Robust, evaluating at a larger testing resolution and/or applying FixRes at a higher resolution can substantially improve the accuracy on the anchor frame and the robustness accuracy for small EfficientNet and ResNet models, but does not help the larger ones. For the BiT models, the resolution suggested by the authors is almost always optimal, except on ObjectNet and ImageNet-A, where changing the preprocessing considerably helps. FixRes arguably does not lead to improvements as it was already applied in BiT as a part of the BiT HyperRule.

Summary These empirical results point to the following conclusion: for models with enough capacity, increasing the amount of pre-training data, with no additional changes, leads to substantial gains in all considered OOD generalization tasks simultaneously. Secondly, resolution adjustments as outlined above can address the considerable distribution shift caused by resolution mismatch.

SI-Score: A fine-grained analysis of robustness to common factors of variation

The results in Section 4 do not reveal the underlying reasons for the success of larger models trained on more data on all robustness metrics. Intuitively, one would expect that these models are more invariant to specific factors of variation, such as object location, size, and rotation. However, a systematic assessment hinges on testing data which can be varied according to these axes in a controlled way. At the same time, the combinatorial nature of the problem precludes any large-scale systematic data collection scheme.

In this work we present a scalable alternative and construct a novel synthetic dataset for fine-grained evaluation: SI-Score (Synthetic Interventions on Scenes for Robustness Evaluation). In a nutshell, we paste a large collection of objects onto uncluttered backgrounds (Figure 6, Figure 14(a)), and can thus conduct controlled studies by systematically varying the object class, size, location, and orientation.The synthetic dataset and code used to generate the dataset are open-sourced on GitHub and are being hosted by the Common Visual Data Foundation.

Synthetic dataset details The foregrounds were extracted from OpenImages using the provided segmentation masks. We include only object classes that map to ImageNet classes. We also removed all objects that are tagged as occluded or truncated, and manually removed highly incomplete or inaccurately labeled objects. The backgrounds were images from nature taken from pexels.com (the license therein allows one to reuse photos with modifications). We manually filtered the backgrounds to remove ones with prominent objects, such as images focused on a single animal or person. In total, we converged to 614 object instances across 62 classes, and a set of 867 backgrounds.

We constructed three subsets for evaluation, one corresponding to each factor of variation we wanted to investigate, as shown in Table 1. In particular, for each object instance, we sample two backgrounds, and for each of these object-background combinations, we take a cross product over all the factors of variation. For the datasets with multiple values for more than one factor of variation, we take a cross product of all the values for each factor of variation in the set (object size, rotation, location). For example, for the rotation angle dataset, there are four object sizes and 18 rotation angles, so we do a cross product and have 72 factor of variation combinations. For the object size and rotation datasets, we only consider images where objects are at least 95% in the image. For the location dataset, such filtering removes almost all images where objects are near the edges of the image, so we do not do such filtering. Note that since we use the central coordinates of objects as their location, at least 25% of each object is in the image even if we do not do any filtering. The results in the following sections are similar when filtering out objects that are less than 50% or 75% in the image.

Learned invariances as a function of scale We study one factor of variation at a time. For example, when studying the impact of changing the location of the object center, we measure the average performance for each location over a uniform grid. Building on our investigation in the previous section, we test whether increasing model size and dataset size improves robustness to these three factors of variation by evaluating the ResNet-50 and ResNet-101x3 models. We observe that the models indeed become more invariant to object location (Figure 6), rotation (Figure 7, left), and size (Figure 7, right) as the pre-training set size increases. Specifically, as we pre-train on more data, the average prediction accuracy across various object locations, sizes, and rotation angles becomes more uniform. Furthermore, the larger ResNet-101x3 model is indeed more robust. Analogous results on the JFT dataset are presented in Appendix D.

Related work

There has been a growing literature exploring the robustness of image classification networks. Early investigations in face and natural image recognition found that performance degrades by introducing blur, Gaussian noise, occlusion, and compression artifacts, but less by color distortions . Subsequent studies have investigated brittleness to similar corruptions , as well as to impulse noise , photometric perturbations , and small shifts and other transformations . CNNs have also been shown to over-rely upon texture rather than shape to make predictions, in contrast to human behavior . Robustness to adversarial attacks is a related, but distinct problem, where performance under worst-case perturbations are studied. In this paper we did not study such adversarial robustness, but have focused on average-case robustness to natural perturbations.

2Several techniques have been shown to improve model robustness on these datasets. Using better data augmentation can improve performance on data with synthetic noise . Auxiliary self-supervision can improve robustness to label noise and common corruptions . Transductive fine-tuning using self-supervision on the test data improves performance under distribution shift . Training with adversarial perturbations improves many robustness benchmarks if one uses separate Batch-Norm parameters for clean and adversarial data . Finally, additional pre-training using very large auxiliary datasets has recently shown significant improvements in robustness. Noisy Student reports good performance on several robustness datasets, while Big Transfer (BiT) reports strong performance on the ObjectNet dataset .

Deep networks are often trained by pre-training the network on a different problem and then fine-tuning on the target task. This pre-training is often referred to as representation learning; representations can be trained using supervised , weakly-supervised , or unsupervised data . Recent benchmarks have been proposed to evaluate transfer to several datasets, to assess generalization to tasks with different characteristics, or those disjoint from the pre-training data . While state-of-the-art performance on many competitive datasets is attained via transfer learning , the implications for final robustness metrics remain unclear.

Creating synthetic datasets by inserting objects onto backgrounds has been used for training and evaluating models , but previous works do not systematically vary object size, location or orientation, or analyze translation and rotation robustness only at the image level .

Given the lack of a consensus on what “natural” perturbations are, there are no established general laws on how models behave under various data shifts. Concurrently, investigated whether higher accuracy on synthetic datasets translates to superior performance on natural OOD datasets. They also identify model size and training data set size as the only technique providing a benefit. In the authors list several of the hypotheses that appear in the literature, and collect new datasets that provide (both positive and negative) evidence for their soundness.

Limitations and future work

We analyzed OOD generalization and transferability of image classifiers, and demonstrated that model and data scale together with a simple training recipe lead to large improvements. However, these models do exhibit substantial performance gaps when tested on OOD data, and further research is required. Secondly, this approach hinges on the availability of curated datasets and significant computing capabilities which is not always practical. Hence, we believe that transfer learning, i.e. train once, apply many times, is the most promising paradigm for OOD robustness in the short term. One limitation of this study is that we consider image classification models fine-tuned to the ImageNet label space which were developed with the goal of optimizing the accuracy on the ImageNet test set. While existing work shows that we do not overfit to ImageNet, it is possible that these models have correlated failure modes on datasets which share the biases with ImageNet . This highlights the need for datasets which enable fine-grained analysis for all important factors of variation and we hope that our dataset will be useful for researchers.

The introduced synthetic data can be used to investigate other qualitative differences between models. For example, when comparing ResNet-50s trained on ImageNet, a ResNet using GroupNorm does better on smaller objects than one with BatchNorm, whereas the model with BatchNorm does better on larger objects (Figure 14(b) in the appendix). While a thorough investigation is beyond the scope of this work, we hope that SI-Score will be useful for such future studies.

Instead of requiring the model to work under various dataset shifts, one can ask an alternative question: assuming that the model will be deployed in an environment significantly different from the training one, can we at least quantify the model uncertainty for each prediction? This important property remains elusive for moderate-scale neural networks , but could potentially be improved by large-scale pretraining which we leave for future work.

References

Appendix A Analysis of existing robustness and transfer metrics

Here, we provide additional details related to the analyses and benchmarks presented in Section 3.

A.2 Dimensionality of the space of robustness metrics

To estimate how many different dimensions are measured by the robustness metrics beyond what is already explained by ImageNet accuracy, we proceeded as follows. For each of the robustness metrics shown in Figure 8 and 10, a linear regression was fit to predict that metric’s value for the 39 models, using ImageNet accuracy as the sole predictor variable. Then, the residuals were computed for each metric by subtracting the linear regression prediction. The plot shows the fraction of variance explained for the first 4 principal components of the space of residuals of the robustness metrics. As a null hypothesis, we assumed that there is no correlation structure in the metric residuals. To construct corresponding null datasets, we randomly permuted the values for each metric independently, which destroys the correlation structure between metrics. Figure 9(a) shows that only the first principal component is significantly above the value expected under the null hypothesis.

A.3 Informativeness of robustness metrics

To estimate how useful different combinations of robustness metrics are for discriminating between model types, we trained logistic regression classifiers to discriminate between the 12 model groups outlined in the main paper. We consider ImageNet accuracy as a baseline metric and therefore compare the performance of a classifier using only ImageNet accuracy as input feature, to a classifier using ImageNet either one (Figure 10, left) or two (Figure 10, right) additional metrics as input features. Figure 10 shows difference in accuracy to the baseline (ImageNet) classifier. These results can serve practitioners with a limited budget as a rough guideline for which metric combinations are the most informative. In our experiments, the most informative combination of metrics in addition to ImageNet accuracy was ObjectNet and YouTube-BB, although other combinations performed similarly within the statistical uncertainty.

A.4 Visual Task Adaptation Benchmark Details

The Visual Task Adaptation Benchmark (VTAB) contains 19 tasks. Either the full dataset or 10001000-example training sets may be used. We use the version with 10001000-example training sets (VTAB-1k).

The tasks are divided into three groups: natural consists of standard natural image classification problems; specialized consists of domain-specific images captured with specialist equipment (e.g. medical images); structured consists of classification tasks that require geometric understanding of a scene. The natural group contains the following datasets: Caltech101 , CIFAR-100 , DTD , Flowers102 , Pets , Sun397 , SVHN . The specialized group contains remote sensing datasets EuroSAT and Resisc45 , and medical image datasets Patch Camelyon and Diabetic Retinopathy . The structured group contains the following tasks: counting and distance prediction on CLEVR , pixel-location and orientation prediction on dSprites , camera elevation and object orientation on SmallNORB , object distance on DMLab and vehicle distance on KITTI .

Appendix B Scale and OOD generalization

The models are firstly pre-trained on ImageNet-21k and JFT, and are then fine-tuned on ImageNet to match the label space for evaluation. We follow the pre-training and BiT-HyperRule fine-tuning setup proposed in .

Specifically, for pre-training, we use SGD with momentum with initial learning rate of 0.1, and momentum 0.9. We use linear learning rate warm-up for 5000 optimization steps and multiply the learning rate by batch size256\frac{\text{batch size}}{256}. We use a weight decay of 0.0001. We use the random image cropping technique from , and random horizontal mirroring followed by resizing the image to 224×224224\times 224 pixels. We use a global batch size of 1024 and train on a Cloud TPUv3-128. We pre-train models for the cross product of the following combinations:

Dataset Size: {1.28M (1×\times ImageNet train set), 2.6M (2×\times ImageNet train set), 5.2M (4×\times ImageNet train set), 9M (7×\times ImageNet train set), 13M (10×\times ImageNet train set)}.

Train Schedule (steps): {113K (90 ImageNet epochs), 229K (180 ImageNet epochs), 457K (360 ImageNet epochs), 791K (630 ImageNet epochs), 1.1M (900 ImageNet epochs)}.

For fine-tuning, we use the BiT-Hyperrule as described in : batch size 512, learning rate 0.003, no weight decay, the classification head initialized to zeros, Mixup with α=0.1\alpha=0.1, fine-tuning for 20 00020\,000 steps with 384×384384\times 384 image resolution.

Additional Results

Here we highlight the results equivalent to Figure 3, with the only difference that we consider subsets of the JFT dataset, instead of ImageNet-21k (Figure 11). We present the results on the synthetic dataset in Appendix D.

Appendix C Effect of the testing resolution

Before applying the respective model, we first resize every image such that the shorter side has length ⌊1.15⋅r⌋\lfloor 1.15\cdot r\rfloor while preserving the aspect ratio and take a central crop of size r×rr\times r. For the widely used 224×224224\times 224 testing resolution, this leads to standard single-crop testing preprocessing, where the images are first resized such that the shorter side has length 256256.

Training details for FixRes

For fine-tuning to the target resolution (FixRes) we use SGD with momentum with initial learning rate of 0.0040.004 (except for the BiT models for which we use 0.00040.0004), and momentum 0.9, accounting for varying batch size by multiplying the learning rate with batch size256\frac{\text{batch size}}{256}. We train for 15 00015\,000⋅batch size2048\cdot\frac{\text{batch size}}{2048}, decaying the learning rate by a factor of 1010 after 1/31/3 and 2/32/3 of the iterations. The batch size is chosen based on the model size to avoid memory overflow; we use 20482048 in most cases. We train on a Cloud TPUv3-64. We emphasize that we did not extensively tune the training parameters for FixRes, but chose a setting that works well across models and data sets.

Additional results

In Figure 12 we provide an extended version of Figure 5 that shows the effect of FixRes for all datasets and models. In Figure 13 we plot the performance of all models and their FixRes variants as a function of the resolution.

Appendix D Additional results on SI-Score, the synthetic dataset

Appendix E Overview of model abbreviations