Self-supervised Pretraining of Visual Features in the Wild
Priya Goyal, Mathilde Caron, Benjamin Lefaudeux, Min Xu, Pengchao Wang, Vivek Pai, Mannat Singh, Vitaliy Liptchinsky, Ishan Misra, Armand Joulin, Piotr Bojanowski
Introduction
A recent trend shows that well-tailored model pretraining approaches (weakly-supervised, semi-supervised, self-supervised) can drastically improve the performance on downstream tasks for most deep learning applications. It has been observed for Natural Language Processing , Speech Recognition and Computer Vision . There are two key ingredients that have contributed towards this success. The first is pretraining on massive datasets: the GPT-3 language model is pretrained on B words, while the speech model Wav2vec2.0 is learned on K hours of audio . The second ingredient is the use of models with massive capacity, even reaching hundreds of billions of parameters for the largest NLP models .
While the benefit of pretraining has been demonstrated in computer vision, it has been in the limited scope of curated datasets originally collected for supervised or weakly supervised learning. These datasets represent only a limited fraction of the general distribution of internet scale images . Prior attempts to train self-supervised models on uncurated data have only used a few millions of images for which using small model architectures is sufficient. This still leaves an open question - can we achieve good performance by pretraining on an extremely large collection of random, uncurated and unlabeled images? Answering this question has important implications. Practically, it may lead to strategies for pretraining that use unlabeled data to achieve state-of-the-art performance on transfer learning, and to create systems that continually learn in a self-supervised manner from an unending data stream.
In this work, we address this question by pretraining high capacity models on billions of random internet images, \ie, completely unconstrained images from the internet. We do not rely on image meta-data or any form of weak/manual annotations to filter data or train the model. Training powerful image representations on unlabeled data has recently been made possible with advances in self-supervised learning . The most recent self-supervised models pretrained on ImageNet even surpass supervised pretrained models on multiple downstream tasks . Recent developments have also shown their potential when trained on random internet images. Further, these methods are amenable to online learning, making them perfect candidates for training large models on unlimited data.
For our analysis, we focus on the RegNet family of architectures and in particular the architecture with M parameters. The RegNet architectures are particularly well suited for this task for two reasons. First, they offer an excellent trade-off of efficiency and performance. Second, they are very flexible for scaling the number of parameters. We train these models online on a dataset of B random internet images using the SwAV self-supervised approach . We use SwAV for this study for its fast convergence to good performance in the online setting with large batch size. To make this study tractable at this scale, we leverage several existing tools to reduce the memory usage of our models, including mixed precision and gradient checkpointing.
The main finding of our study is that our SElf-supERvised (“SEER”) pretrained models are not only good for initializing training on curated dataset like ImageNet, they are also excellent few shot learners, achieving 75.1% with only 10% of ImageNet. Our model also achieves better performance than supervised model trained on ImageNet on several downstream tasks, confirming the benefits of self-supervised pretraining, even when performed on uncurated data.
Related Work
Our work explores the limits of training large architectures on large uncurated datasets with self-supervised learning. We build on prior work from different areas: self-supervised learning, training at scale and large convolutional network architectures.
Self-supervised learning has a long history in computer vision with methods based on autoencoders , clustering or instance-level discrimination . Recently, methods based on contrastive learning have shown that unsupervised pretraining produces features that surpass the supervised feature representations on many downstream tasks . These methods discriminate either between each instance feature or between their cluster assignments . Most works on unsupervised pretraining focus on supervised datasets like ImageNet or curated datasets collected by filtering images related to pre-defined labels . The key takeaway from these works is that supervised labels are not required as long as you trained on the filtered data. Some works have explored unsupervised training in the wild for images and videos . These studies were conducted at a small scale, and there are now evidences that self-supervised pretraining benefits greatly from large archtiectures . Our work builds upon these findings to explore if we can learn good visual representations by training large architectures on large collection of random, uncurated and unlabeled images.
Learning of visual features at scale.
Benefiting from the advances in distributed training , several works have shown the advantages of pretraining on large curated image datasets with weak-supervised learning , semi-supervised learning or supervised training on hundreds of millions of filtered images . Of particular interest, Mahajan et al. show that the pretraining on billions of images significantly improves the performance of large architectures compared to training them from scratch. Most works on training at large data scale rely on a data filtering step to only keep the images associated with targeted concepts. This filtering either uses hastags that are synsets of ImageNet classes , or the predictions from a pretrained object classifier . As opposed to this line of work, we are interested in learning features that cover any available image, and hence, we do not curate our training dataset to match a pre-defined set of concepts.
Scaling architectures for image recognition.
Many works have shown the benefits of training large architectures on the quality of the resulting visual features . Training large architectures is especially important when pretraining on a large dataset, where a model with limited capacity will underfit . This becomes further more important when the pretraining is performed with contrastive learning, where the network has to learn to discriminate between each instance of the dataset in order to learn good visual representations. For instance, Kolesnikov et al. have demonstrated the importance of training wider networks for the quality of visual features learned with self-supervision. More recently, Chen et al. have achieved impressive performance with deeper and wider configurations of ResNet . However, scaling architectures for image recognition goes beyond simply changing the width and the depth of a model, and a large amount of literature is dedicated to building scale efficient models with large capacity . Of particular interest, the RegNets achieve competitive performance on standard image benchmarks while offering an efficient runtime and memory usage making them a candidate for training at scale. In our work, we show the benefits of this model family for large scale self-supervised pretraining.
Method
In this section, we provide a brief overview of the components used in this work to pretrain visual features in the wild. We describe the self-supervised method, SwAV , and the family of convnet architectures, RegNet . We then discuss several technical strategies required to train large models on billions of images with self-supervision.
We pretrain our model with an online self-supervised approach called SwAV that we briefly summarize in this section. We refer to Caron et al. for more details.
SwAV is an online clustering method to train convnets without annotations. It works by training an embedding that yields consistent cluster assignments between multiple views of the same image. By mining clusters invariant to data augmentations, the system learns semantic representations. In practice, SwAV works by comparing the features of different views of the same image using their intermediate cluster assignments. If these features capture the same information, it should be possible to predict the assignment of one from the feature of another view. More precisely, we consider a set of clusters, each associated with a learnable -dimensional prototype vector . Given a batch of images, each image is transformed into two views: and . All views are then featurized with a convnet, resulting in two sets of features and . Each set of features is assigned independently to the cluster prototypes using an Optimal Transport solver. This solver enforces that the features are uniformly split across clusters, avoiding trivial solutions where all representations are mapped to an unique prototype. The resulting assignments are then swapped between the two sets: the cluster assignment of the view has to be predicted from the feature representation of the view , and vice-versa. Formally, the convnet and prototypes weights are trained to minimize the following loss for all examples :
2 Scale efficient model family: RegNetY
Scaling data and model capacity jointly requires using architectures that are efficient in terms of both memory and runtime. RegNets are a family of models designed for this purpose and we briefly describe them in this section. We refer to Radosavovic et al. for more details.
RegNets are a family of architectures defined by a design space of convnets consisting of 4 stages, with each stage containing a series of identical blocks, while keeping the structure of their blocks fixed – namely the residual bottleneck block of He et al. . In this work, we focus on the RegNetY architectures, that add a Squeeze-and-excitation op to the standard RegNets to further improve their performance. The RegNetY model family is parameterized by 5 parameters, allowing the search of a good instance with a certain number of FLOPs with reasonable resources. The models we used were all searched on ImageNet using the same procedure as Radosavovic et al. . We believe our results can further be improved by searching for RegNetYs directly on our self-supervised pre-training task.
Our model of focus is the RegNetY-256GF architecture. Its parametrization is given by the scaling rules of RegNets :
It has 4 stages with stage depths (, , , ) and stage widths (, , , ), leading to a total of M parameters. It takes ms for a single training iteration over images on V100 32GB NVIDIA GPUs. Training this model on billion images requires training iterations for a batch size of images, summing to days of training over GPUs.
3 Optimization and Training at Scale
In this work, we propose several adjustments to the training of self-supervised methods to adapt it to a large scale.
We explore two learning rate schedules: the cosine wave and a simpler fixed learning rate schedule. The cosine wave adapts to the number of updates and we focus on this scheduling for fair comparison between different models. However, it is not adapted to online large scale training because it uses the total of updates for scheduling and it also weighs images differently depending on when they are seen during training. For this reason, we also explore a fixed learning rate schedule. In this scheduling, we keep the learning rate fixed until the loss is non-decreasing, then we divide the learning rate by . Our observation is that this schedule works as well in practice and allows for more flexible training. However, we train our largest model, the RegNetY-256GF with cosine learning rate schedule since we use only 1B images.
Reducing memory consumption per GPU.
We reduce the amount of GPU memory required during training with gradient checkpointing and mixed precision. We use O1 optimization level from NVIDIA Apex libraryhttps://github.com/NVIDIA/apex to perform operations like GEMMs and convolutions in -bits floating-point precision. We use PyTorch’s gradient checkpointing implementation which trades compute for memory. It discards intermediate activations during the forward pass, and recomputes them during the backward pass. In our experiments, using gradient checkpointing, we observe negligible compute overhead in memory-bound settings.
Optimizing Training speed.
Enabling mixed-precision for memory optimization has additional benefits, as modern accelerators take full advantage of the FP reduced size by increasing throughput when compared to FP. This improves memory-bandwidth bottleneck and speeds up training. We also use the optimized SyncBatchNorm implementation with kernels through CUDA/C++ extensions from NVIDIA Apex library. For synchronizing BatchNorm layer across GPUs, we create process groups instead of performing global sync which is slow. Finally, our dataloader pre-fetches more training batches leading to higher data throughput than the default PyTorch dataloader.
Large scale Pretraining data.
For our billion scale pretraining, we consider a dataloader that directly samples random, public, and non-EU images from Instagram. As we train online and in the wild, we do not apply any curation or pre-processing on the images, such as hashtag filtering or de-duplication. This dataset is not static and gets refreshed every days, however, we can confirm that the refreshment doesn’t degrade the model performance.
Implementation details.
We pretrain a RegNetY-256GF with SwAV, using crops per image of resolutions . We follow the same data augmentation as in Caron et al. . During pretraining, we use a 3-layer multi-layer perceptron (MLP) projection head of dimensions , and . We do not use BatchNorm layers in the head. We use K prototypes, temperature set to , the Sinkhorn regularization parameter to and perform iterations of Sinkhorn algorithm. We synchronize BatchNorm stats across gpus and create process groups of size 64 for synchronization. We use a weight decay of , LARS optimizer and O1 mixed-precision optimization from Apex library. We also apply activation checkpointing . We train our model with stochastic gradient descent using a large batch size of different images distributed over NVIDIA V100 32GB GPUs, resulting in different images per GPU. The learning rate is linearly ramped up from to for the first K training updates. After warmup, we follow a cosine learning rate schedule and decay the learning rate to final value . Overall, we train on B images for a total of K iterations.
Main Results
We study the quality of the features generated by our self-supervised pretraining on a variety of downstream tasks and benchmarks. We also consider a low-shot setting with limited access to images and their labels for the downstream task, as well as, standard evaluation using the entire data available for the downstream task. We also compare with prior work trained on large curated datasets.
In this section, we measure the quality of models pretrained in the wild by transferring them to the ImageNet object classification benchmark.
We pretrain RegNet architectures of different capacities, namely RegNetY-{8,16,32,64,128,256}GF, on B random, public and non-EU Instagram images with SwAV. We finetune these models on the task of image classification on ImageNet, using the standard M training images with labels and evaluate on images in the standard validation set. We apply the same data augmentation as in SwAV . We finetune for epochs with SGD, batch size of , learning rate of reduced by factor of after epochs, weight decay of and momentum of . We report top-1 accuracy on validation set using the center crop.
Comparision with other self-supervised pretraining.
In Table 1, we compare our largest pretrained model, a RegNetY-256GF, with existing self-supervised pretrained models. We achieve 84.2% top-1 accuracy on ImageNet, surpassing by +1%, the best existing pretrained model from SimCLRv2 . In the Figure 1, we show the same comparison with different model capacities. The conclusion remains unchanged regardless of the model capacity, showing that combining RegNet with SwAV is a good candidate for pretraining.
Impact of the model capacity.
In Figure 2, we show the impact of model capacity on the performance of pretraining compared to training from scratch. While model capacity benefits both initializations, it has a more significant impact on pretrained models when scaled to hundreds of millions of parameters. A reason is that training these architecture from scratch could overfit on ImageNet which is a relatively small dataset. We confirm that the log-scale performance gain from increasing model capacity also appears in the case where the pretraining data is uncurated.
2 Low-shot learning
In this section, we are interested in evaluating the performance of our pretrained model in the low-shot setting, i.e., with a fraction of data on the downstream task.
We consider two datasets for low-shot learning, namely ImageNet and Places205 . We assume a limited access to the dataset during transfer learning, both in terms of labels and images. This setting differs from the standard setting used in self-supervised learning where the entire datasets is accessible and only the access to labels is limited . For the rest, we follow their experimental setting for finetuning the features.
Results on Places205.
In Figure 3, we show the impact of pretraining on different fractions of the Places205 dataset . We compare to pretraining on ImageNet with supervision with the same RegNetY-128GF architecture. A surprising result is that we observe a stable gain of in top-1 accuracy, regardless of the fraction of training data available to finetune on Places205. The difference between self-supervised and supervised pretraining may be explained by the difference in the nature of training data: features learned from images in the wild may be more suitable to classify scene. Additionally, the non-uniform distribution of underlying concepts in the wild may also provide an advantage to our pretraining on a unbalanced dataset like Places205.
Results on ImageNet.
In Table 2, we show the performance of our self-supervised pretrained model on low-shot learning. For completeness, we report performance of existing semi-supervised and self-supervised methods. We note that all of these methods use the entire set of M images from ImageNet for pretraining and only restrict the access to the labels, while we only see 1% and 10% of the images. This greatly favors these approaches since the network has seen more images from the same distribution during pretraining as the fraction used for transfer. Nonetheless, our approach achieves a top-1 accuracy of 77.9% with only 10% of ImageNet, which is competitive with these methods (2% gap). On 1% of the data, i.e, 10K images, the gap increases significantly but note that the other methods are using the full ImageNet from pretraining.
Impact of the model capacity.
In Figure 4, we explore the impact of model capacity in the different low-shot settings - 1%, 10% and 100% of ImageNet. A first observation is that increasing model capacity gives a higher relative improvement as we decrease the access to both labels and images. This result extends the observation of Chen et al. on the low-shot setting. Interestingly, the relative gains are comparable in both settings ( in of the data), even though low-shot learning is strictly harder.
3 Transfer to Other Benchmarks
In these experiments, we further evaluate our pretrained features by transferring them to other downstream tasks.
In Table 3, we compare the features from our pretrained RegNetY-128GF and RegNetY-256GF with features from the same architecture pretrained on ImageNet with and without supervision. To assess features quality, we freeze the model weights and learn a linear classifier on top of the features using the training set of each downstream task. We consider the following benchmarks: iNaturalist , OpenImages , Places205 and Pascal VOC . We observe that self-supervised features transfer better than supervised features regardless of the pretraining data.
Detection and segmentation.
In Table 4, we evalaute pretrained features on detection and segmentation. We train a Mask-RCNN model on the COCO benchmark with pretrained RegNetY-64GF and RegNetY-128GF as backbones. For both downstream tasks and architectures, our self-supervised pretraining outperforms supervised pretraining by AP points. However, the gap in performances between different architectures is small ( AP) compared to what we observed on ImageNet.
4 Comparing to Weakly-Supervised Pretraining
Many online images have some metadata, e.g., hashtags or geo-localization, that can be leveraged during pretraining. In particular, Mahajan et al. show that pretraining by predicting a curated set of hashtags can greatly improve the quality of the resulting visual features. Their approach requires to filter images and only works in the presence of textual metadata. In Table 5, we compare our self-supervised pretraining on random images to theirs on the same architecture, a ResNeXt101-32x8d, with finetuning. For completeness, we also report their best number with their largest architecture. First, we observe that both pretrainings improve top-1 accuracy over a model trained from scratch, showing in general the benefits of pretraining. Our approach is also in the same ballpark as theirs even though we do not rely on data curation nor supervision. Note that, when the features are frozen, their approach maintains high performance on ImageNet, with top-1 accuracy while our model performance drops significantly – around top-1. This result is not surprising: they pretrain on data that follows the same concepts as ImageNet classes and thus the learned features are more aligned with the target distribution. Since we pretrain our model on random images, we require a full-finetuning step of epochs to adapt to the target distribution. This experiment shows that the benefits of pretraining with finetuning exist even if the features come from a different image distribution.
Ablation Studies
These ablation studies focus on the model architecture, how its performance scales with capacity, and the specificities of our pretraining data and our self-supervised method.
We consider several RegNetY architectures with growing capacity, namely the RegNetY-{8,16,32,64,128}GF. We also consider ResNet-{50, 101} and the ResNeXt architectures with a growing number of parameters, namely RX101-32x{4, 8}d . We refer to the appendix for details about the different architectures. Every model is pretrained for epoch on B random, public and non-EU Instagram (IG) images with SwAV using K prototypes. We use same hyperparameters for pretraining all the ablation models.
Impact of the architecture.
In Figure 5, we measure how changing the architecture affects the quality of pretrained features with a linear evaluation on ImageNet. This evaluation does not favor models that perform well when trained from scratch on ImageNet, and hence, we directly probe the pretrained features. For ResNets and ResNeXts, we observe that the features from the penultimate layer work better in this setting and we report those for fair comparison with RegNets. Overall, RegNets surpass the other architectures, justifying our choice of architecture for our main model. Finally, we observe, regardless of the architecture, increasing model capacity significantly improves the quality of the features with a logarithmic gain in performance.
2 Scaling the Training Data
Pretraining on a larger dataset can improve the quality of the learned visual features for two reasons: more parameter updates and more unique images. In this section, we disentangle these two effects.
On the left panel of Figure 6, we show the performance as we train a RegNetY-128GF model online on 1B images. We observe that the performance steadily increases with the number of updates as expected and the performance does not saturate even after a number of updates corresponding to B images.
Increasing the number of unique images.
On the right panel of Figure 6, we report the performance of two models, RegNetY-8GF and RegNetY-16GF when trained for the same number of updates but with a different number of unique images. We train the models for a number of updates that corresponds to epoch over B unique images, or epochs for M unique images, with a single half-cosine wave learning rate. An interesting observation is that, with this learning rate schedule, the minimum number of unique images required to obtain good performance is greater than the size of ImageNet by only an order of magnitude.
Overall, these experiments show that the number of updates matters more than seeing the same images multiple times. There is thus no need to fix the pretraining dataset, and instead validating their continual online pretraining.
3 Scaling the self-supervised model head
In this section, we study the impact of growing the size of the self-supervised model head during pretraining.
In Table 6, we compare RegNetY-8GF architectures trained with different capacity self-supervised heads. We report top-1 accuracy on ImageNet obtained with a linear classifier trained on frozen features. The models are trained on B images with a cosine wave learning rate schedule. In particular, we show the impact of a larger MLP and more prototype vectors. We adjust the head from 2-layer MLP of dimensions ( and ) to 3-layer MLP of dimensions (, , ), and increase the number of prototypes from K to K. We observe that simply increasing the number of the parameters in the head, and the number of clusters significantly improves the performance of the resulting model (%) with the same model, hence same feature size. The reason is that B random images contain much more concepts that the original SwAV classifier can memorize in its small head, hence the information about the clusters leaks to the features, degrading their performance. Increasing the head reduces this effect at a minimal compute and storage cost.
Conclusion
We show that pretraining features on random images with no annotation achieves competitive performance on any downstream task. This result confirm that the recent progress of self-supervised learning is not specific to curated training set, like ImageNet, and could benefit a large range of applications associated with uncurated data. Our work benefits from the scalability of modern self-supervised learning methods in terms of data, and modern efficient high-capacity architectures. In particular, the scalability of RegNets have played a key role in pushing the limits of self-supervised pretraining, and in the future, we plan to search for larger RegNet architectures suited for this task.
References
Supplementary Material
Model architectures.
We describe below the model architecture settings used for ablation studies. In order to compare the architectures fairly, we follow the same hyperparameters for pre-training. We describe next the setup used for pretraining of ResNet-{50,101}, ResNeXt101-32x{4,8}d and RegNetY-{8,16,32,64,128}GF.
We pretrain standard ResNet-{50,101} from He et al. and standard RX101-32x{4,8}d from Xie et al. with SwAV, using crops per image of resolutions . We follow the same data augmentation as in Caron et al. . During pretraining, we use a 2-layer multi-layer perceptron (MLP) projection head of dimensions and . We do not use BatchNorm layers in the head. We use K prototypes, temperature set to , the Sinkhorn regularization parameter to and perform iterations of Sinkhorn algorithm. We synchronize BatchNorm stats across gpus and create process groups of size 32 for synchronization. We use a weight decay of , LARS optimizer and O1 mixed-precision optimization from Apex libraryhttps://github.com/NVIDIA/apex. We train our model with stochastic gradient descent using a large batch size of different images distributed over NVIDIA V100 32GB GPUs, resulting in different images per GPU. The learning rate is linearly ramped up from to for the first K training updates. After warmup, we follow a half cosine wave learning rate schedule and decay the learning rate from to final value . Overall, we train on B images for a total of K iterations.
2 Pretraining of RegNet architectures.
We train 5 different RegNet architectures namely the RegNetY-{8,16,32,64,128}GF of different capacity. RegNet architectures are generated by following the scaling rules described in Radosavovic et al. . We first share the parametrization used for each of the RegNet architecture below. We then describe how we train these architectures with SwAV for our ablation study in Section 5.
The model has depth = and RegNet parameters:
RegNetY-16GF.
The model has depth = and RegNet parameters:
RegNetY-32GF.
The model has depth = and RegNet parameters:
RegNetY-64GF.
The model has depth = and RegNet parameters:
RegNetY-128GF.
The model has depth = and RegNet parameters:
RegNetY-256GF.
The model has depth = and RegNet parameters:
For pretraining the above RegNetY architectures, we follow the same pretraining hyperparams as ResNet and ResNeXt training with two differences. We use crops per image of resolutions . However, we confirm that the crop resolutions didn’t impact the model performance on ImageNet linear classification task on which we show our ablations in Section 5. Infact, using the bigger resolution crops leads to more GPU memory requirement with no impact on model performance on transfer task. The only other difference is the dimensions of 3-layer MLP in the head. Each RegNetY architecture has difference output channel and hence we adapt 3-layer MLP according to the architecture. More concretely, the head dimensions are: RegNetY-8GF has [, and ], RegNetY-16GF has [, and ], RegNetY-32GF has [, and ], RegNetY-64GF has [, and ], RegNetY-128GF has [, and ]