With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, Andrew Zisserman
Introduction
How does one make sense of a novel sensory experience? What might be going through someone’s head when they are shown a picture of something new, say a dodo? Even without being told explicitly what a dodo is, they will likely form associations between the dodo and other similar semantic classes; for instance a dodo is more similar to a chicken or a duck than an elephant or a tiger. This act of contrasting and comparing new sensory inputs with what one has already experienced happens subconsciously and might play a key role in how humans are able to acquire concepts quickly. In this work, we show how an ability to find similarities across items within previously seen examples improves the performance of self-supervised representation learning.
A particular kind of self-supervised training -- known as instance discrimination -- has become popular recently. Models are encouraged to be invariant to multiple transformations of a single sample. This approach has been impressively successful at bridging the performance gap between self-supervised and supervised models. In the instance discrimination setup, when a model is shown a picture of a dodo, it learns representations by being trained to differentiate between what makes that specific dodo image different from everything else in the training set. In this work, we ask the question: if we empower the model to also find other image samples similar to the given dodo image, does it lead to better learning?
Current state-of-the-art instance discrimination methods generate positive samples using data augmentation, random image transformations (e.g. random crops) applied to the same sample to obtain multiple views of the same image. These multiple views are assumed to be positives, and the representation is learned by encouraging the positives to be as close as possible in the embedding space, without collapsing to a trivial solution. However random augmentations, such as random crops or color changes, can not provide positive pairs for different viewpoints, deformations of the same object, or even for other similar instances within a semantic class. The onus of generalization lies heavily on the data augmentation pipeline, which cannot cover all the variances in a given class.
In this work, we are interested in going beyond single instance positives, i.e. the instance discrimination task. We expect by doing so we can learn better features that are invariant to different viewpoints, deformations, and even intra-class variations. The benefits of going beyond single instance positives have been established in , though these works require class labels or multiple modalities (RGB frames and flow) to obtain the positives which are not applicable to our domain. Clustering-based methods also offer an approach to go beyond single instance positives, but assuming the entire cluster (or its prototype) to be positives could hurt performance due to early over-generalization. Instead we propose using nearest neighbors in the learned representation space as positives.
We learn our representation by encouraging proximity between different views of the same sample and their nearest neighbors in the latent space. Through our approach, Nearest-Neighbour Contrastive Learning of visual Representations (NNCLR), the model is encouraged to generalize to new data-points that may not be covered by the data augmentation scheme at hand. In other words, nearest-neighbors of a sample in the embedding space act as small semantic perturbations that are not imaginary, i.e. they are representative of actual semantic samples in the dataset. We implement our method in a contrastive learning setting similar to . To obtain nearest-neighbors, we utilize a support set that keeps embeddings of a subset of the dataset in memory. This support set also gets constantly replenished during training. Note that our support set is different from memory banks and queues , where the stored features are used as negatives. We utilize the support set for nearest neighbor search for retrieving cross-sample positives. Figure 1 gives an overview of the method.
We make the following contributions: (i) We introduce NNCLR to learn self-supervised representations that go beyond single instance positives, without resorting to clustering; (ii) We demonstrate that NNCLR increases the performance of contrastive learning methods (e.g. SimCLR ) by and achieves state of the art performance on ImageNet classification for linear evaluation and semi-supervised setup with limited labels; (iii) Our method outperforms state of the art methods on self-supervised, and even supervised features (learned via supervised ImageNet pre-training), on out of transfer learning tasks; Finally, (iv) We show that by using the NN as positive with only random crop augmentations, we achieve ImageNet accuracy. This reduces the reliance of self-supervised methods on data augmentation strategies.
Related Work
Self-supervised Learning. Self-supervised representation learning aims to obtain robust representations of samples from raw data without expensive labels or annotations. Early methods in this field focused on defining pre-text tasks -- which typically involves defining a surrogate task on a domain with ample weak supervision labels, like predicting the rotation of images , relative positions of patches in an image , or tracking patches in a video . Encoders trained to solve such pre-text tasks are expected to learn general features that might be useful for other downstream tasks requiring expensive annotations (e.g. image classification).
One broad category of self-supervised learning techniques are those that use contrastive losses, which have been used in a wide range of computer vision applications. These methods learn a latent space that draws positive samples together (e.g. adjacent frames in a video sequence), while pushing apart negative samples (e.g. frames from another video). In some cases, this is also possible without explicit negatives . More recently, a variant of contrastive learning called instance discrimination has seen considerable success and have achieved remarkable performance on a wide variety of downstream tasks. They have closed the gap with supervised learning to a large extent. Many techniques have proved to be useful in this pursuit: data augmentation, contrastive losses , momentum encoders and memory banks . In this work, we extend instance discrimination to include non-trivial positives, not just between augmented samples of the same image, but also from among different images. Methods that use prototypes/clusters also attempt to learn features by associating multiple samples with the same cluster. However, instead of clustering or learning prototypes, we maintain a support set of image embeddings and using nearest neighbors from that set to define positive samples.
Queues and Memory Banks. In our work, we use a support set as memory during training. It is implemented as a queue similar to MoCo . MoCo uses elements of the queue as negatives, while this work uses nearest neighbors in the queue to find positives in the context of contrastive losses. use a memory bank to keep a running average of embeddings of all the samples in the dataset. Likewise, maintains a clustered set of embeddings and uses nearest neighbors to those aggregate embeddings as positives. In our work, the size of the memory is fixed and independent of the training dataset, nor do we perform any aggregation or clustering in our latent embedding space. Instead of a memory bank, SwAV stores prototype centers that it uses for clustering embeddings. SwAV’s prototype centers are learned via training with Sinkhorn clustering and persist throughout pre-training. Unlike , our support set is continually refreshed with new embeddings and we do not maintain running averages of the embeddings.
Nearest Neighbors in Computer Vision. Nearest neighbor search has been an important tool across a wide range of computer vision applications , from image retrieval to unsupervised feature learning. Nearest-neighbor lookup as an intermediate operation has also been useful for image alignment and video alignment tasks. propose a method to learn landmarks on objects in an unsupervised manner by using nearest-neighbors from other images of the same object while show unsupervised learning of action phases by using soft nearest-neighbors across videos of the same action. In our work, we also use cross-sample nearest neighbors but train on datasets with many classes of objects with the objective of learning transferable features. Related to our work, uses nearest neighbor retrieval to define self-supervision for video representation learning across different modalities (e.g. RGB and optical flow). In contrast, in this work we use nearest neighbor retrieval within a single modality (RGB images), and we maintain an explicit support set of prior embeddings to increase diversity.
concurrently propose leveraging nearest-neighbors in embedding space to improve self-supervised representation learning using BYOL . The authors propose using an additional term for the nearest-neighbor to the BYOL loss. Instead, in our formulation all our loss terms use the nearest-neighbor.
Approach
We first describe constrastive learning (i.e. the InfoNCE loss) in the context of instance discrimination, and discuss SimCLR as one of the leading methods in this domain. Next we introduce our approach, Nearest-Neighbor Contrastive Learning of visual Representations (NNCLR), which proposes using nearest-neighbours (NN) as positives to improve contrastive instance discrimination methods.
InfoNCE loss (i.e. contrastive loss) is quite commonly used in the instance discrimination setting. For any given embedded sample , we also have another positive embedding (often a random augmentation of the sample), and many negative embeddings . Then the InfoNCE loss is defined as follows:
where (, ) is the positive pair, (, ) is any negative pair and is the softmax temperature. The underlying idea is learning a representation that pulls positive pairs together in the embedding space, while separating negative pairs.
SimCLR uses two views of the same image as the positive pair. These two views, which are produced using random data augmentations, are fed through an encoder to obtain the positive embedding pair and . The negative pairs (, ) are formed using all the other embeddings in the given mini-batch.
Formally, given a mini-batch of images , two different random augmentations (or views) are generated for each image , and fed through the encoder to obtain embeddings and , where is the random augmentation function. The encoder is typically a ResNet-50 with a non-linear projection head. Then the InfoNCE loss used in SimCLR is defined as follows:
Note that each embedding is normalized before the dot product is computed in the loss. Then the overall loss for the given mini-batch is .
As SimCLR solely relies on transformations introduced by pre-defined data augmentations on the same sample, it cannot link multiple samples potentially belonging to the same semantic class, which in turn might decrease its capacity to be invariant to large intra-class variations. Next we address this point by introducing our method.
2 Nearest-Neighbor CLR (NNCLR)
In order to increase the richness of our latent representation and go beyond single instance positives, we propose using nearest-neighbours to obtain more diverse positive pairs. This requires keeping a support set of embeddings which is representative of the full data distribution.
SimCLR uses two augmentations (, ) to form the positive pair. Instead, we propose using ’s nearest-neighbor in the support set to form the positive pair. In Figure 2 we visualize this process schematically. Similar to SimCLR we obtain the negative pairs from the mini-batch and utilize a variant of the InfoNCE loss (1) for contrastive learning. Building upon the SimCLR objective (2) we define NNCLR loss as below:
where is the nearest neighbor operator as defined below:
As in SimCLR, each embedding is normalized before the dot product is computed in the loss (3). Similarly we apply normalization before nearest-neighbor operation in (4). We minimize the average loss over all elements in the mini-batch in order to obtain the final loss .
Implementation details. We make the loss symmetric by adding the following term to Eq. 3: Though, this does not affect performance emperically. Also, inspired from BYOL , we pass through a prediction head to produce embeddings . Then we use instead of in (3). Using a prediction MLP adds a small boost to our performance as shown in Section 4.4.
Support set. We implement our support set as a queue (i.e. first-in-first-out). The support set is initialized as a random matrix with dimension , where is the size of the queue and is the size of the embeddings. The size of the support set is kept large enough so as to approximate the full dataset distribution in the embedding space. We update it at the end of each training step by taking the (batch size) embeddings from the current training step and concatenating them at the end of the queue. We discard the oldest elements from the queue. We only use embeddings from one view to update the support set. Using both views’ embeddings to update does not lead to any significant difference in downstream performance. In Section 4.4 we compare the performance of multiple support set variants.
Experiments
In this section we compare NNCLR features with other state of the art self-supervised image representations. First, we provide details of our architecture and training process. Next, following commonly used evaluation protocol , we compare our approach with other self-supervised features on linear evaluation and semi-supervised learning on the ImageNet ILSVRC-2012 dataset. Finally we present results on transferring self-supervised features to other downstream datasets and tasks.
Architecture. We use ResNet-50 as our encoder to be consistent with the existing literature . We spatially average the output of ResNet-50 which makes the output of the encoder a 2048-d embedding. The architecture of the projection MLP is fully connected layers of sizes where is the embedding size used to apply the loss. We use in the experiments unless otherwise stated. All fully-connected layers are followed by batch-normalization . All the batch-norm layers except the last layer are followed by ReLU activation. The architecture of the prediction MLP is fully-connected layers of size . The hidden layer of the prediction MLP is followed by batch-norm and ReLU. The last layer has no batch-norm or activation.
Training. Following other self-supervised methods , we train our NNCLR representation on the ImageNet2012 dataset which contains images, without using any annotation or class labels. We train for epochs with a warm-up of epochs with cosine annealing schedule using the LARS optimizer . Weight-decay of is applied during training. As is common practice , we don’t apply weight-decay to the bias terms. We use the data augmentation scheme used in BYOL and we use a temperature of when applying the softmax during computation of the contrastive loss in Equation 3. The best results of NNCLR are achieved with queue size and base learning rate of .
2 ImageNet evaluations
ImageNet linear evaluation. Following the standard linear evaluation procedure we train a linear classifier for epochs on the frozen -d embeddings from the ResNet-50 encoder using LARS with cosine annealed learning rate of with Nesterov momentum of and batch size of .
Comparison with state of the art methods is presented in Table 1. First, NNCLR achieves the best performance compared to all the other methods using a ResNet-50 encoder trained with two views. NNCLR provides more than improvement over well known constrastive learning approaches such as MoCo v2 and SimCLR v2 . Even compared to InfoMin Aug. , which explicitly studies ‘‘good view’’ transformations to apply in contrastive learning, NNCLR achieves more than improvement on top-1 classification performance. We outperform BYOL (which is the state-of-the-art method among methods that use two views) by more than .
We also achieve improvement compared to the state of the art clustering based method SwAV in the same setting of using two views. To compare with SwAV’s multi-crop models, we pre-train for epochs with views (two and six views) using only the larger views to calculate the NNs. In this setting our method outperforms SwAV by in Top-1 accuracy. Note that while multi-crop is responsible for performance improvement for SwAV, for our method it provides a boost of only . However, increasing the number of crops quadratically increases the memory and compute requirements, and is quite costly even when low-resolution crops are used as in .
Semi-supervised learning on ImageNet. We evaluate the effectiveness of our features in a semi-supervised setting on ImageNet and subsets following the standard evaluation protocol . We present these results in Table 2. The first key result of Table 2 is that our method outperforms all the state of the art methods on semi-supervised learning on ImageNet subset, including SwAV’s multi-crop setting. This is a clear indication of good generalization capacity of NNCLR features, particularly in low-shot learning scenarios. Using the ImageNet subset, NNCLR outperforms SimCLR and other methods. However, SwAV’s multi-crop setting outperforms our method in ImageNet subset.
3 Transfer learning evaluations
We show representations learned using NNCLR are effective for transfer learning on multiple downstream classification tasks on a wide range of datasets. We follow the linear evaluation setup described in . The datasets used in this benchmark are as follows: Food101 , CIFAR10 , CIFAR100 , Birdsnap , Sun397 , Cars , Aircraft , VOC2007 , DTD , Oxford-IIIT-Pets , Caltech-101 and Oxford-Flowers . Following the evaluation protocol outlined in , we first train a linear classifier using the training set labels while choosing the best regularization hyper-parameter on the respective validation set. Then we combine the train and validation set to create the final training set which is used to train the linear classifier that is evaluated on the test set.
We present transfer learning results in Table 3. NNCLR outperforms supervised features (ResNet-50 trained with ImageNet labels) on 11 out of the 12 datasets. Moreover our method improves over BYOL and SimCLR on 8 out of the 12 datasets. These results further validate the generalization performance of NNCLR features.
4 Ablations
In this section we present a thorough analysis of NNCLR. After discussing the default settings, we start by demonstrating the effect of training with nearest-neighbors in a variety of settings. Then, we present several design choices such as support set size, varying k in top-k nearest neighbors, type of nearest neighbors, different training epochs, variations of batch size, and embedding size. We also briefly discuss memory and computational overhead of our method.
Default settings. Unless otherwise stated our support set size during ablation experiments is and our batch size is . We train for 1000 epochs with a warm-up of 10 epochs, base learning rate of and cosine annealing schedule using the LARS optimizer . We also use the prediction head by default. All the ablations are performed using the ImageNet linear evaluation setting.
Nearest-neighbors as positives. Our core contribution in this paper is using nearest-neighbors (NN) as positives in the context of contrastive self-supervised learning. Here we investigate how this particular change, using nearest neighbors as positives, affects performance in various settings with and without momentum encoders. This analysis is presented in Table 4. First we show using the NNs in contrastive learning (row 2) is better in Top-1 accuracy than using view 1 embeddings (similar to SimCLR) shown in row 1. We also explore using momentum encoder (similar to MoCo ) in our contrastive setting. Here using NNs also improves the top-1 performance by .
Data Augmentation. Both SimCLR and BYOL rely heavily on a well designed data augmentation pipeline to get the best performance. However, NNCLR is less dependent on complex augmentations as nearest-neighbors already provide richness in sample variations. In this experiment, we remove all color augmentations and Gaussian blur, and train with random crops as the only method of augmentation for 300 epochs following the setup used in . We present the results in Table 5. We notice NNCLR achieves top-1 performance on the ImageNet linear evaluation task suffering a performance drop of only . On the other hand, SimCLR and BYOL suffer larger relative drops in performance, and respectively. The performance drop reduces further as we train our approach longer. With pre-training epochs, NNCLR with all augmentations achieves while with only random crops NNCLR manages to get , further reducing the gap to just . While NNCLR also benefits from complex data augmentation operations, the reliance on color jitter and blurring operations is much less. This is encouraging for adopting NNCLR for pre-training in domains where data transformations used for ImageNet might not be suitable.
Pre-training epochs. In Table 6, we show how our method compares to other methods when we have different pre-training epoch budgets. NNCLR is better than other self-supervised methods when pre-training budget is kept constant. We find that base learning rate of works best for epochs, and works for , and epochs.
Support set size. Increasing the size of the support set increases performance in general. We present results of this experiment in Table 7(a). By using a larger support set, we increase the chance of getting a closer nearest-neighbour from the full dataset. As also shown in Table 7(b), getting the closest (i.e. top-1) nearest-neighbour obtains the best performance, even compared against top-2.
We also find that increasing the support set size beyond doesn’t lead to any significant increase in performance possibly due to an increase in the number of stale embeddings in the support set.
Nearest-neighbor selection strategy. Instead of using the nearest-neighbor, we also experiment with taking one of the top- NNs randomly. These results are presented in Table 7(b). Here we investigate whether increasing the diversity of nearest-neighbors (i.e. increasing ) results in improved performance. Although our method is somewhat robust to changing the value of , we find that increasing the top- beyond always results in slight degradation in performance. Inspired by recent work we also investigate using a soft-nearest neighbor, a convex combination of embeddings in the support-set where each one is weighted by its similarity to the embedding (see for details). We present results in Table 7(e). We find that the soft nearest neighbor can be used for training but results in worse performance than using the hard NN.
Batch size. Batch size has shown to be an important factor that affects performance, particularly in the contrastive learning setting. We vary the batch size and present the results in Table 7(c). In general, larger batch sizes improve the performance peaking at .
Embedding size. Our method is robust to choice of the embedding size as shown in Table 7(d). We vary the embedding size in powers of 2 from 128 to 2048 and find similar performance over all settings.
Prediction head As shown in Table 7(f), adding a prediction head results in a modest boost in the top-1 performance.
Different implementations of support set. We also investigate some variants of how we can implement the support set from which we sample the nearest neighbor. We present results of this experiment in Table 8. In the first row, instead of using a queue, we pass a random set of images from the dataset through the current encoder and use the nearest neighbor from that set of embeddings. This works reasonably well but we are limited by how many examples we can fit in accelerator memory. Since we cannot increase the size of this set beyond , which results in sub-optimal performance. Also using a random set of size is about four times slower than using a queue (when training with a batch size of ) as each forward pass requires four times more samples through the encoder. This experiment also shows that NNCLR does not need features of past samples (akin to momentum encoders ) to learn representations. We also experiment with updating the elements in the support set randomly as opposed to the default FIFO manner. We find that FIFO results in more than better ImageNet linear evaluation Top-1 accuracy.
Compute overhead. We find increasing the size of the queue results in improved performance but this improvement comes at a cost of additional memory and compute required during training. In Table 9 we show how queue scaling with affects memory required during training and number of training steps per second. With a support size of about elements we require a modest 100 MB more in memory.
5 Discussion
Ground Truth Nearest Neighbor. We investigate two aspects of the NN: first, how often does the NN have the same ImageNet label as the query; and second, if the NN is always picked to be from the same ImageNet class (with an Oracle algorithm), then what is the effect on training and the final performance? Figure 3, shows how the accuracy of the NN picked from the queue varies as training proceeds. We observe that towards the end of training the accuracy of picking the right neighbor (i.e. from the same class) is about . The reason that it is not higher is possibly due to random crops being of the background, and thus not containing the object described by the ImageNet label.
We next investigate if NNCLR can achieve better performance if our top-1 NN is always from the same ImageNet class. This is quite close to the supervised learning setup except instead of training to predict classes directly, we train using our self-supervised setup. A similar experiment has also been described as UberNCE in . This experiment verifies if our training dynamics prevent the model from converging to the performance of a supervised learning baseline even when the true NN is known. To do so, we store the ImageNet labels of each element in the queue and always pick a NN with the same ImageNet label as the query view. We observe that with such a setting we achieve accuracy in epochs. With the Top-1 NN from the support set, we manage to get in 300 epochs. This suggests that there is still a possibility of improving performance with a better NN picking strategy, although it might be hard to design one that works in a purely unsupervised way.
Training curves. In Figure 4 we show direct comparison between training using cross-entropy loss with an augmentation of the same view as positive (SimCLR) and training with NN as positive (NNCLR). The training loss curves indicate NNCLR is a more difficult task as the training needs to learn from hard positives from other samples in the dataset. Linear evaluation on ImageNet classification shows that it takes about epochs for NNCLR to start outperforming SimCLR, and remains higher until the end of pre-traininig at epochs.
NNs in Support Set. In Figure 5 we show a typical batch of nearest neighbors retrieved from the support set towards the end of training. Column 1 shows examples of view 1, while the other elements in each row shows the retrieved nearest neighbor from the support set. We hypothesize that the improvement in performance is due to this diversity introduced in the positives, something that is not covered by pre-defined data augmentation schemes. We observe that while many times the retrieved NNs are from the same class, it is not uncommon for the retrieval to be based on other similarities like texture. For example, in row 3 we observe retrieved images are all of underwater images and row 4 retrievals are not from a dog class but are all images with cages in them.
Conclusion
We present an approach to increase the diversity of positives in contrastive self-supervised learning. We do so by using the nearest-neighbors from a support set as positives. NNCLR achieves state of the art performance on multiple datasets. Our method also reduces the reliance on data augmentation techniques drastically.
References
Appendix
Appendix A Pseudo-code
In Algorithm 1 we present the pseudo-code of NNCLR.
It is possible to use momentum encoder with NNCLR training. The pseudo-code when momentum encoder is used is shown in Algorithm 2.
Appendix B Evolution of Nearest-Neighbors
In Figure 6 we show how the nearest-neighbors (NN) vary as training proceeds. We observe consistently that in the beginning of training the NNs are usually chosen on the basis of color and texture. As the encoder becomes better at recognizing classes later in training, the NNs tend to belong to similar semantic classes.
Appendix C SimSiam with Nearest-neighbor as Positive
In this experiment we want to check if it is possible to use the nearest-neighbor in a non-contrastive loss. To do so we use the self-supervised framework SimSiam , in which the authors use a mean squared error on the embeddings of the 2 views, where one of the branches has a stop-gradient and the other one has a prediction MLP. We replace the stop-gradient branch with its nearest-neighbor from the support set. We call this method NNSiam. In Figure 7 we show how NNSiam differs from SimSiam. We also show how they both differ from SimCLR and NNCLR. Note that there is an implicit stop-gradient in NNSiam because of the use of hard nearest-neighbors. For this experiment, we use an embedding size of 2048, which is the same dimensionality used in SimSiam. We train with a batch size of 4096 with the LARS optimizer with a base learning rate of . We find that even with the non-contrastive loss using the nearest neighbor as positive leads to 1.3% improvement in accuracy under ImageNet linear evaluation protocol. This shows that the idea of using harder positives can lead to performance improvements outside the InfoNCE loss also.
Appendix D Experiments with Vision Transformers
Vision Transformers (ViT) are a class of architectures introduced recently to process images using Transformers. We explore the effectiveness of using self-supervised learning methods to train vision transformers. We find the Adam optimizer to be effective for training ViT models. We train for epochs using 2 crops with a base learning rate of with a warmup schedule of 10 epochs followed by cosine decay learning schedule, weight decay of and a stochastic depth dropout of . We use the output corresponding to the [CLS] token of the final layer as the embedding in the loss and as the representation used as input for the linear classifier. We present the results of this experiment in Table 11. In our experiments, we find NNCLR well suited to train Visual Transformers outperforming the supervised learning model (trained without the augmentation introduced in DeIT ) by and the SimCLR model by .
Appendix E Self-supervised Learning as a Pre-training Step for Supervised Learning
Until recently it was believed supervised learning on a particular dataset would always be better than self-supervised learning on that dataset. BYOL showed that some large models can end up achieving higher accuracy than their supervised counterparts. In this experiment, we similarly show that self-supervised learning can serve as a useful pre-training step for supervised learning, alleviating the need for extra data or complex regularization techniques that are used to boost the performance of the final model.
In this experiment we first pretrain a model with NNCLR for 1000 epochs on the ImageNet 2012 dataset. We then proceed to perform regular supervised training for 100 epochs. This is similar to the semi-supervised learning setup but with 100% of the labels available for fine-tuning the pre-trained model. In Table 12 we show results on this experiment. We observe initialization from a self-supervised model improves the performance of ResNet50 from 76.2% to 79.1%. This performance is comparable to training a ResNet-50 model with JFT-300M dataset which is considerably more data than ImageNet. We find this pre-training technique is especially useful with the newly proposed ViT architecture which requires additional data or regularization techniques to achieve good performance. ViT-B/16 pre-trained with NNCLR outperforms DeIT but is only worse than the same architecture trained with the JFT-300M dataset by 0.5%. This experiment highlights another use-case for self-supervised learning: to provide a good initialization for supervised learning.
Appendix F Transfer Learning
In this section we study the performance of Visual Transformers (ViT) in the transfer learning setting. To do this, we repeat the experiment described in Section 4.3 with ViT models. The results of this experiment are presented in Table 13. First, we observe that ViT-B/16 trained with only a supervised learning objective using the data augmentation strategy outlined in DeIT on the ImageNet dataset transfers poorly as compared to ResNet50 architecture trained with the supervised loss. ViT-B/16 is worse on 10 datasets out of the 12 datasets in the transfer learning benchmark. However, ViT-B/16 trained with NNCLR objective outperforms the ResNet50 architecture trained with the same loss on 8 out of the 12 datasets in the benchmark. Additionally, ViT-B/16 trained with NNCLR has better performance on all datasets as compared to the same model trained with just the supervised loss. This shows that the Vision Transformer’s potential for transfer learning is enhanced by using self-supervised training like NNCLR. Of particular note is the increase in performance on Birdsnap (), Sun397 (), Cars (), Aircraft (), DTD () and Flowers (). We also train a ViT-B/8 model that has the same architecture as ViT-B/16 but uses -sized non-overlapping patches in the images to produce the tokens used in the Transformer architecture. We find that ViT-B/8 outperforms ViT-B/16 on all datasets in the transfer learning setup. We also observe that NNCLR brings significant gains over just supervised learning. Of particular note is the increase in performance on Birdsnap (), Sun397 (), Aircraft (), DTD () and Flowers (). Finally, we test the transfer learning performance on models pre-trained with NNCLR and fine-tuned with the supervised loss as outlined in Section E. We find this setup increases the performance over the already strong baseline of ViT-B/8 trained with just NNCLR. We find a boost of 8.3% on Food, 5.3% on Birdsnap, 2.9% on Sun397, 15.8% on Cars, 2.2% on VOC2007, 2.1% on Pets. However, we also observe fine-tuning on the supervised loss also results in degradation in performance on some datasets: 1.1% on DTD and 1.1% on Flowers. Inspite of the degradation in performance, ViT-B/8 pretrained with NNCLR and fine-tuned with the supervised loss on ImageNet is always better than just using the supervised loss only. Overall, we find Vision Transformers are well suited for transfer learning if they have been pre-trained with a self-supervised loss like NNCLR.
Appendix G Visualizations
We visualize ‘‘attention" of various models trained with NNCLR to probe what the model might be focusing on. In order to visualize attention, we need a query embedding and the feature map produced by the image. We convolve the query embedding across the feature map to produce an attention map. In Figure 8, we show examples of attention corresponding to different query embeddings on images from the COCO dataset. We use the original resolution of the images to get a larger feature map that can highlight object boundaries and locations of small objects. Each arrow points to the query embedding in the image and the corresponding attention of that query embedding in the feature map. In each sub-figure’s caption we also mention the object class they are pointing at. In , the authors observe that vision transformer models trained with self-supervised losses like DINO show emergence of properties like objectness and part localization in the attention layer of the transformer. We find these emerging properties also exist in ResNet50 models trained with self-supervised losses like NNCLR.
G.2 Cross-image Attention with NNCLR ResNets
In the previous section we presented attention of a query embedding with parts of the same image. To visualize if models trained with NNCLR encode object categories, we conduct the following experiment. We use a query embedding of an object from one image by average pooling the features in the bounding box of an object denoted by the red box. We convolve this query embedding over the feature maps obtained by passing other images through a ResNet50 trained with NNCLR loss. We show the results of this experiment in Figure 9. We observe that features of the same object in different images are close to each other in the NNCLR feature space. The results of this experiment highlight that the learned features encode semantic similarity beyond color, in order to localize objects of the same category across different images as shown in Figure 9. In the first row, the features are able to localize multiple objects of the same category giraffe. In the second row, we show that while the ground-truth class is that of stop-sign the features are considering different kinds of street signs (hotel name, street name sign, railroad-crossing sign) as similar. This is interesting because the query features are pooled only from the stop-sign region but the retrieved regions of similarity are of different colors. We also observe the model does not activate over the text of the stop-sign. In the last two rows, we show how the model is able to localize small objects (like ball and frisbee) in different images.
G.3 Comparison of Attention in ResNets vs ViT
In this experiment we want to compare the attention maps of 2 architectures, ResNet-50 and ViT-B/16, each with 3 different weights: randomly initialized, ImageNet supervised and ImageNet NNCLR. For ViT we use the average self-attention over all heads in the final layer for the [CLS] token. For ResNet-50 we use the method described in Sec. G.1 and use the average pooled embedding as the query embedding since it is the equivalent of the [CLS] token in ViTs. We show our results in Figure 10. First we observe the attention maps delineate salient objects in the image. Similar to DINO we observe that ViTs trained with just a supervised learning objective do not have objects highlighted in their attention map. However, we do observe that both supervised and self-supervised ResNets have delineated objects in their attention maps. We also note that sometimes randomly initialized ResNets (rows 2 and 5) and ViTs (rows 4 and 7) are able to localize individual objects in their attention map because of similarity in color and texture across the spatial extent of the object. This fact makes it difficult to conclude that a model that produces the well-delineated objects in the self-attention map will necessarily have learned semantically meaningful features. We suggest using a combination of self-attention and cross-attention maps (described in Section G.2) as a more robust visualization technique for interpreting semantic features.