SLIP: Self-supervision meets Language-Image Pre-training

Norman Mu, Alexander Kirillov, David Wagner, Saining Xie

Introduction

Much of the recent progress in deep learning has been driven by the paradigm of pre-training powerful, general-purpose representations that transfer well to a variety of specific applications. Within computer vision, supervised learning on image classification and self-supervised learning on unlabeled images comprise the two primary approaches to representation learning. After AlexNet , researchers soon realized that supervised pre-training yields a generic visual backbone which can be repurposed for many different tasks . Today, most state-of-the-art results still depend on supervised pre-training, and scaling to massive amounts of data, such as Google’s proprietary JFT dataset, remains one of the most reliable methods for improving downstream performance. Self-supervised learning, a form of unsupervised learning, found tremendous success first in the domain of language , but has also made significant recent progress in vision. A major motivation for studying self-supervised learning has been a desire to supersede supervised pre-training and its reliance on labor-intensive human annotation. Indeed, self-supervised pre-training has outperformed supervised learning for some time now on small datasets, but only recently with the development of contrastive methods has it begun to improve performance on larger datasets such as ImageNet.

Both supervised and self-supervised pre-training today rely heavily on ImageNet (i.e. ImageNet-1K) , a highly curated dataset with particular idiosyncrasies and biases . The YFCC100M dataset was released in 2015 and remains the largest publicly-accessible collection of images. To date, the field of representation learning has found much less use for this dataset. On the other hand, the full ImageNet dataset of 14M images (i.e. ImageNet-22K) has become very popular for its role in training Vision Transformer models which require a larger amount of data than ImageNet-1K . Why are uncurated datasets not more common in the study of representation learning? There are a few possible reasons. Most immediately, uncurated datasets also lack labels and so long as supervised pre-training remains the simpler and more accessible option for most researchers, datasets like YFCC100M are a non-starter. As we confirm again later in this work, the standard self-supervised evaluation task of ImageNet classification from frozen features heavily biases results against models not also pre-trained on ImageNet . Finally, while progress on ImageNet has been encouraging, there has not been strong evidence that current self-supervised methods scale well to larger uncurated datasets .

Recently, CLIP introduced an exciting new approach to representation learning. It re-examines language supervision for learning visual representations, and catapults it into contention with label supervision and self-supervision. CLIP requires only images and free-form text captions, thus revitalizing the use of YFCC100M in representation learning. In addition to no longer requiring label annotations, CLIP accuracy also scales well to large datasets and models. The best results for CLIP are achieved with big models on a curated dataset of 400M images and captions, though promising results are also shown on a subset of YFCC100M. CLIP also enables many exciting new applications with its flexible language-guided capabilities.

In this work, we explore whether the momentum of self-supervised learning on images carries into the setting of language supervision. In particular, we investigate whether language supervision in the form of CLIP also benefits from image self-supervision. We note that it is not immediately clear that these two training objectives should be stronger together. The two objectives each require the model to encode qualitatively different and conflicting information about the image, leading to interference.

In order to explore these questions, we introduce SLIP (Self-supervision meets Language-Image Pre-training), a multi-task framework combining language supervision and self-supervision. We pre-train various SLIP models on a subset of YFCC100M, and thoroughly evaluate representation quality under three distinct settings: zero-shot transfer, linear classification, and end-to-end finetuning. We evaluate downstream performance on ImageNet, in addition to a battery of 25 other classification benchmarks. Additionally, we further validate our findings with experiments on different model sizes, training schedules, and pre-training datasets. Our findings conclusively show that SLIP improves performance across most evaluations by a significant margin, an encouraging signal for the general utility of self-supervision in the context of language supervision. Additionally, we analyze various components of our method in further detail such as the choices of pre-training dataset and data processing method. We conclude with a discussion of our evaluations as well as the ethical and practical limitations of this class of methods.

Related Work

Early work explored learning visual representations from image captions, even before the advent of deep learning . DeViSE jointly embeds images and textual class labels within a shared semantic space, allowing the model to recognize classes that were not explicitly trained for. Initial attempts at leveraging the YFCC dataset for representation learning included predicting the bag-of-words representation or n-gram occurrence from images. ICMLM and VirTex showed that language supervision on COCO Captions produced useful visual representations. Prior to CLIP, Multimodal Contrastive Training adds contrastive image-image and language-image losses to VirTex which further improve performance. CLIP quickly garnered significant attention for its simplicity, scale, and strong results. Developed concurrently, ALIGN , uses a larger but noisier uncurated dataset and shows similar results.

Self-supervised learning.

Earlier self-supervised learning methods have shown subpar scaling with dataset size . Contrastive learning methods ushered in rapid progress due to their simplicity and effectiveness. Recent methods for self-supervised learning also propose a variety of alternatives to the contrastive objective such as self-distillation , or input reconstruction .

Multi-modal multi-task learning.

MURAL extends ALIGN to the multi-lingual setting and introduces a cross-lingual objective to improve multi-lingual image and text retrieval. Concurrently to this work, DeCLIP adds several additional training objectives and more data collected in-house to CLIP in order to improve data efficiency.

SLIP Framework

We introduce SLIP, a framework for combining language supervision and image self-supervision to learn visual representations without category labels. During pre-training, separate views of each input image are constructed for the language supervision and image self-supervision branches, then fed through a shared image encoder. Through the course of training, the image encoder learns to represent visual input in a semantically meaningful manner. We can then measure the quality of these learned representations by evaluating their utility in downstream tasks.

Radford et al. demonstrated the ability of contrastive learning (CLIP) on corresponding images and captions to learn powerful visual representations. CLIP first embeds images and text with separate modality-specific models. These vectors are then projected into a shared embedding space and normalized. The InfoNCE loss is computed using these final embeddings, with corresponding images and captions as positive pairs and all non-matching images and captions as negative pairs.

Non-contrastive alternatives for language supervision include predicting the bag-of-words representation of the caption or the original caption from the image. However, these methods appear to yield weaker results than CLIP. The contrastive objective also enables image classification without re-training dataset-specific classification layers (zero-shot transfer).

2 Image Self-Supervision

View-based self-supervised learning, in which models are trained to represent views or augmentations of the same image similarly, has yielded strong results across a variety of different formulations. In this work we primarily use an adaptation of SimCLR , a representative example of these methods, as the self-supervised objective in SLIP. However, other frameworks can be swapped in quite easily, and we explore this in Section 6. We focus on the Vision Transformer architecture for its simplicity and good performance. We follow MoCo v3 on the hyperparameter settings for training self-supervised Vision Transformers, which will be described later in Section 4.1.

3 Our Method

We outline SLIP with SimCLR for self-supervision (i.e. SLIP-SimCLR) in Algorithm 1. During each forward pass in SLIP, all images are fed through the same encoder. The CLIP and SSL objectives are computed on the relevant embeddings and then summed together into a single scalar loss. The two objectives can be balanced differently by rescaling the SSL objective. We find that a scale of 1.0 for the self-supervised objective, i.e. no re-scaling, works well for SimCLR. Unless otherwise noted, we refer to SLIP-SimCLR simply as SLIP.

SLIP increases the number of images processed which results in approximately 3×\times more activations. This expands the model’s memory footprint and slows down the forward pass during training. See Section 7 for further discussion.

Improved Training Procedure

The authors of CLIP focus primarily on training with a large private dataset of 400M image-text pairs, where the large scale of data lessens the need for regularization and data augmentation. While re-implementing CLIP, we found some simple adjustments primarily to data augmentation which significantly improved performance when pre-trained on YFCC15M. Our improved training procedure achieves 34.6% zero-shot transfer to ImageNet with a modifiedThe initial 7×\times7 conv is replaced by three 3×\times3 convs; global average pooling is replaced by a self-attention pooling layer with 14M parameters. ResNet-50, exceeding the original result of 31.3%. Another re-implementation achieves 32.7% accuracy on ImageNet . In our experiments we focus primarily on the Vision Transformer model family for their strong scaling behavior . We train all Vision Transformer models with our improved procedure as well, in order to set strong baselines for comparing our methods.

We focus primarily on a 15M subset of YFCC100M filtered by Radford et al. consisting of English-only titles and descriptions, which we refer to as YFCC15M. We also evaluate on Conceptual Captions 3M (CC3M) and Conceptual Captions 12M (CC12M) .

Data Augmentation.

During training, we randomly sample a valid caption for each image (i.e. title or description for YFCC15M). Images for the CLIP branch are randomly resized and cropped to between 50% and 100% of the original image, which we refer to as global cropping. In the self-supervised branch we sample two views with the augmentation from MoCo v3 .

Architecture.

We use the original ViT-B/16 and ViT-L/16 architectures from the ViT paper for our image encoders, as well as a ViT-S/16 architecture which is comparable to ResNet-50 in FLOPs and parameters. For our text encoders, we use the smallest text Transformer model from CLIP which contains 38M parameters and uses byte-pair encoding with a 49K token vocabulary, and maximum context length of 77.

For the CLIP objective, our model projects the image and caption embeddings into a 512-dim space with separate learned linear projections. In the self-supervised branch, we use the 3-layer MLP projection head with 4096-dim hidden layers to transform the image embeddings into a 256-dim output space.

Training.

We train with a batch size of 4096 and the AdamW optimizer in all our experiments. Following CLIP, we set the β2=0.98\beta_{2}=0.98 to improve training stability, but we keep ϵ=1e−8\epsilon=1e-8. We use a weight decay of 0.5 for CLIP and 0.1 for SLIP. Instead of the custom mixed-precision recipe used in CLIP, we opt for the built-in automatic mixed precision library in PyTorch.

Zero-shot Transfer Evaluation.

We evaluate zero-shot transfer to various classification benchmarks including ImageNet. We perform prompt ensembling by averaging the caption embeddings for each class across the prompt templates. This average caption embedding is then used to compute cosine similarity with the image embeddings. CLIP provides prompt templates and class names for these benchmarks, which we use directly for ease of comparison.

Linear Classification Evaluation.

We use the same setup as MoCo v3 to evaluate linear classification performance. We use SGD w/ momentum and no weight decay. On ImageNet, we use a learning rate of 0.01 and on the other downstream datasets we tune the learning rate and report the best result. We train for 100 epochs and perform standard cropping and flipping augmentations.

End-to-end Finetuning Evaluation.

To finetune our models on ImageNet, we use the training procedure from BeiT . This procedure employs significant regularization and data augmentation, as well as layerwise learning rate decay which exponentially decays the learning rate across layers. We disable relative positional embedding, layer scaling, and average pooling across tokens. For ViT-B and ViT-S we train for 100 epochs, while on ViT-L we train for 50 epochs.

For finetuning on smaller downstream datasets, we use the simpler DeiT training procedure .

Empirical Evaluations

We evaluate performance on ImageNet under three distinct settings: zero-shot transfer, linear classification, and end-to-end finetuning. The zero-shot transfer task evaluates model performance on classification benchmarks directly after pre-training without updating any of the model weights. A model trained with contrastive language supervision can be used as an image classifier by simply selecting the class whose caption embedding aligns most closely with the input image. Linear classification, also called linear probing, is a standard evaluation method used to evaluate unsupervised or self-supervised representations. A randomly initialized final classification layer is trained while all other model weights are frozen. Finally, another way of evaluating representation quality is whether a pre-trained model can improve upon the performance of supervised learning when finetuning the model end-to-end.

One common evaluation setup in the self-supervised learning literature is to train both the model and the linear classifier on ImageNet (i.e. ImageNet-1K), which even without labels is a highly curated and class-balanced dataset. In Table 1 we train ViT-B/16 with SimCLR and MoCo v3 on both YFCC15M and ImageNet. The resulting models are evaluated on ImageNet using linear classification and end-to-end finetuning. Both SimCLR and MoCo v3 experience a more than 10% drop in linear classification accuracy when pretrained on YFCC15M instead of ImageNet, a dramatic degradation in performance. For this reason, the baseline linear results in our experiments are lower than what is typically reported in the self-supervised literature. Similarly, we observe a less severe but consistent degradation for end-to-end finetuning results as well. We argue that training on uncurated data is a more realistic and informative setting, especially given the original motivations of learning vision from less supervision.

In Table 2, we provide evaluation results for CLIP, SimCLR, and SLIP across three sizes of Vision Transformer and on all three ImageNet settings. All models are trained for 25 epochs on YFCC15M. We find that language supervision and image self-supervision interact constructively in SLIP, improving upon the performance of both methods alone.

Self-supervised models do not support zero-shot transfer evaluation since there is no way to directly map the learned representations onto categorical labels. SLIP consistently outperforms CLIP by around +5% on zero-shot transfer across all three model sizes, a very large margin relative to the original number. The gap between SLIP and CLIP does close slightly between ViT-Small (22M params) and ViT-Large (300M params) from +5.6% to +4.8%. This trend suggests that SLIP would continue to yield benefits over CLIP even for the largest Vision Transformer architectures currently in use.

With ViT-Large SLIP achieves 46.2% top-1 accuracy, which is far below the performance achieved by smaller models pre-trained on massive curated datasets. In absolute terms however, this is a very surprising result considering that YFCC15M contains very little data of the specific form seen during zero-shot transfer evaluation (i.e. object-centered images labeled with captions of the form “a photo of a class name”.

Linear Classification.

In this setting we also observe the synergy between language supervision and image self-supervision. CLIP outperforms SimCLR, but by a much smaller margin than SLIP outperforms SimCLR. We see that SLIP significantly outperforms SimCLR in linear classification accuracy across all three model sizes. The gap between SLIP and SimCLR is largest with ViT-L at almost +10%, suggesting that SLIP continues to scale with larger models while SimCLR slightly saturates in performance.

End-to-end Finetuning.

We see in Table 1 that finetuning performance is somewhat less affected by pre-training on YFCC15M than linear performance is affected, possibly because the model is allowed to adapt to the target distribution. Both SimCLR and MoCo v3 experience -0.3% drops in finetuning accuracy when pre-trained on YFCC15M instead of ImageNet, which is still quite significant for this setting. We re-iterate that the results in Table 2 are not directly comparable with methods which are pre-trained on ImageNet-1K.

When finetuning on ImageNet, CLIP is particularly weak: ViT-S and ViT-B performance is below even that of training from a random weight initialization . The performance of CLIP does not scale well with model size either, as CLIP ViT-L performance is only +0.5% above CLIP ViT-B. On the other hand, self-supervised learning does quite well in this setting, especially with the larger models. SimCLR ViT-L enjoys a +3.0% gain in accuracy over CLIP ViT-L, and SLIP ViT-L does slightly better than SimCLR ViT-L, though by a very marginal amount. These results suggest that the subpar finetuning performance of CLIP is mostly solved with self-supervision.

2 Model and Compute Scaling

We also investigate the scaling behavior of SLIP with more compute (longer training) and larger vision models. We note that 100 epochs of training on YFCC15M corresponds to around 1200 epochs of training on ImageNet-1K. In Table 3 we experimented with holding model size fixed (ViT-B/16) and training for longer as well as training different model sizes for an extended training schedule (100 epochs). Our results indicate that SLIP scales well with both longer training and larger models. We show full results simultaneously varying model and compute scaling with SLIP in the appendix.

3 Additional Benchmarks

While evaluating classification performance on ImageNet gives a broad overview of representation quality, it is also informative to measure performance on a variety of narrowly targeted downstream datasets. In Table 4 we evaluate zero-shot transfer on a battery of downstream image classification tasks compiled by . These datasets span many different domains including everyday scenes such as traffic signs, specialized domains such as medical and satellite imagery, video frames, rendered text with and without visual context, and more. We remove Pascal VOC and replace NABirds with CUB-200-2011. To preprocess the datasets into a unified pipeline we use the extra scripts included in VISSL . We catalog chance performance along with short descriptions of the datasets in the appendix.

For a given model size, the relative ranking between methods appears surprisingly inconsistent. On a few datasets such as Rendered SST2, KITTI depth, and PatchCamelyon (PCAM) it appears at first glance that smaller models and less training improve performance. However we note that performance on these datasets is only around chance performance, likely because these datasets have little overlap with the semantic distribution of YFCC15M, and thus an unreliable indicator of representation quality. Performance is stronger on categories are well represented in YFCC15M, such as Food-101, Oxford Pets, Caltech-101, and STL-10. On these datasets we see that larger models and training for longer with SLIP more generally improve zero-shot transfer accuracy. We view these results on tasks with more reasonable representation in YFCC15M as more informative of representation quality.

Zero-shot performance on the low-resolution datasets (MNIST, CIFAR-10, CIFAR-100) is also very poor. On many datasets performance is several multiples of chance performance yet still much lower than what is achievable with a lightweight model trained on a modest amount of application-specific data. This suggests that language supervision alone is an inefficient way of training models for specific tasks of interest.

We also provide linear classification results on these benchmarks in the appendix.

4 Additional Pre-training Datasets

In addition to YFCC15M, we experiment with two additional image-text datasets: CC12M and CC3M. In Table 5, we train ViT-B/16 with both SLIP and CLIP on CC12M and CC3M, and compare against our previous numbers on YFCC15M. SLIP maintains its margin of improvement over CLIP in all ImageNet evaluation settings. Notably, pre-training SLIP on CC12M instead of YCC15M yields lower zero-shot accuracy but actually results in higher linear and finetuning performance. CLIP sees an even more surprising boost to finetuning performance of +1.6%.

Our improved training recipe (see Section 4.1) largely alleviates overfitting by CLIP on YFCC15M and CC12M, but on the smaller CC3M dataset CLIP overfits quite dramatically. This may be due to the hypernymization used in CC3M to make the captions more amenable to image captioning. CLIP reaches its highest zero-shot ImageNet accuracy after just 15 out of 40 epochs of training on CC3M, after which we observe a steady decline in ImageNet accuracy. In contrast, on CC3M SLIP reaches its highest zero-shot ImageNet performance after 35 epochs.

5 Alternative Self-Supervised Frameworks

As noted in Section 3.2, SLIP enables the use of many different self-supervision methods. We ran several experiments on ViT-B/16 with different alternatives to SimCLR, in particular MoCo v3 , BYOL , and BeiT . Similar to how we tuned the hyperparameters for SLIP-SimCLR, we largely keep the original self-supervised hyperparameters and add in the CLIP objective and text encoder. MoCo v3 and BeiT are already designed for ViT, but with BYOL we tuned the learning rate and weight decay while copying the data augmentation and projector/predictor architecture from MoCo v3. We also lightly tune different scaling parameters for the self-supervised loss. All models are trained for 25 epochs on YFCC15M.

Our results in Table 6 show that all three alternatives underperform SLIP-SimCLR, despite being individually stronger self-supervised methods. Most surprising is the result that SLIP-BEiT performs the worst despite BEiT being the strongest self-supervised method tested here. This may be due to a greater input discrepancy between pre-training and deployment stage. Nonetheless, all these suboptimal variants of SLIP still improve performance over CLIP.

Further Analysis

An alternative to SLIP would be to simply initialize the image encoder of CLIP with SSL-trained weights. We tried training CLIP ViT-B/16 under this setting but found worse performance than training jointly with CLIP and SSL. Progress after a few training epochs exceeds that of SLIP at the same point in training, but stalls throughout the rest of training (25 epochs). In 7, we see this approach underperforms SLIP in all three ImageNet evaluation settings.

Is SLIP just CLIP with data augmentation?

We examine the effects of adding further data augmentation to CLIP and whether this explains the performance improvements seen in SLIP. The SimCLR augmentation can be separated into two components: color (jitter or grayscale) + blur, and resize crop + flip. We train CLIP with these two components individually and also with the full SimCLR augmentation. When training with color + blur, we use the original CLIP cropping strategy from in which we resize the shorter side to 224px then perform a random square crop. Our results are shown in Table 8. While augmentation and resize crop + flip hurt performance, color + blur do improve zero-shot transfer performance by +0.8% which is still far below the gain by SLIP.

Can we fully decouple self-supervision from language supervision?

We experimented with a version of SLIP we call SLIP-decoupled in which the self-supervised objective is computed on a disjoint set of 15M images from the YFCC15M images used in the text supervision object. During training, the images are sampled independently from both sets, effectively decoupling the language-image supervision and self-supervision signals. In Table 9, we find that SLIP-decoupled does just as well as SLIP.

Discussion

Our results on ImageNet and other classification benchmarks show that language supervision and self-supervision are indeed highly synergistic. As shown in Table 2, SLIP improves zero-shot ImageNet performance across model sizes by large margins of +4.8% to +5.6%. Similar gains can be seen in the linear classification setting, with consistent but marginal improvements in the end-to-end finetuning setting.

These trends remain consistent on longer training schedules with the exception of linear probe performance on SLIP ViT-L which actually decreases with more training. With SLIP ViT-L pre-trained on YFCC15M for 100 epochs, we achieve our strongest result of 47.9% zero-shot accuracy on ImageNet. SLIP also shows significant improvements on CC3M and CC12M. Finally, we also confirm our findings with zero-shot and linear evaluations on additional downstream benchmarks.

Prior work on representation learning has argued against end-to-end finetuning for its sensitivity to optimization hyperparameters , and against linear classification for being too contrived . We note that zero-shot transfer, along with linear classification and end-to-end finetuning, can be viewed as one cohesive paradigm for evaluating representation quality. Zero-shot transfer represents the strictest setting, where the exemplar vector for each class must be specified through natural language. Linear classification is a relaxation of zero-shot transfer, in which the class exemplars are optimized on training data. Finally, end-to-end finetuning represents a further relaxation of linear classification where all model parameters are allowed to adapt to training data. Representation quality should be assessed by performance across multiple settings, in the same way that an ROC curve offers a more holistic account of model performance than evaluation at a single operating point.

Zero-shot ImageNet monitor.

SLIP may also serve as a useful framework within which to evaluate new methods for self-supervised learning. Training loss on the pre-text task is a poor predictor of downstream performance, so a simple external metric like kNN accuracy is important for quickly estimating performance and diagnosing training issues such as overfitting or instability. However, kNN classification requires encoding and storing every single training image and naive inference requires very expensive matrix multiplications. The memory bank kNN monitor alleviates this cost but is not feasible when pre-training on unlabeled datasets such as YFCC100M. Instead, zero-shot evaluations on ImageNet are virtually as fast as evaluating validation accuracy in the supervised setting.

Ethical considerations.

SLIP faces all of the same ethical considerations as CLIP, both in terms of the harmful applications it may enable, as well as the potential for amplifying and perpetuating problematic behavior in the real world. CLIP’s ability to leverage noisy and minimally filtered data scraped from the open internet has already spurred researchers to begin collecting data in a more careless manner than previously possible for supervised learning . A more cautious and responsible approach to selecting training data may alleviate the most egregious model behaviors.

Practical limitations.

SLIP computes embeddings of image views for both the self-supervised objective and the CLIP objective. This increases the activation count and memory footprint of the model during the forward pass, which results in slower training (30.5 hours for SLIP vs 22.3 hours for CLIP to train ViT-B/16 on 64 V100 GPUs). After pre-training, SLIP incurs no additional cost since its vision backbone can be used in the same manner as CLIP or a self-supervised model.

From the downstream results in Table 4, we note that pre-training on uncurated data alone appears to be an inefficient route to recognizing specific visual concepts, especially concepts unlikely to be widely shared on social media or the broader internet. Even with a massive amount of curated data CLIP’s zero-shot performance on many datasets is still far below what can easily be achieved by finetuning a small pre-trained model on a modest amount of labeled data. This can be easily addressed by simply finetuning CLIP for specific applications or even including more pre-training data from the domain of interest.

References

Appendix A Additional Implementation Details

YFCC15M contains raw HTML captions and titles which we lightly preprocess before training. We unescape the HTML then remove HTML tags and urls with simple regex matching.

CC3M is collected from an initial set of 5B candidate images, of which 99.9% are filtered out according to simple image and text heuristics for quality and content. Many of these filters are relaxed by CC12M in order to collect a bigger and potentially noisier dataset. CC3M also hypernymizes proper nouns, numbers, and infrequent entities to make the dataset more amenable to training and evaluating image captioning systems, the original design for the dataset. In contrast, CC12M only replaces person names for privacy. Our versions of these datasets contain 3.1M and 11.0 M images respectively, due to asset removal.

Pre-training.

During pre-training we use a cosine learning rate decay schedule with 1 epoch (∼\scriptstyle\sim3500 iterations) of linear warmup when training on YFCC15M. When pre-training for 100 epochs we use 2 warmup epochs. On YFCC15M (14.6M images), we train for 25 epochs and on CC12M (11.0M images) we train for 35 epochs. This amounts to approximately the same number of iterations as 300 epochs on ImageNet-1K . Due to the smaller size of CC3M (3.1M images), we train for 40 epochs to reduce overfitting. We trained on up to sixteen 8×\times V100-32GB servers, and to fit SLIP ViT-Large/16 in memory we accumulated gradients over two steps.

End-to-end Finetuning.

We use a similar training recipe for finetuning all models on ImageNet based on the ImageNet finetuning recipe from BeiT using AdamW and a batch size of 1024 with learning rate of 4e-3 and weight decay of 0.05, along with various data augmentations and regularization methods. As we increase model size we also increase regularization. For ViT-S we set drop path to 0 and layer decay to 0.65, for ViT-B we set drop path to 0.1 and layer decay to 0.65, and for ViT-L we set drop path to 0.

Appendix B Full Scaling Results

We include the full results of our scaling experiments in Table 10, in which we simultaneously increase model size and training epochs. As measured by ImageNet classification accuracy under the three settings (zero-shot transfer, linear classification, and end-to-end finetuning), both large models and longer training generally improve performance.

The exception to this trend is the linear classification performance of SLIP ViT-L/16, which degrades slightly with longer training. This behavior also persists across the various other downstream benchmarks, where SLIP ViT-L/16 does worse on average when trained for 100 epochs than when trained for 25 epochs. We note that both the zero-shot transfer and end-to-end finetuning performance of SLIP ViT-L/16 improve with longer training, contrary to the behavior seen with linear classification. Thus we cannot declare this behavior to be a case of simple overfitting, as the representations are still improved for the other evaluation settings.

Appendix C Additional Linear Classification Benchmarks

In Table 11 we show linear classification results on all 26 downstream datasets (including ImageNet). With ViT-B and ViT-S, SLIP pre-training for 100 epochs does best. As with ImageNet, SLIP ViT-L also does worse on average when trained for 100 epochs than when trained for 25 epochs. The dataset average is 0.5 points lower for the 100 epoch model.

As expected, linear classification accuracy is much higher than zero-shot transfer accuracy (shown in Table 4). However, the gap between zero-shot and linear performance varies between datasets. On datasets which are straightforward vision tasks but poorly represented among the YFCC100M imagery, such as Patch Camelyon, MNIST, KITTI distance, and GTSRB, linear classification massively improves accuracy, often from a baseline of around chance performance. On datasets which share more overlap with YFCC100M, such as Food-101, Caltech-101, and Caltech-UCSD Birds 2011, we see significant improvements as well.

However, with HatefulMemes and Rendered SST2, two datasets which require OCR capabilities, the linear classification performance of all models is still around chance. These results suggest, perhaps unsurprisingly, that zero-shot transfer results are much more dependent on what visual and semantic concepts were seen during training than linear classification, since they do not enjoy the benefit of further training examples. We also note that relative rankings within each model size are also quite unstable where the best results alternate between the 25 and 100 epoch models. This is very similar to what we see in the zero-shot transfer evaluations, as discussed in Section 5.