The effectiveness of MAE pre-pretraining for billion-scale pretraining
Mannat Singh, Quentin Duval, Kalyan Vasudev Alwala, Haoqi Fan, Vaibhav Aggarwal, Aaron Adcock, Armand Joulin, Piotr Dollár, Christoph Feichtenhofer, Ross Girshick, Rohit Girdhar, Ishan Misra
Introduction
The pretrain-then-finetune paradigm in visual recognition has enabled high performance visual recognition models across a range of tasks such as image classification , video action recognition , object detection , 3D etc. Typically, pretraining consists of training a model using a pretraining task on large scale data. The resulting pretrained models learn general purpose visual representations that can be used for a range of target tasks, often with limited labeled data, by transfer learning.
In this paper, we show that an initial stage of pre-pretraining before the standard pretraining task can improve vision models across a variety of different tasks. Our method combines two common pretraining tasks in vision: (1) weakly supervised pretraining that uses weak, often noisy, signals such as text or image hashtags as supervision, and (2) self-supervised pretraining that only uses the data without additional supervision. Both forms of pretraining start training with a randomly initialized model and have proven effective at learning general purpose vision models. While there have been attempts to combine both these forms of pretraining , they are typically used independently in the pretrain-then-finetune two stage paradigm .
In this work we explore the combination of self- and weakly-supervised learning in a simple pre-pretraining framework, as follows. We first begin with the Masked Autoencoder (MAE) self-supervised learning technique to pre-pretrain vision models without using any labels. After initializing from the pre-pretrained model, we use standard weakly supervised pretraining on billions of images with noisy labels. We perform a large-scale empirical study to measure the effectiveness of pre-pretraining on 10 different visual recognition tasks spanning image classification, video recognition, object detection, low-shot classification and zero-shot recognition. Our study reveals that pre-pretrain initialization improves the performance for the weakly supervised models, and this improvement holds even at billion scale weakly labeled data, and across vision tasks (Fig. 1). It also improves the model convergence during pretraining, leading to an efficient way of training large scale vision models. Pre-pretrain further enjoys the computational efficiency of the MAE approach, making it simple and scalable. Finally, we show that by using pre-pretraining, both self-supervised learning and weakly supervised learning can be combined for improved model performance for billion-scale data.
Pre-pretrain is related to ‘intermediate finetuning’ which introduces a stage after pretraining to better align the pretrained features with the downstream task using labeled data. In contrast, pre-pretrain serves as a better way to initialize a model before pretraining. Since we leverage MAE for pre-pretraining, we do not need additional information or labels for this stage and can re-use the pretraining data. This makes pre-pretrain convenient and simple to use with existing pretraining datasets.
Our study on large-scale pre-pretraining reveals that model initialization plays a significant role, even for web-scale pretraining, and pre-pretraining is a simple and promising technique in that direction. In particular, we show that (i) MAE not only scales with model size as shown in , but also with the size of the training data (Fig. 2). (ii) Pre-pretraining improves both the model convergence and the final downstream performance for different sized models (millions to billions of parameters) trained on different sized datasets (millions to billions of images). (iii) Using pre-pretraining combines the benefits of both self-supervised learning and large scale weakly-supervised learning, and our models achieve excellent performance on a variety of different visual recognition tasks (Fig. 1). Most prominently, our model sets new state-of-the-art results on image classification on iNaturalist-18 (91.7%) and ImageNet-ReaL (91.1%), 1-shot ImageNet-1k classification (63.6%), and zero-shot transfer on Food-101 (96.2%).
Related Work
Supervised pretraining of transferrable representations on large labeled datasets and employing them for downstream recognition tasks, has emerged as a powerful approach in computer vision. It has spurred rapid progress on various tasks including image classification , object detection/segmentation , image captioning and video action recognition . While useful, such representations are often limited by the scale and diversity of the supervision in the pretraining datasets. Hence, recent work has probed the effectiveness, robustness, and fairness of these representations .
Self-supervised pretraining is a promising alternative to learn these representation without relying on large well-labeled datasets. Initial works focused on reconstructions methods before moving to other pretraining tasks such as solving jigsaw puzzles , constrastive learning or joint embedding approaches . With the advent of Vision Transformers , approaches based on reconstructions such as got renewed interest for their simplicity and state of the art performance. Of particular interest to us is MAE for its state of the art performance on many transfer tasks and its computational efficiency. Given the lack of supervision during pretraining, these representations often require significant finetuning to align to downstream tasks.
Weakly supervised pretraining (WSP) is a middle-ground between supervised and self-supervised pretraining. Instead of ignoring annotations completely as in self-supervised pretraining, or requiring exhaustive labels as in supervised pretraining, WSP relies on the large quantity of “free” annotation available on the internet. These annotations occur as image-text pairs , where the text can additionally be processed to produce pseudo labels. Of particular interest to us is the latter, i.e. approaches which leverage multi-label classification on noisy labels which have shown state of the art fine-tuning performance, and at the same time can be adapted using image-text data to gain zero-shot capabilities . In this work, we explore WSP in conjunction with self-supervised pre-pretraining, and show faster convergence and stronger performance.
Setup
Our goal is to empirically study the effectiveness of self-supervised pre-pretraining as a precursor to billion scale weakly supervised pretraining for representation learning. Given the simplicity and efficiency of Masked AutoEncoding (MAE) , we leverage it as the self-supervised pre-pretraining approach. Our study shows that MAE scales with the size of the pretraining dataset and model size, and combining it with weak supervision improves large scale vision models. Additionally, such a combination leads to faster convergence and is a simple, scalable way to learn visual representations at scale. We describe our setup and the approaches in detail next.
Architecure. We use the Vision Transformer (ViT) architecture as the visual encoder for all our experiments. ViTs employ minimal vision-specific inductive biases combined with the standard transformer architecture , and yet have emerged as an architecture of choice for a wide variety of visual and multimodal recognition tasks . We train ViT models at various scales in terms of number of parameters, including ViT-B (86M), ViT-L (307M), and ViT-H (632M). We also train on larger 1.9B and 6.5B parameter ViT models, which we call ViT-2B and ViT-6.5B, respectively (Table 1). As is common practice , we train models of sizes ViT-B, ViT-L with a patch size of 16 and larger models with a patch size of 14. We pretrain with a resolution for all models.
Pre-pretraining (MAE) learns visual representations from image datasets without using any labels. We choose this approach as it is simple to implement and scales very effectively with large ViT model sizes due to patch dropping as described next. MAE randomly masks 75% of an image and trains the model to reconstruct the masked input image by minimizing the pixel reconstruction error. The target pixel values for a given patch are normalized by the mean and standard deviation of all pixels in it. Coupled with the ViT architecture, MAE can be trained by only processing the 25% unmasked image patches. A separate, smaller, decoder is then used to reconstruct the missing part of the input. This asymmetrical design makes training the encoder extremely efficient, allowing for scaling visual encoder sizes.
Weakly-supervised pretraining (WSP) leverages images with associated ‘weak’ supervision for training models. In particular, we focus on internet images and use their associated text information as supervision. We convert the text into a discrete set of labels, specifically leveraging hash-tag information . We then use a multi-label classification loss to train models. We refer to this method as WSP.
MAEWSP, or MAWS for short, first trains the encoder using the MAE self-supervised method using only the images. This pre-pretraining stage initializes the model while simultaneously being computationally efficient because of the masking used in MAE. In the second stage, we pretrain the encoder using both the image and associated weak supervision. This combination outperforms using either strategy in isolation, i.e., an MAE model or a weakly supervised model trained from scratch.
Experiments
We empirically evaluate and analyze large scale MAE pre-pretraining using Instagram data on a variety of different visual recognition tasks. We describe the datasets used for pretraining and evaluation, followed by analysis of the pretraining design decisions, and finally the downstream transfer evaluation of our learned representation.
Pretraining dataset. We use Instagram-3B (IG-3B) a billion-scale multi-label dataset sourced from Instagram (IG). This multi-label dataset contains 28K classes and 3B unique images, resampled to 5B total images, and was produced by running the dataset generation pipeline from SWAG without modification. Compared to , our version of the dataset has 16% fewer images (3.0B vs. 3.6B), but we were able to reproduce the results from with our version. We obtain labels using an automated process wherein we first obtain hashtags from the associated image captions, and then map the hashtags to WordNet synsets following . After this processing, we get the weakly labeled IG dataset that contains images and their associated labels.
Evaluation datasets. We evaluate MAEWSP on a variety of different downstream visual recognition tasks. To evaluate our model on image classification, we use the standard ImageNet-1k (IN1k) dataset, and also the long-tailed and fine-grained iNaturalist-18 (iNat18) dataset. For object detection and segmentation, we use the popular COCO dataset, and also LVIS , a large vocabulary dataset for long tailed object recognition. We evaluate video classification performance using two popular action recognition datasets, Kinetics-400 (K400) and Something Something-v2 (SSv2). For zero-shot transfer, we evaluate on IN1k and Food-101 (F-101). We also evaluate the robustness of our models on test sets which overlap with IN1k classes, specifically ImageNetv2 (INv2), ImageNet-ReaL (IN-ReaL), and ObjectNet (ON). Please see Table 2 for more details.
MAE pretraining details. We follow to train MAE models on IG-3B without using any labels. We mask 75% of the image for this training and train the model for 1 epoch over the dataset. We follow the same hyperparameters used in for pretraining on IN1k.
Supervised pretraining details. We train with a supervised cross-entropy loss on IG-3B using the hashtags as labels. This model is trained by default with random weight initialization and we use the training hyperparameters from .
Using pre-pretraining. When using pre-pretraining, we first train a model from scratch using MAE on the IG dataset. We then use the weights of the MAE encoder and perform supervised pretraining using the cross-entropy loss as described above. We reuse the same hyperparameters and training details as , i.e. there is no hyperparameter search needed for MAEWSP, and we train for 1 epoch on IG-3B.
Zero-shot training and evaluation details. To impart zero shot understanding capabilities to our models, we use the LiT approach from . For LiT, we use the original (image, caption) pairs from the IG-3B dataset. We freeze the image encoder, and train a text encoder to encode the image captions and match the text embeddings to the associated image embedding using a CLIP loss . We train the text encoder for 1 epoch. For evaluation, we follow – we use the text encoder to compute embeddings from the templated text descriptions of classes and use the cosine similarity of the image and text embeddings as the classification score.
For full training details and hyperparameters, please refer to Appendix A.
2 Scaling MAE pretraining to large data
Since our pre-pretraining uses MAE in the very first stage, we first study how MAE behaves on the large scale IG-3B dataset. We compare the performance of MAE pretraining on the large scale IG-3B with the original MAE models trained on IN1k for 1600 epochs. We train models of varying sizes, from ViT-B to ViT-H as in . To test the scaling behavior further, we also train MAE on ViT-2B and ViT-6.5B, with 2B and 6.5B parameters, respectively. We measure the performance of the resulting models in Fig. 2 on four different vision tasks.
We observe that using the IG-3B data provides consistent gains over IN1k for all vision tasks, and the gain increases for larger models. These experiments show that MAE scales with the size of the pretraining dataset, and benefits from using billions of images from IG-3B. He et al. ’s findings were limited to the fact that MAE scales with the size of the model, and thus our findings on MAE scaling with the size of the pretraining data are complementary to theirs.
Our ViT-2B model pretrained on IG-3B improves upon the best results from on image classification, attaining 87.8% on IN1k (+0.9%) and 85.6% on iNat18 (+2.6%) at 224224 resolution. The gains on detection are equally encouraging, with our ViT-2B reaching 53.6 APbox on LVIS (+2.1 over ) and 59.8 APbox on COCO (+1.2 over ). Our ViT-6.5B model further improves the downstream performance, attaining 88.3% on IN1k and 86.6% on iNat18 at just 224224 resolution.
Lastly, we highlight the simplicity of our setup, since we use the same hyperparameters as to train MAE on IG-3B, even for the and larger ViT-2B and ViT-6.5B models, without requiring extra tweaks. We note that training the ViT-6.5B was sensitive to numerical stability issues, so we trained the model with full precision to avoid divergence.
3 MAE pre-pretraining
Given the promising aspects of MAE as a pretraining approach from § 4.2, specifically that MAE (i) trains with larger models and datasets without needing any tuning (ii) shows gains when scaling model and / or dataset size (iii) is efficient to train, we investigate it as a pre-pretraining approach for supervised pretraining (WSP).
Fig. 1 shows the performance of a ViT-L with MAE pretraining, supervised pretraining (WSP), or MAE pre-pretraining followed by supervised pretraining (MAEWSP). We see that MAE and WSP have different strengths. MAE has strong performance for object detection, and full finetuned image classification. However, MAE underperforms on tasks where the model is not finetuned, such as linear classifiers, zero-shot, or low-shot classification – situations where WSP performs better. For these evaluations MAE lags behind WSP by more than 10 points, which is why the results for MAE are not visible in Fig. 1. For video classification, MAE performs significantly better than WSP on SSv2, but lags behind it on K400.
MAEWSP outperforms either of MAE or WSP pretraining on most evaluations, across image classification, video recognition, zero-shot evaluation, object detection, etc. Given that all baselines are trained at billion-scale, these results show that MAEWSP is a simple yet promising strategy to improve performance while requiring no extra data or tuning. Next, we ablate the key aspects of MAEWSP.
Effect of model size. In Fig. 3 we study the effect of pre-pretraining compared to random initialization for different model sizes. After initialization, all models are pretrained using WSP on the IG-3B dataset, and we measure the transfer performance on IN1k using linear probing. We observe that MAE pre-pretraining gives consistent gains over the WSP baseline across all model sizes, ranging from 86M to 6.5B parameters. The gains over the WSP baseline increase for larger model sizes showing that pre-pretraining shows promising scaling behavior with model sizes. Notably, a 2B MAEWSP model outperforms a larger 6.5B WSP model.
Number of pre-pretraining epochs. We vary the number of MAE pre-pretraining and the number of WSP pretraining epochs to understand their effect on the final recognition performance. We study this in Fig. 4.
Pre-pretraining improves results over the standard pretraining (random initialization, w/o pre-pretraining), and provides large gains with fewer WSP pretraining epochs. Pre-pretraining also leads to faster convergence since even a small amount of pre-pretraining for 0.1 epochs provides improvements. Increasing the epochs of pre-pretraining provide a larger improvement, and the gains saturate at 1 epoch of pre-pretraining. Finally, pre-pretraining’s gains do not diminish even after 4 epochs of WSP (20 billion samples) showing the value of pre-pretraining at scale. We also note that these gains are independent of the evaluation protocol, and we observed them with full finetuning on IN1k as well.
Training efficiency. Fig. 5 shows a comparison between WSP and MAEWSP when comparing training FLOPs. For the same training FLOPs, MAEWSP achieves better transfer performance compared to WSP, and is up to more efficient. Pre-pretraining’s training efficiency holds over a large compute window.
Different datasets for pre-pretraining. We also evaluate the performance of MAEWSP when pre-pretraining MAE on the much smaller ImageNet-1k dataset below, and find that pre-pretraining remains just as effective. This allows reusing pretrained MAE models in practice.
Different datasets for pre-pretraining and pretraining. We investigate the effect of the dataset used in pre-pretraining and pretraining by using IN21k for all methods, including MAE and MAEWSP. Compared to IG-3B, IN21k is more curated, smaller (14M images), and has cleaner labels (21K classes), where each image is labeled with one class from the WordNet synsets . For evaluating zero shot performance, we use the PMD dataset for LiT training. For full details about the hyperparameters, refer to Appendix A.
Fig. 6 compares the performance of MAE, WSP and MAEWSP when pretrained on this IN21k dataset. We notice a similar trend as when pretraining on IG-3B where MAEWSP outperforms both MAE and WSP. This shows that MAE pre-pretraining works with datasets of different scales and distributions.
4 Transfer Evaluation
We compare with state-of-the-art research on image and video classification, detection and segmentation, low-shot image classification, zero shot transfer, robustness analysis. For brevity, we refer to MAEWSP as MAWS in this section.
ImageNet-1k image classification. Table 3 shows the performance of different methods on IN1k. MAWS gets the best performance for a ViT-H sized model (89.3%). Recent methods such as Scale-ViT are better on IN1k and we hypothesize that this gap stems mainly from the differences in the pretraining datasets (IG-3B vs. JFT-3B). We also compute linear performance using frozen features on IN1k at 224px resolution. Our models produce strong features which outperform other methods with WSP objectives. They also surpass the performance of the self-supervised DINOv2 model optimized to produce strong frozen representations:
Robustness for image classification. We evaluate the robustness of our models finetuned on IN1k on additional test sets whose classes overlap with IN1k in Table 3 to evaluate the generalization and robustness of our models to additional test sets. We find that, despite MAWS being 0.4% behind Scale-ViT on IN1k, it is significantly more robust and generalizes better on these additional test sets – MAWS gets the highest reported performance on ImageNetv2, ImageNet-ReaL and ObjectNet for IN1k finetuned models. We also see the benefits of scaling models even up to 6.5B parameters on ImageNetv2 and ObjectNet, where the performance continues to improve significantly with increases in model size.
Generalization in image classification. We evaluate the generalization of our model on additional fine-grained image classification using iNaturalist-18. iNat18 is a challenging long-tailed and fine-grained dataset with images of multiple species of visually similar plants and animals. For the same size, our ViT-H outperforms the previous best result by 3.7%. Our ViT-6.5B sets a new state-of-the-art result on iNat18 (+4.9% over )
Video classification. In Table 4 we investigate how MAWS’s pretraining transfers to video action classification on K400 and SSv2. MAWS is competitive with state-of-the-art methods, including ones that pretrain on videos, whereas our models are only pretrained on images. Specifically, our ViT-L gets the highest reported performance on both video datasets. For all video finetuning, we use relative position embeddings , which improves our performance by 0.6% on K400 for a ViT-L. Overall, the results indicate the promise of MAE pre-pretraining for building strong video understanding models.
Low-shot image classification. We evaluate the label efficiency of our models using a few examples per class for finetuning. We use two datasets, IN1k and iNat18, with shots (labeled examples per class), . For iNat18, as some classes have less than images, we adapt our setting to consider at most shots. For each value of , we generate 5 splits of the original dataset using 5 different random seeds and report the mean top-1 accuracy.
We evaluate two protocols for low-shot finetuning – linear classifiers and Adapters , both of which keep the entire model parameters frozen and introduce a few trainable parameters. We evaluated multiple Adapters proposed for ViTs – LoRA , AdaptFormer , and VPT . We found that VPT performed the best while being robust to the choice of hyperparameters, and outperforms linear classifiers for our models. For other works, we report with the best protocol. Full details in Appendix B.
Table 5 shows a comparison with state-of-the-art methods on low-shot IN1k and iNat18, including foundational and self-supervised models. Our models show impressive low-shot performance on both IN1k and iNat18, reaching 84.6% and 80.9% top-1 accuracy with only 10 labeled examples per class, respectively. They reach the highest reported performance with just one labeled example per class on IN1k of 63.6%.
Zero-shot transfer. Strong foundational models are expected to also have a good open world understanding of visual concepts. To equip our pretrained vision encoders with such capabilities, we utilize LiT . We initialize the text encoder from an XLM-R Large model. Table 6 shows the zero-shot transfer performance of our models on ImageNet-1k, and Food-101. Our ViT-2B attains 82.1% accuracy on IN1k, outperforming a similarly sized OpenCLIP model. Our results lag behind other works which pretrain on datasets such as JFT-3B and ALIGN, highlighting the importance of the pretraining dataset for performance. This data-advantage is also observed in the finetuning results discussed in Table 3, where JFT-3B provides best IN1k accuracy. On Food-101 we attain the highest zero-shot transfer accuracy of 96.2%. The performance on the two datasets is not well correlated, further demonstrating the impact the pretraining dataset can have on a particular zero-shot task.
Detection and segmentation. Next, we evaluate models on detection and instance segmentation, on the LVIS (Table 7) and COCO datasets (Table 8). We use the Cascade Mask R-CNN framework , with the ViTDet architecture, and initialize the backbone with our pretrained models. For finetuning our models, we start with the hyperparameters from and adapt them for our models, full details are in Appendix B. We also perform a system level comparison with other state-of-the-art works, but note that drawing meaningful conclusions out of this is difficult, on account of the multiple differences in the detection frameworks, model architectures, and datasets.
On both benchmarks, MAWS considerably outperforms the weakly supervised SWAG model, demonstrating the benefit of our additional pre-pretraining stage. On the long-tailed LVIS dataset, MAWS outperforms MAE IN1k pretraining on detection AP , but lags slightly behind on COCO. MAE’s strong performance on detection using both IN1k and IG-3B (Fig. 2) can potentially be explained by the fact that it is trained to reproduce images with 75% masking, whereas WSP is only trained to predict one or a few salient objects in an image. Lastly, methods which use additional detection data, such as Objects365 or FLOD-9M , have strong detection performance.
Analyzing detection performance. We further inspect the benefits of pre-pretraining for detection in Fig. 7. Unlike other tasks like image / video classification, scaling model size using WSP pretraining does not improve detection performance. However, adding MAE pre-pretraining provides consistent gains and allows WSP to scale with model size.
We dissect the performance of ViT-L models trained with MAE, WSP, and MAEWSP in Fig. 8 based on the size of the object, or by the frequency of the object’s class. We observe that MAEWSP performs better than WSP on all tasks – detecting rare to frequent classes across small to large object sizes. It improves over MAE at detecting rare objects, presumably because of the diversity in the IG-3B labels.
Performance over model size spectrum. So far we have focused our evaluations on large scale models (ViT-H, ViT-2B, ViT-6.5B). In Table 9 we show the performance of all our models on IN1k, evaluated using linear classifiers as well as with finetuning. We note that our approach works very well for small scale models (ViT-B, ViT-L) as well. MAWS is within 0.3% of the best performance for a ViT-B (86.4% vs. 86.7% for ), and achieves the best performance for ViT-L (88.8%) on IN1k.
Conclusion
We introduced pre-pretraining which is an initial stage in the standard pretrain-then-finetune paradigm. Pre-pretraining uses MAE and thus, does not need additional supervision and can be conveniently added to web-scale training pipelines. We show that pre-pretraining improves downstream performance on multiple different recognition tasks, improves model convergence, and is overall more efficient than standard weakly-supervised pretraining. Our self-supervised pre-pretraining improves results for models trained with billions of labels, showing that it is a scalable technique that matters even at web-scale. The benefits of using pre-pretraining hold across varying model sizes, and different pretraining data distributions showing that it is a robust technique. Finally, pre-pretraining naturally and successfully combines the two most common pretraining strategies – self-supervised and weakly-supervised learning. Our results suggest that model initialization plays a significant role in the final performance, training dynamics etc. even for web-scale training with billions of parameter updates and labels, and should be further investigated.
Acknowledgements
We are grateful to Kaiming He for invaluable discussions and suggestions, and his advice around MAE pre-pretraining. We thank Mary Williamson for her help, guidance and support in project planning, execution, and managing various uncertainties throughout the research project. We thank Devi Parikh for her support around the final parts of the project. We are grateful to Vivek Pai for his help with the training infrastructure. We also thank Yanghao Li for his help with the detection evaluations, Dominik Kallusky for his help with the dataset pipeline, and Tsung-Yu Lin for his help with preparing the image-caption dataset. Lastly, we thank Stephen Roller and Naman Goyal for helpful discussions and feedback.
References
Appendix
Appendix A Pretraining details
MAE pretraining. We make no changes from He et al. for MAE pretraining. We utilize the same decoder dimensions as well – 8 layers, 512 dimension, and 16 heads. Training hyperparameters are shared in Table 10.
WSP and MAEWSP pretraining. We note that large scale WSP pretraining is quite robust to the choice of training hyperparameters, including the choice of (i) a single vs. two layer MLP head, (ii) softmax cross-entropy vs. binary cross-entropy training loss, (iii) additional training augmentations like mixup , CutMix , etc., (iv) learning rate over a range. Most of these choices seem to affect results at small scales, like 0.1 epochs over IG-3B (500 million samples seen), but the differences dissipate when training for a full epoch (5 billion samples seen). Our training hyperparameters are shared in Table 11, and we use these settings for both WSP and MAEWSP. We follow Singh et al. for the training setup and hyperparameters for consistency and easier comparisons. For a label vocabulary , and an image with labels , we utilize softmax cross-entropy, where our label is normalized to a probability distribution, , and the model output is passed to a softmax followed by a cross-entropy loss . We attach a two layer MLP head to the trunk – Linear(embed, embed), Tanh(), Linear(embed, classes).
LiT training. We follow LiT to train a text encoder on Instagram captions. We sanitize the captions and also remove the pound signs in front of hashtags. We use a context length (number of tokens) of 100 per caption. Following CLIP, we chose the embedding dimension for aligning the encoders to be 512, 768 and 1024 for ViT-B, ViT-L and ViT-H, respectively, and set it to 2048 for ViT-2B. We use pretrained XLM-R text encoders, with the Base size (270M parameters) for ablations and Large (550M parameters) for the results in Table 6. Our LiT training hyperparameters are shared in Table 12.
ImageNet-21k pretraining. In IN21k each image is labeled with one class, but the classes are based on WordNet synsets which are hierarchical in nature. Some images in this dataset are duplicated across more than one class – we deduplicate the images by hashing the image contents, and convert the dataset to a multi-label dataset. We disregard the class hierarchy amongst the labels and treat them independently.
For MAE pretraining we again follow and use the hyperparameters in Table 10 and train for 160 epochs over the dataset. For WSP (and MAEWSP) pretraining on ImageNet-21k, we use the hyperparameters from Steiner et al. , with a few minor differences. We train for 90 epochs over the dataset, and select the augmentation setting medium2 of the paper. We train with a softmax cross-entropy loss similar to IG-3B pretraining, but utilize a head with just a single linear layer. Full hyperparameters are in Table 13.
For LiT finetuning of models pretrained on IN21k, we follow a similar setup and hyperparameters used for IG-3B (Table 12), and train for 20 epochs on PMD . Unlike Instagram-3B dataset which has 5 billion image-text pairs, PMD has only about 70 million image-text pairs, so we train for multiple epochs on this dataset.
Appendix B Transfer learning details
Image classification details. We finetune models pretrained with MAE, WSP and MAEWSP by attaching a single linear layer on IN1k and iNat18. We finetune models at either resolution or resolution (or for models which use a patch size of 16) and use the same hyperparameters at all resolutions. The hyperparameters for high resolution finetuning are shared in Table 14 for MAE models and in Table 15 for models pretrained with WSP or MAEWSP.
Video classification details. For video finetuning, we sample 32 frames out of 2.7 second clips for Kinetics-400 and 4 second clips for Something Something-v2, and train the models at resolution. We convert the input into patches of size , akin to MAE approaches applied to video models . We initialize the video models with weights from our inflated pretrained image models. Table 16 contains the finetuning hyperparameters for both K400 and SSv2.
Low-shot image classification details. We adapt all MAEWSP and SWAG models with VPT with the same hyperparameters – 8 tokens of size 192 per self attention layer, as this proved to be a reasonable default setting. The full settings for training our models with VPT are described in Table 17.
We also attempted to train CLIP and OpenCLIP models with Adapters, but however noticed that they do not work as great with VPT as they do with logistic regression, despite sweeping a range of different hyperparameters. We ultimately adopt the logistic regression protocol of MSN for CLIP and OpenCLIP, and also for DINO and MSN. For MAE low-shot evaluations, we noticed that neither Adapters nor logistic regression led to great results, something already seen in MSN . We therefore follow MSN and finetune MAE in low-shot settings, but improve upon the results published in for MAE significantly. All these results are reported in Table 5.
Zero-shot transfer details. For evaluation of our LiT models, we follow the zero-shot evaluation strategy proposed in CLIP . We leverage the prompt templates and class names introduced in CLIP for ImageNet-1k and Food-101. We compute the cosine similarity between the query image and all the generated prompts for each class. While doing so, we also take advantage of the class prompt ensembling approach introduced in CLIP by considering multiple prompt templates for each class.
Detection and segmentation details. We train all of our models with the Cascade Mask-RCNN framework and use the ViTDet architecture to leverage our ViT models within this framework.
For our MAE models trained on IG data, we use the hyperparameters of MAE trained on IN1k from . For the large ViT-2B model, we use settings similar to ViT-H and only change the layer decay to be the same as ViT-L, since both those models have transformer layers.
For WSP and MAEWSP pretraining, we adapt the parameters used for MAE pretraining. The most salient change is a modification of the layer decay scheme of MAE to cap it at a minimum value of , allowing WSP pretrained models to update the initial layers to align better for detection and segmentation tasks. This was particularly useful for instance segmentation.
The hyperparameters used for training our detection and instance segmentation models are available in Table 18 for LVIS and Table 19 for COCO.