EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, Yue Cao
Introduction
Scaling up pre-trained language models (PLMs) has revolutionized natural language processing (NLP) in the past few years. The key to this success lies in the simple and scalable self-supervised learning task of masked signal prediction , with which Transformer models could be scaled up to billions of parameters using nearly unlimited unlabelled data, and generalize well to a wide range of downstream tasks with little tuning. With further scaling on compute, data, and model sizes, PLMs have led to not only continuous performance improvements , but also a surprising emergence of in-context learning capability .
Motivated by the success of model scaling in NLP, it is appealing that we can also translate this success from language to vision, i.e., to scale up a vision-centric foundation model that is beneficial for both vision & multi-modal downstream tasks. Recently, masked image modeling (MIM) has boomed as a viable approach for vision model pre-training and scaling. However, the most competitive billion-sized vision pre-trained models still heavily rely on supervised or weakly-supervised training with hundreds of millions of (often publicly inaccessible) labeled data. MIM is somewhat only adopted as an initialization stage before the heavily supervised pre-training , or a pure MIM pre-trained model could not achieve favorable performance at billion-scale model sizes . We regard this gap stems from the fact that natural images are raw and information-sparse. Meanwhile, an ideal vision pretext task needs the abstraction of not only the low-level geometry & structure information, but also high-level semantics, which is hardly captured by pixel-level recovery tasks .
In this work, we seek a suitable MIM pretext task for large scale vision representation learning and explore its limits at the scale of one billion parameters with tens of millions of unlabeled data. Recently, there are a few trials leveraging the semantic information from image-image or image-text contrastive learning for MIM pre-training , which perform fairly well in vision downstream tasks. However, there remains a debate that (i) tokenized semantic features could provide better supervision signal for masked modeling in vision , and (ii) good performances could be also achieved via a simple post-distillation process without masked prediction tasks . Through a pilot empirical study, we find that simply using image-text aligned (i.e., CLIP ) vision features as the prediction targets in MIM scales up well and achieves satisfactory performances on a broad range of downstream benchmarks. This pre-training task draws the benefits from both the high-level semantic abstraction of image-text contrastive learning as well as the good capture of geometry & structure in masked image modeling, which typically covers the information needed for most visual perception tasks.
Via this MIM pretext task, we can efficiently scale up a vanilla ViT encoder , dubbed EVA, to one billion parameters with strong visual representations that transfers well to a wide range of downstream tasks (Fig. 1). Using 29.6 million public accessible unlabeled images for pre-training, EVA sets new records on several representative vision benchmarks, such as image classification on ImageNet-1K (89.7% top-1 accuracy), object detection and instance segmentation on LVISv1.0 (62.2 AP & 55.0 AP on ) and COCO (64.5 AP & 55.0 AP on , 64.7 AP & 55.5 AP on -), semantic segmentation on COCO-stuff (53.4 mIoU) and ADE20K (62.3 mIoU), and video action recognition on Kinetics-400 (89.7% top-1 accuracy), Kinetics-600 (89.8% top-1 accuracy), Kinetics-700 (82.9% top-1 accuracy). Notably, different from other state-of-the-art billion-scale vision foundation models that demand tens of millions of or even billions of labeled images, such as SwinV2-G using ImageNet-21K-ext-70M and ViT-g/G using JFT-3B , EVA does not need a costly supervised training stage and only leverage images from open-sourced datasets for academic reproducibility.
Moreover, we observe quantitative changes in scaling EVA result in qualitative changes in transfer learning performance that are not observed in other smaller-scale models, e.g., EVA makes a significant breakthrough in the challenging large vocabulary object-level recognition task: our model achieves almost the same performance on LVISv1.0 (55.0 AP on ), an instance segmentation benchmark with more than 1,200 categories, as COCO (55.0 AP on ), which almost shares the same image set as LVISv1.0 but with only 80 categories annotated. This emergent ability well matches the expectation of model scaling , that larger capability of model results in not only predictable performance improvements on standard benchmarks, but also unpredictable phenomenons and capabilities for resolving more challenging tasks.
Going beyond a pure vision encoder, EVA can also serve as a vision-centric, multi-modal pivot that builds a bridge between vision and language. We show that initializing the image encoder via pre-trained EVA in a 1.1 billion parameters CLIP model can outperform the training from scratch counterpart on a broad range of zero-shot image / video classification benchmarks with much fewer samples and less compute. Moreover, EVA can greatly stabilize the giant CLIP’s training & optimization process. Since large CLIP models usually suffer from training instability and inefficiency issues , we hope our solution opens up a new direction for scaling up and accelerating the costly training of multi-modal foundation models.
By scaling up vision-centric foundation models with MIM pre-training to achieve strong performance on broad downstream tasks, we hope EVA would bridge the gap between vision and language with masked signal modeling, and contributes to the big convergence across different modalities.
Fly EVA to the Moon
We first conduct a series of pilot experiments for choosing an ideal vision pretext task in §2.1, then we scale up EVA pre-training via the chosen pre-training objective in §2.2. Finally, we evaluate the pre-trained representation on various downstream tasks in §2.3. Detailed experimental settings and configurations are in Appendix A.
In this section, we seek a MIM vision pretext task with compelling transfer performance. Based on previous literature on vision pre-training, we study two promising candidates: (i) recovering the masked out tokenized semantic vision features , and (ii) feature distillation from strong pre-trained representation as in . Both of them exploit pre-trained image-text aligned vision features (i.e., CLIP vision features). Via a series of pilot experiments shown in Table 2, we find that: (i) the (additional) CLIP feature tokenization process is unnecessary for achieving good downstream performance (ii) feature distillation fails to provide consistent performance gain as the pre-training becomes longer. Instead, we find that simply reconstructing the masked out CLIP vision features conditioned on visible image patches is highly performant, which is chosen for scaling up EVA.
We clarify that this MIM pretext task is not originally proposed by us. Regressing the masked out image-text aligned vision features for MIM pre-training has been studied in MVP and recently has been revisited by MILAN . In this work, we show that this pretext task can scale up to billion-scale parameters and tens of millions of unlabeled images for vision-centric representation learning without (i) semantic feature quantization / tokenization , and (ii) explicitly using image-text paired pre-training data and large corpora as in BEiT-3 .
2 Pre-training
The architecture configurations of EVA are in Table LABEL:tab:_pt_arch. EVA is a vanilla ViT with 1.0B parameters. The shape of her follows ViT giant and the vision encoder of BEiT-3 . We do not use relative positional embeddings and layer-scale during pre-training.
EVA is pre-trained to reconstruct the masked out image-text aligned vision features conditioned on visible image patches. We corrupt the input patches with tokens, and we use block-wise masking with a masking ratio of 40% following . The target for MIM pre-training is from the publicly availableSource: https://github.com/openai/CLIP OpenAI CLIP-L/14 vision tower trained on 224224 pixel images . The output feature of EVA is first normalized and then projected to the same dimension as the CLIP feature via a linear layer. We use negative cosine similarity as the loss function.
The data we used for pre-training EVA are summarized in Table LABEL:tab:_pt_data. For CC12M and CC3M datasets, we only use the image data without captions. For COCO and ADE20K datasets, we only use the train set data. ImageNet-21K and Object365 image data are also used. All these data are publicly accessible. The merged dataset for pre-training has 29.6 million images in total.
The CLIP features we used as MIM prediction targets are trained on a 400 million image-text dataset in a self-supervised manner. So during pre-training EVA also implicitly exploits the knowledge from this dataset to some extent. Meanwhile, these CLIP features is also widely used in other state-of-the-art representation learning & pre-training works such as the BEiT family , AI generated content and large scale dataset filtering .
As shown in Table LABEL:tab:_pt_setting_and_hyperparams, EVA is optimized via Adam with decoupled weight decay of 0.05. The peak learning rate is 1e-3 and decays according to a cosine learning rate schedule. We employed stochastic depth with a rate of 0.1 for regularization and (0.2, 1) for data augmentation. Color jitter is not used.
Some basic pre-training statistics are available in Table LABEL:tab:_pt_stat. The GPU we use is NVIDIA A100-SXM4-40GB. Pre-training code is based on BEiT written in PyTorch . We also adopt optimization library with ZeRO stage-1 optimizer to save memory. We find using format with dynamic loss scaling is stable enough during the whole course of pre-training while using format is unnecessary. Since we use precision, EVA can also be pre-trained using 16 NVIDIA 24GB (32GB) GPUs with (without) gradient checkpointing .
3 Evaluation on Downstream Tasks
In this section, we extensively evaluate pre-trained EVA on several representative benchmarks, such as image classification (§2.3.1), video action recognition (§2.3.2), object detection & instance segmentation (§2.3.3), semantic segmentation (§2.3.4), and contrastive image-text pre-training with zero-shot evaluation (§2.3.5). EVA achieves state-of-the-art performance on a broad range of downstream tasks.
For image classification task, we evaluate EVA on ImageNet-1K (IN-1K) validation set. We also evaluate the robustness & generalization capability of EVA along with our training settings & hyper-parameters using ImageNet-V2 matched frequency (IN-V2) , ImageNet-ReaL (IN-ReaL) , ImageNet-Adversarial (IN-Adv.) , ImageNet-Rendition (IN-Ren.) , ImageNet-Sketch (IN-Ske.) .
Following the conventional setting , we first perform intermediate fine-tuning on ImageNet-21K for 60 epochs with an image resolution of 224, then EVA is further fine-tuned on ImageNet-1K training set for 10 epochs. Different from that use multi-head attention pooling and BEiT-3 that exploits an additional pre-trained giant language tower as the image classification task layer, we simply adopt a linear layer as the classifier . Notice that the supervised intermediate fine-tuning consumes only 1/5 of the time & compute of the MIM pre-training stage. While for other billion-scale vision models such as SwinV2-G-3B, the supervised training phase costs 1.5 resources than the MIM pre-training.
Table 4 compares EVA with some state-of-the-art models on ImageNet-1K validation set. EVA achieves 89.6% top-1 accuracy with 336 inputs, comparable to BEiT-3. Using a larger image resolution of 560 can further boost the top-1 accuracy to 89.7%. Notice that BEiT-3 treats image classification as an image-to-text retrieval task. Therefore they leverage an additional one billion parameters pre-trained language encoder along with 35 million image-text data (21M pairs from CC12M, CC3M, SBU, COCO, VG and 14M pairs from ImageNet-21K) as well as 160GB text data in total. Meanwhile, we simply use a linear classifier on top of EVA with only ImageNet-21K image-tag data used for additional fine-tuning. With only publicly available data, EVA creates a new state-of-the-art image classification result on ImageNet-1K with a much neater architecture.
We evaluate the robustness and generalization capability of EVA trained with an image size of 336 on 6 different ImageNet-1K validation set variants. In Table 5, we compare EVA with some top open-sourced models collected by the librarySource: https://github.com/rwightman/pytorch-image-models/tree/main/results (timestamp: Nov 10, 2022). The detailed model configurations are (arch-model_size-img_resolution-data): ConvNeXt-XL-384px-21K , SwinV2-L-384px-21K , MAE-H-448px-1K , DeiT3-L-384px-21K , EfficientNet-L2&NS-800px-JFT300M , BEiTv2-L-224px-21K , BEiT-L-512px-21K , EVA-g-336px-merged30M&21k. . Following the evaluation procedure in , all these models are first fine-tuned on the original ImageNet-1K training set and then evaluated on different validation sets using the same fine-tuned model without further hyper-parameter selection and specialized fine-tuning.
As shown in Table 5, EVA is the most competitive one in terms of absolute top-1 accuracies. However, these model various in pre-train data (from ImageNet-1K, ImageNet-21K to JFT-300M), input resolutions (from 224 to 800), model sizes (from hundreds of millions to one billion parameters) as well as architectures (ConvNets, vanilla & hierarchical ViTs), etc. Therefore their absolute accuracies are not directly comparable. Instead, we are more interested in the gap between the averaged top-1 accuracy on 6 validation set variants and the original ImageNet-1K validation set top-1 accuracy (the lower the better), i.e., we care about whether a model along with its training settings biases towards the original validation set and generalize well on other variants. From this perspective, EVA not only achieves the highest averaged accuracy, but also has the smallest performance gap, which reflects the excellent robustness and generalization ability of EVA.
3.2 Video Action Recognition
For video action recognition, we evaluate EVA on Kinetics-400 (K-400) , Kinetics-600 (K-600) and Kinetics-700 (K-700) benchmarks. We first conduct intermediate fine-tuning on a merged dataset coined Kinetics-722 (K-722) that integrates videos from K-400, K-600 and K-700. We remove leaked as well as repeated videos in both training and validation sets. After this data de-duplicating process, K-722 has 0.63M training videos in total with 722 action classes.
EVA processes video data simply via spatial-temporal attention as with no specific architectural adaptation for video related tasks. We first train EVA using K-722 training set for 40 epochs with 8 frames and 224 resolution, then we fine-tune EVA on each dataset for only 1 or 2 epochs. We set framecropclip to 1634 for fine-tuning and evaluation for all datasets. The frame resolution is 224.
As shown in Table 6, EVA achieves better performance compared with some recent video-specific or large foundation models in video recognition. For reference, directly adapting image-only pre-trained EVA to K-400 without K-722 intermediate fine-tuning can also achieve a very competitive top-1 accuracy of 88.4%.
3.3 Object Detection & Instance Segmentation
We evaluate the object detection and instance segmentation performance of EVA on both COCO and LVISv1.0 benchmark. COCO is a widely used object-level recognition benchmark consisting of 118k , 5k , and 20k - images respectively, with 80 common object categories. LVISv1.0 is an emerging large-vocabulary object-level recognition benchmark, which has more than 1,200 object categories as well as more than 2 million high quality instance segmentation masks (nearly 2 of COCO instance masks).
Notably, COCO and LVISv1.0 almost use the same set of images, and both and split of LVISv1.0 have a huge overlap with COCO and split. Meanwhile, COCO has much fewer object categories than LVISv1.0 (i.e., 80 v.s.1,200+). Therefore it is interesting and meaningful to evaluate a model’s performance on both COCO and LVIS.
For COCO, we report the standard box AP (AP) and mask AP (AP) on both and - split. For LVISv1.0, we evaluate EVA using AP, AP and AP defined in on the v1.0 set. We also report the experimental “boundary, fixed AP”(AP in Table LABEL:tab:_lvis) used in the LVIS 2021 challengeSource: https://www.lvisdataset.org/challenge_2021 for reference.
EVA uses Cascade Mask R-CNN as the detector and adopts the training settings (e.g., large scale jitter data augmentation ) & architecture configurations (e.g., interleaved window & global attention) of ViTDet . Following the common practice , we first conduct intermediate fine-tuning for the whole detector using Objects365 dataset with a resolution of 1024, then we fine-tune the detector on COCO and LVISv1.0 split respectively with 1280 inputs.
We report single-scale evaluation and multi-scale evaluation / test-time augmentation (tta) results of EVA for comparison. For COCO, Soft-NMS is also applied. For instance segmentation task, the classification score is calibrated via maskness .
The model architecture as well as the hyper-parameters for COCO and LVISv1.0 are almost the same (i.e., the hyper-parameters are nearly “zero-shot” transferred from COCO to LVISv1.0), expect we use federated loss and repeat factor sampling following ViTDet on LVISv1.0.
Perhaps COCO is the most fierce vision benchmark. Table 7 compares EVA with some previous leading approaches on COCO. Our model creates new state-of-the-art results on both object detection and instance segmentation tasks.
Compared with ViTDet-H that uses Cascade Mask R-CNN as well, EVA shows that with a larger model and better encoder & detector pre-training, the performance can be greatly improved with the same detector.
Compared with FocalNet and Group DETRv2 that choose better-established and highly-optimized DINO detector , EVA demonstrates that with sufficient model size, data and pre-training, better performance can be also achieved via the classic R-CNN framework . On the other hand, FocalNet and Group DETRv2 are incapable of instance segmentation due to using DINO.
Compared with SwinV2-Giant and FD-SwinV2-Giant that also adopt a (stronger HTC++ ) detector from the R-CNN family but with 3 model size of EVA, our approach streamlines the pre-training processes and pulls off a “Giant-killing” act via better representations.
Compared with BEiT-3, EVA shows that is possible to build a state-of-the-art object-level recognition system without exploiting (i) semantic feature quantization / tokenization , and (ii) image-text paired pre-training data and large corpora during pre-training.
Table LABEL:tab:_lvis summarizes the results on LVISv1.0 set. EVA achieves state-of-the-art performance under all metrics with single-scale evaluation, outperforming the previous best approaches by a large margin.
It is interesting and valuable to evaluate and analyze a model on both COCO and LVISv1.0 benchmarks, as they share almost the same image set but with different numbers of annotated object categories. Compared with COCO which has only 80 annotated categories, LVIS annotates more than 1,200 object categories and thus naturally features a long-tail distribution, which is more close to the challenging real world scenario . In general, LVIS is considered to be a much harder benchmark than COCO in object-level recognition, and conventional methods usually suffer from a large performance drop on LVIS compared with COCO.
In Table LABEL:tab:_lvis_coco_gap, we analyze the performance gap between LVISv1.0 and COCO benchmark of EVA and other state-of-the-art approaches. For previous leading methods such as ViTDet, the performance gap between AP is around 8, and the gap between AP is around 5. However, using the same detector (Cascade Mask R-CNN) and almost the same settings as ViTDet pre-trained via MAE-Huge (ViTDet-H), EVA not only achieves the state-of-the-art results on both LVIS & COCO benchmarks simultaneously, but also largely closes the performance gap between them, especially for the instance segmentation task, that EVA achieves the same performance on LVIS and COCO with single-scale evaluation. Compared with ViTDet-H, we show a little larger model and stronger representations can take a great leap in the challenging large vocabulary instance segmentation benchmark–with one caveat described next.
Notice that it is inaccurate to say EVA “solves” the LVIS large vocabulary instance segmentation task based on “zero AP gap”, and it is possible to achieve even higher AP on LVIS than COCO (e.g., LVIS has higher quality mask annotations for both training and evaluation). Although the AP of EVA on LVIS is much better than previous approaches, there is still a big gap between the rare categories and common / frequent categories (i.e., AP = 48.3 v.s. AP = 55.5 / AP = 57.4). From another perspective, the gap of AP between LVIS and COCO (1.9 for EVA) may better reflect the overall gap between the large vocabulary and common vocabulary object-level recognition, since this metric eliminates the confounder of instance mask annotation quality.
Nevertheless, EVA makes a significant breakthrough in the challenging large vocabulary object-level recognition task. We believe other large vision foundation models such as SwinV2-G and BEiT-3 also have similar properties on COCO and LVISv1.0 benchmark. We hope our efforts can encourage the community to pay more attention to the relationships between different tasks while scaling up models and chasing the state-of-the-art in each individual task.
3.4 Semantic Segmentation
We evaluate EVA on ADE20K and COCO-Stuff-164K datasets for semantic segmentation task. ADE20K includes 150 semantic categories, and has 20k images for training & 2k images for validation. COCO-Stuff-164K augments 164K complex images from COCO with pixel-level annotations that span over 172 categories including 80 things, 91 stuff, and 1 unlabeled class. Compared with ADE20K, COCO-Stuff is a more challenging but under-explored semantic segmentation benchmark.
We follow the task transfer pipelines of ViT-Adapter +mask2former but with a weakened model adaptation processes due to GPU memory limitation (40GB of VRAM): (i) relative position biases are not applied. (ii) We use 8 decoders in mask2former segmentation head instead of 9. (iii) The feature dimension in mask2former head is 0.6 of EVA encoder.
We compare EVA with other leading semantic segmentation methods in Table 9. EVA achieves strong results in both ADE20K and COCO-Stuff-164K datasets. On the other hand, the segmentation performance of EVA is slightly lower compared with BEiT-3 on ADE20K, we suspect this is partially due to our weakened architectural configurations.
3.5 Contrastive Language-Image Pre-training with Zero-shot Classification Evaluation
CLIP (Contrastive Language-Image Pre-training) is a type of multi-modal foundation model that connects vision and language via contrastive image-text pre-training. CLIP can be applied to any image classification benchmark by simply providing the names of the visual categories to be recognized . Thus the introduction of CLIP essentially reshapes the landscape of visual recognition. Meanwhile, CLIP features also play a central role in representation leaning , AI generated content and large dataset filtering , etc.
In this section and Table 10, we show that EVA is not only a strong encoder for a wide range of vision downstream tasks, but also a multi-modal pivot that builds a bridge between vision and language. To demonstrate that, we train & evaluate EVA as a billion-scale CLIP’s vision tower in various zero-shot image / video classification benchmarks.
We compare our CLIP (dubbed EVA CLIP) with other open-sourced strong CLIP competitors that exploit publicly accessible data / academic resources only. Model configurations and statistics are detailed in Table LABEL:tab:_clip_cfg.
There are two well-known major challenges of CLIP model training and scaling: (i) Large-scale Open CLIP models (e.g., Open CLIP-H & Open CLIP-g ) usually suffer from severe training instability issues and have to use format for optimization. (ii) The training efficiency is low, which may hinder model scaling and downstream performance. For instance, Open CLIP-g is heavily under-trained due to its large compute requirement, and its performance is even worse than the sufficiently-trained Open CLIP-H with a smaller model size.
Compared with our CLIP model, Open CLIP-H & -g are trained from scratch with much more image-text pairs (2.9 and 1.1 of ours) sampled from a much larger dataset (5 of ours) on 3 of GPUs. While by leveraging EVA, billion-scale CLIP model training can be accelerated with improved zero-shot classification performance, described next.
For our CLIP model, we initialize the vision encoder via pre-trained EVA and the language encoder from OpenAI CLIP-L. The pre-training implementation is based on Open CLIP . We also adopt optimization library with ZeRO stage-1 optimizer to save memory. We find using format with dynamic loss scaling is stable enough during the whole course of training while using format is unnecessary. These modifications allow us to train a 1.1B CLIP with a batch size of 41k on 256 NVIDIA A100 40GB GPUs.
We evaluate zero-shot image / video classification performance of each CLIP model on 12 benchmarks and report top-1 accuracy for comparisons.
For zero-shot image classification task, we choose 8 benchmarks, i.e., ImageNet-1K , ImageNet-V2 , ImageNet-Adversarial (ImageNet-Adv.) , ImageNet-Rendition (ImageNet-Adv.) , ImageNet-Sketch (ImageNet-Ske.) , ObjectNet , CIFAR-10 and CIFAR-100 . We are also interested in the robustness of CLIP models, evaluated via the performance gap between the averaged performance of ImageNet-{1K, V2, Adv., Ren., Ske.} & ObjectNet that with natural distribution shifts and the original ImageNet-1K validation accuracy.
For zero-shot video classification task, we choose 4 benchmarks, namely UCF-101 , Kinetics-400 , Kinetics-600 , and Kinetics-700 .
Table LABEL:tab:_clip_result shows the comparison. Our EVA CLIP achieves the highest averaged accuracy, and performs the best in 10 out of 12 zero-shot classification benchmarks. Notably, the ImageNet-1K validation zero-shot top-1 accuracy is 78.2% without using any of its training set labels, matching the original ResNet-101 . Moreover, our model is quite robust and suffers from the smallest performance drop when facing natural distribution shifts in ImageNet.
At last, in Table 11 we provide zero-shot, linear probing & end-to-end fine-tuning top-1 accuracy of EVA-CLIP on ImageNet-1K validation set for reference. Our approach creates the new state-of-the-art results among all existing self-supervised learning methods.
Notice that EVA CLIP’s vision branch learns from OpenAI CLIP-L, while language branch initialized from the same CLIP-L model. Therefore, starting from a CLIP-L with only 430M parameters, we progressively scale up a 1.1B EVA CLIP-g with large performance improvements. This implies that interleaved MIM & image-text contrastive pre-training could be an efficient and scalable CLIP training approach. To our knowledge, EVA CLIP-g is the largest performant CLIP model trained via publicly accessible data and resources. We hope our practice on scaling and improving CLIP can also inspire and transfer to the study of other large scale multi-modal foundation models.
Related Work
learns rich visual representations via predicting masked visual contents conditioned on visible context. ViT and iGPT report the first meaningful MIM pre-training results. The BEiT family greatly improves MIM’s performance via masked visual token prediction. Recent work (re-)explore pixel / feature regression in MIM, but only in a relatively small model and data scales. In this work, we explore the limits of large scale MIM pre-training via masked image-text aligned feature prediction .
ConvNets have long been the de-facto standard visual architecture ab initio. Since AlexNet , ConvNets have rapidly evolved and become deeper, wider and larger . However, at sufficient model and data scales, ConvNets lag behind ViTs due to a lack of scalable pre-training tasks and the built-in inductive biases. Entering the 2020s, large pre-trained ViTs such as SwinV2-G with hierarchical architectures as well as BEiT-3 with multi-modal representations started to demonstrate various vision benchmarks. In this work, we show by leveraging unlabeled images, vanilla ViT can be efficiently scaled up to billion-scale parameters, and stands out in various downstream tasks.
Conclusion
In this work, we launch EVA, a one billion parameters vanilla ViT encoder to explore the limits of masked visual representation learning. We show simple masked feature modeling as a visual learning pretext task scales well on an architecture with minimal vision priors, and attains excellent results in a representative & diverse set of downstream tasks. We hope EVA would bridge the gap between vision and language study via masked modeling, and contributes to the Neon Genesis of vision research.
Acknowledgement
We would like to thank Hanxiao Qu, Yan Tian, Yemin Shi and Xigang Cao for their help on GPU resources. Zhao Xue, Quanyue Ma and Bowen Zhang for their help on datasets and benchmarks, and other colleagues at Beijing Academy of Artificial Intelligence for support throughout this project.
Appendix A Appendix
The MIM pre-training and contrastive language-image pre-training settings are already available in our main submission. Here we summarize the detailed configurations for image classification (§A.1), video action classification (§A.2), object detection & instance segmentation (§A.3), and semantic segmentation (§A.4).
The fine-tuning hyper-parameters for ImageNet-21K and ImageNet-1K are shown in Table 12 and Table 13, respectively.
A.2 Video Action Classification
For video action classification tasks, a two-stage fine-tuning process is adopted. The statistics of video datasets we used are available in Table 14.
In the first stage, we conduct intermediate fine-tuning on a merged dataset coined Kinetics-722 (K-722) that integrates all valid training samples from Kinetics-400 (K-400) , Kinetics-600 (K-600) and Kinetics-700 (K-700) . The input video resolution is 224 with 8 frames. Notably, for a fair and legal comparison, we removed leaked videos in all validation sets and duplicated videos in all training sets based on the videos’ “youtube id”. Accordingly, the cleaned K-722 contains 0.63M training videos, covering 722 human action classes. Table 15 lists the detailed settings & hyper-parameters for fine-tuning on this dataset.
In the second stage, we further fine-tune on each dataset using more input video frames of 16 with a resolution of 224. For the frame sampling, we adopt the sparse sampling strategy . During testing, we follow the common practice of multi-view inference with 4 temporal clips and 3 spatial crops. The final prediction is the ensemble of all trials. Table 16 lists the detailed hyper-parameters for fine-tuning on K-400, K-600 and K-700.
A.3 Object Detection & Instance Segmentation
The detailed hyper-parameters are shown in Table 17 and Table 18. For intermediate fine-tuning on Objects365 , the model is trained with a batch size of 128 for 380k iterations. To accelerate the training process, we use a smaller input resolution of 1024 for the first 320k iteration. Afterward, the input resolution is lifted to 1280 for a better adaptation to the fine-tuning of COCO and LVIS.
For fine-tuning COCO and LVIS, the learning rate is initialized as 2.5e-5 and step by a factor of 10 for the last 5k iterations. As shown in Table 18, we use almost identical hyper-parameters for training COCO and LVIS. Except for the commonly used repeat factor sampling and federated loss that are specialized for long-tailed recognition, the only difference in training is that we train the model for 45k steps on COCO, while a longer 75k step on LVIS, since the tail classes generally take a longer schedule to converge .
A.4 Semantic Segmentation
Detailed configurations about semantic segmentation are available in Table 19. Our settings basically follow ViT-Adapter with Mask2Former as the segmentation head. For ADE20K, we use COCO-Stuff pre-trained weights as initialization.