EVA-02: A Visual Representation for Neon Genesis
Yuxin Fang, Quan Sun, Xinggang Wang, Tiejun Huang, Xinlong Wang, Yue Cao
Introduction
Recent research advancements have led to a surge of interest in scaling up vision as well as vision-language representations. These efforts are driven by the belief that increasing the number of parameters, data, and compute budgets will ultimately result in improved performance .
However, there is an increasing gap in computer vision between large-scale models that achieve state-of-the-art performance and models that are affordable for the wider research community. Training, tuning, and evaluating very large vision models requires significant computational resources, which can be prohibitively expensive and time-consuming. This usually leads to large-scale visual representations being trained in a few-shot or even single-shot manner, limiting the ability to fully optimize the entire process. In addition, the study of state-of-the-art representations is frequently conducted using huge amounts of infrastructure and web-scale private training data , which makes it difficult to evaluate the effects of modeling advancements in a feasible and transparent way, and restricts access to a broad range of researchers and practitioners. These challenges highlight a pressing need for a more efficient and accessible approach of training and evaluating state-of-the-art vision as well as vision-language representations.
In this work, we present EVA-02, a series of robustly optimized plain Vision Transformers (ViTs) with moderate model sizes that are equipped with transferable bidirectional visual representations learned from a strong CLIP vision encoder via masked image modeling (MIM) pre-training . Compared with current leading vision models with billions of parameters , these EVA-02 variants require far fewer compute budgets and resources to investigate, allowing for an in-depth exploration of often-overlooked aspects.
Our empirical investigation indicates that the smaller-sized plain ViTs are highly capable, and their potential has been significantly underestimated. By leveraging the latest plain Transformer architecture design borrowed from language models, as well as thorough MIM pre-training from a publicly available giant EVA-CLIP vision encoder, EVA-02 is able to achieve superior performance compared to prior state-of-the-art approaches with much larger model sizes on various visual tasks.
Remarkably, using exclusively 38 million publicly accessible data, the small-sized variant of EVA-02 with only 22M parameters achieves 85.8 fine-tuning top-1 accuracy on ImageNet-1K (IN-1K) val set , while the large model with only 304M parameters achieves an outstanding 90.0 fine-tuning top-1 accuracy. Moreover, we show that initializing the image encoder of a CLIP via MIM pre-trained EVA-02 representations can reach up to 80.4 zero-shot top-1 on IN-1K val, outperforming the previous largest & best open-sourced CLIP-Giant with only 1/6 parameters and 1/6 image-text training data. EVA-02 also achieves state-of-the-art performances on other representative vision tasks such as object detection and instance segmentation on LVIS (65.2 AP & 57.3 AP on ) and COCO (64.5 AP & 55.8 AP on -), as well as semantic segmentation on COCO-stuff-164K (53.7 mIoU) and ADE20K (61.7 mIoU and 62.0 mIoU). For a quantitative summary of EVA-02’s performance, please refer to Table 1.
The proposed EVA-02 series offers a diverse range of model sizes, ranging from 6M to 304M parameters, each demonstrating exceptional performance. The aim of this work is not necessarily to propose a novel method, but strive to identify a robust and effective recipe for making state-of-the-art models more affordable in practice. By providing a more accessible and performant option, EVA-02 democratizes access to state-of-the-art vision models, allowing researchers as well as practitioners to conduct high-quality research without the need for extensive infrastructure or resources. We hope our efforts enable a broader range of the research community to advance the field in a more efficient and equitable manner.
Approach
The aim of EVA-02 is to introduce a next-generation Transformer-based visual representation that achieves strong performances with moderate model sizes. To achieve this goal, our representation instrumentality project consists of two parts: architectural improvements made to the plain ViT in §2.1, as well as our MIM pre-training strategy in §2.2.
At a high level, plain ViT along with its variants comes with interleaved multi-head self-attention (MHSA) layers for global spatial information aggregation & position-wise feedforward networks (FFNs) for feature transformation, without downsampling layers and multi-stage design . This makes it an ideal testbed for representation learning due to its minimal visual structure prior and biases, as well as its natural compatibility with masked modeling, which is proven to be a simple, strong, and scalable pre-training approach . Pre-trained plain ViT can also be successfully adapted to challenging vision tasks that require high-resolution inputs & multi-scale representations with feasible costs .
Although the inner-block micro architecture of plain ViT has continuously evolved since its inception in the year 2020 , we notice that some significant architectural advances in language models have not yet been explored in the context of visual representation learning. These include gated linear unit with sigmoid linear unit (SiLU) / swich activation (SwiGLU) as the feedforward network, sub-LN as the normalization layer, and 2D rotary position embedding (RoPE) for positional information injection.
In Table 2 we conduct a series of pilot experiments studying these architectural modificationsMore technical details can be found in the Appendix. All these modifications do not bring additional parameters as well as FLOPs.. The pretext task is to regress the masked-out EVA-CLIP vision features conditioned on visible image patches using IN-1K training images for 300 epochs, and the evaluation is done by fine-tuning the pre-trained base-sized models on IN-1K. Starting with the baseline ViT configurations used in the original BEiT series pre-training ( in Table 2), we progressively refine the model design and make the following observations: (i) The performance of SwiGLU FFN is mediocre with the random weight initialization method used in BEiT, but works quite well with weight initialization (+1.1). (ii) sub-LN slightly improves the performance compared with pre-LN (+0.2). (iii) 2D RoPE can improve the performance (+0.4), while the standard relative position embedding suffers from unstable pre-training with other configurations unchanged.
The final model configuration ( in Table 2), called Transform Vision (TrV, Fig. 2b), aligns with the model architecture in current leading language models , and achieves a favorable accuracy with an overall improvement of 1.6 points compared to the original configurations (i.e., from 84.0 to 85.6), with one caveat that will be described next.
2 Pre-training Strategy
In the previous section, we choose to use features from a giant CLIP vision encoder with one billion parameters as the target representation for our MIM pretext task. However, we have not yet explained the rationale behind this choice. Although similar pre-training strategies have been well-studied in recent literature and shown to be effective, they typically use vision features from much smaller CLIP models. Choosing the 1B-parameter EVA-CLIP is based on our assumption that larger CLIP will provide more robust and transferable target representations for MIM, and will ultimately lead to better pre-trained models. In Table 3, we study the impact of target representations produced by different-sized CLIPs.
At first glance, compared with the smaller VQKD-B and CLIP-B as MIM teachers, the accuracy degenerate (i.e., from 85.0 to 84.0) with EVA-CLIP target when the students are base-sized plain ViT in with 300 epochs pre-training on IN-1K ( in Table 2 and Table 3). The architectural modifications from TrV compensate for this to some extent, resulting in a modest total improvement of 0.6-point ( in Table 2 and Table 3).
We conjecture that as the teacher becomes stronger, it becomes harder for the students to learn robust and transferable representations in a crash course. Consequently, more extensive pre-training is required for the students to fully master the teacher’s knowledge. As we extend the pre-training schedule to 1600 epochs (1M steps), TrV with EVA-CLIP as the MIM teacher yields a 1.3-point non-trivial improvement over BEiTv2 . Furthermore, with 150 epochs (1M steps) pure MIM pre-training on ImageNet-21K (IN-21K, 14.2M images) , our base-sized TrV achieves 87.0 top-1 accuracy, even outperform BEiTv2 with 1600 epochs (1M steps) MIM pre-training on IN-1K plus an additional 90 epochs intermediate fine-tuning on IN-21K with labels.
Ulteriorly, in Table 4 we show that scaling model size, resolution as well as injecting labels via intermediate fine-tuning can further boost the performance, reaching up to 90.0 top-1 accuracy on IN-1K with only a 304M-parameter EVA-02. Notably, our pure MIM pre-trained representations can achieve very competitive performance without additional intermediate fine-tuning.
From now on, we denote TrV with sufficient MIM pre-training from EVA-CLIP as EVA-02. In the rest of this section, we present some technical details of MIM pre-training before go into the performance evaluation in §3.
We provide four variants, i.e., EVA-02-Ti (6M), -S (22M), -B (86M) and -L (304M), as detailed in Table 5. The marco architecture (e.g., model depth, width, #head) of EVA-02 variants follows the canonical plain ViT configurations in . The inner-block modifications are detailed in §2.1.
is similar to EVA , which is to regress the masked-out image-text aligned vision features conditioned on visible image patches only. We corrupt the input patches with tokens, and we use block-wise masking with a masking ratio of 40% following . The target representation for MIM pre-training is from the publicly accessible EVA-CLIP vision tower with one billion parameters. The output feature of EVA-02 is first normalized and then projected to the same dimension as the EVA-CLIP’s vision feature via a linear layer. We use negative cosine similarity as the loss function.
For EVA-02-Ti, -S and -B, we use images from IN-21K for pre-training. For EVA-02-L, we use a merged dataset consisting of IN-21K, CC12M , CC3M , COCO , ADE20K , Object365 and OpenImages . For CC12M and CC3M, we only use the image data without captions. For COCO and ADE20K, we only use the training set images. The merged dataset for pre-training EVA-02-L has 38 million images in total (denoted as Merged-38M). All these datasets are publicly accessible.
generally follow the BEiT series . The optimizer is Adam with decoupled weight decay / of 0.05 / 0.98 . The peak learning rate / batch size is 3e-3 / 4k for tiny- and small-sized models, and 1.5e-3 / 2k for base- and large-sized models. We train tiny- and small-sized models for 0.8M steps, and base- and large-sized models for 1M steps.
The pre-training code is based on the open-sourced EVA implementation . We adopt with ZeRO stage-0 / -1 optimizer and precision with dynamic loss scaling . All MHSA operations are accelerated by . Although our MIM teacher comes with one billion parameters, the wall-clock pre-training time is 10% shorter than the official BEiT series implementations .
Experiments and Evaluation
In this section, we present a comprehensive evaluation of our approach on representative vision tasks and benchmarks, including image classification in §3.1, contrastive image-text pre-training (CLIP) with zero-shot evaluation in §3.2, object detection & instance segmentation in §3.3.1, and semantic segmentation in §3.3.2. We conduct experiments mainly using base-sized (86M) and large-sized (304M) pre-trained representations. Our results demonstrate that EVA-02 is capable of outperforming larger counterparts and achieving state-of-the-art performance without or with only minimal additional intermediate fine-tuning. Additional details and results can be found in the Appendix.
For image classification, we mainly evaluate EVA-02 on IN-1K . We also evaluate the robustness & generalization capability of EVA-02 along with our training settings & hyper-parameters using some IN-1K validation set variants, including ImageNet-V2 matched frequency (IN-V2) , ImageNet-ReaL (IN-ReaL) , ImageNet-Adversarial (IN-Adv.) , ImageNet-Rendition (IN-Ren.) , ImageNet-Sketch (IN-Ske.) , as well as ObjectNet (ObjNet) , following the settings in .
To fully unleash the potential of EVA-02, we optionally perform intermediate fine-tuning following for base- / large-sized model on IN-21K for 40 / 30 epochs in Table 7. The final IN-1K fine-tuning for all-sized models (including EVA-02-Ti and -S) can be done without using strong regularization such as cutmix , mixup and random erasing . In the Appendix, we show that our pre-trained representations are robust enough that can be fine-tuned using various numerical precisions (e.g., and ) and optimizers (e.g., Lion , AdamW , and SGD ). Remarkably, the fine-tuning can be done even using the SGD optimizer with only 0.1-point performance drop.
Table 7 compares EVA-02 with some state-of-the-art models on IN-1K val set. Our base-sized model, trained with ImageNet data only, outperforms several strong competitors and achieves the same performance with a ViT-B distilled from a 4B-parameter teacher using large-scale in-house training data . Furthermore, EVA-02-L with only 304M-parameter can achieve a phenomenal 90.0 fine-tuning top-1 accuracy, outperforms several state-of-the-art larger models trained with more (often publicly inaccessible) data, including its fine-tuned EVA-CLIP MIM teacher, which distinguishes MIM from knowledge distillation .
It is commonly believed that plain ViTs perform mediocrely due to the lack of inductive biases in light-weight settings. However, compared with specialized light-weight networks with strong visual structure prior in Table 8, EVA-02 as a plain ViT variant equipped with extensive MIM pre-training can trump inductive biases, and achieve favorable performance with tiny and small models.
We evaluate the robustness and generalization capability of EVA-02 on several IN-1K val set variants. Following the evaluation procedure in , all these models are first fine-tuned on the original IN-1K training set, and then directly evaluated on different val sets using the same fine-tuned model without further hyper-parameter selection and specialized fine-tuning.
In Table 6, we compare EVA-02 with some top open-sourced models. EVA-02 is the most competitive one in terms of top-1 accuracies. Besides the absolute performance, we also care about whether a model along with its training settings biases towards the original validation set and generalizes well on others. From this perspective, EVA-02 not only achieves the highest averaged accuracy, but also has the smallest performance gap (as measured by the difference between the averaged accuracy of val set variants and the original IN-1K val set accuracy), which reflects the excellent robustness and generalization ability of EVA-02.
2 Contrastive Language-Image Pre-training and Zero-shot Evaluation
Contrastive Language-Image Pre-trained (CLIP) model is a kind of foundation model that aligns vision and natural language through contrastive image-text pre-training . Its impact on the field of representation learning has been significant, making it a powerful engine for both recognition and generation tasks, as well as uni-modal and multi-modal applications .
In this section, we thoroughly demonstrate the efficacy of initializing EVA-02 as the CLIP vision encoder following the settings in . The resulting model, referred to as EVA-02-CLIP, significantly improves zero-shot performance, sample efficiency, and training speed.
We present CLIP model configurations and IN-1K zero-shot accuracies in Table 10. To train EVA-02-CLIP, we merge the data from the publicly accessible LAION-2B and COYO-700M , which results in a dataset of 2 billion image-text pairs (we only have 1.6B / 400M valid samples from LAION-2B / COYO-700M datasets). Leveraging MIM pre-trained EVA-02 representations, our CLIP model significantly outperforms previous approaches in IN-1K zero-shot classification, achieving an outstanding 74.7 / 80.4 top-1 accuracy with base- / large-sized models.
In Table 9, we further demonstrate the efficacy and robustness of our approach on 26 additional zero-shot classification benchmarks. Notably, our EVA-02-CLIP-L model, which only has 1/2 of the model size and 1/5 image-text pairs, achieves a 1.2-point non-trival averaged improvement over OpenCLIP-H.
Finally, in Table 11 we show that EVA-02-CLIP is also quite effective in zero-shot video recognition benchmarks.
Table 12 comprehensively reports the zero-shot image and text retrieval results on Flickr30K and COCO . EVA-02-CLIP outperforms all the competitors with the same model size. While the zero-shot retrieval performance of EVA-02-CLIP is not as significant as classification compared to OpenCLIP-H, the results are still competitive. We speculate that the main reason for this difference is that retrieval tasks depend more on the capacity and capability of the language encoder compared to classification tasks.
3 Object Detection and Segmentation
In this section, we evaluate the transfer learning performance of EVA-02 to mainstream object-level and pixel-level recognition benchmarks, namely, object detection and instance segmentation on COCO and LVIS in §3.3.1, as well as semantic segmentation on COCO-Stuff-164K and ADE20K in §3.3.2.
For thoroughly evaluating the performance of EVA-02 on object detection and instance segmentation tasks, we adopt the canonical Cascade Mask R-CNN as the task layer. This choice is motivated by its versatility in simultaneously performing both tasks, as well as its robustness and accuracy. To ensure fair comparisons with existing state-of-the-art methods, we essentially follow the training settings and architecture configurations of ViTDet , which includes large-scale jittering (LSJ) data augmentation and interleaved windowed and global attention mechanisms.
The model architecture as well as the hyper-parameters for COCO and LVIS are almost the same, except we use federated loss and repeat factor sampling following ViTDet on LVIS. For LVIS, we use the IN-21K MIM pre-trained checkpoints of EVA-02 for all experiments, as the COCO training images in the Merged-38M dataset include 10k images in LVIS setIn the Appendix, we show including unlabeled images from development / test set for MIM pre-training does not improve the final performance..
In the rest of this section, we evaluate EVA-02 under three different transfer learning settings in Table 13 and Table 14, including (i) a sanity check, (ii) a system-level comparison without using additional detection data, and (iii) a system-level comparison with additional intermediate detection fine-tuning.
We first use the same open-sourced architectural configurations as ViTDet (LSJ with 1024 crops, 4global attention blocks) to perform a head-to-head comparison. In general, both EVA-02-B in Table 13a and EVA-02-L in Table 14a can outperform the same- / larger-sized ViTDet w/ Cascade Mask R-CNN counterparts by a large margin, especially on LVIS.
In Table 13b and Table 14b, we explore the limits of pure MIM pre-trained EVA-02-B and -L representations in object detection and instance segmentation tasks. To fully unleash the potential of EVA-02, we use an improved ViTDet configuration (LSJ with 1536 crops, windowed attention with a size of 32, and 6 / 8global attention blocks for base- / large-sized models). Soft-NMS is also applied. For instance segmentation task, the classification score is calibrated via maskness . The baselines we compared also adopt improved settings such as larger input resolution, Soft-NMS, etc., and RevCol initializes HTC++ as the task layer, which is an improved version of Cascade Mask R-CNN we used.
Our experiments demonstrate that EVA-02 significantly outperforms the same- and larger-sized counterparts, particularly on LVIS. These findings are consistent with our previous results in Table 13a and Table 14a. We also encourage future work in representation learning to conduct more in-depth investigations on the original pre-trained representations before adding more intermediate processes to chase the absolute performance.
3.2 Semantic Segmentation
We comprehensively evaluate the semantic segmentation performance of EVA-02-B and -L models using two different task layers: UperNet and Mask2Former on two widely adopted benchmarks: ADE20K and COCO-Stuff-164K . Notably, unlike previous mainstream approaches that involve additional fine-tuning such as using IN-21K intermediate fine-tuned models for semantic segmentation, we primarily evaluate pure MIM pre-trained representations of EVA-02.
As shown in Table 15, both pure MIM pre-trained EVA-02-B and -L models with UperNet segmenter significantly outperform the same-sized BEiTv2 models without or with the additional 90-epoch IN-21K intermediate fine-tuning. Furthermore, our representation can outperform larger pre-trained counterparts such as ConvNeXt V2, InternImage, etc., and achieves up to 60.1 mIoU with single-scale evaluation.
Table 16 shows the state-of-the-art model comparisons on COCO-Stuff-164K and ADE20K benchmarks. Models for ADE20K segmentation are initialized from COCO-Stuff-164K pre-trained representations as long as the COCO-Stuff-164K results are reported. BEiTv2-L and EVA also utilize ViT-Adapter (ViT-Ada. in Table 16) for architectural improvements.
Compared with larger models using the Mask2Former task layer, our approach is still quite performant, and creates new state-of-the-art results with large-sized models on both COCO-Stuff-164K and ADE20K semantic segmentation benchmarks.
4 Summary of All Evaluations
In §3, we demonstrate the excellent transfer learning ability of pre-trained EVA-02 representations on a large diversity of downstream tasks. Although all tasks / benchmarks we evaluated are at the core of computer vision, here we would like to (re-)emphasize the importance of the ones related to EVA-02-CLIP: not only for the promising zero-shot transferability, but also because the vision features from EVA-02-CLIP are well aligned with natural language that comes with much broader supervision than pure vision signals / features as well as fixed set of pre-determined label sets. Therefore, we hope EVA-02-CLIP can serve as a basic building blocks and provide more robust vision features for future multi-modal systems.
Related Work
Some previous advancements in representation learning do not necessarily come with entirely new ideas or novel approaches. The GPT series achieve quantitative changes that transform the landscape of scientific research by continuously scaling the simplest language modeling. RoBERTa present a detailed replication study of BERT pre-training that carefully measures the impact of many key hyper-parameters, training data and objectives, which results in greatly improved bidirectional language representations. DeiT and RSB closely evaluate the training recipe for smaller-sized plain ViTs and ResNets respectively, while ConvNeXt collectively examines previous architectural advancements for the next-gen ConvNets model design. empirically shows that a robust and effective recipe of knowledge distillation makes state-of-the-art large-scale image classification models affordable in practice.
Inspired by the spirits of these works, this paper provides a thorough evaluation of MIM visual representation learning that significantly bridge the gap between large-scale visual representations that achieve state-of-the-art performance and models that are affordable and accessible for the wider research community.
Discussion and Conclusion
In this work, we aim to contribute to the ongoing research on visual and vision-language representation learning. Instead of proposing an entirely new architecture or method, we present an in-depth evaluation of the existing MIM pre-training with CLIP vision features as the pretext task’s targets. Our experiments demonstrated that if robustly optimized, this approach is capable of producing highly performant, affordable, and transferable representations that outperform larger state-of-the-art specialized models.
Our analysis has revealed that base- & large-sized EVA-02 models can be effectively leveraged to obtain compact and expressive CLIP representations, which have the potential to facilitate modular, reusable, and scalable model design in the future . Our findings on moderate-sized models can also serve as a valuable reference for future research on model and representation scaling.
Furthermore, in combination with EVA , we demonstrate that alternate training of the pure MIM visual representations as well as vision-language CLIP representations can improve both MIM and CLIP performances in a bootstrapped manner (Fig. 3). This suggests a promising and scalable approach for pre-training both vision and vision-language representations of various sizes, which warrants further exploration in future research.
Acknowledgement
We would like to thank Hanxiao Qu, Yan Tian, Yemin Shi and Xigang Cao for their help on GPU resources. Zhao Xue, Quanyue Ma and Bowen Zhang for their help on datasets and benchmarks, and other colleagues at Beijing Academy of Artificial Intelligence for support throughout this project. We thank Wen Wang for constructive discussions on object detection & instance segmentation tasks, and Qiang Chen for constructive discussions on model weight initialization.
Appendix A Appendix
The position-wise feedforward network (FFN) in the original ViT design is a multi-layer perceptron (MLP) contains two layers (represented by the weight matrices and , biases are omitted) with a GELU activation function, denoted as . Formally,
SwiGLU FFN replace the first transformation of the original ViT’s FFN with a variant of the Gated Linear Unit (GLU) with a SiLU () activation function , Formally,
where is the element-wise product.
To keep the number of parameters and the amount of computation constant, we reduce the hidden units (the output dimension of and and the input dimension of ) of by a factor of 2/3 when comparing these layers to the original .
We use sub-LN (We find the inner attention LN unnecessary so we drop it) as the default normalization scheme for EVA-02-B and -L blocks. For the tiny- and small-sized model, we find using the default pre-LN configuration following is sufficient.
is a type of position embedding that unifies absolute as well as relative potential representations, and is widely adopted in state-of-the-art language models . For a detailed description of RoPE, please refer to . Our implementation is based on the open-sourced .
In brief, RoPE twists / rotates the input embedding (without changing the norm) such that the attention of a token at position to a token at position is linearly dependent on . Notably, unlike the conventional relative position representations that inject the positional information into the attention matrix, RoPE only manipulates , vectors. So RoPE is naturally compatible with off-the-shelf fused high-performance MHSA operators such as .
is a type of position embedding that unifies absolute as well as relative potential representations, and is widely adopted in state-of-the-art language models . For a detailed description of RoPE, please refer to . Our implementation is based on the open-sourced .
In brief, RoPE twists / rotates the input embedding (without changing the norm) such that the attention of a token at position to a token at position is linearly dependent on . Notably, unlike the conventional relative position representations that inject the positional information into the attention matrix, RoPE only manipulates , vectors. So RoPE is naturally compatible with off-the-shelf fused high-performance MHSA operators such as .
We use to initialize all weights in TrV blocks. The weight matrices in MHSA and FFN are sampled from , where std is .
A.2 Additional Results for Image Classification
In Table 17, we show that sufficiently pre-trained pure MIM EVA-02 representations (w/o IN-21K intermediate fine-tuning) outperform some previous leading approaches (even w/ intermediate fine-tuning).
In Table 18, we show that sufficiently pre-trained EVA-02 representations are robust enough that can be fine-tuned using various numerical precisions (e.g., and ) and optimizers (e.g., Lion , AdamW , and SGD ). Remarkably, the fine-tuning can be done using the SGD optimizer with only little performance drop.
Table 19 distinguishes MIM from conventional knowledge distillation in the context of “pre-training & fine-tuning” paradigm.
A.3 Data Contamination in MIM Pre-training: A Case Study
We provide a case study about the impact of data contamination in MIM pre-training when transferred to object detection and instance segmentation tasks. In short, we find the impact is minor.
We pre-train two EVA-02-L models, one uses the Merged-38M unlabeled images for MIM pre-training, and the other uses the images from IN-21K as the pre-training data. Both models are pre-trained with 1M steps with a batch size of 2k. Other settings & configurations are the same. Notice that the Merged-38M unlabeled images contain all Object365 (O365) set images, as well as 15k out of 20k LVIS set images (the Merged-38M images contain all the COCO training images, and LVISv1.0 split also contains 15k images from the COCO training set).
We study the transfer learning performance in three different settings:
(i) Directly transfer pure MIM pre-trained EVA-02 representations to O365 (MIM to O365), The performance is evaluated using the O365 setThe set images are publicly available. We have permission to access the annotations.(the set is a very large and challenging benchmark with 200k images and 2.5M instances in 365 different categories).
(ii) Directly transfer pure MIM pre-trained EVA-02 representations to LVIS (MIM to LVIS, Table 14a). The performance is evaluated using LVIS set (the set is a long-tail, large-vocabulary challenging benchmark with 20k images and 0.25M federated annotated instances in more than 1.2k different categories).
(iii) Transfer the EVA-02 representations with additional O365 intermediate fine-tuning to LVIS (MIM to O365 to LVIS, Table 14d). The performance is evaluated using LVIS set.
The results are summarized in Table 20. Overall, we find including unlabeled images from the development / test set for MIM pre-training has little impact on the final performance.
These experiments are motivated by our initial use of the Merged-38M pre-trained representation for LVIS set evaluation, which resulted in unintended use of unlabeled images from the development / test set for MIM pre-training, similar to the issue raised in . also reports a small percentage of images from IN-1K along with its variants, Flickr30K and COCO were detected in the LAION-400M dataset. This data contamination issue raises concerns about the validity of downstream benchmarks when a large number of unlabeled images are used for pre-training. While it is possible to identify and remove all duplicates for existing benchmarks, it may be infeasible to do so on already pre-trained models for future benchmarks or in real-world applications. Nonetheless, we believe that this issue should not hinder progress in data scaling for future representation learning studies.
A.4 Implementation Details
In this section, we summarize the training / evaluation settings, configurations, and hyper-parameters.