EVA-CLIP: Improved Training Techniques for CLIP at Scale

Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, Yue Cao

Introduction

CLIP (Contrastive Language-Image Pre-training) is a powerful vision-language foundation model that leverages large-scale datasets to learn rich visual representations by bridging vision and language via contrastive image-text pre-training. The CLIP model exhibits robust zero-shot transferability , and has the potential to enhance both multimodal and unimodal visual tasks, such as AI-generated content applications . Despite its significance, training CLIP models remains an inevitable challenge due to its high computational cost and training instability issues when scaling up.

In this paper, we propose EVA-CLIP, a family of models that provides a feasible, efficient, and effective solution for training CLIP models. Our approach incorporates several techniques that can significantly reduce training costs, stabilize the training process and improve zero-shot performance, including initialize CLIP with pre-trained EVA representations, the LAMB optimizer, randomly dropping input tokens , and a speedup trick named flash attention . With these techniques, we are able to greatly stabilize the training of CLIP models at scale with less computational costs and outperform the training-from-scratch counterpart with much fewer samples on a broad range of zero-shot benchmarks. Our largest 5.0B-parameter EVA-02-CLIP-E/14+ with only 9 billion seen samples achieves 82.0% zero-shot top-1 accuracy on ImageNet-1K val. A smaller EVA-02-CLIP-L/14+ with only 430 million parameters and 6 billion seen samples achieves 80.4% zero-shot top-1 accuracy on ImageNet-1K val.

Approach

Training CLIP models is very hard and costly. The need for a large batch size and scaling up CLIP models can lead to significant computational resource requirements and even instability training problem. Fortunately, EVA-CLIP offers a highly efficient and effective solution that significantly reduces the computational cost while achieving superior zero-shot performance across a broad range of benchmarks.

. To improve feature representation and expedite convergence of CLIP models, we adopt pre-trained EVA models which combines high-level semantics of image-text contrastive learning with geometric and structural capture from masked image modeling. We use pre-trained EVA weights to initialize the image encoder of EVA-CLIP. Our empirical study demonstrates that pre-trained EVA models not only help EVA-CLIP achieve superior performance on various zero-shot benchmarks but also expedites and stabilizes the training process.

Optimizer

. We use the LAMB optimizer for training our EVA-CLIP models. LAMB optimizer is specifically designed for large-batch training, and its adaptive element-wise updating and layer-wise learning rates enhance training efficiency and accelerate convergence rates. Given the exceptionally large batch sizes used for training CLIP models, the original CLIP model uses a batch size of 32,768 and some open-sourced CLIP models even use an extremely large batch size of more than 100k. EVA-CLIP shows that LAMB optimizer is the preferred optimizer for training large-scale CLIP models.

FLIP [33]

has shown promising results in large-scale settings. In this work, we leverage FLIP to improve the time efficiency of training CLIP models. Specifically, we randomly mask 50% image tokens during training esulting in a significant reduction of time complexity by half. This approach also allows for a 2×\times increase in batch size, without any additional memory costs.

Experiments

. In our experiments, we initialized the vision encoder with pre-trained weights from EVA and the text encoder with pre-trained weights from either OpenAI CLIP or OpenCLIP . Specifically, the vision encoder of EVA-01-CLIP is initialized from EVA-01 , and the vision encoder of EVA-02-CLIP is initialized from EVA-02 . We adopt the LAMB optimizer with β1\beta_{1} = 0.9, β2\beta_{2}=0.98, and a weight decay of 0.05. We applied different learning rates and layer decay rates to the vision encoder and text encoder to ensure optimal training. For example, we set the learning rate to 2e-4 for the vision encoder and 2e-5 for the text encoder of EVA-01-CLIP-g during the first 2000 warming-up steps. Afterward, we decayed the learning rate linearly to 0 for the remainder of the training steps. To further improve the training process, we used the DeepSpeed\mathtt{DeepSpeed} optimization library with ZeRO stage-1 optimizer , gradient checkpointing and flash attention to save memory and accelerate training process. We found that using the fp16\mathtt{fp16} precision with dynamic loss scaling was sufficiently stable throughout the EVA-01-CLIP-g training process, while bfloat16\mathtt{bfloat16} format was necessary to stabilize the training process of EVA-02-CLIP-E+.

To construct our training dataset, Merged-2B, we merged 1.6 billion samples from LAION-2B dataset with 0.4 billion samples from COYO-700M .

System-level Comparison

. We present CLIP model configurations and zero-shot accuracies on ImageNet variants and ObjectNet in Table 1. EVA-02-CLIP-E/14+ achieves the highest zero-shot top-1 accuracy of 80.9% averaged accross all 6 benchmarks, with the smallest performance drop (1.1% top-1 accuracy gap). Notably, this result surpasses the previous largest and best open-sourced OpenCLIP-G/14 by 1.9% on ImageNet and 4.7% on the average accuracy of 6 benchmarks. Remarkably, with those powerful techniques, the large size EVA-02-CLIP-L can even reach up to 80.4% zero-shot top-1 on ImageNet, outperforming OpenCLIP-G/14 with only ∼\scriptstyle\sim1/6 parameters and ∼\scriptstyle\sim1/6 image-text training samples.

In Table 2, we further demonstrate the efficacy and robustness of our approach on all 27 zero-shot image classification benchmarks. Our largest EVA-02-CLIP-E/14+ achieves averaged 77.5% on all 27 benchmarks. Notably, our EVA-02-CLIP-L/14+ model, which only has ∼\scriptstyle\sim1/2 of the model size and ∼\scriptstyle\sim1/5 image-text pairs, achieves a 1.2-point averaged improvement over OpenCLIP-H/14.

For video classification, we sample only a single center frame in each video, making it an image classification task. Following the conventional settings, we report the top-1 accuracy for UCF-101 and the mean of top-1 and top-5 accuracy for Kinetics-400 , Kinetics-600 and Kinetics-700 . In Table 3 we show that EVA-CLIP is also quite effective in zero-shot video recognition benchmarks.

Table 4 reports the zero-shot image and text retrieval results on Flickr30K and COCO . EVA-CLIP outperforms all the competitors at the base and large model size. While the zero-shot retrieval performance of EVA-02-CLIP-E/14 is slightly lower than OpenCLIP-G/14, the results are still competitive. We speculate that the main reason is that retrieval tasks depend more on the capacity of the text encoder and the number of training samples, and in comparison, the capacity of the text encoder in EVA-02-CLIP-E/14 is smaller and the number of training samples is less than that of OpenCLIP-G/14. To this end, we have trained EVA-02-CLIP-E/14+ with a larger capacity text encoder and more training samples. The results show that this improved model can substantially improve retrieval performance and outperform OpenCLIP-G/14 on zero-shot text retrieval tasks.

Ablation Study

. We first ablate the EVA-CLIP design in Table 5. The image encoder is ViT-B/16 model and text encoder is CLIP-B-16. We conducted experiments with a batch size of 32k and evaluated zero-shot accuracy on the ImageNet-1K validation set. It is important to note that we used a shorter training schedule than the final models.

We trained our model on the LAION-400M dataset using the AdamW optimizer. Compraing to training from scratch, EVA initialization resulted in a 1.8% increase in zero-shot top-1 accuracy on ImageNet with only ∼\scriptstyle\sim1/2 seen samples.

Furthermore, we experimented with using the LAMB optimizer instead of AdamW with EVA initialization on the LAION-400M dataset. This resulted in a 0.7% increase in zero-shot top-1 accuracy on ImageNet with the same seen samples. When 50% masking was applied, the accuracy decreased by 0.7% while enjoying a speedup of 2×\times. These results highlight the significance of LAMB optimizer in training high-performing models and the strategy of masking image tokens in faster training without marginally decreasing accuracy.

We also conducted experiments with the LAION-2B dataset using EVA initialization, LAMB optimizer, and 50% masking, which increased the accuracy by 0.7% compared to LAION-400M. Only half samples were required to achieve the same top-1 accuracy when using the merged-2B dataset. It demonstrates the importance of dataset sizes and the significant convergence speed through merging the two datasets.

Computation Costs

. In Table 6, we present the memory and time cost of our implementation. As shown, masking 50% of image tokens can accelerate training time by 2×\times and using flash atttention can reduce additional 15% training time.

Using all these techniques, we can train EVA-CLIP with a lower budget than other counterpart CLIP models. For instance, EVA-CLIP-B/16 can be trained on a batch size of 32k and converges within 300 hours using on 16 NVIDIA 40GB-A100 GPUs. Similarly, the billion-scale EVA CLIP-g/14 can be trained on a batch size of 65k and requires less than 25 days to train 12B samples using 64 NVIDIA 40G-A100 GPUs. These results demonstrate the scalability and effectiveness of our method in achieving state-of-the-art results while maintaining an optimal balance between training time and GPU memory utilization.

References