CLIP Itself is a Strong Fine-tuner: Achieving 85.7% and 88.0% Top-1 Accuracy with ViT-B and ViT-L on ImageNet
Xiaoyi Dong, Jianmin Bao, Ting Zhang, Dongdong Chen, Shuyang Gu, Weiming Zhang, Lu Yuan, Dong Chen, Fang Wen, Nenghai Yu
Introduction
CLIP is undoubtedly of paramount importance in advancing the zero-shot performance for multiple computer vision tasks, especially for zero-shot image classification. By leveraging web-scale noisy image-text pairs crawled from the Internet and simply aligning the image features with text features via contrastive learning, CLIP demonstrates strong transferable ability and has been widely explored as cross-modal guidance in various multi-modal tasks .
Recently, masked image modeling (MIM) has emerged as a successful visual pre-training paradigm and shows remarkable fine-tuning performance on downstream tasks. MIM essentially follows a mask-then-predict idea that masks a large ratio of the input image and forces the model to predict the masked part. The prediction target is a vital component and has attracted a lot of research efforts exploiting different targets such as discrete dVAE code , pixels , perceptual codebook , HOG features , online features and so on. Until recently, it has been well acknowledged in many works that utilizing CLIP features as prediction target of MIM leads to superior fine-tuning performance than other prediction targets.
Observing the success of distilling CLIP features to MIM models, we naturally ask a question: “If CLIP is a good teacher, why CLIP itself cannot be a strong fine-tuner?” However, we notice that CLIP ImageNet fine-tuning results reported in previous works are significantly different from each other. As a result, there is still not a clear answer to the question. With this in mind, we investigate what is the real best performance that CLIP can achieve by directly fine-tuning itself on the ImageNet dataset (the CLIP image encoder to be more specific).
In this paper, we justify that CLIP itself, regardless of its impressive zero-shot capability, is super powerful under fine-tuning setting, achieving unprecedented performance and even outperforming all existing methods that use CLIP as the teacher. We provide concrete fine-tuning details in depth to ensure reproducibility. Meanwhile, we also present a comprehensive analysis about each fine-tuning technique including the initial learning rate, the simultaneously updated EMA model, the layer-wise learning rate strategy, and fewer training epochs with weak input augmentation. As shown in Fig.1, starting from a commonly used ImageNet-1K fine-tuning setting, we eventually end up with a proper fine-tuning strategy with several off-the-shelf fine-tuning techniques that significantly improve the CLIP fine-tuning performance.
The final results are surprisingly inspiring. With input resolution, the CLIP ViT-Base model gets Top-1 accuracy on ImageNet-1K, even surpassing the state-of-the-art CLIP-targeted MIM method by . With a larger input resolution of , the CLIP ViT-Base model achieves top-1 accuracy, comparable with the strong competitive with the same model size and same input resolution but is supervisedly trained over JFT-300M or even JFT-3B . We also provide fine-tuning results over a larger backbone ViT-Large, obtaining and top-1 accuracy with and input resolution respectively. We think these results are encouraging that, with the noisy web image-text data, CLIP could learn a very strong visual representation, which is comparable with the model learned from carefully annotated classification data at a similar scale, at least on the classification task.
In summary, this paper builds a strong and reproducible baseline for directly fine-tuning CLIP model itself. Our paper demonstrates that the fine-tuning strategy is of crucial importance and justifies CLIP for ImageNet-1K fine-tuning. It will also motivate researchers in this field to rethink the latest proposed improvements upon CLIP.
Experiments
We first report the baseline results. The backbone is initialized from the CLIP pretraining model and we add a new LayerNorm layer and a fully-connected layer as the classification head. We use the average pooling feature of the backbone output as the classification head input. We finetune the model for 100 epochs, with the most commonly used augmentations and regularization, as shown in Table.1. We get a baseline result with top-1 ImageNet-1K classification accuracy. In the following, We gradually add or remove some of the finetune configuration components to further improve the performance.
Adjust Learning Rate. We first study the influence of different learning rates. As shown in the following Table 2, we find that the fine-tuning need a quite small learning rate or , about smaller than our default setting.
Exponential Moving Average (EMA). Transferring a model pre-trained on a large dataset to a small new dataset always suffers from the overfitting problem, and the EMA is a commonly used method to relieve it. The EMA is realized by keeping a moving average of the weight of all the model parameters. We set the EMA momentum factor as and report its influence with five different learning rates in Table 3. We find with the EMA, most of the settings get better results, especially when the learning rate is not proper, the EMA reduces its gap toward the best setting.
Layer-wise Learning Rate Decay (LLRD). The Layer-wise learning rate decay is a fine-tuning technique widely used in the Nature Language Processing (NLP) area to fine-tune the BERT model. It has been used in BEiT and shows significant performance.
The LLRD assigns different learning rates for each layer of the model backbone. It sets a large learning rate for the top layer and uses a multiplicative decay rate to decrease the learning rate layer-by-layer from top to bottom. With a large learning rate, the feature of the top layers changes more and could adapt to new tasks. On the contrary, the bottom layers have a small learning rate, so the strong feature learned from the pre-training is preserved.
We first pick the base learning rate (the learning rate of the last layer) from and increase the LLDR from to . As shown in Table 4, we find a large base learning rate with a small LLDR performs better.
Following the observation above, we pick LLDR from a small range and increase the learning rate from to . As shown in Table 5, we find with a proper LLDR, all the results are better than our baseline result of . With LLDR and learning rate , we reach top-1 accuracy, better than the EMA baseline.
Training Length. We show the epoch-accuracy curve in Fig.2, and we find that with 100 epoch fine-tuning, the model converges quite fast. It reaches the best performance before the 50th epoch and tends to overfit the training set with the rest epochs.
When we reduce the training length to 50 epochs with 10 epochs warmup, and keep the rest setting unchanged. We find that the model gets a similar top-1 accuracy of and it seems under-fitting, as the online accuracy (line w/o EMA in the figure) is still increasing. The shortened training length also reduces the fine-tuning cost to half, so we conduct the following experiments with the 50-epoch setting.
Data Augmentation. In our default setting, we use two kinds of augmentations: content modification augmentations and content diversification augmentations. The former changes the image content with other images (Mixup, CutMix), or removes part of the content (Random Erase). The latter applies some affine transformation (rotate, shear, translate, etc. in the Random Augmentation) or some statistics changes (contrast, brightness, color, etc. in the Random Augmentation) to the image.
We first study the influence of Mixup and CutMix in Table 6. We find that under the 50 epoch setting, removing the Mixup/Cutmix gets the best results. We suspect that as the feature learned by the CLIP model is good enough, it only needs some weak augmentations to transfer to a new dataset.
Then we follow the setting above and study the influence of Random Erase in Table 7. We find that the influence of Random Erase is quite minor and the default setting works well.
We further study the influence of Random Augmentation in Table 8. It has two main hyper-parameter, for the augmentation magnitude (severity) and for the number of transformations selected per image. When we fix and change from to , we find or augmentations for each image work better.
Finally we fix and change the from to (10 is the max level). As shown in Table 9 We find the influence of the augmentation strength is quite small.
Final Results. With the above attempts, we improve the CLIP-Base/16 fine-tuning accuracy from to , with improvement. We find all the successful attempts follow a similar idea: adapt the CLIP to new tasks while keeping its representation changes as slight as possible. For data augmentation, we remove the strong MixUp and CutMix which is totally different from the pre-training data, which may force the model to change more to adapt to it. For model parameter update, we shorten the training epochs, and use both layer-wise learning rate decay and EMA to slow down the change of the CLIP representation, especially its bottom layers representation.
In Table 11, we compare our results with previous supervised pretraining methods, self-supervised MIM methods and CLIP-based MIM methods (i.e., use CLIP as the teacher). We find without any additional design nor training of a new model, the CLIP model with a proper fine-tuning setting outperforms these methods, surpassing the SOTA method by .
We further apply the fine-tuning strategy to a larger resolution and get top-1 accuracy. Compared with the supervised training baseline with different data scales(results are copied from and appendix of , selecting best results for each setting), we find that the CLIP with 400M web image-text pairs shows similar performance with the carefully annotated JFT dataset with similar or even larger data scale.
We also apply the fine-tuning strategy to the CLIP-Large/14 model. With the configurations shown in Table.10, we get a strong result, top-1 accuracy with resolution input. For the ViT-Large model, the CLIP use ViT-L/14 with resolution input while the previous works report results of ViT-L/16 with a larger input . Here we show their FLOPs for clear comparison. We find that with half of the FLOPs, the CLIP with ViT-L/14224 shows comparable results with the JFT-300M pre-trained model. When we further increase the input resolution to (190.6G FLOPs, similar to the ViT-L/16384 with 190.7G FLOPs), we get top-1 accuracy, better than the model trained on JFT-300M result and slightly worse than that on much larger data JFT-3B.
We think these results are encouraging that only with the noisy image-text data collected from the Internet, the CLIP could learn a very strong visual representation, which is comparable with the model learned from carefully annotated classification data at a similar scale, at least on the classification task. This shows the strong capability and potential of contrastive image-language learning.
2 More Ablations.
Partial fine-tuning. The layer-wise learning rate decay is used with an idea that the early layers learn quite good features from the large-scale pre-training so it only needs a small learning rate and changes slightly. Here we use the partial fine-tuning design to study the feature learned from different layers. In practice, we fine-tune part of the model layers and freeze the rest layers. As shown in Fig.3, with only one layer fine-tuning, the model gets top-1 accuracy, this shows the feature learned by CLIP is strong. With half of the layers frozen, the model gets top-1 accuracy, close to the fully fine-tuning results. This also proves the necessity of the layer-wise learning rate decay that the bottom layers only need small changes.
Architecture Modification. CLIP uses the vanilla Vision Transformer architecture with absolute position encoding (APE), and there are several advanced techniques to further improve the model performance, such as the relative position encoding (RPE) and layer scale (LS). The RPE provides more inductive bias with a relative position table, and LS adds a learnable scale factor for each block, improving the model capability. Here we further add them to the fine-tuning model and study their influence in Table 12. In practice, we initialize the RPE table with and the LS factor with . However, their influence is quite minor. This may be because the model already learns enough inductive bias through large-scale pre-training. This is also consistent with the observation in that a vanilla ViT trained with large-scale data could learn a similar pattern to a hybrid ViT that uses conv-stem to attend local manually.
Data Augmentations. In our main experiment, we study the influence of commonly used data augmentations. Here we further study two other augmentations in Table 13: 3Aug and the AugMix. The 3Aug also proposes using single resize crop (SRC), instead of the default random resize crop (RRC). We find that the 3Aug+RRC performs similarly to the RandAug, while the 3Aug+SRC and AugMix perform worse.
In the random resize crop, an important factor is the crop ratio, it controls the cropped region size of the input image. Its upper bound is 1 and the default lower bound is 0.08. A larger lower bound leads to a weaker augmentation that the cropped results contain more area of the input. Here we study the influence of the lower bound. We find the influence is minor and a smaller crop ratio works better.
Model regularization. Here we study the influence of different regularization methods.
Drop path rate (DPR) is a commonly used regularization to ease the over-fitting problem. For vision transformer architecture, it randomly drops the MHSA or FFN results in each block and only output the shortcut result. In most cases, we set the transfer learning DPR slightly larger than it was used in the pretraining to get better results. The CLIP is pre-trained with DPR and here we study the fine-tuning performance with a larger DPR. As shown in the following Table 15, the fine-tuning performance decrease with using DPR. This may be because the DPR is not used in the CLIP pre-training so the model is not aware of such regularization.
Label smoothing is a regularization used in the cross-entropy loss to solve the overconfident problem of the classification model. As shown in Table 16, we find that removing it leads to overfitting and worse performance.
Weight Decay is an additional loss calculated in the optimizer to prevent overfitting. In Table 17, we find that its influence is quite minor.
Conclusion
In this paper, we study how far the performance can be achieved by directly fine-tuning the CLIP on the ImageNet-1K dataset. We find that CLIP itself is a super strong finetuner, even outperforming existing methods that use CLIP as the teacher. We hope that our study could be a new baseline for the following works, raise attention to the strong recognition capability of the CLIP model, and rethink recent improvements based on CLIP. In the future, we will transfer our strategy to other multimodal vision foundation models Florence and OmniVL.