MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image Pretraining

Xiaoyi Dong, Jianmin Bao, Yinglin Zheng, Ting Zhang, Dongdong Chen, Hao Yang, Ming Zeng, Weiming Zhang, Lu Yuan, Dong Chen, Fang Wen, Nenghai Yu

Introduction

Vision-language (VL) contrastive learning has shown remarkable success in pretraining for various tasks. With large-scale image-text pairs available on the Internet, the model composed of a simple dual encoder design learns strong semantic prior by aligning between image and text. The resulting visual encoder not only exhibits excellent linear probing and finetuning performance, but also enables impressive zero-shot performance with the guidance of the language encoder, showing the generality of natural language and its ability to supervise a wide range of visual concepts.

Nonetheless, the associated language description, though providing richer information than mere class labels, still can hardly describe all the information in the corresponding image, as images are continuous signals with fine-grained details and complex semantics. As a result, the VL contrastive by aligning global representations may only focus on the text-described objects and ignore the rest which might be useful for downstream tasks.

In this paper, we are interested in how to fully leverage the image itself to facilitate the VL contrastive to further improve the transfer capability. (1) Firstly, the learned feature representation shall characterize local patches, serving as a complementary for global representation in VL contrastive. Inspired by the recent success of masked image modeling in learning patch representations, we also randomly mask the input image with a large portion to force the visual encoder to focus on the remaining visible patches. (2) Secondly, the learned representation for local patches shall possess semantic meanings, being consistent with the global representation receiving semantic text supervision. We bring mean teacher self-distillation to supervise the learned patch representations with the visual feature representations, enabling implicit supervision from natural language. The resulting objective is denoted as masked self-distillation where the student model and the teacher model come from the same neural networks and the knowledge is distilled from the full image (fed to the teacher model) to the masked image (fed to student model). To this end, we introduce MaskCLIP by incorporating masked self-distillation into VL contrastive to advance the transferable visual encoder.

There are several recent attempts also exploring the capability of the visual encoder under natural language supervision. The common approach is to introduce contrastive learning or masked image modeling on the vision side together with contrastive language-image pretraining. However, the performance indeed improves based on CLIP but does not as well as our masked self-distillation. We argue that (1) the contrastive learning objective based on central crop augmentation actually learns global representations for salient objects while lack of attention on the surrounding backgrounds ; and (2) masked image modeling usually needs to remap the learned representation to pixels or discrete tokens . Such low-level prediction target is inefficient for semantic feature learning and thus also conflicts with high-level language supervision in VL contrastive. A brief illustration is presented in Figure 1. In the experiments, we conduct comprehensive ablations to analyze the difference and provide numerical and visual evidence for better understanding.

Symmetrically, we argue that local semantic supervision on the text branch is also helpful for the text encoder and eventually beneficial for zero-shot performance. So we introduce the same mask-data-modeling format supervision into the text branch as well. Different from images where the pixel is low-level signal, the words crafted by human beings are already highly semantic, so we use the tokenized word piece as the prediction target directly, following the well-studied mask language modeling method BERT. Meanwhile, to reduce the output conflicts between contrastive learning and mask language modeling, we introduce a small decoder for the mask language modeling branch.

We train our MaskCLIP on a subset of a publicly available image-text pairs dataset, YFCC , and thoroughly evaluate the transfer ability of visual representations on several vision benchmarks: ImageNet-1K for classification, ADE20K for semantic segmentation, MS-COCO for detection and segmentation, as well as a batch of other classification benchmarks. When it comes to ImageNet-1K classification, MaskCLIP achieves +6.9%+6.9\%, +7.2%+7.2\%, +1.3%+1.3\% higher than CLIP for zero-shot transfer, linear probing, and finetuning respectively. For vision downstream tasks, we reach +2.7+2.7 mIoU on ADE20K and +1.8+1.8 APb, +1.4+1.4 APm on MS-COCO . For vision-language tasks, MaskCLIP achieves +6.1%+6.1\% average zero-shot accuracy on 20 datasets, and +17.2%+17.2\%, +12.8%+12.8\% rank@1 improvement on the Flickr30K image-test retrieval. In the recent Image Classification in the Wild challenge academic track, our MaskCLIP gets the 1st1_{st} result with 48.9%48.9\% TOP-1 average accuracy, surpassing the second team with 3.4%3.4\%.

In summary, the major contributions of this work are:

We present a novel vision-language pretraining framework MaskCLIP, by introducing masked self-distillation objective to facilitate VL contrastive for better transferable visual models.

We present extensive ablation studies on MaskCLIP variants and provide in-depth analysis numerically and visually to help understand how the proposed masked self-distillation assists VL contrastive.

We demonstrate our MaskCLIP on tens of benchmarks, showing the superiority under all three settings: zero-shot, linear probing, and finetuning.

Related Work

Vision-language pretraining Recent years have seen rapid progress made in vision-language pretraining . Several multiple cross-modality loss functions have been proposed for the training objective, such as image-text matching , masked language modeling , masked image modeling , contrastive loss . These objects are often mixed with each other to form a compound objective. While a variety of approaches have been proposed, few works investigate the performance on visual representation learning for image classification. Recently, CLIP and ALIGN show that the image-text contrastive learning objective achieves promising performance for visual representation learning. There are many following works proposed to further improve the pretraining performance, DeCLIP , SLIP , COTS , ViCHA , CYCLIP use additional uni/multi-modality supervision to improve the model capability, and PyramidCLIP , KLITE , IDEA seek to external knowledge from pre-trained models or datasets as the additional guidance. FILIP and LOUPE introduce fine-grained alignment to the model. Focusing on this research direction, we analyze the desired properties of supervision which could be complementary to CLIP, and propose the masked self-distillation objective incorporated with the image-text contrastive loss to further improve pretraining performance for various visual understanding tasks.

Self-supervised learning Self-supervised visual representation learning has attracted increasing attention over the past few years. The objective of the self-supervised learning is mainly divided into two categories: contrastive and generative . The contrastive methods, such as MOCO , SimCLR , BYOL , SimSiam , and DINO measure the similar and dissimilar samples by contrastive loss. Their success heavily depends on the strong data augmentation. The generative methods, such as BEiT , MAE , PeCo , BEVT , BootMAE and MaskFeat leverage masked image modeling to reconstruct the remaining masked part of its original input from the given visible parts. The generative methods show more promising transfer performance than the contrastive methods, as generative objective learns patch representations while contrastive objective focuses on learning centric global representations .

Self-knowledge distillation Self-knowledge distillation aims to distill the knowledge in a model itself and uses it for training the model. Instead of distilling knowledge from a pretrained teacher model , self-knowledge distillation regards a temporal ensemble of the student model as the teacher. It means that a student model becomes a teacher model itself, which gradually utilizes its own knowledge for softening the hard targets to be more informative during training. Self-knowledge distillation has been explored in semi-supervised learning , contrastive learning , self-supervised learning . In this paper, we use visual features supervised by natural language for guidance in masked self-distillation which naturally fit VL contrastive to learn more transferable visual representations.

MaskCLIP

We introduce MaskCLIP, a novel framework that learns visual representations. The core part of MaskCLIP is its backbone image encoder, denoted by EIE_{I} as shown in Figure 1. It obtains the transferable capability during pretraining that could benefit downstream vision tasks. Following recent self-supervised approaches , we implement the backbone EIE_{I} as a Vision Transformer (ViT) . The prediction results from EIE_{I} given an input image II then should be a collection of visual feature tokens, represented as

Here cls is short for class token. 1,…,N1,\dots,N are the indexes of the non-class tokens.

The rest of this section starts with the utilization of language supervision. More shall be emphasized on the masked self-distillation, which we deem crucial for visual pretraining.

Following , we introduce a Transformer-based text encoder ETE_{T} to leverage language knowledge. It aims to align the global feature representations of an image and a text with respect to some forms of similarity. Precisely, consider a given image-text pair {I,T}\{I,T\}, besides extracting the visual feature representation EI(I)E_{I}(I) using the vision backbone as shown by Equation 1, we additionally use the text encoder ETE_{T} to extract linguistic features from the text TT.

The mean feature of the two branches are regarded as the global representations and are fed into a projection head (implemented as a fully-connected layer) respectively to obtain the metric embeddings eTe^{T} and eIe^{I}. Image-text contrastive loss is employed to align them during pretraining. The loss can be formulated as LT+LI\mathcal{L}_{T}+\mathcal{L}_{I}, with

where BB stands for the number of image-text pairs within a training mini-batch, i,ji,j are indexes within the batch; σ\sigma stands for the temperature for the loss functions, which is learned together with all other parameters during training.

2 Masked Self-distillation for Visual Encoder

Knowledge distillation is a learning paradigm where a student model is trained to match the output of a given teacher model, so that the student model can be improved by the teacher. Instead of bringing in an external teacher, self-distillation methods such as proposes using a mean teacher model that is derived from the student itself. In specific, the teacher shares the same structure with the student, while the parameters of the teacher are exponential moving averages (EMA) of the parameters from the student. In the following, we would use the term “EMA model” to represent such mean teacher model constructed from the student.

MaskCLIP leverages the mean teacher self-distillation to enhance its vision representations. Let EˉI\bar{E}_{I} be the EMA model of the backbone encoder EIE_{I}. θt\theta_{t} and θˉt\bar{\theta}_{t} are the parameters of EIE_{I} and EˉI\bar{E}_{I} at training step tt. θˉt\bar{\theta}_{t} is updated with

where α\alpha is a hyper-parameter for smoothing updates. We propose to incorporate masked image modeling into self-distillation, resulting in masked self-distillation with asymmetric input for student model and teacher model.

In specific, considering a given input image II, we first feed it to the EMA model EˉI\bar{E}_{I} (teacher model) to obtain the distillation targets. These target features can be represented as

In the meantime, we randomly mask a large portion of the input image patches and then feed it into the original backbone EIE_{I} (student model). Following , we only feed the visible (unmasked) patches, denoted by I′I^{\prime}, into the original backbone EIE_{I} to speed up computation and save memory. Let M\mathcal{M} be the indexes of all the masked tokens. These encoded features corresponding to visible tokens can then be denoted as EI(I′)={fcls′}⋃{fk∉M′}E_{I}(I^{\prime})=\{f^{\prime}_{\textit{cls}}\}\bigcup\left\{f^{\prime}_{k\not\in\mathcal{M}}\right\}. They are then joined with a shared and learnable feature vector, denoted as mm, that represents mask tokens, to form a complete set of features {fcls′,f1′,f2′,…,fN′}\{f^{\prime}_{\textit{cls}},f^{\prime}_{1},f^{\prime}_{2},\dots,f^{\prime}_{N}\}, with fi∈M′=mf^{\prime}_{i\in\mathcal{M}}=m. We attach positional embeddings onto all these tokens, and append a small Transformer DD as a decoder to predict features of the masked region from the visible tokens, which could be formulated as

Inspired by , we use an online quantizer h()h() to transform the output features into a soft codewords distribution, and minimize the cross-entropy between the target features and the predicted features. Formally,

here the parameter of the teacher quantizer hˉ()\bar{h}() is also EMA updated by the online quantizer, similar to the teacher model.

3 Local Semantic Learning for Text Encoder

Besides the local semantic supervision for the visual encoder, we argue it is also helpful for the text encoder. So we introduce the BERT pretraining into the text branch. For the text T={tsos,t1,t2,...,tM,teos}T=\{t_{sos},t_{1},t_{2},...,t_{M},t_{eos}\}, we denote the masked input as T′={tsos′,t1′,t2′,...,tM′,teos′}T^{\prime}=\{t_{sos}^{\prime},t_{1}^{\prime},t_{2}^{\prime},...,t_{M}^{\prime},t_{eos}^{\prime}\}, where ti∈MT′=mtt_{i\in\mathcal{M}_{T}}^{\prime}=m_{t} and ti∉MT′=tit_{i\notin\mathcal{M}_{T}}^{\prime}=t_{i}, and MT\mathcal{M}_{T} be the indexes of all the masked text tokens. The output feature of the encoder is ET(T′)E_{T}(T^{\prime}).

To reduce the output conflict between the global image-text contrastive learning and the local mask language modeling, we further introduce a small text decoder, which shares the same architecture as the encoder but with only a few layers. So that the global prediction and local prediction are conducted at different layers. We denote the output feature as: (DT∘ET)(T′)={tsos′′,t1′′,t2′′,...,tM′′,teos′′}(D_{T}\circ E_{T})(T^{\prime})=\{t_{sos}^{\prime\prime},t_{1}^{\prime\prime},t_{2}^{\prime\prime},...,t_{M}^{\prime\prime},t_{eos}^{\prime\prime}\} and the loss could be formulated as:

4 Overall Loss Functions

Finally, we pretrain MaskCLIP with all these losses combined:

with λ,β\lambda,\beta being the hyper-parameter weighting between VL contrastive loss and self-supervised learning loss. All the components of MaskCLIP are trained from scratch, including the visual backbone EIE_{I}, the visual decoder DD, the text encoder ETE_{T}, as well as the text decoder DTD_{T}.

Experiments

Model architecture. Our framework consists of the visual encoder EIE_{I}, the text encoder ETE_{T}, the visual decoder DD, and the text decoder DTD_{T}. We adopt the widely used Transformer ViT-B/16 for a fair comparison. It is composed of 12 layers, 768 width, and 12 head. The input image is 224×224224\times 224 resolution and is further split into 14×1414\times 14 patches with size 16×1616\times 16. A learnable cls token is prepended to the 196 embeddings. For the text encoder, we adopt a 12-layer, 512-width, and 8-head Transformer following CLIP , and the text decoder has 4 layers. The number of text tokens is fixed to 77 with necessary truncations or paddings. For the image decoder, we directly use a one-layer Vision Transformer.

Pretraining details. We train our proposed MaskCLIP from scratch for 25 epochs, the batch size is fixed to 4096 for all the experiments. The masks used in the mask self-distillation branch and mask language modeling branch are random mask with a mask ratio of 75% and 20%. We pretrain all the models with the commonly used YFCC15M dataset, which is flited from the YFCC100M dataset by .

Downstream details. We evaluated MaskCLIP on several downstream datasets, including ImageNet-1K , ADE20K , MS-COCO , Flicr30K et al. For ImageNet-1K, we report zero-shot, linear probing, and finetuning performance. The zero-shot is conducted following the label prompt setting in SLIP . For linear probing, we fix the backbone and train a new linear classifier for 90 epochs. For finetuning, we follow the setting in BEiT and finetune the model for 100 epochs with a layer-decayed learning rate. See supplemental materials for more details.

2 Analysis

We first present our analysis by studying different ways of boosting CLIP. The baseline is CLIP trained on the YFCC-15M. Besides the introduced masked self-distillation, we consider two other popular methods: (1) SimCLR , a representative contrastive method; and (2) MAE the state-of-the-art masked image modeling approaches. All the compared methods are trained on the YFCC-15M for a fair comparison. We have the following observations.

Vision self-supervision helps VL contrastive. We evaluate the models on both vision task ImageNet-1K classification and vision-language task image-text retrieval on Flicker30K and present the comparison in Table 1. All the added vision self-supervision, regardless of contrastive or generative, improves the baseline CLIP. Among them, our proposed MaskCLIP achieves the best results in terms of all the evaluation metrics, outperforming CLIP with +6.9%, +7.2%, + 1.3% on ImageNet-1K classification for zero-shot, linear probing, and finetuning respectively, and +17.2%, +12.8% on Flicker30K for image-to-text retrieval and text-to-image retrieval. We also report the training GPU memory usage and time-consuming cost in Table 1. It is worth noting that the contrastive model (CLIP+SimCLR) compares two additional views of the input image, resulting in larger GPU memory usage and longer training time.

Masked image modeling is able to learn representations for local patches. We argue that the image encoder only pays attention to the text-described objects under VL contrastive due to sparse text description and to the centric objects under image contrastive due to central-crop augmentation. In contrast, masked image modeling forces the image encoder to focus on local patches using token-wise objectives by mandatorily masking a large portion of patches. Here, we provide numerical comparisons for evidence. We conduct an “Annotation-free zero-shot segmentation” experiment to test the zero-shot segmentation. The results on such a dense prediction task would better reveal the ability of local patch representations than global classification. Following the design in DenseCLIP , we use the prompted label feature as the linear classification weight to realize segmentation, without any training procedure. We evaluate the performance on two widely used datasets: ADE20K and Pascal Context . The results are shown in Table.2. We can see that equipped with masked image modeling, our MaskCLIP as well as CLIP+MAE achieves better results than CLIP and CLIP+SimCLR, validating our hypothesis.

Masked self-distillation learns semantic representations for local patches. Our masked self-distillation predicts visual features dynamically outputted by the visual encoder and thus implicitly gets supervision from the text side via VL contrastive. While MAE predicts fixed low-level pixels, making it inefficient to learn semantic representations (as the objective may force the representation to memorize low-level details) and thus causing conflict with VL contrastive. To show this, we select images from MS-COCO and calculate the feature similarity between image features and their corresponding caption features. We also select objects in the caption, prompt it to a new caption, such as “a photo of teddy bears”, and calculate the similarities. An example is shown in Figure 5 (More can be found in the supplementary material). Comparing MaskCLIP with CLIP+MAE in the fourth column, we can see that CLIP+MAE uses color as evidence and fails to distinguish the white teddy bear from the white snow. While our MaskCLIP successfully differentiates the two objects, suggesting ours learn more semantic features. On the other hand, the superior results of MaskCLIP shown in Table 1 and Table 2 also validate this. It is worth mentioning that CLIP and CLIP+SimCLR fail to have a correct response partition for different single objects like MaskCLIP, further justifying our second observation.

3 Comparison with Previous Methods

To show the effectiveness of MaskCLIP as a general vision-language pretrain method, we conduct experiments on both vision tasks and vision-language tasks. For vision tasks, we report results on ImageNet-1K classification, MS-COCO object detection, and ADE20K semantic segmentation. For vision-language tasks, we report zero-shot results on recent challenging ICinW 20 datasets benchmark and image-text retrieval results on Flickr30K and MS-COCO . In the following, we compare with the supervised baseline DeiT , self-supervised methods SimCLR and MAE , and vision-language methods CLIP and SLIP . For a fair comparison, we train SimCLR and MAE on YFCC-15M with the same epochs.

Classification on ImageNet-1K. As shown in Table 3, MaskCLIP benefits from the advantages of both VL pretraining and image mask self-distillation that shows strong performance on all the metrics. For zero-shot tasks, MaskCLIP outperforms CLIP by +6.9%+6.9\% with 25 epoch training and achieves +1.7%+1.7\% higher than the recent work SLIP. When it comes to finetune, MaskCLIP reaches 83.6%83.6\% top-1 accuracy, and outperforms CLIP by +1.3%+1.3\%.

Semantic segmentation on ADE20K. Then we apply our MaskCLIP to the semantic segmentation task. Here we use the UperNet framework with 512×512512\times 512 input and end-to-end training for 160K iterations. The evaluation metric is the mean Intersection of Union (mIoU) and we report single-scale evaluation results here. The results are given in Table 3. Our method achieves 50.5 mIoU, +2.7+2.7 mIoU than our baseline method CLIP, and +2.0+2.0 mIoU than SLIP. This verifies the effectiveness of our introduced incorporation.

Object detection and instance segmentation on MS-COCO. We further investigate our transfer performance on object detection and instance segmentation in Table.3. Here we use Mask-RCNN framework with single-scale input and 1×1\times schedule (12 epochs). Our method achieves 45.445.4 box AP and 40.940.9 mask AP, +1.8/1.4+1.8/1.4 better than CLIP, and +1.4/0.6+1.4/0.6 better than SLIP.

Zero-shot on small datasets. We also report zero-shot performance on 20 small datasets under the ICinW setting (see the introduction below) in Table 4. We find that all the methods perform poorly on some datasets such as Aircraft(1% acc for random guessing, we omit the description in the following), Fer(24.7%), Country211(0.5%), GTSRB(5.9%), Cars(0.8%). This might be caused by the data domain gap that the YFCC-15M contains few related images and descriptions. For the rest of the datasets, all the methods get reasonable performance and our MaskCLIP gets the best performance on most datasets.

Image Classification in the Wild (ICinW) Challenge The ICinW challenge is a newly proposed visual pretraining benchmark, which contains 20 diverse downstream classification datasets, measuring the ability of pre-training models on both the prediction accuracy and their transfer efficiency in a new task. The pretraining is limited to three datasets: YFCC-15M , GCC3M +12M and ImageNet-21K (ImageNet-1K data is excluded). We pretrain our MaskCLIP on it and get the 1st result in the zero-shot track (we submit the results anonymously). As shown in Table 4, the 2nd2_{nd} team KLITE uses a strong Swin-B as the backbone and additional knowledge from GPT-3 and Wiktionary , and the 4th4_{th} use the strong Focal-B as the backbone, while our MaskCLIP greatly outperforms these methods with a simple ViT-B backbone and no additional knowledge.

Zero-shot on text-image retrieval. We further report the zero-shot text-image retrieval results on two benchmark datasets, Flicr30K and MS-COCO . We find that the text without any prefixes or suffixes works well for all the models. Table 5 shows the results. We can see that MaskCLIP exhibits a strong zero-shot performance. For example, with 25 epochs training, MaskCLIP reaches 41.4% Rank@1 image-to-text accuracy on MS-COCO, outperforming CLIP with 13.9%, and 25.5% Rank@1 text-to-image accuracy, +7.8% higher than CLIP.

4 Ablations

We compare our default settings with other alternatives to justify the efficacy of our model designs.

Training objectives ablation. As shown in Table.LABEL:tab:loss_ablation, when we remove the mask language modeling loss LMLM\mathcal{L}_{\text{MLM}}, the performance of the image-text task drops, including the zero-shot accuracy and retrial performance. While benefiting from the distillation loss, the finetuning performance on ImageNet-1K is not influenced. When we remove the distillation loss LDis\mathcal{L}_{\text{Dis}}, we observe a performance drop on all tasks, especially the finetuning results.

Distillation loss format. Different from previous methods that calculate the per-element distance as the loss function, we use an online tokenizer to map the feature to soft codewords and use the cross-entropy loss as the supervision. Here we study their difference in Table.LABEL:tab:ce_loss. We find that although they get similar fine-tuning performance, the CE loss gets better zero-shot and linear probing performance. The reason may be that the per-element MSE loss leads the model to fit some unnecessary details of the target feature, while the CE loss with soft tokenizer helps the model to focus more on the important feature.

Distillation & MLM loss weight. Here we set the loss weight of the CLIP branch as 1 and study the loss weight of the two additional branches. As shown in Table.LABEL:tab:dis_loss_weight and Table.LABEL:tab:mlm_loss_weight, setting λ=1\lambda=1 or β=1\beta=1 emphasize too much on new tasks, which mislead the model to a wrong converge direction, resulting in poor performance. When we reduce the loss weight by 10×10\times, the two additional tasks are helpful for the model and show a consistent gain on all the metrics. We suspect this is because the CLIP loss requires two different capabilities: understanding the input content and aligning them into a shared feature space. And the goal of the two additional self-supervised learning tasks is to facilitate understanding.

Image & Text decoder depth. Then we study the influence of the decoder depth for both image and text decoders. As shown in Table.LABEL:tab:image_decoder, we find the image decoder with only one layer works well, increasing the decoder depth leads to worse performance on all metrics. Similarly, Table.LABEL:tab:text_decoder shows that the text branch benefits from a shallow decoder design. We argue that a too-deep decoder would make the encoder lazy, relying on the strong decoder to resolve the challenging mask feature/language modeling tasks. And the different depth choice between the image and text branches is caused by the framework difference: the image branch sees the mask tokens at the decoder, while the text branch takes the mask tokens as the encoder input. Note that if we remove the text decoder, the performance gets worse. We think this is largely caused by the output conflict that the global recognition feature aggregation and local word prediction are conducted at the same layer.

Single-Stage v.s. two-Stage. Our MaskCLIP learns the VL contrastive and masked self-distillation simultaneously and jointly in a single stage. One possible variant is to first train CLIP and then use CLIP feature from the first stage to train masked image modeling as in . We report results on three datasets in Table 7. We can see that the second stage achieves better finetuning results compared with results from the stage one, showing the effectiveness of masked image modeling. Nonetheless, such two-stage training requires longer training time and loses the transfer capability in a zero-shot setting. In contrast, our MaskCLIP achieves superior results under all settings with fewer epochs.

Conclusion

We present MaskCLIP, a new VL pretraining framework that incorporates masked self-distillation into VL contrastive. We point out that masked self-distillation learns local semantics, fitting nicely to the VL contrastive that aims to learn global semantics, and this is supported with comprehensively designed experiments. We also utilize mask language modeling to enhance the text encoder which is critical for zero-shot performance. The resulting visual encoder shows strong transfer capability across widely adopted benchmarks for linear probing, fine-tuning, and also zero-shot evaluation.

References

Appendix A More Experiment

Comparison over small model and small dataset. As some baselines report ViT-B/32 instead of ViT-B/16, in order to compare, we further experiment MaskCLIP with a smaller model ViT-B/32 and report the zero-shot performance on ImageNet-1K. As shown in Table 8 left, our MaskCLIP outperforms the combination of two recent strong methods DeCLIP and FILIP . We also investigate the performance on a smaller dataset CC3M (we use ViT-B/16 here in coherency with previous experiments). Table 8(right part) shows that MaskCLIP achieves consistent gain.

Ablation on distillation loss. Here we further study the effectiveness of each component in the distillation loss. We start from CLIP+MAE and add three components of the distillation loss one by one. We find that 1) using the feature as the prediction target improves all metrics; 2) using EMA model gets better performance; 3) the MLM loss improves all the vision-language tasks.

Appendix B Experiment detail

Pre-training We train our proposed MaskCLIP from scratch and training for 25 epochs, the batch size is fixed to 4096 for all the experiments. We use 32 V100 for training with 128 samples per GPU. We use the AdamW optimizer with weight decay 0.1. The learning rate is set to 1e−31e^{-3} with one epoch warm-up and decay to 1e−51e^{-5} followed by a cosine schedule. The masks used in the mask self-distillation branch are random mask with a mask ratio of 75%. The EMA weight is set to 0.999 and linearly increases to 0.9999 during the training. We pretrain all the models with the commonly used YFCC15M dataset, which is flited from the YFCC100M dataset by .

For the ICinW academic track experiment, we pretrain the model with three datasets: YFCC-15M , GCC3M +12M and ImageNet-21K (ImageNet-1K data is excluded). Here we use the UniCL to utilize the ImageNet-22k dataset in the pretraining with a unified format. We train the model for 32 epochs and 16384 batch size, the rest settings are the same as the YFCC15M setting.

Zero-shot ImageNet-1K classification. For zero-shot on ImageNet-1K, we follow the prompt setting in to convert the labels to text features, which contains 7 prompt templates and we use the average feature as the final label feature. We calculate the similarity between image feature and all the label features to get its zero-shot classification result.

Linear-probing ImageNet-1K classification. For linear probing, we fix the backbone and train a new linear classifier for 90 epochs. Following the setting in MAE , we add a batch-norm layer without learnable affine parameters before the classifier to avoid adjusting the learning rate for each model. We set the batch size to 16384 and use the LARS optimizer with weight decay 0 and momentum 0.9. The learning rate is set to 6.4 and decays to 0 following the cosine schedule.

Fine-tuning ImageNet-1K classification. When fine-tuning on the ImageNet-1K dataset, we average pool the output of the last transformer of the encoder and feed it to a softmax-normalized classifier. We fine-tune 100 epochs for all the experiments, the learning rate is warmed up to 0.0006 for 20 epochs and decay to 1e−61e^{-6} following the cosine schedule. Similar to recent works, we also apply the layer decayed learning rate used in and we set the decay factor as 0.7. Note that we use the pure ViT architecture, without the techniques used in , such as layer scale and relative position embedding. The evaluation metric is top-1 validation accuracy of a single 224×224224\times 224 crop.

Zero-shot Semantic segmentation. Here we follow the setting in DenseCLIP based on the implementation from mmsegmentaion . For ADE20K and MS-COCO, we report the single-scale test result with 512×512512\times 512 input. For Pascal Context, we use 480×480480\times 480 input. To avoid the influence of position embedding caused by changing input size, we use sliding inference with 224×224224\times 224 input and stride 112112. To convert the labels to text embedding, we use 85 prompt templates and use the average feature as the final label feature.

ADE20K Semantic segmentation. Here we use: UperNet based on the implementation from mmsegmentaion . For UperNet, we follow the settings in and use AdamW optimizer with initial learning rate 2e−42e^{-4}, weight decay of 0.05 and batch size of 16 (8 GPUs with 2 images per GPU) for 160K iterations. The learning rate warmups with 1500 iterations at the beginning and decays with a linear decay strategy. We use the layer decay for the backbone and we set it as 0.6. As the ViT architecture outputs features with the same size, here we add four different scale FPNs to scale the feature map into different size. Specifically, we upsample the output feature of the 4th4th block 4×4\times, upsample the output feature of the 6th6th block 2×2\times, keep the output feature of the 8th8th block unchanged and downsample the output feature of the 12th12th block 2×2\times. We use the default augmentation setting in mmsegmentation including random horizontal flipping, random re-scaling (ratio range [0.5, 2.0]) and random photo-metric distortion. All the models are trained with input size 512×512512\times 512. The stochastic depth is set to 0.1. When it comes to testing, we report single-scale test result.

COCO Object Detection and Instance Segmentation. We use the classical object detection framework Mask R-CNN based on the implementation from mmdetection . We train it the 1×1\times schedule with single-scale input (image is resized so that the shorter side is 800 pixels, while the longer side does not exceed 1333 pixels) for 12 epochs. We use AdamW optimizer with a learning rate of 1e−41e^{-4}, weight decay of 0.05 and batch size of 16. We also use the layer decay for the backbone and we set it as 0.75. The learning rate declines at the 8th8th and 11th11th epoch with decay rate being 0.1. The stochastic depth is set to 0.1. Similar to the implementation of semantic segmentation above, we also use four different scale FPNs to scale the feature map into different size.

Appendix C More visualization results.

Here we provide more visualization results on the MS-COCO val set. In most cases, our MaskCLIP gets a better feature alignment performance between image and text.

Appendix D Societal impacts

MaskCLIP is an improvement of CLIP, so it has the same societal impacts of CLIP, including some malicious usages and positive applications. Meanwhile, CLIP and MaskCLIP may suffer from some unwanted data bias, as the data used for training are roughly collected from the Internet.