MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image Pretraining
Xiaoyi Dong, Jianmin Bao, Yinglin Zheng, Ting Zhang, Dongdong Chen, Hao Yang, Ming Zeng, Weiming Zhang, Lu Yuan, Dong Chen, Fang Wen, Nenghai Yu
Introduction
Vision-language (VL) contrastive learning has shown remarkable success in pretraining for various tasks. With large-scale image-text pairs available on the Internet, the model composed of a simple dual encoder design learns strong semantic prior by aligning between image and text. The resulting visual encoder not only exhibits excellent linear probing and finetuning performance, but also enables impressive zero-shot performance with the guidance of the language encoder, showing the generality of natural language and its ability to supervise a wide range of visual concepts.
Nonetheless, the associated language description, though providing richer information than mere class labels, still can hardly describe all the information in the corresponding image, as images are continuous signals with fine-grained details and complex semantics. As a result, the VL contrastive by aligning global representations may only focus on the text-described objects and ignore the rest which might be useful for downstream tasks.
In this paper, we are interested in how to fully leverage the image itself to facilitate the VL contrastive to further improve the transfer capability. (1) Firstly, the learned feature representation shall characterize local patches, serving as a complementary for global representation in VL contrastive. Inspired by the recent success of masked image modeling in learning patch representations, we also randomly mask the input image with a large portion to force the visual encoder to focus on the remaining visible patches. (2) Secondly, the learned representation for local patches shall possess semantic meanings, being consistent with the global representation receiving semantic text supervision. We bring mean teacher self-distillation to supervise the learned patch representations with the visual feature representations, enabling implicit supervision from natural language. The resulting objective is denoted as masked self-distillation where the student model and the teacher model come from the same neural networks and the knowledge is distilled from the full image (fed to the teacher model) to the masked image (fed to student model). To this end, we introduce MaskCLIP by incorporating masked self-distillation into VL contrastive to advance the transferable visual encoder.
There are several recent attempts also exploring the capability of the visual encoder under natural language supervision. The common approach is to introduce contrastive learning or masked image modeling on the vision side together with contrastive language-image pretraining. However, the performance indeed improves based on CLIP but does not as well as our masked self-distillation. We argue that (1) the contrastive learning objective based on central crop augmentation actually learns global representations for salient objects while lack of attention on the surrounding backgrounds ; and (2) masked image modeling usually needs to remap the learned representation to pixels or discrete tokens . Such low-level prediction target is inefficient for semantic feature learning and thus also conflicts with high-level language supervision in VL contrastive. A brief illustration is presented in Figure 1. In the experiments, we conduct comprehensive ablations to analyze the difference and provide numerical and visual evidence for better understanding.
Symmetrically, we argue that local semantic supervision on the text branch is also helpful for the text encoder and eventually beneficial for zero-shot performance. So we introduce the same mask-data-modeling format supervision into the text branch as well. Different from images where the pixel is low-level signal, the words crafted by human beings are already highly semantic, so we use the tokenized word piece as the prediction target directly, following the well-studied mask language modeling method BERT. Meanwhile, to reduce the output conflicts between contrastive learning and mask language modeling, we introduce a small decoder for the mask language modeling branch.
We train our MaskCLIP on a subset of a publicly available image-text pairs dataset, YFCC , and thoroughly evaluate the transfer ability of visual representations on several vision benchmarks: ImageNet-1K for classification, ADE20K for semantic segmentation, MS-COCO for detection and segmentation, as well as a batch of other classification benchmarks. When it comes to ImageNet-1K classification, MaskCLIP achieves , , higher than CLIP for zero-shot transfer, linear probing, and finetuning respectively. For vision downstream tasks, we reach mIoU on ADE20K and APb, APm on MS-COCO . For vision-language tasks, MaskCLIP achieves average zero-shot accuracy on 20 datasets, and , rank@1 improvement on the Flickr30K image-test retrieval. In the recent Image Classification in the Wild challenge academic track, our MaskCLIP gets the result with TOP-1 average accuracy, surpassing the second team with .
In summary, the major contributions of this work are:
We present a novel vision-language pretraining framework MaskCLIP, by introducing masked self-distillation objective to facilitate VL contrastive for better transferable visual models.
We present extensive ablation studies on MaskCLIP variants and provide in-depth analysis numerically and visually to help understand how the proposed masked self-distillation assists VL contrastive.
We demonstrate our MaskCLIP on tens of benchmarks, showing the superiority under all three settings: zero-shot, linear probing, and finetuning.
Related Work
Vision-language pretraining Recent years have seen rapid progress made in vision-language pretraining . Several multiple cross-modality loss functions have been proposed for the training objective, such as image-text matching , masked language modeling , masked image modeling , contrastive loss . These objects are often mixed with each other to form a compound objective. While a variety of approaches have been proposed, few works investigate the performance on visual representation learning for image classification. Recently, CLIP and ALIGN show that the image-text contrastive learning objective achieves promising performance for visual representation learning. There are many following works proposed to further improve the pretraining performance, DeCLIP , SLIP , COTS , ViCHA , CYCLIP use additional uni/multi-modality supervision to improve the model capability, and PyramidCLIP , KLITE , IDEA seek to external knowledge from pre-trained models or datasets as the additional guidance. FILIP and LOUPE introduce fine-grained alignment to the model. Focusing on this research direction, we analyze the desired properties of supervision which could be complementary to CLIP, and propose the masked self-distillation objective incorporated with the image-text contrastive loss to further improve pretraining performance for various visual understanding tasks.
Self-supervised learning Self-supervised visual representation learning has attracted increasing attention over the past few years. The objective of the self-supervised learning is mainly divided into two categories: contrastive and generative . The contrastive methods, such as MOCO , SimCLR , BYOL , SimSiam , and DINO measure the similar and dissimilar samples by contrastive loss. Their success heavily depends on the strong data augmentation. The generative methods, such as BEiT , MAE , PeCo , BEVT , BootMAE and MaskFeat leverage masked image modeling to reconstruct the remaining masked part of its original input from the given visible parts. The generative methods show more promising transfer performance than the contrastive methods, as generative objective learns patch representations while contrastive objective focuses on learning centric global representations .
Self-knowledge distillation Self-knowledge distillation aims to distill the knowledge in a model itself and uses it for training the model. Instead of distilling knowledge from a pretrained teacher model , self-knowledge distillation regards a temporal ensemble of the student model as the teacher. It means that a student model becomes a teacher model itself, which gradually utilizes its own knowledge for softening the hard targets to be more informative during training. Self-knowledge distillation has been explored in semi-supervised learning , contrastive learning , self-supervised learning . In this paper, we use visual features supervised by natural language for guidance in masked self-distillation which naturally fit VL contrastive to learn more transferable visual representations.
MaskCLIP
We introduce MaskCLIP, a novel framework that learns visual representations. The core part of MaskCLIP is its backbone image encoder, denoted by as shown in Figure 1. It obtains the transferable capability during pretraining that could benefit downstream vision tasks. Following recent self-supervised approaches , we implement the backbone as a Vision Transformer (ViT) . The prediction results from given an input image then should be a collection of visual feature tokens, represented as
Here cls is short for class token. are the indexes of the non-class tokens.
The rest of this section starts with the utilization of language supervision. More shall be emphasized on the masked self-distillation, which we deem crucial for visual pretraining.
Following , we introduce a Transformer-based text encoder to leverage language knowledge. It aims to align the global feature representations of an image and a text with respect to some forms of similarity. Precisely, consider a given image-text pair , besides extracting the visual feature representation using the vision backbone as shown by Equation 1, we additionally use the text encoder to extract linguistic features from the text .
The mean feature of the two branches are regarded as the global representations and are fed into a projection head (implemented as a fully-connected layer) respectively to obtain the metric embeddings and . Image-text contrastive loss is employed to align them during pretraining. The loss can be formulated as , with
where stands for the number of image-text pairs within a training mini-batch, are indexes within the batch; stands for the temperature for the loss functions, which is learned together with all other parameters during training.
2 Masked Self-distillation for Visual Encoder
Knowledge distillation is a learning paradigm where a student model is trained to match the output of a given teacher model, so that the student model can be improved by the teacher. Instead of bringing in an external teacher, self-distillation methods such as proposes using a mean teacher model that is derived from the student itself. In specific, the teacher shares the same structure with the student, while the parameters of the teacher are exponential moving averages (EMA) of the parameters from the student. In the following, we would use the term “EMA model” to represent such mean teacher model constructed from the student.
MaskCLIP leverages the mean teacher self-distillation to enhance its vision representations. Let be the EMA model of the backbone encoder . and are the parameters of and at training step . is updated with
where is a hyper-parameter for smoothing updates. We propose to incorporate masked image modeling into self-distillation, resulting in masked self-distillation with asymmetric input for student model and teacher model.
In specific, considering a given input image , we first feed it to the EMA model (teacher model) to obtain the distillation targets. These target features can be represented as
In the meantime, we randomly mask a large portion of the input image patches and then feed it into the original backbone (student model). Following , we only feed the visible (unmasked) patches, denoted by , into the original backbone to speed up computation and save memory. Let be the indexes of all the masked tokens. These encoded features corresponding to visible tokens can then be denoted as . They are then joined with a shared and learnable feature vector, denoted as , that represents mask tokens, to form a complete set of features , with . We attach positional embeddings onto all these tokens, and append a small Transformer as a decoder to predict features of the masked region from the visible tokens, which could be formulated as
Inspired by , we use an online quantizer to transform the output features into a soft codewords distribution, and minimize the cross-entropy between the target features and the predicted features. Formally,
here the parameter of the teacher quantizer is also EMA updated by the online quantizer, similar to the teacher model.
3 Local Semantic Learning for Text Encoder
Besides the local semantic supervision for the visual encoder, we argue it is also helpful for the text encoder. So we introduce the BERT pretraining into the text branch. For the text , we denote the masked input as , where and , and be the indexes of all the masked text tokens. The output feature of the encoder is .
To reduce the output conflict between the global image-text contrastive learning and the local mask language modeling, we further introduce a small text decoder, which shares the same architecture as the encoder but with only a few layers. So that the global prediction and local prediction are conducted at different layers. We denote the output feature as: and the loss could be formulated as:
4 Overall Loss Functions
Finally, we pretrain MaskCLIP with all these losses combined:
with being the hyper-parameter weighting between VL contrastive loss and self-supervised learning loss. All the components of MaskCLIP are trained from scratch, including the visual backbone , the visual decoder , the text encoder , as well as the text decoder .
Experiments
Model architecture. Our framework consists of the visual encoder , the text encoder , the visual decoder , and the text decoder . We adopt the widely used Transformer ViT-B/16 for a fair comparison. It is composed of 12 layers, 768 width, and 12 head. The input image is resolution and is further split into patches with size . A learnable cls token is prepended to the 196 embeddings. For the text encoder, we adopt a 12-layer, 512-width, and 8-head Transformer following CLIP , and the text decoder has 4 layers. The number of text tokens is fixed to 77 with necessary truncations or paddings. For the image decoder, we directly use a one-layer Vision Transformer.
Pretraining details. We train our proposed MaskCLIP from scratch for 25 epochs, the batch size is fixed to 4096 for all the experiments. The masks used in the mask self-distillation branch and mask language modeling branch are random mask with a mask ratio of 75% and 20%. We pretrain all the models with the commonly used YFCC15M dataset, which is flited from the YFCC100M dataset by .
Downstream details. We evaluated MaskCLIP on several downstream datasets, including ImageNet-1K , ADE20K , MS-COCO , Flicr30K et al. For ImageNet-1K, we report zero-shot, linear probing, and finetuning performance. The zero-shot is conducted following the label prompt setting in SLIP . For linear probing, we fix the backbone and train a new linear classifier for 90 epochs. For finetuning, we follow the setting in BEiT and finetune the model for 100 epochs with a layer-decayed learning rate. See supplemental materials for more details.
2 Analysis
We first present our analysis by studying different ways of boosting CLIP. The baseline is CLIP trained on the YFCC-15M. Besides the introduced masked self-distillation, we consider two other popular methods: (1) SimCLR , a representative contrastive method; and (2) MAE the state-of-the-art masked image modeling approaches. All the compared methods are trained on the YFCC-15M for a fair comparison. We have the following observations.
Vision self-supervision helps VL contrastive. We evaluate the models on both vision task ImageNet-1K classification and vision-language task image-text retrieval on Flicker30K and present the comparison in Table 1. All the added vision self-supervision, regardless of contrastive or generative, improves the baseline CLIP. Among them, our proposed MaskCLIP achieves the best results in terms of all the evaluation metrics, outperforming CLIP with +6.9%, +7.2%, + 1.3% on ImageNet-1K classification for zero-shot, linear probing, and finetuning respectively, and +17.2%, +12.8% on Flicker30K for image-to-text retrieval and text-to-image retrieval. We also report the training GPU memory usage and time-consuming cost in Table 1. It is worth noting that the contrastive model (CLIP+SimCLR) compares two additional views of the input image, resulting in larger GPU memory usage and longer training time.
Masked image modeling is able to learn representations for local patches. We argue that the image encoder only pays attention to the text-described objects under VL contrastive due to sparse text description and to the centric objects under image contrastive due to central-crop augmentation. In contrast, masked image modeling forces the image encoder to focus on local patches using token-wise objectives by mandatorily masking a large portion of patches. Here, we provide numerical comparisons for evidence. We conduct an “Annotation-free zero-shot segmentation” experiment to test the zero-shot segmentation. The results on such a dense prediction task would better reveal the ability of local patch representations than global classification. Following the design in DenseCLIP , we use the prompted label feature as the linear classification weight to realize segmentation, without any training procedure. We evaluate the performance on two widely used datasets: ADE20K and Pascal Context . The results are shown in Table.2. We can see that equipped with masked image modeling, our MaskCLIP as well as CLIP+MAE achieves better results than CLIP and CLIP+SimCLR, validating our hypothesis.
Masked self-distillation learns semantic representations for local patches. Our masked self-distillation predicts visual features dynamically outputted by the visual encoder and thus implicitly gets supervision from the text side via VL contrastive. While MAE predicts fixed low-level pixels, making it inefficient to learn semantic representations (as the objective may force the representation to memorize low-level details) and thus causing conflict with VL contrastive. To show this, we select images from MS-COCO and calculate the feature similarity between image features and their corresponding caption features. We also select objects in the caption, prompt it to a new caption, such as “a photo of teddy bears”, and calculate the similarities. An example is shown in Figure 5 (More can be found in the supplementary material). Comparing MaskCLIP with CLIP+MAE in the fourth column, we can see that CLIP+MAE uses color as evidence and fails to distinguish the white teddy bear from the white snow. While our MaskCLIP successfully differentiates the two objects, suggesting ours learn more semantic features. On the other hand, the superior results of MaskCLIP shown in Table 1 and Table 2 also validate this. It is worth mentioning that CLIP and CLIP+SimCLR fail to have a correct response partition for different single objects like MaskCLIP, further justifying our second observation.
3 Comparison with Previous Methods
To show the effectiveness of MaskCLIP as a general vision-language pretrain method, we conduct experiments on both vision tasks and vision-language tasks. For vision tasks, we report results on ImageNet-1K classification, MS-COCO object detection, and ADE20K semantic segmentation. For vision-language tasks, we report zero-shot results on recent challenging ICinW 20 datasets benchmark and image-text retrieval results on Flickr30K and MS-COCO . In the following, we compare with the supervised baseline DeiT , self-supervised methods SimCLR and MAE , and vision-language methods CLIP and SLIP . For a fair comparison, we train SimCLR and MAE on YFCC-15M with the same epochs.
Classification on ImageNet-1K. As shown in Table 3, MaskCLIP benefits from the advantages of both VL pretraining and image mask self-distillation that shows strong performance on all the metrics. For zero-shot tasks, MaskCLIP outperforms CLIP by with 25 epoch training and achieves higher than the recent work SLIP. When it comes to finetune, MaskCLIP reaches top-1 accuracy, and outperforms CLIP by .
Semantic segmentation on ADE20K. Then we apply our MaskCLIP to the semantic segmentation task. Here we use the UperNet framework with input and end-to-end training for 160K iterations. The evaluation metric is the mean Intersection of Union (mIoU) and we report single-scale evaluation results here. The results are given in Table 3. Our method achieves 50.5 mIoU, mIoU than our baseline method CLIP, and mIoU than SLIP. This verifies the effectiveness of our introduced incorporation.
Object detection and instance segmentation on MS-COCO. We further investigate our transfer performance on object detection and instance segmentation in Table.3. Here we use Mask-RCNN framework with single-scale input and schedule (12 epochs). Our method achieves box AP and mask AP, better than CLIP, and better than SLIP.
Zero-shot on small datasets. We also report zero-shot performance on 20 small datasets under the ICinW setting (see the introduction below) in Table 4. We find that all the methods perform poorly on some datasets such as Aircraft(1% acc for random guessing, we omit the description in the following), Fer(24.7%), Country211(0.5%), GTSRB(5.9%), Cars(0.8%). This might be caused by the data domain gap that the YFCC-15M contains few related images and descriptions. For the rest of the datasets, all the methods get reasonable performance and our MaskCLIP gets the best performance on most datasets.
Image Classification in the Wild (ICinW) Challenge The ICinW challenge is a newly proposed visual pretraining benchmark, which contains 20 diverse downstream classification datasets, measuring the ability of pre-training models on both the prediction accuracy and their transfer efficiency in a new task. The pretraining is limited to three datasets: YFCC-15M , GCC3M +12M and ImageNet-21K (ImageNet-1K data is excluded). We pretrain our MaskCLIP on it and get the 1st result in the zero-shot track (we submit the results anonymously). As shown in Table 4, the team KLITE uses a strong Swin-B as the backbone and additional knowledge from GPT-3 and Wiktionary , and the use the strong Focal-B as the backbone, while our MaskCLIP greatly outperforms these methods with a simple ViT-B backbone and no additional knowledge.
Zero-shot on text-image retrieval. We further report the zero-shot text-image retrieval results on two benchmark datasets, Flicr30K and MS-COCO . We find that the text without any prefixes or suffixes works well for all the models. Table 5 shows the results. We can see that MaskCLIP exhibits a strong zero-shot performance. For example, with 25 epochs training, MaskCLIP reaches 41.4% Rank@1 image-to-text accuracy on MS-COCO, outperforming CLIP with 13.9%, and 25.5% Rank@1 text-to-image accuracy, +7.8% higher than CLIP.
4 Ablations
We compare our default settings with other alternatives to justify the efficacy of our model designs.
Training objectives ablation. As shown in Table.LABEL:tab:loss_ablation, when we remove the mask language modeling loss , the performance of the image-text task drops, including the zero-shot accuracy and retrial performance. While benefiting from the distillation loss, the finetuning performance on ImageNet-1K is not influenced. When we remove the distillation loss , we observe a performance drop on all tasks, especially the finetuning results.
Distillation loss format. Different from previous methods that calculate the per-element distance as the loss function, we use an online tokenizer to map the feature to soft codewords and use the cross-entropy loss as the supervision. Here we study their difference in Table.LABEL:tab:ce_loss. We find that although they get similar fine-tuning performance, the CE loss gets better zero-shot and linear probing performance. The reason may be that the per-element MSE loss leads the model to fit some unnecessary details of the target feature, while the CE loss with soft tokenizer helps the model to focus more on the important feature.
Distillation & MLM loss weight. Here we set the loss weight of the CLIP branch as 1 and study the loss weight of the two additional branches. As shown in Table.LABEL:tab:dis_loss_weight and Table.LABEL:tab:mlm_loss_weight, setting or emphasize too much on new tasks, which mislead the model to a wrong converge direction, resulting in poor performance. When we reduce the loss weight by , the two additional tasks are helpful for the model and show a consistent gain on all the metrics. We suspect this is because the CLIP loss requires two different capabilities: understanding the input content and aligning them into a shared feature space. And the goal of the two additional self-supervised learning tasks is to facilitate understanding.
Image & Text decoder depth. Then we study the influence of the decoder depth for both image and text decoders. As shown in Table.LABEL:tab:image_decoder, we find the image decoder with only one layer works well, increasing the decoder depth leads to worse performance on all metrics. Similarly, Table.LABEL:tab:text_decoder shows that the text branch benefits from a shallow decoder design. We argue that a too-deep decoder would make the encoder lazy, relying on the strong decoder to resolve the challenging mask feature/language modeling tasks. And the different depth choice between the image and text branches is caused by the framework difference: the image branch sees the mask tokens at the decoder, while the text branch takes the mask tokens as the encoder input. Note that if we remove the text decoder, the performance gets worse. We think this is largely caused by the output conflict that the global recognition feature aggregation and local word prediction are conducted at the same layer.
Single-Stage v.s. two-Stage. Our MaskCLIP learns the VL contrastive and masked self-distillation simultaneously and jointly in a single stage. One possible variant is to first train CLIP and then use CLIP feature from the first stage to train masked image modeling as in . We report results on three datasets in Table 7. We can see that the second stage achieves better finetuning results compared with results from the stage one, showing the effectiveness of masked image modeling. Nonetheless, such two-stage training requires longer training time and loses the transfer capability in a zero-shot setting. In contrast, our MaskCLIP achieves superior results under all settings with fewer epochs.
Conclusion
We present MaskCLIP, a new VL pretraining framework that incorporates masked self-distillation into VL contrastive. We point out that masked self-distillation learns local semantics, fitting nicely to the VL contrastive that aims to learn global semantics, and this is supported with comprehensively designed experiments. We also utilize mask language modeling to enhance the text encoder which is critical for zero-shot performance. The resulting visual encoder shows strong transfer capability across widely adopted benchmarks for linear probing, fine-tuning, and also zero-shot evaluation.
References
Appendix A More Experiment
Comparison over small model and small dataset. As some baselines report ViT-B/32 instead of ViT-B/16, in order to compare, we further experiment MaskCLIP with a smaller model ViT-B/32 and report the zero-shot performance on ImageNet-1K. As shown in Table 8 left, our MaskCLIP outperforms the combination of two recent strong methods DeCLIP and FILIP . We also investigate the performance on a smaller dataset CC3M (we use ViT-B/16 here in coherency with previous experiments). Table 8(right part) shows that MaskCLIP achieves consistent gain.
Ablation on distillation loss. Here we further study the effectiveness of each component in the distillation loss. We start from CLIP+MAE and add three components of the distillation loss one by one. We find that 1) using the feature as the prediction target improves all metrics; 2) using EMA model gets better performance; 3) the MLM loss improves all the vision-language tasks.
Appendix B Experiment detail
Pre-training We train our proposed MaskCLIP from scratch and training for 25 epochs, the batch size is fixed to 4096 for all the experiments. We use 32 V100 for training with 128 samples per GPU. We use the AdamW optimizer with weight decay 0.1. The learning rate is set to with one epoch warm-up and decay to followed by a cosine schedule. The masks used in the mask self-distillation branch are random mask with a mask ratio of 75%. The EMA weight is set to 0.999 and linearly increases to 0.9999 during the training. We pretrain all the models with the commonly used YFCC15M dataset, which is flited from the YFCC100M dataset by .
For the ICinW academic track experiment, we pretrain the model with three datasets: YFCC-15M , GCC3M +12M and ImageNet-21K (ImageNet-1K data is excluded). Here we use the UniCL to utilize the ImageNet-22k dataset in the pretraining with a unified format. We train the model for 32 epochs and 16384 batch size, the rest settings are the same as the YFCC15M setting.
Zero-shot ImageNet-1K classification. For zero-shot on ImageNet-1K, we follow the prompt setting in to convert the labels to text features, which contains 7 prompt templates and we use the average feature as the final label feature. We calculate the similarity between image feature and all the label features to get its zero-shot classification result.
Linear-probing ImageNet-1K classification. For linear probing, we fix the backbone and train a new linear classifier for 90 epochs. Following the setting in MAE , we add a batch-norm layer without learnable affine parameters before the classifier to avoid adjusting the learning rate for each model. We set the batch size to 16384 and use the LARS optimizer with weight decay 0 and momentum 0.9. The learning rate is set to 6.4 and decays to 0 following the cosine schedule.
Fine-tuning ImageNet-1K classification. When fine-tuning on the ImageNet-1K dataset, we average pool the output of the last transformer of the encoder and feed it to a softmax-normalized classifier. We fine-tune 100 epochs for all the experiments, the learning rate is warmed up to 0.0006 for 20 epochs and decay to following the cosine schedule. Similar to recent works, we also apply the layer decayed learning rate used in and we set the decay factor as 0.7. Note that we use the pure ViT architecture, without the techniques used in , such as layer scale and relative position embedding. The evaluation metric is top-1 validation accuracy of a single crop.
Zero-shot Semantic segmentation. Here we follow the setting in DenseCLIP based on the implementation from mmsegmentaion . For ADE20K and MS-COCO, we report the single-scale test result with input. For Pascal Context, we use input. To avoid the influence of position embedding caused by changing input size, we use sliding inference with input and stride . To convert the labels to text embedding, we use 85 prompt templates and use the average feature as the final label feature.
ADE20K Semantic segmentation. Here we use: UperNet based on the implementation from mmsegmentaion . For UperNet, we follow the settings in and use AdamW optimizer with initial learning rate , weight decay of 0.05 and batch size of 16 (8 GPUs with 2 images per GPU) for 160K iterations. The learning rate warmups with 1500 iterations at the beginning and decays with a linear decay strategy. We use the layer decay for the backbone and we set it as 0.6. As the ViT architecture outputs features with the same size, here we add four different scale FPNs to scale the feature map into different size. Specifically, we upsample the output feature of the block , upsample the output feature of the block , keep the output feature of the block unchanged and downsample the output feature of the block . We use the default augmentation setting in mmsegmentation including random horizontal flipping, random re-scaling (ratio range [0.5, 2.0]) and random photo-metric distortion. All the models are trained with input size . The stochastic depth is set to 0.1. When it comes to testing, we report single-scale test result.
COCO Object Detection and Instance Segmentation. We use the classical object detection framework Mask R-CNN based on the implementation from mmdetection . We train it the schedule with single-scale input (image is resized so that the shorter side is 800 pixels, while the longer side does not exceed 1333 pixels) for 12 epochs. We use AdamW optimizer with a learning rate of , weight decay of 0.05 and batch size of 16. We also use the layer decay for the backbone and we set it as 0.75. The learning rate declines at the and epoch with decay rate being 0.1. The stochastic depth is set to 0.1. Similar to the implementation of semantic segmentation above, we also use four different scale FPNs to scale the feature map into different size.
Appendix C More visualization results.
Here we provide more visualization results on the MS-COCO val set. In most cases, our MaskCLIP gets a better feature alignment performance between image and text.
Appendix D Societal impacts
MaskCLIP is an improvement of CLIP, so it has the same societal impacts of CLIP, including some malicious usages and positive applications. Meanwhile, CLIP and MaskCLIP may suffer from some unwanted data bias, as the data used for training are roughly collected from the Internet.