FILIP: Fine-grained Interactive Language-Image Pre-Training

Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, Chunjing Xu

Introduction

Large-scale Vision-Language Pre-training (VLP) models like CLIP (Radford et al., 2021) and ALIGN (Jia et al., 2021) have recently demonstrated success across various downstream tasks. They learn visual and textual representations from millions of image-text pairs collected from the Internet and show superior zero-shot ability and robustness. The core technique of these models lies in the global contrastive alignment of the images and texts through a dual-stream model. Such architecture is inference-efficient for downstream tasks like retrieval because the encoders for the two modalities can be decoupled and the image or text representations can be pre-computed offline. However, CLIP and ALIGN model the cross-modal interaction via solely the similarity of the global feature of each modality, lacking the ability of capturing finer-level information like the relationship between visual objects and textual words. In this paper, we develop a simple yet efficient cross-modal finer-grained interaction mechanism for large-scale VLP.

To achieve finer-grained cross-modal interaction, previous methods mainly exploited two kinds of methods. (1) One line of work (Chen et al., 2020; Li et al., 2020b; Dong et al., 2021; Li et al., 2021b; Zhang et al., 2021; Zhan et al., 2021) uses a pre-trained object detector to extract region-of-interest (ROI) features from images, and then fuses it with the paired text through a VLP model. This design complicates the pre-training due to pre-computing and storing a large number of ROI features. In addition, the zero-shot ability of these approaches is usually limited by the predefined number of classes and their performance is also restricted by the quality of the detector. (2) Another line of work (Li et al., 2021a; Kim et al., 2021) enforces the token-wise or patch-wise representations from both modalities into the same space and models these finer-grained interactions via cross-attention (Li et al., 2021a) or self-attention (Kim et al., 2021). However, these methods are usually less efficient in terms of both training and inference. In particular, during training, cross-attention in (Li et al., 2021a) requires to be performed in an encoder-decoder structure, while the complexity of the self-attention (Kim et al., 2021) grows quadratically with the length of the prolonged concatenated sequences of both modalities. During inference, the data from both modalities are intertwined to compute the cross-attention or self-attention, and can not be pre-computed offline as dual-stream models like CLIP and ALIGN. This can be less efficient for downstream tasks like image/text retrieval and image classification.

In this paper, we propose a large-scale Fine-grained Interactive Language-Image Pre-training framework named FILIP to address these limitations. Inspired by Khattab & Zaharia (2020), we model the fine-grained semantic alignment through a novel cross-modal late interaction mechanism in the contrastive loss, instead of using cross or self-attention. Specifically, our fine-grained contrastive learning uses a token-wise maximum similarity between visual and textual tokens to guide the contrastive objective. In this way, FILIP successfully leverages the finer-grained expressiveness among image patches and textual words while simultaneously gaining the ability to pre-compute image and text representations offline. Unlike Khattab & Zaharia (2020), we discard the padded tokens and use average instead summation of token-wise maximum similarities when computing the image-text alignment, which enhances the cross-modal representation learning and stabilizes training. Furthermore, we construct a large-scale pre-training dataset named FILIP300M from the Internet. Data cleaning and image-text data augmentation are also explored and proved useful in this work.

Extensive experiments show that by effectively learning fine-grained representations, FILIP achieves state-of-the-art performance on multiple downstream vision-language tasks, including zero-shot image classification and image-text retrieval. For example, FILIP reaches 77.1% top-1 accuracy for zero-shot ImageNet classification, surpassing CLIP with less training data. Visualizations on word-patch alignment further show that FILIP learns meaningful finer-grained features with promising localization ability.

Related Work

The pre-train-and-fine-tune scheme has achieved great success in the domains of natural language processing (Devlin et al., 2019; Brown et al., 2020) and computer vision (Dosovitskiy et al., 2020). It is then naturally extended to a joint cross-modal domain of Vision-and-Language Pre-training (VLP). The pre-training datasets of recent VLP models include publically available datasets like YFCC100M (Thomee et al., 2016) and CC12M (Changpinyo et al., 2021), as well as larger-scale datasets with more than 100M samples in CLIP (Radford et al., 2021) and ALIGN (Jia et al., 2021), which are shown to be even more powerful. The pre-training tasks of VLP models can be categorized into two categories: image-text contrastive learning task and Language Modeling (LM) based tasks: (i) CLIP (Radford et al., 2021), ALIGN (Jia et al., 2021) and UNIMO (Li et al., 2021b) make use of cross-modal contrastive learning which aligns the textual and visual information into a unified semantic space; (ii) VisualBERT (Li et al., 2019), UNITER (Chen et al., 2020), M6 (Lin et al., 2021), and DALL-E (Ramesh et al., 2021) employ LM-like objectives, including both masked LM (e.g., Masked Language/Region Modeling), and autoregressive LM (e.g., image captioning, text-grounded image generation). On the other hand, some methods rely on a pre-trained object detection model such as Faster-RCNN (Ren et al., 2015) to extract image regional features offline, which requires extra labeled bounding-box data and makes the approach less scalable. Recent efforts such as SOHO (Huang et al., 2021) and SimVLM (Wang et al., 2021) try to eliminate this burden via visual dictionary or PrefixLM (Raffel et al., 2020). In this paper, we directly learn fine-grained vision-language representations in an end-to-end and simpler manner while maintaining the benefit of inference efficiency.

The core of vision-language pre-training models lies in modeling the interaction between the two modalities. There are mainly two types of cross-modal interaction architectures: single-stream and dual-stream models. Single-stream models like VisualBERT (Li et al., 2019) and ViLT (Kim et al., 2021) directly concatenate the patch-wise or regional visual features and textual embeddings and feed them to the transformer-based model. Dual-stream models such as ViLBERT (Lu et al., 2019) and CLIP (Radford et al., 2021) have separate encoders for different modalities. This allows flexible use of different models for different modalities, and efficient inference for downstream tasks like image-text retrieval, through the ability of decoupling the encoders and pre-compute image/text features offline. In this paper, while following the dual-stream approach for its flexible and efficient inference, we further propose a new multi-modal interaction mechanism to capture the fine-grained representations.

Method

In this paper, we propose a new cross-modal pre-training model that excels in fine-grained interaction between image encoder and text encoder for mining more detailed semantic alignment, named as FILIP, as shown in Figure 1. Particularly, FILIP is a dual-stream model with Transformer-based image and text encoders. For the visual modality, the image encoder is a Vision Transformer (Dosovitskiy et al., 2020) which takes the concatenation of an extra [CLS] token embedding and linearly projected image patches as input. For the textual modality, following Radford et al. (2021), we use the lower-cased byte pair encoding (BPE) (Sennrich et al., 2016b) with a vocabulary size of 49,408 to tokenize the text. Each text sequence starts with [BOS] token and ends with [EOS] token. After the word embedding layer, the token embeddings are fed into a modified decoder-only Transformer model as in (Radford et al., 2019). On top of the image and text encoders, the representations of textual tokens and visual tokens are linearly projected to the multi-modal common space, and are separately L2-normalized. Different from existing dual-stream models (e.g., CLIP and ALIGN) which models cross-modal interaction via only the global features of the entire image and text sequence, we introduce a novel fine-grained contrastive learning objective equipped with cross-modal late interaction which takes into account the fine-grained interaction between image patches and textual tokens, detailed in Section 3.1.

Contrastive representation learning has recently been found to learn better representations than its predictive counterpart in both visual (Tian et al., 2020) and vision-language cross-modal pre-training (Radford et al., 2021). Under a general formulation of cross-modal contrastive learning (Radford et al., 2021), we want to learn encoders fθf_{\theta} for image data I\mathcal{I} and gϕg_{\phi} for text data T\mathcal{T} such that, given an image xI∈I{\bm{x}}^{I}\in{\mathcal{I}}, and a text xT∈T{\bm{x}}^{T}\in{\mathcal{T}}, the encoded representations fθ(xI)f_{\theta}({\bm{x}}^{I}) and gϕ(xT)g_{\phi}({\bm{x}}^{T}) are close if they are related and far apart if not, under a distance metric. In each training batch, we sample bb image-text pairs {xkI,xkT}k=1b\{{\bm{x}}^{I}_{k},{\bm{x}}^{T}_{k}\}_{k=1}^{b}. For image xkI{\bm{x}}^{I}_{k} in image-text pair {xkI,xkT}\{{\bm{x}}^{I}_{k},{\bm{x}}^{T}_{k}\}, xkT{\bm{x}}^{T}_{k} is its positive, while the other texts will be used as in-batch negatives. The image-to-text contrastive loss LkI{\mathcal{L}}^{I}_{k} for xkI{\bm{x}}^{I}_{k} can then be formulated as

where sk,jIs_{k,j}^{I} denotes the similarity of the kk-th image to the jj-th text. Similarly, the text-to-image contrastive loss for xkT{\bm{x}}^{T}_{k} is

The total loss of this mini-batch can be represented by

neglecting finer-grained interactions (e.g., word-patch alignment) between the two modalities. To alleviate this problem, while simultaneously maintain the training and inference efficiency of dual-stream models, we apply a cross-modal late interaction inspired by Khattab & Zaharia (2020) to model the token-wise cross-modal interaction.

as its token-wise maximum similarity with xjT{\bm{x}}^{T}_{j}. We then use the average token-wise maximum similarity of all non-padded tokens in the image (resp. text) as the similarity of an image to a text (resp. a text to an image). The similarity of the ii-th image to the jj-th text can thus be formulated as:

where mkI=arg⁡max⁡0≤r<n2[fθ(xiI)]k⊤[gϕ(xjT)]rm_{k}^{I}=\arg\max_{0\leq r<n_{2}}[f_{\theta}({\bm{x}}^{I}_{i})]_{k}^{\top}[g_{\phi}({\bm{x}}^{T}_{j})]_{r}. Similarly, the similarity of the jj-th text to the ii-th image is

where mkT=arg⁡max⁡0≤r<n1[fθ(xiI)]r⊤[gϕ(xjT)]km_{k}^{T}=\arg\max_{0\leq r<n_{1}}[f_{\theta}({\bm{x}}^{I}_{i})]_{r}^{\top}[g_{\phi}({\bm{x}}^{T}_{j})]_{k}. Note that si,jI(xiI,xjT)s_{i,j}^{I}({\bm{x}}^{I}_{i},{\bm{x}}^{T}_{j}) in Equation (4) does not necessarily equal si,jT(xiI,xjT)s_{i,j}^{T}({\bm{x}}^{I}_{i},{\bm{x}}^{T}_{j}) in Equation (5).

Intuitively, the token-wise maximum similarity in Equation (3) means that for each image patch, we find its most similar textual token. Similarly, for each textual token, we also find its closest image patch. By applying this to the similarity calculation in (4) and (5) for contrastive loss (1), the dual-stream model learns fine-grained alignment between image patches and textual tokens.

The original late interaction mechanism in (Khattab & Zaharia, 2020) computes the relevance score of a document to a query padded with mask tokens, as a sum of token-wise maximum similarities, and is optimized via a pairwise softmax cross-entropy loss. Though inspired from Khattab & Zaharia (2020), our proposed cross-modal late interaction differs in several aspects. Firstly, we exclude the padded textual tokens when computing the similarity, as they harm the performance. We speculate that this is because these padded tokens also learn textual representations and will mislead the model to align image patches to these meaningless padded tokens rather than meaningful non-padded words. Secondly, when computing similarities (4) and (5), we use the average of the token-wise maximum similarities instead of summation in (Khattab & Zaharia, 2020). This is because the number of non-padded tokens varies from text to text, and this summation over all non-padded tokens can have quite different magnitudes, leading to less stabilized training and worse final performance. Thirdly, we optimize the late interaction mechanism via a contrastive loss (1) which is found powerful vision-language pre-training (Radford et al., 2021) instead of the original pairwise loss in (Khattab & Zaharia, 2020).

Training Efficiency. Though the cross-modal late interaction is able to capture finer-grained features compared with the original loss, it relies on the token-wise representations of both modalities, and can be inefficient in terms of communication, memory and computation, especially when the batch size is large. To alleviate this problem, we utilize several methods. Firstly, we reduce the embedding size to 256. Besides, we reduce the precision of the last-layer features of both modalities from fp32 to fp16 before node communication in a distributed learning setting, and perform the multiplication in Equations (4) and (5) under the reduced precision. In addition, since the complexity of similarity calculation scales with the sequence length of textual tokens and image patches, for each image (resp. text), we select the 25% tokens with the highest token-wise maximum similarity score (Equation (3)) among all texts (resp. images) in the same local worker before node communication, based on the intuition that each sample can be represented by a few of the most representative tokens. Effects of these modifications are studied in Section 2.

1.2 Prompt Ensemble and Templates

Due to the problem of polysemy and inconsistency with the pre-training process, following Radford et al. (2021), we also use prompt templates to augment the original label for some downstream tasks. For visualizations, for simplicity, we use only one prompt template across the paper, i.e. “a photo of a {label}.” as Radford et al. (2021). For other experiments, we report results using prompt ensemble following Radford et al. (2021). When multiple prompts are allowed, the token-wise representations of different prompt templates for the same class label are different, and can not be summed together to form a mean textual representation as in (Radford et al., 2021). Thus, instead of ensembling different prompt templates by their mean textual representation, we ensemble them by their mean token-wise similarity. Specifically, suppose there are CC prompt templates, each label is augmented to CC different texts x1T,x2T,⋯ ,xCT{\bm{x}}^{T}_{1},{\bm{x}}^{T}_{2},\cdots,{\bm{x}}^{T}_{C}. The similarity between an image xI{\bm{x}}^{I} and this label is computed as 1C∑c=1Cs⋅,⋅I(xI,xcT),\frac{1}{C}\sum_{c=1}^{C}s_{\cdot,\cdot}^{I}({\bm{x}}^{I},{\bm{x}}^{T}_{c}), where s⋅,⋅Is_{\cdot,\cdot}^{I} is defined in Equation (4).

We use a unified rule-based method inspired by Radford et al. (2018) to construct prompt templates for image classification tasks. Specifically, each template consists of four components:

Here, the “[prefix]” is an in-context description like “a photo of a” similar as Radford et al. (2021); “label” is a class label of the dataset; “[category description]” describes the category which is found helpful for some fine-grained image classification datasets (Radford et al., 2021), e.g., “ a type of pet” for dataset Oxford-IIIT Pets. An interesting finding is that, adding a suffix that includes the reference word “it” (e.g., “I like it.”) at the end of the prompt empirically improves the zero-shot classification performance of the proposed model. We speculate this is because the reference word “it” strengthens the fine-grained cross-modal alignment, as it can also be aligned to image patches of the target object. Detailed prompt templates for different datasets can be found in Appendix A.4.

2 Image and Text Augmentation

To obtain better generalization and data-efficiency of the model, we perform data augmentation on both images and texts during the pre-training phase to construct more image-text pairs. We apply AutoAugment (Krizhevsky et al., 2012; Sato et al., 2015; Cubuk et al., 2019; Hoffer et al., 2020) for image augmentation, following the SOTA vision recognition methods (Touvron et al., 2021; Xie et al., 2020b). To ensure the augmented texts are semantically similar as the original one, for text augmentation, we rewrite the original text using back-translation (Xie et al., 2020a; Sennrich et al., 2016a). Specifically, the texts are first translated to the target language and then translated back to the source language. We choose German and Russian as the target language and get extra two texts for each image-text pair. When constructing a batch of image-text pairs during the pre-training, the text of each image-text pair is randomly sampled from the three candidate texts, i.e., the original text and two back-translated texts.

3 Pre-training Dataset

A sufficiently large image-text dataset is a prerequisite for vision-language pre-training. Recent CLIP (Radford et al., 2021) and ALIGN (Jia et al., 2021) construct datasets with 400M and 1800M image-text pairs, respectively. In this work, we also construct a large-scale dataset called FILIP300M, which consists of 300M image-text pairs and covers board vision and language concepts. Specifically, we collect image-text pairs from the Internet, and apply the following image- and text-based filtering rules to clean data. For image-based filtering, we remove the images whose shorter dimension is smaller than 200 pixels and the aspect ratio is larger than 3. For text-based filtering, we keep only English texts, and exclude the meaningless ones, e.g., img_0.jpg. We also discard image-text pairs whose texts are repeated for over 10 times. Besides, we also use 3 public datasets, including Conceptual Captions 3M (CC3M) (Sharma et al., 2018), Conceptual 12M (CC12M) (Changpinyo et al., 2021) and Yahoo Flickr Creative Commons 100M (YFCC100M) (Thomee et al., 2016). We apply the same filtering rules on YFCC100M. Finally, we use about 340M image-text pairs for pre-training. Despite using a smaller training dataset than CLIP and ALIGN, our models still outperform them in most down-steam tasks (see Section 4).

Experiments

Model Architectures. We train two models from scratch, i.e., FILIPbase\text{FILIP}_{\text{base}} and FILIPlarge\text{FILIP}_{\text{large}}. The model architectures follow CLIP (Radford et al., 2021), i.e., the image encoder is ViT-B/32 for FILIPbase\text{FILIP}_{\text{base}} and ViT-L/14 for FILIPlarge\text{FILIP}_{\text{large}}. More details can be found in Appendix A.2.

Pre-training Details. To save memory and scale up the batch size, automatic mixed-precision (Micikevicius et al., 2018) and gradient checkpoint (Griewank & Walther, 2000; Chen et al., 2016) are used The input images are resized to 224×224224\times 224 resolution during pre-training and the maximum length of the text is limited to 7777 tokens following Radford et al. (2021). The training is mainly conducted on Nvidia V100 GPUs and Ascend Cards. FILIPbase\text{FILIP}_{\text{base}} is trained on 128 cards about 9 days and FILIPlarge\text{FILIP}_{\text{large}} takes about 24 days to train on 192 cards. Unless otherwise specified, we use FILIPlarge\text{FILIP}_{\text{large}} to compare with other methods and FILIPbase\text{FILIP}_{\text{base}} for ablation. We train both models using the LAMB optimizer (You et al., 2020) and cosine learning rate schedule (Loshchilov & Hutter, 2016) with a linear warmup. Weight decay regularization is applied to all parameters except bias, layer normalization, token embedding, positional embedding and temperature in contrastive loss. Detailed values of hyperparameters for different datasets and models can be found in Appendix A.2.

2 Zero-Shot Image Classification

In this section, we evaluate our proposed FILIP on the zero-shot image classification task. We compare our FILIP with CLIP (Radford et al., 2021) on 12 downstream classification datasets, using the same evaluation setting as in CLIP. As described in Section 3.1.2, we apply a set of prompts for each dataset and ensemble them to get the final results, see Appendix A.4 for details. We only compare the zero-shot performance with CLIP here as ALIGN does not release its model and the related performances are not reported in their paper.

Table 1 shows the results on 12 datasets. Despite using less training data (340M vs. 400M), both FILIPbase\text{FILIP}_{\text{base}} and FILIPlarge\text{FILIP}_{\text{large}} considerably outperform their CLIP counterparts in terms of average top-1 accuracy over 12 datasets, i.e., achieving absolute improvements of 5.6% and 3.0%, respectively. In particular, our FILIP surpasses CLIP on ImageNet, the largest dataset among 12 datasets. FILIP also achieves substantial performance gains on some domain-specific datasets, e.g., for Aircrafts, the two FILIP models reach a 30% improvement over CLIP on average. We speculate this is because, unlike CLIP which aggregates the information of the whole image into the representation of the [CLS] token, our proposed FILIP model focuses more on the target object by directly aligning the image patches corresponding to the target object with the textual tokens corresponding to the class label (visualizations of word-patch alignment are in Section 4.5).

3 Image-Text Retrieval

Image-text retrieval consists of two sub-tasks: image-to-text retrieval and text-to-image retrieval. We evaluate our FILIP model on two retrieval benchmark datasets: Flickr30K (Plummer et al., 2015) and MSCOCO (Lin et al., 2014), under both zero-shot and fine-tuned settings. More details of experimental setting can be found in Appendix A.2.

Tables 2 and 3 show the results of zero-shot and fine-tuned image-text retrieval, respectively. We compare our FILIP model against methods with complex attention layers including Unicoder-VL (Li et al., 2020a), ImageBERT (Qi et al., 2020), UNITER (Chen et al., 2020), VILLA (Gan et al., 2020), ERNIE-ViL (Yu et al., 2021), Oscar (Li et al., 2020b), VinVL (Zhang et al., 2021), ALBEF (Li et al., 2021a), and methods trained on larger-scale image-text datasets including CLIP (Radford et al., 2021) and ALIGN (Jia et al., 2021). As we can see, FILIP achieves state-of-the-art performances under all metrics on both Flickr30K and MSCOCO datasets, except for zero-shot text-to-image retrieval on Flickr30K, where FILIP achieves competitive performance with SOTA. For zero-shot image-to-text retrieval on MSCOCO dataset, the absolute R@1 of our proposed FILIP is 2.7% higher than ALIGN, which is trained on a much larger dataset.

4 Ablation Study

Effectiveness of Each Component. We study the effectiveness of each component in FILIP, i.e., image/text augmentations and cross-modal late interaction. Experiments are conducted on FILIPbase\text{FILIP}_{\text{base}}, with a filtered subset of YFCC100M as the training dataset (as described in Section 3.3), on both zero-shot retrieval and classification tasks. We measure models’ performance on MSCOCO zero-shot image-text retrieval and ImageNet zero-shot classification, which are two effective indicators for the quality of the learned vision-language representations.

Table 4 reports the results. As can be seen, all three components are beneficial for both tasks. Despite the simple design, cross-modal late interaction brings significant performance improvements over the baseline (the vanilla CLIP ViT-B/32), with an absolute R@1 gain of 5.5% (resp. 3.8%) for image-to-text (resp. text-to-image) retrieval on MSCOCO and an absolute top-1 accuracy gain of 3.9% for zero-shot classification on ImageNet. Further improvements are observed when all components are combined together.

Efficiency Study of Cross-modal Late Interaction. Since the late interaction mechanism in Section 3.1.1 requires to calculate the similarity between all visual and textual tokens, its efficiency can be a problem when employed in large-scale distributed training. As described in Section 3.1.1, we make several attempts to address the issue. Table 5 shows the efficiency improvement on zero-shot classification on ImageNet when these attempts are applied. As can be seen, these attempts improve the efficiency of late interaction without accuracy drop. Combining all three attempts achieves only slightly slower training and larger memory consumption than the original loss in CLIP.

5 Visualization of Fine-grained Alignment

In this section, we visualize FILIP’s capability of capturing fine-grained cross-modal correspondence using the method of word-patch alignment. To make a fair comparison, we use our FILIPbase\text{FILIP}_{\text{base}} trained on YFCC100M and CLIP’s ViT-B/32, which are of the same size, for visualization. Each image is patchified to 7×77\times 7 image patches. More visualization results can be found in Appendix A.3.

Visualization Method. The word-patch alignment is performed based on the token-wise similarity between the image patches and textual tokens. Specifically, for the kk-th image patch, the location index of textual token with the largest similarity with it (mkIm_{k}^{I} in Equation (4)) is considered as its predicted label, and is placed at the center of it. Take class “balloon” as an example. There are 8 tokens in the tokenized textual sequence “[BOS] a photo of a balloon. [EOS]”, and the location index of the class label “balloon” is “5”. Note that one class label may be tokenized to more than one token. Location indices of textual tokens corresponding to the class label are highlighted in red, while the others are marked in white. A desired model that learns fine-grained representations would predict image patches of the target object to red indices.

Observations. Figure 2 shows the word-patch alignment results for FILIP and CLIP on 4 classes from the ImageNet dataset. As can be seen, FILIP exhibits the finer-grained understanding of an image in the following aspects. (i) A single object: From the visualization of class “small white butterfly”, the image patches covering the object are all classified correctly; (ii) Same object in different shapes: From the visualizations of class “balloon” and “lifeboat”, image patches corresponding to all target objects with different shapes and locations are correctly classified; (iii) Key Components of an object: For class “electric locomotive”, there are two key components crucial to correctly classifying the image, i.e., “electric” and “locomotive”, whose corresponding textual token indices are “5” and “6”, respectively. As can be seen, image patches matching these two key components are respectively correctly classified. On the other hand, CLIP can not correctly align image patches with corresponding textual tokens. Compared with Kim et al. (2021) which uses an extra optimal transport to align the textual word and image patch distributions, the word-patch alignment can be simply automatically learned by our method.

Conclusion and Future Work

This paper introduces FILIP, a simple yet generic framework towards fine-grained vision-language pre-training. By using a token-wise maximum similarity, our method learns fine-grained representation for patches in the images and words in the sentences. While it achieves competitive results against several large-scale multi-modal pre-training on various downstream tasks, both its architecture and training procedure can still be optimized to improve its performance. In the future, a more advanced image encoder as well as a well-designed interaction layer can be used to boost the performance. Furthermore, we can further add more masked language/image loss to support more generation tasks. To this end, we hope to extend FILIP as a generic and unified interface for solving a large variety of vision-language tasks.

References

Appendix A Appendix

Table 6 shows the number of image-text pairs of each datasets used in different pre-training methods.

A.2 Detailed Experimental Settings

We follow the same architecture design as CLIP, for both FILIPbase\text{FILIP}_{\text{base}} and FILIPlarge\text{FILIP}_{\text{large}}, except that we reduce the embedding dimension from 512/768 to 256 for the efficiency of loss computation. Table 7 describes the details of architectures.

For the implementation of the contrastive loss, following CLIP (Radford et al., 2021) and ALIGN (Jia et al., 2021), we also set the temperature in the softmax function to be a learnable parameter and initialize it as 0.07. For the pre-training, we use the LAMB optimizer implemented by the cybertronai’s open-source repository (https://github.com/cybertronai/pytorch-lamb). For the learning rate scheduler, we first assign a base learning rate and then linearly warm it up to the peak learning rate according to the effective total batch size by a square root strategy, peak_lr=base_lr×total_bs512peak\_lr=base\_lr\times\sqrt{\frac{total\_bs}{512}}. We note that a large weight decay is crucial to stabilize training and improve generalization. Specifically, we found that the training stability is a challenging issue when applying mix-precision training to large-scale models, i.e., the training is extremely unstable and the NaN loss easily happens. Recent works DALL-E (Ramesh et al., 2021) and Cogview (Ding et al., 2021) also notice this issue and provide their solutions. However, we found that simply increasing the weight decay and applying the trick of removing the weight decay of specific parameters as described in Section 4.1 work for our case. The base learning rate and weight decay are selected manually via observing the performance at the early training stage. Table 8 summarizes the common hyperparameters and Table 9 shows the model- and dataset-specific hyperparameters for FILIP pre-training.

Following previous works (Jia et al., 2021; Li et al., 2021a), for Flickr30K, we test on the 1K test set with or without fine-tuning on the 30K training set, while for MSCOCO, we test on the 5K test set with or without fine-tuning on the 113K training set. We use the similarity between image and text for ranking and use the contrastive loss for fine-tuning. Since there are multiple texts for each image in these two datasets, we change the ground-truth label of contrastive loss to consider multiple positives, by assigning a probability of 1/#positive to each positive following ALBEF (Li et al., 2021a). Besides, we also use prompts during evaluation for both datasets, see Appendix A.4 for details. Table 10 shows the hyperparameters for image-text retrieval fine-tuning.

A.3 More visualizations of Word-patch Alignment and Grad-cam Heatmaps

In Figure 3, we visualize the cross-modal alignment of the proposed method for more images, in terms of both word-patch alignment as described in Section 4.5 and Grad-CAM heatmaps (Selvaraju et al., 2017). We compute the Grad-CAM heatmaps based on the average self-attention maps over the image patches classified to targeted textual tokens (i.e., the textual token(s) corresponding to the class label in the ImageNet dataset) in the last layer of the image encoder. We average the heatmaps over all attention heads. As can be seen, our proposed model learns meaningful alignment between image patches and textual tokens.

A.4 Prompt Templates for downstream tasks

Table 11 shows the prompt templates for different image classification datasets in the form of “ [prefix] {label}, [category description]. [suffix]. ” in Equation (6). There are three components to be determined in the template, i.e., the prefix, the category description and the suffix. For each component, we select several well-performed ones for each dataset. Then we use the full combinations of all three components as the set of prompt templates for ensemble. For instance, we use 5 prefixes, no category descriptions, and 6 suffixes for dataset ImageNet. Then the total number of prompt templates for this dataset is: 5×1×6=305\times 1\times 6=30.

Following CLIP (Radford et al., 2021), we use prompt in zero-shot image-text retrieval for both Flickr30K and MSCOCO datasets. The prompt is selected by the same rule as described in Section 3.1.2, except that we do not use “[category description]” here. Table 12 shows the prompt templates for zero-shot image-text retrieval on Flickr30K and MSCOCO datasets.