X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval

Yiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan, Ji Zhang, Rongrong Ji

Introduction

Video-text retrieval (VTR) is a multi-modal task, which aims to find the most relevant video/text based on the text/video query. With the explosive growth of videos on the Internet, VTR has attracted increasing interests and served as an important role in people’s daily life. Recent years have witnessed the rapid development of VTR, which is supported by a series of pre-training multi-modal models (Radford et al., 2021; Lei et al., 2021; Bain et al., 2021), innovative retrieval methods (Yu et al., 2018; Zhu and Yang, 2020; Lei et al., 2021; Liu et al., 2019; Gabeur et al., 2020; Dzabraev et al., 2021; Mithun et al., 2018; Zhang et al., 2018; Liu et al., 2021; Dong et al., 2019; Bertasius et al., 2021; Arnab et al., 2021; Luo et al., 2021; Jin et al., 2021; Yang et al., 2020; Wang et al., 2021) and video-text benchmarks (Caba Heilbron et al., 2015; Xu et al., 2016; Chen and Dolan, 2011; Rohrbach et al., 2015; Anne Hendricks et al., 2017).

Recently, with great success in large-scale contrastive language-image pre-training, VTR has also achieved great progress. Specifically, with 400M image-text pairs for training, CLIP (Radford et al., 2021) can embed the images and sentences into the shared semantic space for similarity calculation. Furthermore, CLIP4Clip (Luo et al., 2021) transfers the image-text knowledge of CLIP to the VTR task, resulting in significant performance improvements on several video-text retrieval datasets. However, CLIP and CLIP4Clip embed the whole sentence and image/video into textual and visual representations, thus lacking the ability to capture fine-grained interactions. To this end, some previous works (Yao et al., 2021; Lee et al., 2018) propose fine-grained contrastive frameworks, which consider the contrast between each word of the sentence and each frame of the video. Moreover, TACo (Yang et al., 2021) introduces token-level and sentence-level loss to consider both fine-grained and coarse-grained contrast. Although they have shown promising advances on the VTR task, cross-modality semantic contrast still needs to be systematically explored.

As shown in Fig. 1, a video is composed of multiple frames, and a sentence consists of several words. Video and sentence are usually redundant, which may contain some unnecessary frames or unimportant words. Concretely, given a specific video or sentence query, unnecessary frames or unimportant words refer to the candidates with low relevance to the query (i.e., light-colored frames and words in Fig. 1). However, most current works mainly focus on coarse-grained contrast (Radford et al., 2021; Luo et al., 2021), fine-grained contrast (Yao et al., 2021; Lee et al., 2018) or both (Yang et al., 2021), which are inefficient in filtering out these unnecessary frames and words. Specifically, coarse-grained contrast calculates the similarity between video-level and sentence-level features, and fine-grained contrast calculates the similarity between frame-level and word-level features. To this end, we ask: How to effectively filter out unnecessary information during retrieval? To answer this question, we propose the cross-grained contrast, which calculates the similarity score between the coarse-grained features and each fine-grained feature. As shown in Fig. 1, with the help of the coarse-grained feature, unimportant fine-grained features will be filtered out and important fine-grained features will be up-weighted. However, challenges in cross-grained contrast arise from aggregating similarity matrices to instance-level similarity scores. A naive and easy method is to use Mean-Max strategy (Yao et al., 2021; Khattab and Zaharia, 2020; Santhanam et al., 2021; Khattab et al., 2021) to calculate the instance-level similarity score after obtaining the similarity matrix. However, the conventional Mean-Max strategy is not conducive to filtering out the unnecessary information in videos and sentences during retrieval. On one hand, Mean applies the same weight to all frames and words, so the contrast between unnecessary frames and unimportant words may harm the retrieval performance. On the other hand, Max only considers the most important frame and word, ignoring other critical frames and words.

Based on the above analysis, in this paper, we propose an end-to-end multi-grained contrast model, namely X-CLIP, for video-text retrieval. Specifically, X-CLIP first adopts modality-specific encoders to generate multi-grained visual and textual representations and then considers multi-grained contrast of features (i.e., video-sentence, video-word, sentence-frame, and frame-word) to obtain multi-grained similarity scores, vectors, and matrices. To effectively filter out the unnecessary information and obtain meaningful instance-level similarity scores, the AOSM module of X-CLIP conducts the attention mechanism over the similarity vectors/matrices. Different from the conventional Mean-Max strategy, our proposed AOSM module dynamically considers the importance of each frame in the video and each word in the sentence, so the adverse effects of unimportant words and unnecessary frames on retrieval performance are reduced.

To validate the effectiveness of our proposed X-CLIP, we conduct extensive experiments on five widely-used video-text retrieval benchmarks and achieve significantly better performance than previous approaches. Specifically, our X-CLIP achieves 49.3 R@1 on MSR-VTT (i.e., 6.3% relative improvement, 2.9% absolute improvement over the previous state-of-the-art approach). Besides, our proposed X-CLIP achieves 50.4 R@1, 26.1 R@1, 47.8 R@1, 46.2 R@1 on the MSVD, LSMDC, DiDeMo and ActivityNet datasets, respectively, which outperforms the previous SOTA method by +6.6% (+3.1%), +11.1% (+2.6%), +6.7% (+3.0%), +3.8% (+1.7%) on relative (absolute) improvement.

Related Works

With the success of self-supervised pre-training such as BERT (Devlin et al., 2018) in NLP, vision-language pre-training on large-scale unlabeled cross-modal data has attracted growing attention (Lu et al., 2019; Xu et al., 2021; Tan and Bansal, 2019; Li et al., 2020b; Yu et al., 2020; Li et al., 2021; Radford et al., 2021; Jia et al., 2021; Sun et al., 2019b; Li et al., 2020a). One line of work such as LXMERT (Tan and Bansal, 2019), OSCAR (Li et al., 2020b) and ALBEF (Li et al., 2021) focuses on pre-training on enormous image-text pairs data, and obtains significant improvement in a variety of vision-and-language tasks. To better cope with the image-text retrieval tasks, contrastive language-image pre-training methods such as CLIP (Radford et al., 2021), ALIGN (Jia et al., 2021) and WenLan (Huo et al., 2021) have been proposed, by leveraging billion-scale image-text pairs data from the web with a dual-stream Transformer. Due to the great advantage of CLIP for visual representation learning, some recent work such as CLIP4Clip (Luo et al., 2021) has also begun to transfer the knowledge of CLIP to video-text retrieval tasks and obtained new state-of-the-art results. The other line of work such as VideoBERT (Sun et al., 2019b), HERO (Li et al., 2020a) and Frozen in Time (Bain et al., 2021) directly collects video-text pairs data for video-language pre-training, by further considering the temporal information in videos. However, the scale of the video-language pre-training dataset is much smaller than image-text pre-training since the process of video-text dataset collection is much more expensive. In this work, we follow the line of CLIP4Clip (Luo et al., 2021), which enhances video-text retrieval by borrowing the ability of visual representation learning from contrastive image-text pre-training. Different from CLIP4Clip (Luo et al., 2021), we design a multi-grained video-text alignment function to better align the video-text semantics.

2. Video-Text Retrieval

Video-text retrieval is a popular but challenging task, which involves cross-modal fusion of multiple modalities and additional understanding of temporal information in videos. Traditional video-text retrieval methods tend to design task-specific or modality-specific fusion strategies for cross-modal learning from offline extracted video and text features (Yu et al., 2017; Gabeur et al., 2020; He et al., 2021; Liu et al., 2021; Patrick et al., 2020; Jang et al., 2017; Le et al., 2020), including face recognition/object recognition/audio processing. However, they are limited by the pre-extracted single modal features, since these features are not properly learnt for the target downstream tasks. Recently, the paradigm of end-to-end video-text retrieval by training models directly from raw video/text has gained large popularity. For example, MIL-NCE (Miech et al., 2020) adopts Multiple Instance Learning and Noise Contrastive Estimation for end-to-end video representation learning, which addresses visually misaligned narrations from uncurated videos. ClipBERT (Lei et al., 2021) proposes to sparsely sample video clips for end-to-end training to obtain clip-level predictions, while Frozen in Time (Bain et al., 2021) uniformly samples video frames and conducts end-to-end training on both image-text and video-text pairs data. CLIP4Clip (Luo et al., 2021) transfers the knowledge of CLIP to end-to-end video-text retrieval and investigates three similarity calculation approaches for video-sentence contrastive learning. However, cross-grained (i.e., video-word and sentence-frame) contrast is also critical, which has rarely been explored in previous works. We propose the first work of multi-grained contrastive learning for end-to-end video-text retrieval, by considering all the video-sentence, video-word, sentence-frame, and frame-word contrasts.

3. Multi-Grained Contrastive Learning

Recently, contrastive learning (Chen et al., 2020b; He et al., 2020; Chen et al., 2020a; Chen et al., 2021) has been a popular topic in deep learning community. CLIP (Radford et al., 2021) implements the idea of contrastive learning based on a large number of image-text pairs, achieving outstanding performance on several multi-modal downstream tasks (Zhang et al., 2021; Ji et al., 2022; Ma et al., 2022; Zhu et al., 2022; Ji et al., 2021; He et al., 2022). To achieve fine-grained contrastive learning, FILIP (Yao et al., 2021) contrasts the patch in the image with the word in the sentence, achieving fine-grained semantic alignment. TACo (Yang et al., 2021) proposes token-level and sentence-level losses to include both fine-grained and coarse-grained contrasts. Although contrastive learning has been widely used in multi-modal pre-training, cross-grained contrast has rarely been explored in previous works, which is also critical for semantic alignment. Therefore, we propose a multi-grained contrastive learning method for video-text retrieval, which aims to achieve multi-grained semantic alignment.

METHODOLOGY

In this section, we elaborate each component of our proposed X-CLIP, whose architecture is shown in Fig. 2. Specifically, we first introduce how to extract the multi-grained visual and textual representations in Sec. 3.1. We then explain the multi-grained contrastive learning based on these feature representations in Sec. 3.2, which aims to obtain multi-grained contrast scores, vectors, and matrices. We also introduce how to aggregate the similarity vectors/matrices to the instance-level similarity score in Sec. 3.3. Finally, we describe the similarity calculation and objective function for video-text retrieval in Sec. 3.4 and 3.5, respectively.

For a video v^i∈V^\hat{v}_{i}\in\mathbf{\hat{V}}, we first sample video frames using the sampling rate of 1 frame per second (FPS). Frame encoder is used to process these frames to obtain frame-level features, which is a standard vision transformer (ViT) with 12 layers. Following the previous work (Luo et al., 2021), we initialize our frame encoder with the public CLIP (Radford et al., 2021) checkpoints. The architecture of ViT is the same as the transformer (Vaswani et al., 2017) encoder in natural language processing (NLP), except ViT introduces a visual tokenization process to convert video frames into discrete token sequences. The discrete token sequence, which is prepended with a [CLS] token, is then fed into the Transformer of ViT. The [CLS] tokens from the last layer are extracted as the frame-level features vˉ(i,j)∈Vˉi\bar{v}_{(i,j)}\in\mathbf{\bar{V}}_{i}.

1.2. Visual Representation

However, vˉ(i,j)∈Vˉi\bar{v}_{(i,j)}\in\mathbf{\bar{V}}_{i} are extracted from separate frames, without considering the interaction among frames. Therefore, we further propose a temporal encoder with temporal position embedding P\mathbf{P}, which is a set of predefined parameters, to model the temporal relationship. To be specific, the temporal encoder is also a standard transformer with 3 layers, which can be formulated as:

1.3. Textual Representation

2. Multi-Grained Contrastive Learning

Previous VTR works (Luo et al., 2021; Lee et al., 2018) focus on fine-grained and coarse-grained contrastive learning, which include video-sentence and frame-word contrasts. However, as explained in Sec. 1, cross-grained (i.e., video-word and sentence-frame) contrast is explicit to filter out the unnecessary information in the video and sentence. Therefore, different from previous works (Luo et al., 2021; Lee et al., 2018; Yao et al., 2021), which only focus single-grained contrast, X-CLIP is a multi-grained contrastive framework for VTR.

2.2. Video-Word Contrast

2.3. Sentence-Frame Contrast

2.4. Frame-Word Contrast

The fine-grained similarity matrix between word representations and frame representations can be also obtained using the matrix multiplication:

3. Attention Over Similarity Matrix (AOSM)

To obtain the instance-level similarity, we fuse the similarity vector/matrix in Eq. 4, Eq. 5 and Eq. 6. As discussed in Sec. 1, Mean-Max strategies (Yao et al., 2021; Khattab and Zaharia, 2020; Santhanam et al., 2021; Khattab et al., 2021) ignore the importance of different frames and words. To address this issue, we propose the Attention Over Similarity Matrix (AOSM) module, where scores in similarity vectors/matrices will be given different weights during aggregation.

where τ\tau is the temperature parameter of Softmax.

4. Similarity Calculation

The similarity score s(vi,tj)s(v_{i},t_{j}) measures the semantic similarity between the two instances. Different from the previous work (Luo et al., 2021) that only consider the coarse-grained contrast, our proposed X-CLIP adopt multi-grained contrast during retrieval. Therefore, the final similarity score s(vi,tj)s(v_{i},t_{j}) of X-CLIP contains multi-grained contrastive similarity scores, which can be represented as follows:

5. Objective Function

During training, given a batch of BB video-text pairs, the model will generate a B×BB\times B similarity matrix. We adopt the symmetric InfoNCE loss over the similarity matrix to optimize the retrieval model, which can be formulated as:

EXPERIMENTS

MSR-VTT (Xu et al., 2016) is a popular video-text retrieval dataset, which contains 10,000 videos and 200,000 captions. The length of videos in this dataset ranges from 10 to 32 seconds. In this paper, we adopt the widely-used ‘Training-9K’ split, where 9,000 videos and 180,000 captions are used for training and the rest are used for testing.

MSVD (Chen and Dolan, 2011) contains 1,970 videos, the duration of which vary from 1 to 62 seconds. Each video is annotated with 40 English captions. We use 1,200, 100, 670 videos for training, validating, and testing.

LSMDC (Rohrbach et al., 2015) is a dataset that contains 118,081 videos and captions. The duration of each video ranges from 2 to 30 seconds. We adopt 109,673, 7,408, and 1,000 videos for training, validating, and testing.

DiDeMo (Anne Hendricks et al., 2017) contains 10,000 videos and 40,000 captions. Following previous works (Liu et al., 2019; Lei et al., 2021; Bain et al., 2021), all captions of a video are concatenated together during video-paragraph retrieval.

ActivityNet (Caba Heilbron et al., 2015) contains 20,000 YouTube videos, which are annotated temporally. Following previous works (Luo et al., 2021; Sun et al., 2019a; Gabeur et al., 2020), all captions of a video are also concatenated together during video-paragraph retrieval for fair comparison.

2. Experimental Settings

We conduct the experiments on 4 NVIDIA Tesla V100 32GB GPUs using the PyTorch library. Following the previous work (Luo et al., 2021), the text encoder and frame encoder of X-CLIP are initialized by the public CLIP checkpoints. We use the Adam optimizer (Kingma and Ba, 2015) to optimize the X-CLIP and decay the learning rate using a cosine schedule strategy (Loshchilov and Hutter, 2016). Since the parameters of the text encoder and frame encoder are initialized from the public CLIP checkpoints, we adopt different learning rates for different modules. Specifically, the initial learning rate for text encoder and frame encoder is 1e-7, and the initial learning rate for other modules is 1e-4. We set the max token length, max frame length, batch size, and the training epoch to 32, 12, 300, and 3 for MSR-VTT, MSVD, and LSMDC datasets. Since videos and captions in DiDeMo and ActivityNet are longer and more complex, we set the max token length, max frame length, and the training epoch to 64, 64, and 20. Due to the limitation of GPU memory, we also reduce the batch size of DiDeMo and ActivityNet to 64. We conduct ablation, quantitative and qualitative experiments on the MSR-VTT dataset, it is more popular and competitive compared with other datasets. The base model of X-CLIP is ViT-B/32 if not specified. In order to enhance the expression ability of the model, we adopt linear embedding during calculating the video-sentence and frame-word similarity scores, which are initialized with the identity matrices. Besides, we also use the FC layers which are initialized with the identity matrices on similarity scores to enhance the modeling ability of the model.

2.2. Evaluation Protocols

To evaluate the retrieval performance of our proposed model, we use recall at Rank K (R@K, higher is better), median rank (MdR, lower is better), and mean rank (MnR, lower is better) as retrieval metrics, which are widely used in previous retrieval works (Yu et al., 2018; Zhu and Yang, 2020; Lei et al., 2021; Liu et al., 2019; Gabeur et al., 2020; Dzabraev et al., 2021; Mithun et al., 2018; Zhang et al., 2018; Liu et al., 2021; Dong et al., 2019; Bertasius et al., 2021; Arnab et al., 2021; Luo et al., 2021).

3. Performance Comparison

We compare X-CLIP against the previous works on MSR-VTT, MSVD, LSMDC, DiDeMo, and ActivityNet. X-CLIP achieves the SOTA results on all five datasets with significant improvements.

For the MSR-VTT dataset, the performance comparison is shown in Tab. 1. By analyzing the table, we gain the following observations:

Benefiting from the large-scale image-text pre-training, both CLIP4Clip and our model X-CLIP can obtain significant gains in performance compared with all the baselines. The consistent improvements verify that it is important to adopt end-to-end finetuning to realize the full potential of the image-text pre-trained model on video-text retrieval.

Compared with the strongest competitor (i.e., CLIP4Clip-seqTransf), X-CLIP obtains 49.3 R@1 (6.3% relative improvement, 2.9% absolute improvement) in the text-to-video retrieval task and 48.9 R@1 (7.7% relative improvement, 3.5% absolute improvement) in the video-to-text retrieval task by employing CLIP(ViT-B/16) as pre-trained model. This can be attributed to that our proposed cross-grained contrast and the AOSM module are critical to reducing the bad effects of unnecessary frames and unimportant words.

Compared to all the other state-of-the-arts, our model with ViT-B/16 achieves the best performance in all metrics. Surprisingly, our model with the ViT-B/32 can even achieve comparable performance to CLIP4Clip with ViT-B/16, which again demonstrates the effectiveness and superiority of multi-grained contrast and the AOSM module.

We also further validate the generalization of X-CLIP on MSVD, LSMDC, DiDeMo and ActivityNet in Tab. 2 - 5. It is worth noting that, in all variants of CLIP4Clip, we only report the performance of CLIP4Clip-MeanP and CLIP4Clip-seqTranf, because they perform better than the other two variants in consideration of experience in the previous work (Luo et al., 2021) and performance comparison in Tab. 1. By analyzing these tables, we can observe that X-CLIP also achieves significant improvement on these datasets for text-to-video and video-to-text retrieval tasks. Specifically, for the text-to-video retrieval task, X-CLIP outperforms the CLIP4Clip with ViT-B/16 on R@1 by +6.6% (+3.1%), +11.1% (+2.6%), +6.7% (+3.0%), +3.8% (+1.7%) relative (absolute) improvement on aforesaid four datasets respectively. For the video-to-text retrieval task, X-CLIP obtains +5.7% (+3.6%), +12.9% (+3.0%), +1.3% (+0.6%), +5.2% (+2.3%) relative (absolute) improvement on R@1. This demonstrates that our proposed X-CLIP can achieve consistent performance improvement on several video-text retrieval datasets. More experimental results are in the supplementary materials.

4. Ablation Study

To fully examine the impact of different contrastive modules, we conduct an ablation study to compare different variants of X-CLIP. As shown in Tab. 6, we gain two important observations:

With the number of contrastive modules increasing, the retrieval performance tends to be higher. When X-CLIP is equipped with all contrastive modules, the best retrieval performance can be achieved. This may be because each contrastive module plays a different role in the retrieval task and different contrast modules can promote each other to achieve better retrieval results.

Our proposed cross-grained contrast can assist fine-grained contrast or coarse-grained contrast to achieve better performance in the retrieval task. Specifically, X-CLIP with the sentence-video contrast module (i.e., Exp1) only achieves 43.0 R@1 in the text-to-video retrieval task. However, when X-CLIP is additionally equipped with cross-grained contrast modules (i.e., Exp8 and Exp9), the performance gets obvious absolute improvements of 2.4% and 1.0% respectively. Similarly, when X-CLIP is only equipped with fine-grained and coarse-grained contrast modules (i.e., Exp10), it achieves 44.8 R@1 in the text-to-video task. However, when it is additionally equipped with cross-grained contrast modules (i.e., Exp13 and Exp14), 1.0% and 0.7% absolute improvement of R@1 can be achieved. Therefore, we conclude that the performance improvement of cross-grained contrast modules in the retrieval task does not conflict with that of coarse-grained and fine-grained contrast modules.

To justify the effectiveness of the proposed AOSM module, we compare our method with the conventional Mean-Max and other variants (i.e., Max-Max, Max-Mean and Mean-Mean). As shown in Tab. 7, we observe that the Mean-Mean strategy performs worst. This may be because the Mean-Mean strategy, which applies the same weight to all similarity scores during aggregating, can not eliminate the adverse effects of unnecessary frames and unimportant words on the retrieval results. The Max-Mean, Mean-Max and Max-Max strategies perform better than the Mean-Mean strategy. This can be attributed to that these strategies adopt the highest similarity during aggregation, so contrast scores between unnecessary frames and unimportant words will be filtered out. However, since these strategies adopt the top-1 similarity score, some important similarity scores will also be ignored. To address this issue, we propose the AOSM module, where all similarity scores will be applied with different weights during aggregation. From Tab. 7, we observe that compared with other strategies, our proposed attention mechanism achieves better performance.

To explore the impact of the temporal encoder module in X-CLIP, we also conduct an ablative study to compare the X-CLIP with and without the temporal encoder. As shown in Tab 8, based on either ViT-B/32 or ViT/16, X-CLIP with temporal encoder consistently outperforms X-CLIP without temporal encoder. This may be because the temporal encoder is used to model the temporal relation of different frames in a video. Therefore, X-CLIP without temporal encoder can not understand and perceive the information that requires a combination of multiple frames, e.g., action. Based on the above analysis, we conclude that temporal modeling is also a key to improving the performance of retrieval tasks.

5. Effect of Temperature Parameter

To explore the effect of different τ\tau in the AOSM module, we also designed a group of experiments by setting different temperature parameters τ\tau in Softmax. From Tab. 9, we observe that the retrieval performance first improves before reaching the saturation point (i.e., τ=0.01\tau=0.01), and then begins to decline slightly. The main reason may be that when τ\tau is large, too many noisy similarity scores are considered. On the contrary, if the τ\tau is small, some important similarity scores may be ignored. Besides, our proposed attention mechanism with different τ\tau consistently performs better than the Mean-Mean strategy, and the attention mechanism with the optimal τ\tau outperforms other strategies in all evaluation protocols. This justifies that our proposed attention mechanism helps to strengthen the influence of important similarity scores and weaken the influence of noisy similarity scores, thus achieving better retrieval performance.

6. Qualitative Analysis

To qualitatively validate the effectiveness of our proposed X-CLIP, we show some typical video-to-text and text-to-video retrieval examples in Fig. 3 and Fig. 4, respectively. From these retrieval results, we find that X-CLIP could accurately understand the content of sentences and videos. Meanwhile, it is robust for X-CLIP to comprehend complex and similar sentences and videos, which is mainly attributed to the multi-grained contrast of our proposed model. To be specific, as shown in the first example in Fig.3, although the top-3 retrieved sentences are similar, our proposed X-CLIP can still choose the correct sentence by understanding the details of sentences and videos. Similarly, as shown in the first example in Fig.4, all top-3 retrieved videos describe the same cartoon, while “squid” does not appear in the second and third videos. Due to the multi-grained contrast, X-CLIP performs well in visual and textual content understanding, so it can retrieve the correct video.

Conclusion

In this paper, we present X-CLIP, a novel end-to-end multi-grained contrastive model for video-text retrieval, which first encodes the sentences and videos into coarse-grained and fine-grained representations, and conducts fine-grained, coarse-grained, and cross-grained contrasts over these representations. The multi-grained contrast and the AOSM module of X-CLIP help to reduce the negative effects of unnecessary frames and unimportant words during retrieval. Significant performance gains on five popular video-text retrieval datasets demonstrate the effectiveness and superiority of our proposed model.

References

Appendix

To verify the effectiveness of our method, we display the detailed comparison between our proposed X-CLIP and all variants of CLIP4Clip on different backbones (i.e., ViT-B/32 and ViT-B/16). As shown in Tab. 10 - Tab. 13, our proposed X-CLIP outperforms all variants of CLIP4Clip. Notably, X-CLIP with a weak backbone (i.e., ViT-B/32) even achieves comparable performance to CLIP4Clip with a strong backbone (i.e., ViT-B/16). This may be because our proposed cross-grained contrast is conducive to removing the noise information in the videos and sentences and capturing the important information. The outstanding performance again proves the importance and effectiveness of multi-grained contrast and the AOSM module.

2. Effect of training dataset size on contrastive modules

To gain deep insight into our four contrastive modules, we conduct the experiment to validate the X-CLIP with a single contrast module on the training datasets of different sizes. As illustrated in Fig. 5, when the training data is sufficient (i.e., 9k), the video-to-text and text-to-video retrieval performance of four variants is similar. When the size of the training dataset is reduced to 3k, the performance differences of different variants begin to appear and the word-frame contrastive module performs worse than other modules. Furthermore, when the size of the training dataset is reduced to 0.1k, other contrastive modules perform better than the word-frame contrastive module by a significant margin. The main reason can be that compared with other modules, the word-frame contrastive module is more complex, so it is difficult to optimize this module on a small amount of training data.

3. Effect of the AOSM module and Transformer modeling

To demonstrate the superiority and effectiveness of our proposed AOSM module, we also try to use a Transformer module to model the relationship of multi-grained features, which introduces more computation and parameters. The architecture of the new model is shown in Fig. 6. As shown in Tab. 14, our proposed AOSM module performs better than the Transf module. The performance gain can result from two aspects:

1) The new Transformer architecture introduces too many parameters, which makes it hard to be optimized with the limited amount of data. Our proposed AOSM is a well-designed module, where the importance of each frame and word is explicitly calculated. Thus, noise information in the video and sentence can be removed in X-CLIP. Besides, compared with Transformer, our proposed AOSM module contains fewer parameters, so it is easy to optimize the AOSM module.

2) The similarity scores of Transformer are obtained by a Linear layer, while the similarity scores of our proposed AOSM are obtained by the dot product. Notably, the dot product is the conventional approach for similarity calculation in the CLIP. However, Linear is a new approach, which does not carry any prior knowledge. Therefore, the prior knowledge of CLIP has little gain in Transformer, but our X-CLIP retains this prior knowledge well.