CenterCLIP: Token Clustering for Efficient Text-Video Retrieval

Shuai Zhao, Linchao Zhu, Xiaohan Wang, Yi Yang

Introduction

Text-video retrieval is less studied than the commonly known text-image retrieval task as the intricate context of the video, especially when the length of the video is very long or the temporal variation of the video is large. With the explosive growth of video content on mobile phones and Internet during the past decade, text-video retrieval becomes increasingly popular. People also desire a better text-video retrieval system as searching for videos of interest already becomes a part of daily lives of most people.

Recently, with the success of large-scale contrastive language-image pre-training methods like CLIP (Radford et al., 2021), text-video retrieval also has made great progress. To be specific, CLIP4clip (Luo et al., 2021) transfers the knowledge of CLIP to text-video retrieval tasks, surpassing the previous state-of-the-art methods by a large margin (e.g., more than 30% improvement of the recall metric on ActivityNet (Fabian Caba Heilbron and Niebles, 2015)). This demonstrates the power of billion-scale image-text pairs pre-training via contrastive learning. In CLIP, a vision transformer (Vaswani et al., 2017; Dosovitskiy et al., 2021) is adopted for visual representation learning. Typically, in vision transformer, visual tokenization, i.e., linear projection of non-overlapped image patches to an embedding space, is a necessary component to produce discrete visual token sequences. Then token sequences can be processed by the multi-head self-attention (MHSA) in transformer blocks as the same manner of dealing with text sequences in the original transformer (Vaswani et al., 2017).

When the input of the vision transformer becomes videos, the visual tokenization procedure produces many homogeneous tokens due to the redundancy nature in continuously changing frames. In Figure 1, we extract the token embedding of CLIP from different frames in the same video and visualize them by t-SNE (van der Maaten and Hinton, 2008). From the visualization, we can see those token embeddings from different frames form many tight clusters. Image patches with similar texture features correspond to immediate data points within a certain cluster. It is also clear that the number of clusters and the average number of tokens in clusters are not small, i.e., there are many similar token embedding in high-dimensional space. As a result, repeated computation of these homogeneous tokens in CLIP inevitably introduces a lot of unnecessary computation costs and hinders the training and deployment of video retrieval models in web applications. To resolve the above problem, in this work, we propose to distinguish the most representative tokens, i.e., the center token of each cluster in Figure 1, and only use these typical tokens for visual representation learning as these tokens contribute most to the discriminative feature representation learning.

We introduce a multi-segment token clustering algorithm to find the most representative tokens to reduce computation costs, and achieve segment-level semantic alignment of video and text representation. An input video is divided into multiple temporal segments. Each segment contains the same number of consecutive frames. Given the token embeddings of these frames, a clustering algorithm is performed on each segment independently. After clustering, only center tokens of clusters are reserved and non-center tokens are dropped to avoid duplicated computation of similar tokens. This significantly reduces computation costs. Center tokens from the same temporal segment are then concatenated into a new visual sequence and arranged according to their original spatial-temporal positions, i.e., tokens whose image patches occur earlier in the video would appear at the earlier position of the new visual sequence. Then the new visual sequence is processed by the standard transformer blocks. This enables the model to learn segment-level video representation via attention among tokens from within-segment frames. These segment-level video representations are aligned with the text through contrastive learning.

In this work, we introduce two instances of clustering algorithm in multi-segment token clustering. One is k-medoids equipped with a deterministic centroids initialization method, i.e., KKZ initialization (Katsavounidis et al., 1994; Su and Dy, 2007), to ensure the clustering results are consistent through multiple runs. A good initialization also helps the clustering algorithm converge fast. The other is spectral clustering which suits high-dimensional data points clustering. With our multi-segment token clustering algorithm, CenterCLIP achieves state-of-the-art performance on four common benchmarks: MSR-VTT (Xu et al., 2016), MSVD (Chen and Dolan, 2011), LSMDC (Rohrbach et al., 2015), and ActivityNet (Fabian Caba Heilbron and Niebles, 2015). We achieve significant improvement of retrieval metrics on all these four datasets compared to the baseline. At the same time, we achieve a decent reduction in memory cost and speed up the inference process. Specifically, on ActivityNet, we achieve a 35% reduction in memory cost and 14% speedup of inference speed compared to the baseline.

Related works

Contrastive Vision-Language Pre-Training. Since the success of derivative works of Contrastive Language-Image Pre-Training (CLIP) (Radford et al., 2021) in different areas (Luo et al., 2021; Patashnik et al., 2021; Bommasani et al., 2021; Shen et al., 2021; Xu et al., 2021), visual representation learning under text supervision attracts widespread attention. Huge models pre-trained on billion-scale image-text pairs from web like WenLan (Huo et al., 2021), Google’s ALIGN (Jia et al., 2021), and Microsoft’s Florence (Yuan et al., 2021) emerged. In the language-video understanding area, there are similar works like Frozen in Time (Bain et al., 2021) and HowTo100M (Miech et al., 2019). However, the scale of language-video pre-training is much smaller than language-image pre-training as the former is much more expensive. Following CLIP4clip (Luo et al., 2021), we transfer the knowledge of CLIP to the text-video retrieval task in this work.

Text-video Retrieval. Text-video retrieval is more complex than commonly studied text-image retrieval as the additional temporal dimension introduces complex context information. Previously, standard language-video learning methods tend to design dedicated fusion manners for cross-model learning from offline extracted video and text features (Yu et al., 2016, 2017; Le et al., 2020; Jang et al., 2017; Kaufman et al., 2017; Xu et al., 2017). Recently, the paradigm of end-to-end large-scale pre-training plus task-specific finetune becomes more and more popular for language-video understanding, e.g., HowTo100M (Miech et al., 2019), MIL-NCE (Miech et al., 2020), ActBERT (Zhu and Yang, 2020), VideoBERT (Sun et al., 2019), MMT (Gabeur et al., 2020), and HERO (Li et al., 2020). These methods achieve promising results on many language-video tasks and demonstrate the effectiveness of pre-training. Our work is also in this line, the difference is that we inherit the knowledge from CLIP (Radford et al., 2021), which is pre-trained on image-text pairs rather than video-text pairs.

Efficient Transformer. Recently, transformer becomes the unified model for many vision and text tasks (Vaswani et al., 2017; Radford et al., 2019; Dosovitskiy et al., 2021; Jaegle et al., 2021; Liu et al., 2021). However, there are many time-consuming operations in transformers such as self-attention and softmax operations. Some works try to reduce the complexity of self-attention for very long sequences or remove the softmax operation, e.g., Performer (Choromanski et al., 2020), Linear Transformer (Katharopoulos et al., 2020), Linformer (Wang et al., 2020), Reformer (Kitaev et al., 2020), Sparse Transformer (Child et al., 2019), Routing Transformer (Roy et al., 2021), Longformer (Beltagy et al., 2020), and Galerkin Transformer (Cao, 2021). Very recently, in computer vision, people also notice that not all tokens matter for the final performance of the model. To reduce the computation cost, researchers try to learn a few most representative visual tokens (Ryoo et al., 2021; Wu et al., 2020), learn to rank all tokens and select the most important ones (Wang et al., 2021a), and learn to mask the unimportant tokens (Yin et al., 2021; Rao et al., 2021). Compared to these mentioned methods, we are parameter-free and introduce segment level semantic alignment of text and video representation for text-video retrieval.

Methods

Given a video set V\mathcal{V} and text set T\mathcal{T}, the goal of text-video retrieval is to learn a score function f\mathit{f}, which gives a high similarity score f(vi,ti)\mathit{f}(v_{i},t_{i}) if a video vi∈Vv_{i}\in\mathcal{V} and a text ti∈Tt_{i}\in\mathcal{T} are highly relevant and a low similarity score for an irrelevant video-text pair. Then we can rank videos according to the query text (text to video retrieval) or rank texts according to the query video (video to text retrieval).

In training, a video-text pair (viv_{i}, tit_{i}) is treated as the positive if viv_{i} and tit_{i} are corresponded. All other instances of video or text in the mini-batch are treated as the negative. The text and video encoders are optimized in an end-to-end manner via normalized softmax loss (Zhai and Wu, 2019). The overall loss L\mathcal{L} is the average of video-to-text classification loss (Lv2t\mathcal{L}_{v2t}) and text-to-video classification loss (Lt2v\mathcal{L}_{t2v}):

where NN is the mini-batch size and τ\tau is the temperature to scale the logits. It is worth noting that τ\tau is crucial because both h(vi)\mathit{h}(v_{i}) and g(ti)\mathit{g}(t_{i}) are normalized. We set it as a trainable parameter following the CLIP model. During training, our model is initialized from the pre-trained weight of CLIP. We describe the details of the text encoder and video encoder below.

We instantiate the text encoder using the text model of CLIP. It is a transformer (Vaswani et al., 2017) with the architecture modifications described in BERT (Radford et al., 2019), i.e., only encoder and no decoder. A transformer model typically consists of repeated blocks (layers) of multi-head self-attention (MHSA) and feed-forward networks (FFN). We use a transformer with 12 layers and 512 width with 8 attention heads, where the width is the dimension of the query, key, and value feature. The text tokenizer is a lower-cased byte pair encoding (BPE) (Sennrich et al., 2016) with a 49 152 vocab size. The text sequence is padded with [SOS] and [EOS] tokens. [SOS] and [EOS] is padded at the beginning and end of the text sequence, respectively. The final text feature representation is the activation from the last layer of the transformer that corresponds to the [EOS] token. This text representation is later normalized by layer normalization and linearly projected into the joint video-text embedding space.

1.2. Video encoder

Our video encoder is a vision transformer (ViT), which first successfully applied transformers in vision tasks. The architecture of ViT is the same as the transformer in natural language processing, except ViT introduces an additional visual tokenization process to convert images into discrete sequences. When feeding images or videos into a ViT, we first convert the non-overlapped image patches into visual sequences, where a [CLASS] token is prepended to the beginning of sequences as BERT (Radford et al., 2019). Then the output of [CLASS] token at the final layer is extracted as the visual representation. In this work, we adopt a 2D linear projection to project image patches of different frames into an embedding space independently following the practice of CLIP4clip (Luo et al., 2021). For convenience, we name this linear transformation process as visual tokenization. Generally, we use a ViT-B/32 model (Dosovitskiy et al., 2021) with 12 layers and 512 width with 8 attention heads. ViT-B/32 means the non-overlapped input image patch size is 32×3232\times 32.

When applying the visual tokenization process to videos, it inevitably produces many redundant tokens as shown in Figure 1. Generally, an input video viv_{i} consists of many temporal related frames: vi={vi1,vi2,…,vi∣vi∣}v_{i}=\{v_{i}^{1},v_{i}^{2},\ldots,v_{i}^{|v_{i}|}\}, where ∣vi∣|v_{i}| is the number of frames in viv_{i}. After visual tokenization, if each frame produces LL tokens, the number of visual tokens is L∣vi∣L|v_{i}| (do not consider [CLASS] token). It shows that the number of visual tokens is linear to the number of tokens per frame (LL) and the video length. Given an input frame with a size of 224×224224\times 224, L=49L=49 for the ViT-B/32 model and L=196L=196 for the ViT-B/16 model. With a larger LL, the number of visual tokens becomes much larger. When performing text-video retrieval on long videos, the total number of tokens for a video is large. For example, videos in the ActivityNet (Fabian Caba Heilbron and Niebles, 2015) dataset usually have a few minutes duration. In this case, L∣vi∣L|v_{i}| will be easily larger than 1 000.

The redundant tokens considerably increase computation costs. To make the training and inference of text-video retrieval models more efficient, we propose to use clustering algorithms to find the most representative token embeddings. This process significantly reduces the number of tokens while maintaining the most valuable information of original tokens. After clustering, we only reserve the center tokens and remove other non-center tokens. The reserved tokens contain most of the information about the video and it is sufficient for text-video retrieval. We describe our multi-segment token clustering in the next section.

2. Multi-segment Token Clustering

The overall framework of our video encoder can be found in Figure 2. We perform a multi-segment clustering strategy on visual tokens from a certain temporal segment. This is based on the assumption that neighbor frames are more likely to be the same; then tokens of these similar neighbor frames are more possible to be redundant. Our multi-segment token clustering method empowers the model to achieve segment-level semantic alignment of text and video representations. Previously, CLIP and CLIP4clip adopt the average of frame features or the late fusion of frame features as the video representation. However, the former loses the temporal information, and the latter is a post-processing step and loses the detail temporal variations at the early stage of the transformer. By clustering across multiple frames within a segment at an early or middle stage, image patches from different temporal positions can interact with each other via the self-attention mechanism.

Specifically, a video sequence {vi1,vi2,…,vi∣vi∣}\{v_{i}^{1},v_{i}^{2},\ldots,v_{i}^{|v_{i}|}\} is divided into SS segments {si1,si2,…,siS}\{s_{i}^{1},s_{i}^{2},\ldots,s_{i}^{S}\}. Each segment contains ∣vi∣S\frac{|v_{i}|}{S} frames and L∣vi∣S\frac{L|v_{i}|}{S} tokens. Then we perform token clustering on these L∣vi∣S\frac{L|v_{i}|}{S} tokens segment-wise, namely, clustering for each segment independently. Then the centers of all clusters from one segment, i.e., center tokens, are selected and other non-center tokens are simply dropped. These center tokens are concatenated and arranged according to their original relative spatial-temporal position. Center tokens from the upper-left position and early frames are at the beginning of the new token sequence. Center tokens from the bottom-right position and late frames are at the rear-end of the new token sequence.

Multi-segment token clustering algorithm makes our vision model achieve segment-level temporal modeling and be able to capture the detailed temporal variation of video frames. This allows our methods to achieve segment-level alignment of the text tit_{i} and the video viv_{i} consisted of segments {si1,si2,…,siS}\{s_{i}^{1},s_{i}^{2},\ldots,s_{i}^{S}\}:

The multi-segment token clustering method has at least two advantages: (1) reducing computation costs by cutting down the number of tokens; (2) achieving segment-level semantic alignment of text and video representations via attention among tokens from different frames within the same temporal segment. As shown in Figure 2, assuming we perform token clustering right after the BB-th transformer block and the number of clusters is KK (ignore [CLASS] token after pooling), this means the length of the input sequence length of the following (12−B)(12-B) transformer blocks become KK. Generally, L∣vi∣S>>K\frac{L|v_{i}|}{S}>>K, obviously, computational costs are largely reduced. It is worth noting that the clustering module can be inserted at any place of ViT and the clustering procedure can be performed for any times. Clustering at an early stage reduces more computation costs. Next, we discuss two clustering methods used in the multi-segment token clustering algorithm.

For every ii, set pi≔arg min⁡j∥xi−μj∥22p_{i}\coloneqq\operatorname*{arg\,min}_{j}\lVert x_{i}-\mu_{j}\rVert_{2}^{2};

For every jj, set μj≔∑i=1m1{pi=j}xi∑i=1m1{pi=j}\mu_{j}\coloneqq\frac{\sum_{i=1}^{m}\mathbf{1}\{p_{i}=j\}x_{i}}{\sum_{i=1}^{m}\mathbf{1}\{p_{i}=j\}}; 1{⋅}\mathbf{1}\{\cdot\} equals to 11 if and only if the inner condition is true;

2.2. Spectral clustering

Construct similarity graph. Let WW be its weighted adjacency matrix, DD be the degree matrix;

Compute normalized Laplacian Lsym=D−12(D−W)D−12L_{sym}=D^{-\frac{1}{2}}(D-W)D^{-\frac{1}{2}};

Compute the first KK eigenvectors μ1,…,μk\mu_{1},\ldots,\mu_{k} of LsymL_{sym} which correspond to the first KK least eigenvalues;

Consider each row of UU as a new data point, apply k-means to these data points.

Experiments

Datatest. We validate our model on four datasets: MSR-VTT (Xu et al., 2016), MSVD (Chen and Dolan, 2011), LSMDC (Rohrbach et al., 2015), and ActivityNet (Fabian Caba Heilbron and Niebles, 2015). To save computational costs, the shorter side of videos are resized to 224 and the frame per second (fps) is set to 3. (a) MSR-VTT contains 10 000 videos with a length ranges from 10 ~ 32 seconds and 200 000 captions. We use two types of data splits, training-7K and training-9K, to compare with baselines. The training-7K follows the data splits from HowTo100M (Miech et al., 2019) and the training-9K follows the data splits from (Gabeur et al., 2020). The test data in both splits is ‘test 1k-A’, which contains 1 000 video-text pairs following JSFusion (Yu et al., 2018). If we do not specify, we use training-9K as the default. (b) MSVD contains 1 970 videos with a duration ranges from 1 ~ 62 seconds. Train, validation, and test splits contain 1 200, 100, and 670 videos, respectively. Each video has approximately 40 associated sentences in English. (c) LSMDC is comprised of 118 081 videos that ranges from 2~30 seconds. The videos were extracted from 202 movies. The validation set contains 7 408 videos. The 1 000 videos in the test set are from movies independent from the training and validation splits. (d) ActivityNet (Heilbron and Niebles, 2014; Fabian Caba Heilbron and Niebles, 2015) consists of 20 000 YouTube videos, and some of them are minutes long. We follow (Zhang et al., 2018; Gabeur et al., 2020) to concatenate all the descriptions of a video to form a paragraph and evaluate the model with video-paragraph retrieval on the val1 split.

Learning strategies. We apply warm up and cosine learning rate decay policy (Goyal et al., 2017; He et al., 2018). If the initial learning rate is lrlr and current epoch is epochepoch, for the first slow_epoch steps, the learning rate is lr×epochslow_epochlr\times\frac{\textit{epoch}}{\textit{slow\_epoch}}; for the rest epochs, the learning rate is 0.5×lr×(1+cos⁡(π×epoch−slow_epochmax_epoch−slow_epoch))0.5\times lr\times(1+\cos(\pi\times\frac{\textit{epoch}-\textit{slow\_epoch}}{\textit{max\_epoch}-\textit{slow\_epoch}})). Generally, lrlr is 1e-5 for ActivityNet and 5e-6 for other datasets; max_epoch is 8 for ActivityNet and 5 for other datasets; slow_epoch=0.1×max_epoch\textit{slow\_epoch}=0.1\times\textit{max\_epoch}. AdamW (Loshchilov and Hutter, 2018) optimizer is adopted with decoupled weight decay value 0.2.

Sequence length and batch size. For ActivityNet, the maximal text sequence length is 77 and the frame length is 60. For other datasets, the maximal text sequence length is 32 and the frame length is 12. The total batch size is always 128. Experiments with ViT-B/32 for ActivityNet are done on 8 NVIDIA Tesla V100 GPUs. Experiments with ViT-B/32 for other datasets need at least 2 RTX 3090 GPUs. All experiments are done with mixed precision (Micikevicius et al., 2018).

Frame sampling. We adopt a sparse sampling strategy following TSN (Wang et al., 2016). During training, video frames are divided into NinN_{in} segments, and we randomly sample one frame from each segment. During the evaluation, NinN_{in} frames are uniformly sampled from the video. As the above said, Nin=60N_{in}=60 for ActivityNet and Nin=12N_{in}=12 for the other datasets. These NinN_{in} frames are further divided into SS segments during token clustering.

Setting of CenterCLIP. We use (Ba−S,K)\bm{(B_{a}-S,K)} to represent the setting. It means we perform token clustering right after the aa-th transformer block, the number of temporal segments is SS, and the number of clusters/centers are constant KK. Generally, we construct KNN graph with Gaussian similarity function between two points when applying spectral clustering: exp⁡(−∥xi−xj∥2/(2σ2))\exp(-\lVert x_{i}-x_{j}\rVert^{2}/(2\sigma^{2})). The neighbours of one vertex is 5×5\times {the number of frames in a segment} for ViT-B/32, and plus an additional 5 for ViT-B/16. The variance of the Gaussian function σ\sigma is simply set to 2.0. No normalization is applied for token embeddings before performing clustering. Baselines in the experiments use the same setting as CenterCLIP.

2. Results on Common Benchmarks

As shown in Table 1, Table 2, Table 3, and Table 4. We achieve SOTA performance on all four datasets. Moreover, we also achieve decent memory usage reduction in all cases and obvious speedup of evaluation in some cases. Specifically, for CenterCLIP (ViT-B/32), we achieve a 32% reduction in memory cost and accelerate the model by 6% of the original speed for MSR-VTT, MSVD, and LSMDC in the best situation. For ActivityNet, the reduction in memory cost is 35% and the speedup of evaluation speed is 14% in the best case. These numbers verify the efficiency of our method. For CenterCLIP (ViT-B/16), as the patch size decreases, the number of tokens increases (4 ×\times as the number of tokens of ViT-B/32). In this work, the clustering complexity is at least linearly related to the number of data points. Therefore, CenterCLIP does not gain speedup for ViT-B/16. However, CenterCLIP also achieves a 32% reduction in memory cost. In future work, we will introduce faster clustering algorithms to speed up the whole model.

Compared to the baseline, CenterCLIP achieves significant improvement on recall. When using ViT-B/32, for MSVD, the maximal gain of text→\rightarrowvideo R@1 is 1.7%; for MSR-VTT (training-9K), the number is 1.2%; for LSMDC, it is 1.8%; for ActivityNet, it achieves 2.1% improvement of text→\rightarrowvideo R@1. When using ViT-B/16, for MSVD, the numbers are 1.0%, 2.8%, and 0.1% for text→\rightarrowvideo R@1, 5.7%, 5.2%, and 2.0% for video→\rightarrowtext R@1. CenterCLIP gains more improvement of video→\rightarrowtext retrieval performance in this case. All these results demonstrate the effectiveness of our clustering strategy. It aligns segment semantics of text and video.

It is worth noting that spectral clustering and k-medoids++ achieve similar performance in most cases. This is somehow counter-intuitive as spectral clustering should be more suitable for clustering high-dimensional points. This is possible because the data shape of clusters of token embeddings in high dimensional space is nearly spherical. Spectral clustering does achieve better performance in terms of some metrics, e.g., better R@5 and R@10 on MSR-VTT and ActivityNet, and produces the best video→\rightarrowtext results on MSVD.

3. Diagnostic Experiments

In this section, we will analyze CenterCLIP thoroughly. All diagnostic experiments are taken with CenterCLIP (ViT-B/32).

We provide four more strong baselines: 1) pooling of nearby tokens in a temporal segment, after pooling, we get KK average tokens for one segment; 2) sparse sampling of tokens in a temporal segment, namely, randomly sample KK tokens from a temporal segment during training and uniformly sample KK tokens during validation; 3) temporal shift described in TSM (Wang et al., 2016), here we apply temporal shift to the tokens except [CLASS] embedding; 4) token shift described in (Zhang et al., 2021), the method only shift the [CLASS] embedding. The shift is performed twice in each transformer block, right before MHSA and FFN. Results are shown in Table 5. Shifting all image patches does not work here. sparse sampling produces a little better results than baseline on MSR-VTT and LSMDC. However, CenterCLIP is much better than sparse sampling, this demonstrates the necessity of selecting representative tokens. Tokenshift achieves pretty good performance on short videos, nevertheless, it does not reduce any computational costs.

3.2. The place of performing token clustering

The influence of places of token clustering (k-medoids++) is shown in Figure 3. The smaller the BB, the lower the memory cost. The performance of the whole model will also decrease along with the decreasing or increasing of BB. A good trade-off between memory cost and performance achieves at B=6B=6, this is also our default setting.

We can also take multiple times of clustering. For instance, firstly perform clustering with (S=6,B=4)(S=6,B=4) and then with (S=3,B=8)(S=3,B=8). The results are shown in Table 6. Such progressive clustering strategy achieves pretty good R@5, R@10, MdR, and memory cost reduction. However, performing multiple times will increase the time complexity and this is not suitable for large amounts of tokens. Thus we generally perform clustering once in this work.

3.3. The number of cluster K𝐾K and segment S𝑆S

We perform experiments on LSMDC and ActivityNet. The results including R@1, memory cost, and inference time are shown in Figure 4. Along with the increase of KK, the performance increases, and computation costs also increase. At the same time, a small segment number SS does not always achieve better performance, e.g., S=1S=1 on LSMDC and S=6S=6 on ActivityNet. A small segment number means more tokens are dropped. This will cause the loss of more information. When S=1S=1 on LSMDC and S=6S=6 on ActivityNet, the number of tokens in a segment is large, i.e., 12×4912\times 49 and 10×4910\times 49, this leads to more computational costs of clustering as shown in Figure 4(c) and Figure 4(d). Thus a moderate segment number SS is usually adopted.

We change the number of input video frames and take experiments with CenterCLIP (B6−15,49B_{6}-{15},49) on ActivityNet. The results are shown in Table 7. The large the number of input frames NinN_{in}, the more computation costs, and a small number of frames will lead to worse performance. Similar ablations about the input of frames on short video datasets like MSR-VTT can be found in CLIP4clip (Luo et al., 2021). When the number of segments SS is fixed, a large number of input frames NinN_{in} will also increase computation costs of the clustering process as the number of tokens in one temporal segment increases.

3.5. Normalization of token embeddings

3.6. Learning rate and training epochs

The original CLIP4clip uses a learning rate of 1e-7. When setting lrlr = 1e-7 on MSR-VTT (training-7K), we get 39.7 T→\rightarrowV R@1 with mixed precision (Micikevicius et al., 2018) and 41.7 T→\rightarrowV R@1 without mixed precision on the split ‘test 1k-A’. The corresponding result of CLIP4clip baseline is 42.1 T→\rightarrowV R@1. When increasing the learning rate to 5e-6, we get 42.4 T→\rightarrowV R@1 with mixed precision on ‘test 1k-A’. As the mixed precision training saves a lot of GPU memory and accelerates the training procedure, we are stuck with mixed precision and use lr=lr= 5e-6 for short video datasets. For ActivityNet, we found a large learning rate with more training epochs brings a better result. This is shown in Table 9. It is possibly because of the large number of different video frames in long videos in ActivityNet and the sparse sampling strategy we used during training. The model needs more training epochs to learn good representations of videos.

3.7. Visualization of image patches of center tokens after clustering

We further display visualization results of center tokens after multi-segment clustering with different numbers of video frames within a temporal segment. The results are shown in Figure 5. It is clear that the clustering algorithm reserves the most representative tokens, for example, in the second and third row of Figure 5, tokens of the foreground animal are selected and only part of the tokens of the similar background remains. This verifies our beginning motivation that using a few typical tokens is already enough for learning discriminative features for video representation.

Conclusion

In this work, we propose a multi-segment clustering algorithm to reduce the number of redundant tokens of continuous video frames, and achieve segment-level alignment of video and text representations for text-video retrieval task. Our method, named CenterCLIP as we only reserve center tokens of token clusters and drop non-center tokens, is based on the knowledge of large-scale image-text pairs pre-trained model – CLIP. We take extensive experiments on four common text-video multi-modal datasets: MSR-VTT, MSVD, LSMDC, and ActivityNet. CenterCLIP achieves state-of-the-art performance on all these four datasets and surpass the old SOTA by a large margin. At the same time, CenterCLIP realizes a decent reduction in memory costs and speedup of inference time.

Acknowledgments

This work is supported by National Key R&D Program of China under Grant No. 2020AAA0108800. Thanks Naiyuan Liu for his helpful discussions.

References