HiT: Hierarchical Transformer with Momentum Contrast for Video-Text Retrieval

Song Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen, Wenkui Ding, Zhongyuan Wang

Introduction

Cross-modal Retrieval has attracted the increasing attention with the aim to search the semantic similar samples from different modalities. Specially, the explosive growth of video contents on the internet has brought great challenges to accurate video-text retrieval. In this paper, we focus on the learning of video-text retrieval and also hope to inspire other cross-modal tasks.

Recent works have shown that transformer can learn high level video representations, which capture semantically meaningful and temporally long-range structures for videos. Notably, existing approaches for cross-modal learning can be roughly categorized as two-stream, single-stream and dual stream architectures. Two-stream architecture, as shown in Figure 2-(a), utilizes a vision transformer and a text transformer to learn visual and textual representations independently, then introduces a multi-modal transformer to achieve cross-modal information exchange. Singe-stream architecture , as shown in Figure 2-(b), fuses visual and textual representations at the initial stage of the transformer model. However, these two architectures are not suitable for large-scale cross-modal retrieval tasks, due to the requirement of pairwise inputs and O(MN)\mathcal{O}(MN) time complexity of intra-model information exchange. Approaches with dual-stream architecture and our method, as shown in Figure 2-(c), have become a recent trend for cross-modal retrieval with better efficiency, requiring a time complexity of O(M+N)\mathcal{O}(M+N). In the line of dual-stream architecture, this paper proposes a novel transformer based method to achieve video-text retrieval, namely Hierarchical Transformer (HiT), where two contributions are jointly performed:

Hierarchical Cross-modal Contrastive Matching. According to the attention allocated characteristics of different layers in transformer architectures, the features in different layers focus on different views for samples . For example, the features in lower layers tend to encode more local contents with basic syntactic representations. Higher layers capture more complex semantics and usually produce higher-level semantic representations, as recent works performed. Based on these specialities, we propose Hierarchical Cross-modal Contrastive Matching to achieve multi-view and comprehensive video-text retrieval hierarchically, which is designed as Figure 1.

Momentum Cross-modal Contrast. Recently, a class of self-supervised methods for unsupervised visual representation learning emphasize the necessity of large-scale negative samples. Inspired by these works, we argue that large-scale negative sample interactions in the training process have been neglected in cross-modal contrastive learning. In this paper, we introduce MoCo into HiT to enable large-scale negative sample interactions on-the-fly. We name it as Momentum Cross-modal Contrast (MCC). In MCC, we build several memory banks to save a rich set of negative representations, which help broader negative sample interactions during training. However, if we utilize video and text encoders that are updated dramatically by gradient descent to generate representations for memory banks, it would result in the representation inconsistency in memory banks, thus largely affect the retrieval performance. Hence, key encoders for two modalities with momentum update (updated more smoothly) are required to maintain representation consistency.

Contributions: We propose Hierarchical Transformer (HiT) with Momentum Contrast for Video-Text Retrieval, which jointly performs Hierarchical Cross-modal Contrastive Matching and Momentum Cross-modal Contrast. Extensive experiments demonstrate the advantages of the proposed methods on three benchmarks, including MSR-VTT, ActivityNet and LSMDC.

Related Work

Video-Text Retrieval has received wide attention with the exploitation of the huge multimedia data and rich application scenarios. Several excellent works are introduced to address this task. JSFusion proposes a joint sequence fusion model for sequential interaction of videos and texts. Dual Encoding consists of mean pooling, biGRU and CNN models to encode sequential videos and texts in multiple levels. PVSE presents a polysemous instance embedding network to learn multiple and diverse representations of videos and texts for the polysemous problem. A graph-based framework is proposed in for matching between movie segments and synopsis paragraphs, which takes into account both the flow of events and the interactions among characters. HGR is a Hierarchical Graph Reasoning model, which decomposes video-text matching into global-to-local levels and disentangles texts into a hierarchical semantic graph with three levels of events, actions and entities.

2 Video-Text Learning with Transformer

BERT is a transformer based representation model for natural language process tasks. It evolves a line of works that learn a universal language encoder by pre-training with language modeling objectives. Recently, several attempts have been made which utilize BERTs and transformers as the backbones for cross-modal tasks. In video-text learning tasks, VideoBERT transforms a video into spoken words paired with a series of images and applies a transformer to learn joint representations. ActBERT learns a joint video-text representation that uncovers global and local visual clues from paired video sequences and text descriptions. Both the global and the local visual signals interact with the semantic stream mutually. MMT proposes the multi-modal transformer which processes features extracted from different modalities at different moments in videos, such as video, audio and speech. COOT proposes a hierarchical model that exploits long-range temporal context producing the video/text embeddings based on hierarchically interactions between local and global contexts. Support-set incorporates a auxiliary generative task, i.e., cross-captioning task, to alleviate mismatching problems existed in recent works. Very recently, T2VLAD uses a paradigm of global-local alignment to perform video retrieval. It obtains the global similarities by calculating the similarities multiple times between video-related and text features. For obtaining local similarities, they need to cluster the local features into several shared centers firstly, and calculate the similarities between local features and cluster centers. Though it also performs hierarchical matching as HiT, it performs their idea in a more complicated way.

3 Contrastive Learning

Contrastive Learning has made the remarkable progress in unsupervised visual representation learning. We introduce several representative contrastive learning mechanisms that benefit from the optimization with negative samples. End-to-end mechanism uses samples in the current mini-batch, where one can use its augmented views as positive samples and consider other samples in the current batch as negatives. Memory bank mechanism uses the representations sampled from a memory bank to conduct broader negative sample learning. However, the representations in the memory bank are from very different encoders all over the past epoch and they are less consistent. MoCo improves the memory bank mechanism by using a momentum-updated key encoder to generate the large-scale negative representations for the memory bank which can maintain better representations’ consistency. SimCLR shows that contrastive learning in unsupervised visual representation learning benefits from large batch size negatives, stronger data augmentation and introducing the learnable nonlinear transformation, i.e., using projection heads. Though recent works show that contrastive learning can achieve decent performance even without negatives by using a momentum encoder or stop gradient operation to prevent collapse solutions, our HiT in video-text retrieval and in visual representation learning indeed benefit from the large-scale negative sample learning. The effects of cross-modal learning without negatives are not involved in this paper.

Problem Definition

For the video-text retrieval task, we are given MM videos V={Vi}i=0M−1V=\{V_{i}\}_{i=0}^{M-1} and NN captions T={Ti}i=0N−1T=\{T_{i}\}_{i=0}^{N-1}. Each video has several kinds of expert embeddings to represent videos in multiple views, e.g., motion, appearance and audio. Each caption is represented by the natural language in English. Formally, the target of our methods for video-text retrieval is to obtain two query encoders ff: V→Z={Zi}i=1LV\rightarrow\textbf{Z}=\{Z_{i}\}_{i=1}^{L} and gg: T→Z={Zi}i=1LT\rightarrow\textbf{Z}=\{Z_{i}\}_{i=1}^{L} jointly, where ff and gg are for video and text domains respectively, and Z consists of LL common embedding spaces. In the common embedding spaces, cross-modal samples are represented by a series of compact embeddings. Meanwhile, the distance among similar cross-modal samples are smaller than that of among dissimilar cross-modal samples in the common embedding spaces. The constraint can be formulated as follows:

where d(⋅,⋅)d(\cdot,\cdot) is the distance measurement. The overall similarity between two cross-modal samples is decided by hierarchical contrastive matching results.

Hierarchical Transformer

Figure 3 illustrates the structure of the Hierarchical Transformer (HiT) for video-text retrieval. For video encoding, there are Query Video Encoder and Key Video Encoder. Both two video encoders utilize the same architecture. For text encoding, there are Query Text Encoder and Key Text Encoder that adopt the same architecture. Notably, Siamese encoders, a.k.a., key encoders, are shown for the utilization of Momentum Cross-modal Contrast (MCC), which will be discussed later. There are only two query encoders left if we remove MCC, as shown in Figure 1.

The video encoders, including query and key video encoders, are designed as transformer based architectures. We transform the raw visual features into a discrete sequence of tokens as inputs. To this end, we generate a sequence of pre-trained video-related features, including motion, appearance and audio features, to obtain Visual Embeddings Fv\textbf{F}_{v} as the inputs. Visual Segment Masks Mv\textbf{M}_{v} and Visual Position Embeddings Pv\textbf{P}_{v} are needed to indicate the real numbers and positions of input features respectively. We append Expert Embeddings E to identify the attending expert. The final visual input V can be formulated as follows, also shown in Figure 4:

2 Text Encoders

We leverage BERT-base-uncased as the text encoders and fine-tune it. It’s worth noting that the video features are generated by pre-trained deep neural networks and already have higher level semantic representation ability. While the text modality has different inherent complexity from the video modality and needs more transformer blocks to model semantic relations among words. Thus, text encoders are deeper than video encoders.

Each word in a caption will be embedded into a word embedding vector and we obtain Token Embeddings Ft\textbf{F}_{t}. [CLS] and [END] are embedded into the first and last positions. Text Segment Mask Mt\textbf{M}_{t} is needed to indicate the real length of the input sequence. Text Position Embedding Pt\textbf{P}_{t} is used to represent the word indexes of the input sequence in text encoders. The final input for text encoders is defined as:

3 Momentum Cross-modal Contrast

The end-to-end training mechanism as most methods implemented largely limits the negative sample interactions. To enable large-scale negative sample interactions for generating more precise and discriminative representations, Momentum Cross-modal Contrast (MCC) is proposed. Four memory banks are firstly built as queues for saving negative representations dynamically.

∙\bullet Text Memory Banks. Text memory banks, including BTwB_{T}^{w} for saving key text word-level features and BTsB_{T}^{s} for saving key text semantic-level features, are built as two queues. In each training iteration, the current mini-batch key text representations twkt_{w}^{k} and tskt_{s}^{k} encoded by the key text encoder will be enqueued into BTwB_{T}^{w} and BTsB_{T}^{s}, and the oldest mini-batch will be dequeued. The key text representations in BTwB_{T}^{w} and BTsB_{T}^{s} will be used to calculate the loss with the current mini-batch video representation vfqv_{f}^{q} and vsqv_{s}^{q} encoded by the query video encoder.

∙\bulletVideo Memory Banks. Similarly, video memory banks BVfB_{V}^{f} for saving key video feature-level features vfkv_{f}^{k}, and BVsB_{V}^{s} for saving key video semantic-level features vskv_{s}^{k} are built.

Moreover, to maintain the representation consistency in the memory banks, two key encoders, which perform momentum update , are required. We denote θqv\theta_{q}^{v} and θkv\theta_{k}^{v} as the parameters of the query and key video encoders. θqt\theta_{q}^{t} and θkt\theta_{k}^{t} are the parameters of the query and key text encoders. We formulate the momentum update for θkv\theta_{k}^{v} and θkt\theta_{k}^{t} as:

where m∈[0,1)m\in[0,1) is a momentum coefficient, which is a relatively large value. We set m=0.999m=0.999 in this paper. The parameters θqv\theta_{q}^{v} and θqt\theta_{q}^{t} are updated by back-propagation. The momentum update makes θkv\theta_{k}^{v} and θkt\theta_{k}^{t} evolve more smoothly than θqv\theta_{q}^{v} and θqt\theta_{q}^{t}. As a result, though the key representations in the memory banks are encoded by different encoders (in different mini-batches), the difference among these encoders will be small.

4 Hierarchical Cross-modal Contrastive Matching

We propose hierarchical cross-modal contrastive matching for video-text retrieval learning. Specifically, we utilize video feature-level features and text word-level features for feature-level contrastive matching. The video and text semantic-level features are used for semantic-level contrastive matching.

Feature-level Contrastive Matching. For the view of retrieving texts with videos, we get positive similarity svt+s^{vt+} by calculating cosine similarity between vfqv_{f}^{q} and twkt_{w}^{k}. Then, we obtain negative similarity Svt−={sivt−}i=1KtS_{vt-}=\{s_{i}^{vt-}\}_{i=1}^{K_{t}} by calculating cosine similarity among vfqv_{f}^{q} and all key text representations in BTwB_{T}^{w}. Thus, we achieve Svt={svt+}∪Svt−={sivt}i=11+KtS_{vt}=\{s^{vt+}\}\cup S_{vt-}=\{s^{vt}_{i}\}_{i=1}^{1+K_{t}}, where KtK_{t} is the queue size of BTwB_{T}^{w}. Similarly, for the view of retrieving videos with texts, we get Stv={stv+}∪Stv−={sitv}i=11+KvS_{tv}=\{s^{tv+}\}\cup S_{tv-}=\{s^{tv}_{i}\}_{i=1}^{1+K_{v}}, where KvK_{v} is the queue size of BVfB_{V}^{f}. The InfoNCE , a form of contrastive loss functions, is adopted as our objective function for feature-level contrastive matching:

where γ\gamma is a temperature hyper-parameter, which is set to 0.07 in this paper.

Semantic-level Contrastive Matching. Similarly, we achieve positive and negative similarity Cvt={cvt+}∪Cvt−={civt}i=11+KtC_{vt}=\{c^{vt+}\}\cup C_{vt-}=\{c^{vt}_{i}\}_{i=1}^{1+K_{t}} and Ctv={ctv+}∪Ctv−={citv}i=11+KvC_{tv}=\{c^{tv+}\}\cup C_{tv-}=\{c^{tv}_{i}\}_{i=1}^{1+K_{v}}. The objective function of semantic-level contrastive matching is defined as:

Thus, the overall objective function is L\mathcal{L}:

where α\alpha and β\beta are two hyper-parameters to balance two objectives. We set both α\alpha, β\beta to 1 in our experiments.

Experiments

We adopt video-text retrieval experiments on three datasets. Pre-training experiments are conducted on HowTo100M .

∙\bullet MSR-VTT contains 10,000 videos, where each video is annotated with 20 captions in English. We follow the training protocol defined in to evaluate on text-to-video and video-to-text retrieval tasks on the 1k-A testing split with 1,000 video or text candidates defined by .

∙\bullet ActivityNet Captions consists of 20K YouTube videos temporally annotated with sentence descriptions. We follow the approach of , where all the descriptions of a video are concatenated to form a paragraph. The training set has 10,009 videos. We evaluate our video-paragraph retrieval on the “val1” split (4,917 videos).

∙\bullet LSMDC contains 118,081 short video clips (∼\sim45s) extracted from 202 movies. Each clip is annotated with a caption, extracted from either the movie script or the audio description. The testing set is composed of 1,000 videos, from movies not present in the training set.

∙\bullet Metric. We measure the retrieval performance with common metrics in information retrieval, including Recall at K (R@K and K=1, 5, 10), and Median Rank (MedR). R@K is the percentage of test queries that at least one relevant item is found among the top-K retrieved results. The MedR measures the median rank of correct items in the retrieved ranking list, where lower score indicates a better model. We also take the sum of all R@K as rsum to reflect the overall retrieval performance.

2 Implementation Details

∙\bullet Pre-trained Features. We follow MMT to conduct pre-trained feature extraction. Motion features are extracted from S3D trained on the Kinetics action recognition dataset. Audio features are extracted from VGGish model trained on YT8M. Appearance features are extracted from the final global average pooling layer of SENet-154 trained on ImageNet.

For MSRVTT and LSMDC, we use all motion, appearance and audio experts. We employ 30 features for each type of visual features as the visual input, and the 25 first words from captions as the text input. For HowTo100M and ActivityNet, we use motion and audio experts, each of which has 100 features as the visual input, and the first 100 words as the text input.

∙\bullet Backbone. For text encoders, we use 12-layer BERT-base-uncased and fine-tune it. Video encoders have 4 transformer layers with 4 attention heads. The hidden size and the intermediate size are set to 512 and 3,072, respectively. We set the hidden size of projection heads to 8,192. DvD_{v} and DtD_{t} are both set to 2,048. The ReLU is used as the activation function and BN layers are used in hidden layers.

∙\bullet Optimization. The initial learning rate is set to 2e-5 and the network is optimized by AdamW optimizer. The 10% proportion of warm up and cosine decay are used for scheduling the learning rate. The batch size is 128 and we train 40 epochs. All experiments are conducted on NVIDIA 3090Ti GPUs.

∙\bullet KvK_{v} and KtK_{t} in MCC . For MSR-VTT, we report retrieval results when we set KvK_{v} and KtK_{t} to 4,096. KvK_{v} and KtK_{t} in ActivityNet are set to 512. In LSMDC, KvK_{v} and KtK_{t} are 1,024. We set KvK_{v} and KtK_{t} to 8,192 in HowTo100M. These numbers should vary with the batch size.

3 Compare to state of the art

The Table 1-3 present the retrieval results of HiT on MSR-VTT, ActivityNet Captions and LSMDC. We also compare HiT with other state-of-the-art methods.

As shown in the results, HiT outperforms all comparison methods by a clear margin. For MSR-VTT, we report video-to-text retrieval and text-to-video retrieval results. In particular, our retrieval performance at rsum is 320.3, exceeding recent state-of-the-art methods by a margin of 19.7. It well reflects the overall retrieval quality of HiT. With pre-training on HowTo100M, HiT further boosts the retrieval performance. For ActivityNet Captions and LSMDC, we report the retrieval performance in terms of text-to-video retrieval. HiT still outperforms comparison methods. We find that the growth of retrieval performance benefits from the proposed components, including Hierarchical Cross-modal Contrastive Matching and Momentum Cross-modal Contrast. To demonstrate the effectiveness and robustness of two components, we exhaustively and comprehensively ablate our method in the following sections.

Ablation Study

Hierarchical Cross-modal Matching. As mentioned above, we use token features from the first layers to perform Feature-level Contrastive Matching while token features from the last layers are adopted for Semantic-level Contrastive Matching. In this section, we design several variants to verify the impacts of Hierarchical Cross-modal Contrastive Matching. Note that we do not perform MCC for efficiency in this .

∙\bullet HiT-sl. We only implement semantic-level matching while feature-level matching is removed.

∙\bullet HiT-fl. Only feature-level matching is implemented.

∙\bullet HiT-4-level. To investigate the potential of hierarchical matching for transformer architectures, contrastive matching with respect to more levels is conducted. Since a text encoder has 12 transformer blocks and a video encoder has 4 blocks, except feature-level (between layer-1 in text encoder and layer-1 in video encoder) and semantic-level (between layer-12 in text encoder and layer-4 in video encoder), we append contrastive matching with more levels between layer-5 in text encoder and layer-2 in video encoder, layer-9 in text encoder and layer-3 in video encoder.

∙\bullet HiT-3-level-a. We append contrastive matching between layer-9 in text encoder and layer-3 in video encoder.

∙\bullet HiT-3-level-b. Contrastive matching between layer-5 in text encoder and layer-2 in video encoder is appended.

Table 5 presents the ablation results on MSR-VTT in text-to-video retrieval. We find that using more levels to conduct contrastive matching is able to obtain clear improvements. However, n-level matching requires n times retrieval during inference. In addition, significant improvements are not shown in 3-level and 4-level matching results. For the sake of retrieval efficiency and efficient training with Momentum Cross-modal Contrast, we select 2-level matching in this paper to report the main results.

Momentum Cross-modal Contrast. To explore the impacts of the memory bank size, sufficient experiments are conducted. The results are shown in Table 4. We vary the queue size of KvK_{v} and KtK_{t} from 0 to 8,192, and evaluate R@K and rsum. As shown in the results, it deserves attention that the introduction of large-scale negatives for similarity learning indeed achieves considerable performance improvements, in which we attribute it to broader negative sample interactions for obtaining more precise and discriminative representations. In addition, with the growth of queue size KvK_{v} and KtK_{t}, retrieval performance is slightly degraded after the growth which is probably due to some positive samples are misclassified as negative samples.

Momentum Encoders. For maintaining representation consistency in memory banks, we introduce two key encoders with momentum update for two modalities to generate representations. In this section, we abate two momentum encoders to explore their effectiveness in terms of maintaining representation consistency by evaluating the retrieval performance. We achieve the ablation by directly using query encoders to produce representations for memory banks. Table 6 presents the ablation results. We can find that it shows the performance degradation when we do not use momentum encoders. Particularly, it degrades performance at R@5 to 48.4%, which clearly demonstrates the necessity of momentum encoders.

Contrastive Loss. In Equation 5 and 6, InfoNCE is adopted as the Contrastive loss to perform common space learning. In this section, we use another commonly used loss function, i.e., Triplet Ranking Loss, as the objectives and present the retrieval performance for MSR-VTT in Table 7. Though it exists the difficulty in tuning the appropriate combination of temperature and batch size, we find that InfoNCE achieves better performance than Triplet Ranking Loss in HiT.

The temperature γ\gamma in InfoNCE is a sensitive parameter. To show how γ\gamma affects retrieval performance, the impacts of γ\gamma with regard to rsum are presented in Table 8. We can observe that the best performance can be achieved when we set γ\gamma to 0.07. A number with the same magnitude as 0.07 won’t change the performance obviously.

Expert Utilization. In MSR-VTT, we use three types of expert embeddings as the visual input, including motion, appearance and audio features. The ablation of the different experts are in Table 9.

From the results, we find that the motion expert achieves the best results when we only use one of three experts. Using audio features solely shows the worst performance. When using two experts, the combination of motion and audio experts achieves best results. As analysed in , we also note that audio features contribute the most when being used together with others, which indicates that they provide many complementary cues.

Feature Aggregation. As illustrated in Section 4.1 and 4.2, we leverage Average Pooling to produce aggregated features before projection heads, in the sense of capturing important features from all tokens. Alternately, we evaluate three more aggregation methods, including Max Pooling, 1D-CNN (kernel sizes: ) and using a [CLS] aggregated token. To obtain aggregated visual features from [CLS] token, similar to the text inputs, here we need to embed [CLS] and [END] tokens into the first and last positions of the visual input. We initialize them with random vectors. Table 10 presents comparison results in terms of text-video retrieval. Note that the decent results are not presented in [CLS]. We suppose the reason is that the features are not well aggregated in the [CLS] at feature-level.

Conclusion

We summarize our paper in two aspects: 1) In Hierarchical Cross-modal Contrastive Matching, we show that taking advantage of feature hierarchies in transformers can achieve decent performance gains. 2) Momentum Cross-modal Contrast demonstrates that cross-modal learning can benefit from large-scale negative sample learning. For future: work: 1) To facilitate the exploitation of feature hierarchies in transformers, we can design the fusion modules to utilize hierarchical features more effectively and efficiently. 2) To improve Momentum Cross-modal Contrast, some feature-level operations can be applied in memory banks, such as data mixing, hard negative selection, etc.

References