HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training
Qinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu, Qi Qian, Ji Zhang, Fei Huang
Introduction
Vision and language are two primary signals that constitute the real-world perception of humanity. With the success of image-language pre-training , video-language pre-training has recently received increasing attention. Large-scale video-language pre-training helps the model to learn effective multi-modal representation, which has shown significant improvement on a variety of video-language downstream tasks, such as video-text retrieval, video question answering and video captioning .
Inspired by the success of image-language pre-training paradigm, various methods have been proposed to adapt it to video-language pre-training. ClipBERT and Singularity directly build on representations from image encoders and aggregate them via score aggregation function and temporal encoder. Furthermore, MIL-NCE and Frozen switch image encoder to video encoder for spatio-temporal video representation learning and align the video with corresponding text. In addition, some advanced pre-training tasks are designed through modeling entity , reconstructing masked patches and predicting frame order . Despite their promising performance on downstream tasks, they treat video within global perspective illustrated in Figure 1(a), thus failing to consider fine-grained temporal information and relations which are essential to video-language pre-training.
Since untrimmed video contains various temporal details, directly treating the video globally has two main limitations: (1) Less effective in modeling the fine-grained moment information including atomic actions and moments. As illustrated in Figure 1(b), we vary time resolutions and generate two views (long & short) for the input video. As a result, the shot-view video clip tends to represent the moment information and the long-view video may express more event-level information. For example, the short-view video clip in Figure 1(b) only describes the moment of "lick fingers" rather than "eating ice cream". Such fine-grained moment information is hard to be captured by the long-view video under global event perspective; (2) Ignoring the temporal relations implicitly existed in the video. Knowing the event expressed by the text, the moment "eating ice cream" can be inferred from the moment "lick fingers" shown by short-view video. However, the implicit temporal relations between the moments are rarely explored in previous works.
To address these problems, we propose a Hierarchical Temporal-Aware video-language pre-training framework, HiTeA, for both multi-modal understanding and generation. Except for the standard pre-training tasks, HiTeA introduces two novel temporal-aware video-language pre-training tasks, named cross-modal moment exploration (CME) and multi-modal temporal relation exploration (MTRE), which not only model the fine-grained moments with partial cross-modal alignment but also capture temporal relations between multi-modal pairs hierarchically. Specifically, we first generate the long-view and short-view videos with different time resolutions to build hierarchy of the input video. Then, based on the similarities of words and short-view video, we select the most relevant words as positive and leave the rest of the words as hard negatives. The CME pre-training task is applied to align the positive words and shot-view video representations in the same embedding space. Moreover, to capture association between moments and the event, we match different views for the same video. However, directly matching two views visually would be noisy due to the background similarity . To this end, we perform multi-modal alignment between video-text pairs via the MTRE pre-training task. More specifically, the shot-view video guided by most relevant words and the long-view video guided by text will be aligned. Empowered by above two novel temporal-aware video-language pre-training tasks, HiTeA captures both fine-grained moment information and temporal relations between different views of video.
In spite of a good performance, recent studies reveal most video-language downstream datasets are biased towards still objects, scenes, etc., while the temporal dynamics are negligible. To evaluate the temporal performance of the video-language pre-training model and temporal reliance of downstream datasets, we introduce temporal shuffling test for these datasets. This enables a more comprehensive evaluation of temporal modeling capability in the video-language pre-training field. Besides, our method achieves significant improvement on the datasets with heavy temporal reliance.
In summary, our key contributions are the followings:
We propose a novel hierarchical temporal-aware video-language pre-training framework with both video-language understanding and generation capabilities.
We introduce additional temporal-aware pre-training tasks by performing cross-modal and multi-modal alignment hierarchically, which not only model moment information with fine-grained semantics but also capture temporal relations between moments and event.
Extensive experiments demonstrate the effectiveness of HiTeA, and it achieves state-of-the-art performance on 15 video-language downstream datasets including video-text retrieval, video question answering, and video captioning, especially on temporal-oriented datasets (e.g., SSv2-Template and SSv2-Label) with 8.6% and 11.1% improvement respectively.
Related Work
Benefiting from a large number of image/video-text pairs, video-language pre-training (VLP) exhibits superior capabilities on various video-text benchmarks. The method of VLP is constantly evolving. Traditional approaches leverage offline-extracted dense video features for pre-training to circumvent the expensive computation overhead. In contrast, ClipBERT suggests that sparse sampling can enable affordable end-to-end learning and improve performance simultaneously. Recent emerging approaches adopt this strategy and propose new model architectures and pre-training tasks. Frozen trains jointly on image and video datasets via video-text contrastive learning (VTC). ALPRO proposes a new visually-grounded pre-training task combined with VTC, video-text matching (VTM) and masked language modeling (MLM) to learn fine-grained region-entity alignment. LAVENDER formulates all pre-training and downstream tasks as MLM so that a unified architecture can used for all video-text tasks. Apart from above representative works, frame order modeling (FOM) and masked video modeling (MVM) are designed for VLP. However, the temporal characteristic of video still remains largely unexplored. To this end, we introduce a novel hierarchical temporal-aware VLP framework which not only models the fine-grained moment information but also captures their correlations with different temporal granularities.
The temporal characteristic acts as a vital role in VLP since it provides the model with the capabilities of reasoning and understanding causality. Previous efforts in this field can be roughly divided into three categories. First, several methods directly transfer image-text models to video-text tasks by simply concatenating video frame or building a additional temporal encoder . Second, some works switch the image encoder to video encoder for learning spatio-temporal contexts within videos. Third, HERO and MERLOT design FOM task to explicitly recover the correct temporal order of shuffled frames. Nonetheless, ATP and Singularity reveal the existence of a static appearance bias in popular video-language datasets, and they develop single-frame models to achieve surprisingly strong performance, comparable or even better than above methods with explicit temporal modeling. Therefore, they recommend SSv2 and NExT-QA datasets to test the temporal ability of VLP models. Different from previous approaches, we vary the temporal resolutions and generate two views of video so as to construct the temporal hierarchy, which equips the model with the ability to learn both fine-grained moment information and temporal relations at the same time.
Method
Figure 2 sketches the overview of the HiTeA. In concrete, our model consists of two unimodal encoders for encoding video and text separately, a multi-modal encoder for video and text interaction, and a text decoder for generation which is omitted here for simplicity and detailed in Appendix.
In order to take full advantage of the different view of the video, we introduce cross-modal moment exploration (CEM) to explore the proper words or phrases from input text to align the short-view video with for capturing the moment information in Section 3.2. Furthermore, to model the relations between the short-view video containing moment information and the long-view video with event information, we propose multi-modal temporal relation exploration (MTRE) to align the multi-modal representation of short-view and long-view videos by in Section 3.3. Lastly, we introduce the overall pre-training objective for training the model in Section 3.4.
2 Cross-Modal Moment Exploration
To understand the fine-grained moment information, the video with short temporal range (i.e. short-view of video) should be aligned with the corresponding text. However, since the video is partially aligned with the paired text which describes the whole video, directly aligning short-view of video with paired text would bring noise to model learning and degrade the performance. Therefore, we propose a novel pre-training task named cross-modal moment exploration (CME), which enables the model to understand fine-grained moment information.
Formally, we first discover the possible positive words for the video in short-view by computing the cosine similarity of the word embedding sequence from text encoder and the short-view video representation from video encoder as:
where is a permutation function for ranking such that , and is the set of selected word indices, is the number of possible selected words, and represents the cosine similarity between and . After obtaining the words for the video in short-view as the positive pair, the cross-modal moment exploration loss is computed with negative pairs from other words in the input text, which is defined as:
where is the learnable temperature hyper-parameter that controls the sharpness of the output distribution, and it is initialized as 0.07. As a consequence, the model is able to understand moment information via the proposed cross-modal exploration scheme.
3 Multi-Modal Temporal Relation Exploration
While the video encoder has demonstrated its effectiveness in learning temporal representation implicitly , it remains a challenge to discover the inherent temporal relations. As a result, the limited capabilities in temporal modeling deteriorate the downstream task in temporal reasoning. This is in particular a missing point for the existing video-language pre-training paradigm , which usually focuses on bridging video and text neglecting the function of text for guiding the video context representation learning thus losing the temporal cues.
where represents the multi-modal encoder with video features and text features . However, since the short-view of the video is partially aligned with the text, using the whole text is not reasonable for generating accurate text-guided video feature for short-view. Meanwhile, improper video-text pairs would yield noisy multi-modal representation thus degrading the performance of the model. Therefore, thanks to the positive words mined by cross-modal moment exploration, we can calibrate representation for short-view video by:
where is the index of the set for selected positive words. Then, we aim to match the representation of produced text-guided video features in different granularities in order to enable the model to predict the past and the future from the short-view of video, which benefits for capturing the general structure of the video. Specifically, we adopt the SimSiam framework for minimizing their negative cosine similarity:
where and . The and are projection MLP head and prediction MLP head . Minimizing is equivalent for minimizing the mean square error between and , which encourages the videos in different temporal magnitudes to be similar. Following , we defined a symmetrized loss as:
where is the stop-gradient operation that prevents the model from collapse during training .
4 Pre-training Objectives
Apart from the two proposed temporal-aware pre-training tasks, we follow proven video-text pre-training approaches to adopt the standard pre-training tasks including video-text contrastive (VTC), video-text matching (VTM), masked language modeling (MLM), and prefix language modeling (PrefixLM) described in the related work. Precisely, VTC and VTM align the video and text from the global perspective, while MLM and PrefixLM contribute to multi-modal understanding and generation capabilities of the model. Details of these objectives are described in the Appendix. We simply combine these as the base training objective for our model. Therefore, the full pre-training objective is computed as:
Experiments
Following the recent work , we pre-train our model on a webly-sourced video dataset WebVid-2M with 2.5M video-text pairs and a image-text dataset Google Conceptual Captions (CC3M) with 3M image-text pairs. Unlike previous methods, we do not pre-train our model on the large-scale video-text datasets like HowTo100M with 136M video-text pairs and YT-Temporal-180M due to the heavy computation. For scaling up, we also trained our model on the widely used image-text pre-training datasets including MS COCO , Visual Genome , SBU Captions and Conceptual 12M , we refer this setting as 17M corpus.
We evaluate our pre-trained model on 18 video-language benchmarks including video-text retrieval, video question answering, and video captioning tasks. Specifically, video question answering (VideoQA) can be categorized as Multiple-Choice (MC) and Open-Ended (OE) settings. The evaluation datasets are briefly summarized in below. The details can be found in the Appendix.
Video-Text Retrieval: MSRVTT , DiDeMo , LSMDC , ActivityNet Caption , SSv2-Label , and SSv2-Template ;
VideoQA (MC): TGIF-Action, TGIF-Transition , MSRVTT-MC , LSMDC-MC , and NExT-QA ;
VideoQA (OE): TGIF-Frame , MSRVTT-QA, MSVD-QA , LSMDC-FiB and ActivityNet-QA .
Our implementation of HiTeA is based on PyTorch . In detail, we instantiate the video encoder with MViT-Base model pretrained on ImageNet-21K . The text encoder is initialized from first six layers of pre-trained BERT-Base , and the multi-modal encoder is initialized with last six layers of pre-trained BERT-Base. We pre-train HiTeA for 10 epochs, using a batch size of 16 on 8 NVIDIA A100 GPUs. We use AdamW optimizer with a weight decay of 0.02 and betas (0.9, 0.98). The learning rate is first warmed up to 5e-5 in the first 1000 iterations, and decays following a cosine schedule. During pre-training, we sparsely sample 8 frames for short and long view while preserving their order in-between and resize them to 224 224. The duration of short view is restricted as the 1/8 of the whole video duration. is empirically set to 5. The MLM mask ratio is set to . Details of fine-tuning stage are described in Appendix.
2 Comparison to Prior Arts
In this section, we compare HiTeA with numerous state-of-the-art video-language pre-training methods on several downstream datasets under fine-tuning setting.
Table 1 summarizes the results on MSRVTT , DiDeMo , LSMDC , and ActivityNet Caption under fine-tuning settings. Our method outperforms all of the existing video-language pre-training model by a large margin under the same data scale. In particular, our method yields 6.6% lift in terms of R@1 on MSRVTT dataset while only exploiting 5M video-text pairs. Note that we also include the comparison with the recent works that utilize the powerful encoder from CLIP , our method still can be comparable with them even surpass them, which shows the validness of the proposed method. Besides, we can notice that our method achieves the best result among all of listed methods on LSMDC dataset, which proves that our model can leverage the various moments presented in fruitful movie clips with cross-modal moment exploration.
2.2 Video Question Answering
Table 2 lists the results of HiTeA and current state-of-the-art approaches on nine VideoQA datasets. It can be noticed that our method achieves the best performance in most of VideoQA datasets even with less pre-training data. Specifically, it achieves absolute improvement 1.1% on TGIF-FrameQA, 2.2% on MSRVTT-MC, 1.5% on MSRVTT-QA, 0.2% on MSVD-QA, and 3.3% on ActivityNet-QA. We believe the moments learned by the cross-modal exploration are useful for finding the clue of answers in VideoQA.
2.3 Video Captioning
Table 3 compares HiTeA with existing mthods on video captioning datasets MSRVTT and MSVD. As shown in the table, although we use less pre-training data than compared approaches, HiTeA still obtains significant improvement compared to those large-scale pre-trained models. On MSRVTT Caption, our method surpasses SoTA method MV-GPT by 2.5% CIDEr. Note that MV-GPT is pre-trained for multi-modal video captioning and it leverages the ASR transcripts from audio as the additional input. By contrast, our method only utilizes video as the input during generation.
3 Discussion
In this section, we discuss the temporal characteristics of our model and the datasets.
We investigate the contribution of individual loss terms and the results are shown in Table 4. It can be observed that the combining both and improves the performance of text-to-video retrieval and video question answering by at least 1.7% and 2.9% in Average Recall and Average accuracy respectively. In addition, we also find that the performance of surpasses that of on MSRVTT retrieval dataset that largely dominated by the appearance information. This can be explained that the cross-modal moment exploration loss not only select the positive verbs for the video from the text but also choose the acting object for alignment, which can boost the retrieval performance.
Lei et al. reveal that the previous four retrieval datasets are prone to being biased for appearance while rarely relying on temporal information, thus introducing Something-to-Something v2 (SSv2) Template and SSv2 Label retrieval datasets to test models’ true temporal modeling capability. In particular, SSv2 Template retrieval task requires a deeper understanding of the moment and temporal relation since no objects information are presented. The performance on these datasets are summarized in Table 5. It can be observed that HiTeA achieves significant improvement with gains in terms of R@1 on these two temporal-oriented text-to-video retrieval datasets, which demonstrates the effectiveness of our proposed method through exploring fine-grained moment information and modeling temporal relation. In addition, we evaluate our model on NExT-QA dataset that explicitly designed for temporal and causal understanding. As presented in Table 6, our method significantly surpasses its competitive counterparts, even those methods equipped with powerful image-text pre-trained encoders. Quantitatively, HiTeA obtains an absolute improvement on the causality split with the help of intrinsic temporal relation. Recently, Buch et al. filter out the trivial question for the dataset, and build the hard split for causality and temporal related questions for evaluate the causality and temporal of the model. As we can see in the table, even for the questions that heavily rely on causality, our model can still achieves a relative gain of 4.1% on the model with specific design for VideoQA, which indicates that our model do not solely depend on static appearance.
Previous methods only evaluate the performance of models on the existing datasets to demonstrate the superiority of the methods. However, Buch et al. and Lei et al. reveal that the most of the evaluation are biased towards the static concepts. Here, we investigate the temporal reliance for the evaluated datasets by introducing the temporal shuffling test. Specifically, we compute the performance changes between running inference on the ordered video versus its shuffled version. The large performance drop indicates the dataset has less spatial bias and needs for temporal information. Table 7 and Table 8 conclude the performance gap between ordered and shuffled input video for text-to-video retrieval and VideoQA datasets. For text-to-video retrieval task, SSv2 Template shows the large performance drop after shuffling the input video, which demonstrates that it depends most on the dynamic information thus verifying our assumption. On the contrary, the performance on ActivityNet Caption dataset is barely affected (-0.8 on Mean Recall) since the text almost describes the static objects without relying on temporal information. For video question answering dataset, we observe that the MSVD-QA and ActivityNet-QA are less sensitive to the order of video frames. This is because these two datasets contain more questions requiring frame-region information, such as object categories, scenes, and species. We believe this can be used to evaluate the temporal reliance of the datasets as well as the utilization for temporal cue by models in the future work.
To verify that our model can capture the motion information with respected to the given text rather than inferring from the static signal, we present the query text in SSv2 Template dataset which has masked all of the object information, and also visualize the query in MSRVTT dataset. As we can see in the Figure 3, the attention map of atomic action "talking" mainly focuses on the mouse of the cartoon characters while the baseline largely focusing on the characters, which indicates that our method can understand the moment better when adopting the temporal-aware pre-training tasks. In another example, the word "spins" can reveal the trajectory of the object showing that our method is able to capture the temporal motion presented in the video.
4 Zero-shot Generalizability
To demonstrate the generalizability of proposed video-text pre-trained model, we perform zero-shot evaluation on video-language downstream tasks. Table 9 summarizes the performance of our model and compared approaches on text-to-video retrieval. We can observe that our model yields more than 3.4% lift in R@1 on MSRVTT dataset while exploiting less video-text pairs. Besides, our method surpasses all of the compared models in terms of LSMDC dataset showing the superiority of our model’s generalizability. We also evaluate the zero-shot performance on VideoQA task in Table 10. Our method attains competitive zero-shot performance on MSRVTT-QA and MSVD-QA datasets even without help of audio signal supervision or additional generated video question pairs . In particular, less pre-training data (i.e. 5M 69M) are used while our method can still outperform other SoTA approaches. We also evaluate the zero-shot performance of models supervised on VQA v2 . We can find that our method surpasses the powerful multi-modal SoTA methods (e.g. mPLUG ) with only 5M pre-training data showing the better generalization ability of HiTeA.
Conclusion
In this work, we introduce HiTeA, a novel hierarchical temporal-aware video-language pre-training framework with both understanding and generation capabilities. We vary the video with different views and model cross-modal alignment between moments and texts as well as their temporal relations in a hierarchical way. Specifically, a cross-modal moment exploration pre-training task is proposed to explore the alignment between the text and video moment, which helps to overcome the partially semantic alignment between video and text. Moreover, multi-modal pairs are constructed to learn temporal relations between moments and the event presented by the video with multi-modal temporal relation exploration pre-training task. Even pre-trained on less data, HiTeA still achieves state-of-the-art performance on a wide range of video-language downstream datasets, which clearly shows the superiority of our method.
References
Appendix A Additional Experimental Results
In this section, we provide more experimental results for completeness of our proposed method.
Since images can be viewed as the single-frame videos, we evaluate the proposed method on image-text tasks including image-text retrieval and visual question answering.
We perform the Image-to-Text and Text-to-Image retrieval on COCO datasets, and the results are summarized in Table 11. We can observe that our method surpasses Singularity with same amount of pre-train data, especially 1% improvement on Recall@1 for Text-to-Image retrieval task. Moreover, although some methods leverage 4M dataset which contains the COCO dataset as a part of the pre-training dataset, HiTeA can still attain comparable results showing the good generalization ability.
We also evaluate our method on visual question answering task. Table 12 concludes the image question answering results on VQAv2 datasets. We observe that HiTeA demonstrates competitive performance on the VQA tasks. It is worthwhile noting that our method achieves the better performance compared to Singularity same pre-training datasets, which indicates the video-text pre-training would boost the performance of image-text downstream tasks. However, we still see a gap with state-of-the-art image-text pre-trained models since our method do not use the in-domain data (e.g. COCO) during pre-training, thus leading to the gap with SoTA performance. One future direction is to use more image-text data during video-text pre-training for better generalization.
A.2 Additional Ablation Studies
We investigate the effect of choosing different positive words size during cross-modal moment exploration. As depicted in Figure 4, it can be observed that with the increment of , the performance on each dataset is increasing then start to decrease. In addition, there is a trade-off between the choice of and performance with respected to different datasets, and gives relative good results among these datasets. It also suggests that the small would give more deterministic results since the model would only select the word with the largest similarity, thus more focusing on the single action or object. Then, as number of positive words increased, more accurate words are selected to align with the short-view of video. However, the model no longer benefits from cross-modal moment exploration when is large enough (i.e., or ) due to the increased noise in the selected candidate words.
To further validate the temporal dependency for the proposed method, we adopt the shuffling test for models with different loss terms, as shown in Table 13. Table 13 shows that our loss terms contribute more significantly when the dataset requires more temporal understanding. In concrete, and consistently improve the performances of Original and Gap on more temporal relied datasets (i.e. SSv2-Template and SSv2-Label). For example, model with two loss terms largely surpasses the baseline model in the metric of Gap by achieving 4.4 and 0.5 improvement on SSv2-Template and SSv2-Label, respectively.
Table 14 shows that our proposed method is generalizable to different vision backbones. In details, we instantiate the video encoder with TimeSformer pretrained on ImageNet-21K . It can be observed both CME and MTRE consistently improve the model performance across the video backbones considered showing the generalization of proposed hierarchical temporal-aware pre-training framework. It is worth noting that, TimeSformer generates long video tokens compared to that of Multi-scale ViT , which brings extra memory cost for the multi-modal encoder and decoder since the computation of self-attention is quadratic. This makes TimeSformer expensive to scale to more input frames with longer sequences.
We investigate the influence of language for multi-modal temporal relation exploration. Instead of utilizing the language signals, we directly adopt the video representation from the video encoder during the learning. The results are sketched in Figure 5. It can be observed that the model trained with multi-modal pairs attains better performance than the model without text. In concrete, it achieves 2.1% gains on SSv2-Template which mainly depends on the understanding of actions, which indicates that our method can better understanding the actions via multi-modal temporal relation exploration. Besides, we also notice that performance of the model trained with correct multi-modal pairs surpasses that of model trained by multi-modal pairs with same text, which indicates that improper video-text pair yields noisy multi-modal representation thus degrading the performance of the model.
Appendix B Discussion
We sample some videos and corresponding texts and compute similarities between words and videos in Figure 6. As we can see in the figure, our model can effectively capture the moments such as "spreading", "moving", and "preparing" etc. in the video, which is essential for understanding videos. Besides, we can notice that the video would also attend the object appeared in the video, showing the capability for modeling fine-grained moment information.
B.2 Connection to Other Fine-Grained Methods
Some efforts have been made to learn the fine-grained correlation and alignment between two modalities by leveraging the token-wise similarities in vision-language pre-training. FILIP and TERAN aggregate the maximum token similarity scores and assign the optimal patch-word transport matrix. SCAN utilizes the similarity scores to attend each tokens for soft fine-grained alignment. These approaches are originally tailored for image-text pre-training, which aims to locate the fine-grained static object. However, different from image-text pre-training, video-text pre-training needs to understand the correlation between words and moments, which not only contains static objects but also consists of atomic actions. Our proposed cross-modal moment exploration leverages the short-view of video to reflect the moment information and discover the relationship between short-view videos and words, which results in fine-grained moment representations for video-language pre-training.
B.3 Limitations and Boarder Impact
Despite the effectiveness of the proposed method on various downstream tasks, our method still has some limitations that would make for promising directions for future work. (1) Currently, we only pre-train our model on 5M data with the base-size encoders, and the scalability of the model is not explored which deserves more in-depth investigation in the future. (2) Our method shares similar risks like other pre-training methods that the pre-training data might consist bias and unsafe content which requires further analysis before the deployment.
Appendix C Implementation Details
We include some of previous models with their parameter counts (which were reported in the original paper or calculated by follow-up work), and we compare them with HiTeA in Table 15. Compared with other models, our model is of comparable model size and requires less pair of video-text pre-training data to achieve better performance in terms of both video-language understanding and generation.
C.2 Model Architecture
C.3 Pre-training Objectives
During pre-training, we also perform four pre-training tasks including Video-Text Contrastive Learning (), Video-Text Matching (), Masked Language Modeling (), and Prefix Language Modeling (). The VTC task first is applied to align the unimodal representation of video and text. And the multi-modal representation can be learned by VTM and MLM tasks. Upon on the video-language representations obtained from multi-modal encoder, the decoder is trained by PrefixLM loss with text completion task.
Following , we align the unimodal encoders via this task. Specially, the softmax-normalized video-to-text and text-to-video similarities are computed, and we employ memory queues in MoCo to increase the number of negative samples during learning. Formally, the video-text contrastive loss is calculated as:
where and are the projected representations of and for -th video-text pair in the batch.
This task aims to predict whether a video and a text is paired or not based on the multi-modal representation. As suggested in , hard negative video-text pairs are selected based on the similarity of video and text during contrastive learning. Formally, the video-text matching loss is calculated as:
where denotes the word tokens, and denotes the video features of long-view video.
The setup of this pre-training task is same as that used in BERT , where 15% of tokens in the text are randomly masked, and the model needs to predict the masked tokens based on the multi-modal representation. Formally, the masked language modeling loss is calculated as:
where denotes the masked word token.
This pretext task requires model to complete the truncated texts based on given videos and prefix sequence of truncated texts . The model can be trained by maximizing the likelihood of the truncated text in an auto-regressive manner. Formally, the prefix language modeling loss is calculated as:
where denotes the total number of words in the text, and is the length of a prefix sequence of tokens which is randomly selected.
C.4 Downstream Task Implementation Details
We evaluate HiTeA on various downstream video-language tasks, including Text-to-Video Retrieval, Open-ended VideoQA, Multiple Choice VideoQA, and Video Captioning. The fine-tuning procedures are described as follows:
For retrieval tasks, we jointly optimize the VTC loss and VTM loss for video-text alignment during fine-tuning. During inference, we first select top-k candidates by computing the dot-product similarity between the video and text features, and then reranking the selected candidates based on their VTM scores. is set to 128 by default.
For open-ended VideoQA, we first generate video features and text features with two unimodal encoders, and then fuse them with multi-modal encoder. The output of multi-modal features are fed to text decoder for answer generation. We use the language modeling loss to optimize the model. During inference, the answer would be generated by the text decoder.
For multiple choice VideoQA, we treat the problem as the text-to-video retrieval task where the correct answer should have the highest matching probability. During training, we compute the VTM scores for each candidate answer and video, then optimize the model with cross entropy loss. During the inference, the answer with highest VTM score is the prediction answer.
For Video Captioning, we use the video features from video encoder and directly feed it into text decoder for caption generation. The language modeling loss is utilized for model optimization.
For all above video-language downstream tasks, we resize video frames to . During fine-tuning, following , we randomly sample 12 frames for text-to-video retrieval, 16 frames for video question answering and video captions. We perform uniform sampling during inference. We use RandomCrop with minimum ratio 0.5 and HorizontalFlip with 0.5 probability for data augmentation. The hyperparameters that we used for fine-tuning on downstream tasks are summarized in Table 16. For the video caption task, we use a prefix prompt “A video of” to improve the quality of generated captions.
C.5 Datasets Description
In this section, we describe all of the downstream video-language datasets used during evaluation. The details of the datasets are represented below:
We evaluate HiTeA on 6 popular text-to-video retrieval datasets including MSRVTT , DiDeMo , LSMDC , ActivityNet Caption , SSv2 Template , and SSv2 Label . Details of these datasets: MSRVTT contains 10K YouTube sourced videos with 200K text descriptions. Following , we train the video on 9K videos and evaluate on the rest 1K video. DiDeMo contains of 10K videos from Flickr and 4 descriptions for each video. Following , we concatenate all of the given descriptions from the same video as a paragraph, and evaluate the paragraph-to-video retrieval performance. The number of video in training set is 8K, leaving 1K for validation set and 1K for test set. LSMDC consists of 118K video clips from 202 movies, and each clip is accompanied with a caption from video scripts. It has 101K video clips for training and 1K clips for testing. We use the standard splits from . ActivityNet Caption is built on 20K YouTube videos with 100K captions. We use the train split with 10K videos for training, and report the performance on the val1 split with 4.9K videos. SSv2-Template and SSv2-Label contain 169K videos for training and 2K videos for testing. The text queries in SSv2-Template are templates without object information (e.g. "Throwing [something] in the air and catching it"). By contrast, SSv2-Label contains annotated text queries with specific object information (e.g. "Throwing keys in the air and catching it"). Therefore, SSv2-Template mainly focuses on temporal understanding of actions, while SSv2-Label needs a more comprehensive understanding of both appearance and temporal dynamic.
Five datasets are evaluated for multiple-choice video question answering tasks. TGIF-Action and TGIF-Transition are adopted to evaluate model’s capability to recognize the repeated actions and state transitions in short GIFs. Each video and question is equipped with 5 candidate answers. We concatenate the question and answer as the text and use the highest similarity among the video and candidate texts. TGIF-Action contains 18K GIFs for training and 2K for testing. TGIF-Transitions has 47K GIF-question pairs for training and 6K for testing. MSRVTT-MC and LSMDC-MC are originally retrieval task, but reformulated as the multiple choice video QA task. It requires the model to find the optimal caption that describes the video out of 5 candidate texts. NExT-QA is explicitly designed for temporal and causal understanding. Questions in the dataset are categorized into three types: Descriptive, Temporal, and Causal. Each question in the dataset are paired with 5 candidate answers. Therefore, this dataset is able to evaluate model’s ability in video question answering in different aspects.
For open-ended video QA, we evaluate the model on five datasets. MSRVTT-QA is composed of 243K open-ended questions over 10K videos, while MSVD-QA consists 2K videos with 47K questions. TGIF-Frames collects the answerable with just a single frame in the video, and is divided into training set with 35K questions and test set with 14K questions.. For LSMDC-FiB , the model needs to predict a correct word for the blank with a given video and a sentence with blank. It contains 297K sentences for training and 30K sentences for testing. ActivityNet-QA .
We use MSRVTT and MSVD for video captioning evaluation. As described before, MSRVTT is composed of 10K videos with 20 captions per video, and MSVD contains 2K videos with around 40 captions per video. We follow the standard splits from . During inference, we generate the caption with beam search until the model outputs a [SEP] that indicates the end of sentence or when it reaches the maximum generation step 40.