UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation
Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Jason Li, Taroon Bharti, Ming Zhou
Introduction
With the recent advances of self-supervised learning, pre-training techniques play a vital role in learning visual and language representation. The paradigm is to pre-train the model on a large scale unlabeled data and fine-tune the downstream tasks using task-specific labeled data. Inspired by the BERT Devlin et al. (2019) model’s success for NLP tasks, numerous multimodal image-language pre-training models Lu et al. (2019); Li et al. (2019a, b) have been proposed. Their results have demonstrated the effectiveness of pre-training on various visual and language tasks such as visual question answering. Different from previous text pre-training or image-language pre-training, we focus on video-linguistic pre-training in this paper.
Videos contain rich visual, acoustic, and language information for people to acquire knowledge or learn how to perform a task. This motivates researchers to investigate whether AI agents can learn task completion from videos like humans with both low-level visual and high-level semantic language signals. Therefore, multimodal video-language tasks are of great importance to investigate for both research and applications. In this work, we first propose to pre-train a unified video-language model using video and acoustic speech recognition (ASR) transcript in instructional videos to learn a joint representation of both video and language. Then, we fine-tune this model on five typical multimodal tasks, including understanding and generation targets. Figure 1 presents a showcase of our pre-training and fine-tuning flow. Take multimodal video captioning as an example. The model inputs video and ASR transcript and predicts a captioning sentence.
VideoBERT Sun et al. (2019b) and CBT Sun et al. (2019a) are the first pioneers to investigate video-language pre-training with regard to video representation on instructional videos. They have demonstrated the effectiveness of the BERT based model for capturing video temporal and language sequential features. Besides the above two works, there is some concurrent progress to our model. ActBERT Zhu and Yang (2020) leverages global action information to catalyze mutual interactions between linguistic texts and local regional objects. Moreover, a transformer block is introduced to encode global actions, local regional objects, and linguistic descriptions. HERO Li et al. (2020) hierarchically encodes multimodal inputs. Furthermore, two new pre-training tasks, video-subtitle matching and frame order modeling, are designed to improve the representation learning. VideoAsMT Korbar et al. (2020) takes a generative modeling approach that poses the objective as a translation problem between modalities.
However, most of previous models only pre-train the model on understanding tasks. In this paper, we pre-train on both understanding and generation tasks through an encoder-decoder paradigm. Although the concurrent work VideoAsMT has a similar encoder-decoder as ours, it is not flexible for downstream tasks with only one single unified framework. In this paper, we develop a flexible approach to learn video and language joint representation and adapt downstream multimodal tasks.
We propose UniVL: a Unified Video and Language pre-training model for multimodal understanding and generation. Our UniVL model adopts Transformer Vaswani et al. (2017) as the backbone and has four components, including two single-modal encoders, a cross encoder, and a decoder. In detail, we first encode the text and visual separately by two single-modal encoders. A video-text joint objective performs on these two encoders, which aims to learn better representation for each modality before fusing them. Such a two-stream design is natural to retrieval tasks due to its scalability to very large datasets. The proposed representation can be indexed and has linear complexity in the number of videos. Then we adopt the Transformer based encoder-decoder model to perform the understanding and generation pre-training by four tasks: conditioned masked language model (CMLM for language corruption), conditioned masked frame model (CMFM for video corruption), video-text alignment, and language reconstruction.
Furthermore, we design two pre-training strategies, including stage by stage pre-training strategy (StagedP) and Enhanced video representation (EnhancedV), to promote the UniVL pre-training. The StagedP has two parts in our setting. We only pre-train the text encoder and video encoder by the video-text joint objective for the first stage. Then all modules will be pre-trained under the whole objectives in the second stage. Besides, we adopt an entire masking strategy EnhancedV on text to enhance video representation.
Our contributions are summarized as follows:
1) We propose a multimodal video-language pre-training model trained on a large-scale instructional video dataset. It is a flexible model for both video-language understanding and generation tasks.
2) The pre-training consists of five objectives, including video-text joint, conditioned masked language model, conditioned masked frame model, video-text alignment, and language reconstruction. Two pre-training strategies are proposed to make these objectives work harmoniously.
3) We fine-tune our pre-trained model on five typical multimodal video-language tasks: text-based video retrieval, multimodal video captioning, action segmentation, action step localization, and multimodal sentiment analysis. Extensive experiments demonstrate our model’s effectiveness on downstream tasks and achieve state-of-the-art results.
Related Works
Self-supervised representation learning has been shown to be effective for sequential data, including language and video. Language pre-training models, including BERT Devlin et al. (2019), GPT Radford et al. (2018), RoBERTa Liu et al. (2019), XLNet Yang et al. (2019), MASS Song et al. (2019), UniLM Dong et al. (2019), and BART Lewis et al. (2019), have achieved great success on NLP tasks. BERT Devlin et al. (2019) is a denoising auto-encoder network using Transformer with MLM (masked language model) and NSP (next sentence prediction) as pre-training tasks. It has a strong performance for understanding tasks. MASS Song et al. (2019) focuses on pre-training for generation tasks. UniLM Dong et al. (2019) and BART Lewis et al. (2019) continuously study a unified pre-training model for both understanding and generation tasks.
Video representation learning mostly focuses on the video sequence reconstruction or future frames prediction as pre-training (pretext) tasks. Early works like Mathieu et al. (2015); Srivastava et al. (2015); Han et al. (2019) aim to synthetic video frames through the image patches. Similarly, Wang and Gupta (2015) adopt a siamese-triplet network to rank continuous patches more similar than patches of different videos. Other works predict the feature vectors in latent space using auto-regressive models with the noise-contrastive estimation (NCE) Lotter et al. (2016); Oord et al. (2018). Sun et al. (2019a) adopt NCE to predict corrupted (masked) latent space using the auto-encoder model.
2 Multimodal Pre-Training
Recently, numerous visual-linguistic pre-training models are proposed for multimodal tasks. For image and text pre-training, ViLBERT Lu et al. (2019), LXMERT Tan and Bansal (2019) adopt two separate Transformers for image and text encoding independently. Other models like Visualbert Li et al. (2019b), Unicoder-VL Li et al. (2019a), VL-BERT Su et al. (2020), UNITER Zhou et al. (2019) use one shared BERT model. These models employ MLM and image-text matching as pre-training tasks which are effective for downstream multimodal tasks. VLP Zhou et al. (2019) proposes a unified image-language model for understanding and generation tasks. Different from these works, we focus on video and text pre-training for universal representation.
For video and text pre-training, VideoBERT Sun et al. (2019b) and CBT Sun et al. (2019a) are the first works to explore the capability of pre-training models. Although VideoBERT and CBT pre-train the model on multimodal data, the downstream tasks mainly take video representation for further prediction. ActBERT Zhu and Yang (2020) leverages global action information to catalyze mutual interactions between linguistic texts and local regional objects, and introduces a transformer block to encode global actions, local regional objects, and linguistic descriptions. HERO Li et al. (2020) encodes multimodal inputs in a hierarchical fashion. Besides, two new pre-training tasks, video-subtitle matching and frame order modeling, are designed to improve the representation learning. However, ActBERT and HERO are only pre-train the models on understanding tasks. VideoAsMT Korbar et al. (2020) takes a generative modeling approach that poses the objective as a translation problem between modalities. The difference between our work with VideoAsMT is that our model contains two more separate encoders instead of one unified encoder-decoder, while VideoAsMT is inflexible for downstream tasks due to one single unified framework.
We summarize three pre-training paradigms to cover the previous vision-text pre-training model considering different encoder architecture in literature, as presented in Figure 2. Unicoder-VL Li et al. (2019a), VL-BERT Su et al. (2020), UNITER Zhou et al. (2019), VLP Zhou et al. (2019), VideoBERT Sun et al. (2019b), ActBERT Zhu and Yang (2020), and VideoAsMT Korbar et al. (2020) belong to share-type in Figure 2(a), where the text and vision sequences are combined as the input of one shared Transformer encoder. ViLBERT Lu et al. (2019) and LXMERT Tan and Bansal (2019) are cross-type shown in Figure 2(b). CBT Sun et al. (2019a) and HERO Li et al. (2020) are joint-type shown in Figure 2(c). The cross-type and joint-type architectures have two-stream input, and the difference is the interaction across both modalities. Compared with the single-stream input in the share-type, the two-stream input can accommodate each modality’s different processing needs and interact at varying representation depths Lu et al. (2019). Besides, the joint-type structure has one cross-modal encoder for full interaction between the two streams comparing with the cross-type. We adopt the joint-type as our encoder in this paper.
Method
The problem is defined as: given the input video and the corresponding ASR transcript pairs, pre-train a model to learn joint video and text representation with the self-supervision approach, and fine-tune downstream tasks. In this section, we describe the architecture and pre-training tasks in detail.
Figure 3 presents the UniVL as an encoder-decoder architecture. First, the model extracts representations of the input text tokens and the video frame sequences using various feature extractors. A text encoder then adopts the BERT model to embed the text, and a video encoder utilizes the Transformer encoder to embed the video frames. Next, we employ a Transformer based cross encoder for interacting between the text and the video. Finally, a Transformer decoder is used to reconstruct the input text.
We first pre-process video and language before feeding to the UniVL. For the input text, we tokenize all words by WordPieces Wu et al. (2016) following the pre-processing method in BERT to obtain the token sequence \mathbf{t}=\big{\{}t_{i}|i\in[1,n]\big{\}}, where is the -th token and is the length of the token sequence. For each video clip, we sample a frame sequence \mathbf{v}=\big{\{}v_{j}|j\in[1,m]\big{\}} and adopt them to extract features, where is the -th group of video frames and is the group length of the frame sequence.
1.2 Single Modal Encoders.
where is the hidden size of text representation.
1.3 Cross Encoder.
where denotes the combination operation. It is noted that the combination is operated along with the dimension of sequence, not the dimension of hidden size. One reason is that the text length and video clip length are always different. Another reason is that the semantic between text and video are not absolutely aligned. People are likely to describe an event after or before performing it in the video Miech et al. (2020).
1.4 Decoder.
2 Pre-training Objectives
We have five pre-training objectives: 1) video-text joint, 2) conditioned masked language model (for text corruption), 3) conditioned masked frame model (for video corruption), 4) video-text alignment, and 5) language reconstruction.
As our text encoder, the BERT-Base uncased model is a robust extractor of text representation. So, we utilize a video-text joint objective to enhance the capability of the video encoder. It seems a retrieval orienting operation, which is to align the space of representation between text and video. Considering the misalignment between the text and video clip in narrated videos, we adopt MIL-NCE Miech et al. (2020) on and as our joint objective,
2.2 CMLM: Conditioned Masked Language Model.
Following BERT, we also randomly mask 15% tokens with the special token [MASK] in the sentence and re-produce the masked tokens under the condition of video input and known tokens. This loss function is defined on the feature matrix of the text part in as:
where means the contextual tokens surrounding the masked token , is the trainable parameters.
2.3 CMFM: Conditioned Masked Frame Model.
Similarly, we also propose a masked frame model to predict the correct frames given contextual frames and the input text for semantic constraints. However, it is hard to reconstruct the original RGB frame. We adopt the contrastive learning method to maximize the MI (Mutual information) between the masked output features and the original features. This loss function is NCE Sun et al. (2019a). We randomly mask 15% vectors (also 15% frames) with zeros. The objective is to identify the correct frame compared to negative distractors. The loss is defined as:
2.4 Video-Text Alignment.
We use the fused representation that corresponds to the special token [CLS] to predict scores for the video-text alignment, which is similar to the BERT sentence pair classification task. We adopt the NCE loss to learn to discriminate against the positive from negative video-text pairs. To enhance this capability, we not only randomly sample negative cases but also re-sample video clips from the same video Han et al. (2019). The reason is that the frames inside the same video are more similar than frames of different videos. This loss function is defined as follows,
where means two linear layers with a activation function between them, which is performed on the first hidden state of . We take other video clips in the same batch as negative cases .
2.5 Language Reconstruction.
To reconstruct the input sentence to endow the pretrained model with the generation capability, we employed an auto-regressive decoder with reconstruction objective, and the loss function is,
It is noted that is the masked version of ground-truth text when pre-training. As shown in BART Lewis et al. (2019), pre-training decoder benefits generation tasks.
We jointly optimize our model by a weighted loss:
3 Pre-training Strategies
We develop two pre-training strategies to train the UniVL model effectively.
The UniVL can benefit from the pre-trained BERT-Base uncased model in the text encoder module. The natural idea is to train a peer to peer video encoder as the BERT-Base. We adopt a two-stage training fashion. For the first stage, we only preserve the text BERT and video Transformer to learn the weights using the Video-Text Joint loss Eq. (5). Next, we decrease the learning rate and continue to further pre-train the UniVL by all five objectives. One advantage is to fasten the pre-training speed, and the other advantage is to make the pre-training progress more smoothing on weights.
3.2 EnhancedV: Enhanced Video Representation.
To further enhance the video representation, we adopt a masked modality strategy to make the video to generate transcripts without text input. Specifically, we mask the whole text tokens with a 15% possibility. In other words, there are 15% text-video pairs with entire text tokens masked in each mini-batch, and the model utilizes the video information to complete generation. Such a strategy is a more challenging task for the model to learn a better video representation.
Experiments
We first pre-train our model on the large scale dataset. We download videos with ASR transcripts from Howto100M dataset Miech et al. (2019)https://www.di.ens.fr/willow/research/howto100m/. After filtering the unavailable ones, we get 1.2M videos for pre-training our model. On average, the duration of each video is 6.5 minutes with 110 clip-text pairs.
Then, we fine-tune our pre-trained model on five diverse downstream tasks using five datasets, including text-based video retrieval, multimodal video captioning, action segmentation, action step localization, and multimodal sentiment analysis.
Youcook2 Zhou et al. (2018a) contains 2,000 cooking videos on 89 recipes with 14K video clips. The overall duration is 176 hours (5.26 minutes on average). Each video clip is annotated with one captioning sentence. We evaluate both text-based video retrieval and multimodal video captioning task on this dataset.
For the text-based video retrieval task, we follow the same experimental setting in Miech et al. (2019), and use the captions as the input text queries to find the corresponding video clips. For the video captioning task, we use the same setting as in Shi et al. (2019). We filter the data and make sure there is no overlap between pre-training and evaluation data. In all, we have 1,261 training videos and 439 test videos, that is, 9,776 training clip-text pairs and 3,369 test clip-text pairs.
1.2 MSR-VTT
MSR-VTT Xu et al. (2016) is the open-domain dataset for video retrieval tasks. It has open domain video clips, and each clip has 20 captioning sentences labeled by human. In all, there are 200K clip-text pairs from 10K videos in 20 categories including sports, music, etc. Following JSFusion Yu et al. (2018), we randomly sampled 1,000 clip-text pairs as test data to evaluate the performance of our model on text-based video retrieval task.
1.3 COIN
COIN Tang et al. (2019) is to evaluate action segmentation task, which contains 180 different tasks and 11,827 videos. Each video is labeled with 3.91 step segments. In total, the dataset contains videos of 476 hours, with 46,354 annotated segments.
1.4 CrossTask
CrossTask Zhukov et al. (2019) is to evaluate the action step localization task. It contains 83 different tasks and 4.7k videos. For each task, an ordered list of steps with manual descriptions are provided.
1.5 CMU-MOSI
Multimodal Opinion Sentiment and Emotion Intensity Zadeh et al. (2016) is sentence-level sentiment analysis and emotion recognition in online videos. CMU-MOSI contains 2,199 opinion video clips, each annotated with real-valued sentiment intensity annotations in the range [-3, +3]. We evaluate the performance of our model on multimodal sentiment analysis.
2 Experimental Details
For text encoding, we apply WordPiece embeddings Wu et al. (2016) with a 30,000 token vocabulary to input to BERT model. We exploit the BERT-base model Devlin et al. (2019) with 12 layers of Transformer blocks. Each block has 12 attention heads and the hidden size is 768.
For video encoding, we first extract the 3D feature from video clips using the S3D model pretrained by Miech et al. (2020). The basic visual feature can significantly affect the results from our preliminary experiments. The fps of the 3D feature extractor is 16 and the dimension is 1,024. We then employ Transformer Encoder with 6 layers to capture the sequential information on the 3D feature. Each block has 12 attention heads and the hidden size is 768.
The model consumes the clip-text pairs. The maximal input tokens of text is 32 and the maximal number of video features is 48. For short sentence and clip, we concatenate contextual tokens and frames. For cross encoder and decoder, we use a 2 layers Transformer Encoder as the encoder and a 3 layer Transformer Decoder as the decoder with 12 heads and 768 hidden size. For generation task during the inference stage, we use the beam search with the size of 5. As previously mentioned, the generated sequence is the ground-truth input transcripts in the pre-training phase. Its target is to sequentially learn full information from the masked transcripts and video features.
We pre-train our model on 8 NVIDIA Tesla V100 GPUs. There are two sets of hyper-parameters considering the stage by stage pre-training strategy. In the first stage, the batch size is set to 600 and the model is trained 50 epochs for 1.5 days. In the second stage, the batch size is set to 48 and the model is trained 50 epochs for 12 days. We use the Adam optimizer Kingma and Ba (2015) with an initial learning rate of 1e-3 in the first stage and 1e-4 in the second stage, and employ a linear decay learning rate schedule with a warm-up strategy.
3 Main Results
Text-based video retrieval is defined to retrieve a relevant video/clip given an input text query. As shown in Figure 3 (retrieval block), the model encodes the input text query and candidate video clips through the text encoder and video encoder respectively. Then calculate the matching scores using two different approaches: one is UniVL (FT-Joint), which calculates the score through dot product as in Eq. (6), and use as the loss during the fine-tuning stage; the other is UniVL (FT-Align), which feeds the encodings to both single encoders and the cross encoder to get unified representation and predict the match score through in Eq. (12) on the first token ‘[CLS]’. During the fine-tuning stage, the loss is . We use the Adam optimizer with an initial learning rate of 3e-5 and a batch size of 32 video-caption pairs for Youcook2, an initial learning rate of 5e-5 and a batch size of 128 video-caption pairs for MSR-VTT as hyper-parameters to fine-tune for 5 epochs.
We fine-tune our pre-trained model for text-based video retrieval task on both Youcook2 and MSR-VTT datasets. The evaluation metrics are Recall@n (R@n) and Median R. Tables 1 and 2 list the retrieval results of all baselines and our model on Youcook2 and MSR-VTT separately. We can see that our model achieves the best performance over all baselines to a large extent. We present several baseline methods with or without pre-training. Our model outperforms the Howto100M and VideoAsMT models pre-trained on the same dataset on all metrics. Besides, the experimental results present the a large performance gain with pre-training.
We also notice that UniVL (FT-Align) performs better than UniVL (FT-Joint), which demonstrates that fusion representation generated by the cross encoder is better. Nevertheless, the UniVL (FT-Joint) inference speed is 50 times for Youcook2 and 10 times for MSR-VTT faster than that of the UniVL (FT-Align). Therefore, it is a trade-off between performance and efficiency in practical applications. In the following ablation experiment, we exploit UniVL (FT-Joint) in the retrieval task.
3.2 Multimodal Video Captioning.
Multimodal video captioning aims to generate a sequence of descriptive sentences. As shown in Figure 3 (caption block), the model encodes the input video frames as well as transcripts inside the clips through the video encoder and text encoder respectively, then feeds the encodings to the cross encoder to get unified representation, and finally generates token sequence by the decoder. We use as the loss during the fine-tuning stage. The hyper-parameters are an initial learning rate of 3e-5, a batch size of 32 samples, and fine-tune for 5 epochs.
Table 3 lists the caption results of all baselines and our models on Youcook2. This generation task adopts the corpus-level generation evaluation metric using the pen-source toolhttps://github.com/Maluuba/nlg-eval, including BLEU (BLEU-3, B-3; BLEU-4, B-4) Papineni et al. (2002), METEOR (M) Banerjee and Lavie (2005), ROUGE-L (R-L) Lin and Och (2004), and CIDEr Vedantam et al. (2015). We compare our pre-trained model with several baseline methods. We classify the methods with the setting that the input is video-only or video+transcript. Zhou et al. (2018a) propose an end-to-end model for both procedural segmentation and captioning. Sun et al. (2019b, a); Zhu and Yang (2020); Korbar et al. (2020) adopt the pre-training strategy and evaluate the captioning with the only video as input. Shi et al. (2019) and Hessel et al. (2019) discuss the multimodal input with both video and transcript. Our pre-trained model achieves state-of-the-art results and outperforms the existing pre-trained models, even only considering video as input.
3.3 Action Segmentation.
We fine-tune our pre-train model on action segmentation task using COIN dataset, which is to predict one pre-defined label for each frame of the given video. As shown in Figure 3 (action tasks block), the model encodes the input video frames through the video encoder, followed by a linear classifier upon the output encodings for frame labeling. We do not use the text encoder due to no text description in the dataset. The evaluation metric is frame-wise accuracy (FA). The hyper-parameters are an initial learning rate of 3e-5, a batch size of 32 samples, and fine-tune for 5 epochs. The results are shown in Table 4. The UniVL significantly outperforms the baselines with more than 14% improvements. It shows that the pre-trained UniVL actually learns a good visual representation, even absent of linguistic descriptions.
3.4 Action Step Localization.
We evaluate the action step localization on CrossTask dataset. As shown in Figure 3 (action tasks block), the model encodes the step description (action) and video clip through the text encoder and the video encoder respectively. And then calculate the relevance scores through dot product similar to the retrieval task. To fairly compare to Miech et al. (2019, 2020); Zhu and Yang (2020), we do not fine-tune on the CrossTask dataset. We perform the evaluation protocol by reporting the average recall (CTR) metric for the localization taskThe result is generated following the evaluation process of official project: https://github.com/DmZhukov/CrossTask. The results are shown in Table 5. Our results are even better than the supervised baseline, which demonstrates our UniVL model can learn better joint text-video representation.
3.5 Multimodal Sentiment Analysis.
We evaluate the multimodal sentiment analysis on CMU-MOSI dataset, the goal of which is to identify the sentiment of speaker based on the speaker’s display of verbal and nonverbal behaviors. We employ video and corresponding transcripts to accomplish this task. As shown in Figure 3 (multimodal classification block), the model encodes the input video frames as well as transcripts inside the clips through the video encoder and text encoder, respectively. Then feeds the encodings to the cross encoder to get unified representation, and finally predicts the sentiment score by a linear on the first token ‘[CLS]’. The hyper-parameters are an initial learning rate of 1e-5, a batch size 32, and fine-tune for 3 epochs.
The results are shown in Table 6. Following Zadeh et al. (2019), the evaluation metrics are binary accuracy (BA), F1 score, Mean-Absolute Error (MAE), and Pearson Correlation Coefficient (Corr). Compared with the baseline using video, transcript, and audio inputs, our model trained with video and language still achieves the best results without audio information.
4 Ablation Studies
We analyze the effectiveness of our model design on pre-training objectives and strategies through ablation studies over text-based video retrieval and multimodal video captioning tasks. We also discuss the effectiveness of various visual features.
Table 7 shows the effectiveness of each objective or strategy on the retrieval task. The results are reported on both Youcook2 and MSR-VTT datasets. Simultaneously, Table 8 demonstrates the effectiveness of each objective or strategy on the caption task. For the retrieval task, we exploit UniVL (FT-Joint) fine-tuning strategy to study the objectives: Joint loss, Alignment loss, and Decoder loss, and the strategies: StagedP and EnhancedV show consistent improvement. From the result, we can see that the cross encoder and decoder modules can promote the joint representation of video and text. For the caption task, we find that the decoder module shows great advantage and achieves more than 3 points gain on the BLUE-4 metric. Another finding is that the Joint loss decreases the generation task a little, although it performs well in the retrieval task. Excessive emphasis on coarse-grained matching can affect the fine-grained description at the generation task.
4.2 Visual Features.
We compare the S3D video feature pre-trained on Howto100M and ResNet-152 plus ResNeXt-101 pre-trained on labeled ImageNet and Kinetics respectively. The ResNet-152 (RS152) and ResNeXt-101 (RX101) are used to extract 2D and 3D features from video clips respectively similar to Miech et al. (2019)’s work.
As shown in Table 9 and Table 10, the visual feature is important in our pre-training model and the downstream tasks. It is worth studying an end to end training from raw videos instead of extracted fixed video features in the future. However, the time-cost and the memory-cost are enormous. The key bottleneck is visual representation, and we propose two possible approaches: designing a lightweight training scheme, e.g., training on keyframes of video, using a small feature dimension size.
Conclusion and Discussion
This paper proposes UniVL with self-supervised learning for video and language representation on large scale videos. The UniVL is designed with four modules and five objectives for both video-language understanding and generation tasks. It is a flexible model for most of the multimodal downstream tasks considering both efficiency and effectiveness. We conduct extensive experiments on evaluating our model for five downstream tasks, e.g., text-based video retrieval and multimodal video captioning. The experimental results demonstrate that our pre-trained model can improve the performance to a large extent over the baseline models and achieve state-of-the-art results on five typical multimodal tasks. Besides, we will investigate our model’s performance on more massive datasets and more downstream tasks for future work.