All in One: Exploring Unified Video-Language Pre-training
Alex Jinpeng Wang, Yixiao Ge, Rui Yan, Yuying Ge, Xudong Lin, Guanyu Cai, Jianping Wu, Ying Shan, Xiaohu Qie, Mike Zheng Shou
Introduction
The pre-train-and-then-fine-tune scheme has become a standard paradigm to learn transferable video-language representations for a wide range of downstream video-text tasks, for example, text-video retrieval , video-question answering , multiple choice and visual commonsense reasoning . In recent years, there has been tremendous progress in the development of video-language pre-training (VLP) models , where joint representations are generally produced with a multimodal fusion Transformer network after extracting the visual and language features through unimodal encoders.
Mainstream VLP methods attempt to boost the pre-training in two ways: i. adopting more expensive video/text encoders to obtain more powerful unimodal features ii. designing heavier fusion networks to enhance the associations between modalities . Despite their advanced performance, they suffer from the increasing parameters, leading to significant computational inefficiency in downstream tasks.
To this end, we aim to design the simplest and most lightweight video-language model that gathers all capabilities in one, that is, learning video-language representations from their raw inputs in an end-to-end manner. In this way, we do not need any extra unimodal encoders (e.g., object detector in or ResNet visual encoder in ), and embed visual and text signals in a shared and unified model, termed as All-in-one Transformer in our paper. A recent study, ViLT , accomplishes such end-to-end joint learning in image-text pre-training under the presumption that the Transformer can process images in the same way as it processes text. However, we observe that it is non-trivial to embed videos using a unified Transformer that is also applied to process textual signals due to the unique challenge of modeling temporal information in video.
Existing works encode temporal representations in video-language pre-training via designing temporal attention layers or using temporal-aware visual encoders (e.g., 3D convnets in or video Transformer in ), which are all infeasible to be applied in our All-in-one Transformer as they are modality-dependent. To tackle the challenge, we introduce a novel, effective and flexible method, the temporal token rolling operation, to properly and gradually capture temporal representations in a non-parametric manner. Specifically, a proportion of the visual tokens in each frame of a sparsely sampled video clip are cyclic scrolling from frame to frame. Thus, visual tokens of a certain frame and its corresponding text tokens can “view” temporal dynamics from the rolling tokens of other frames through self-attention layers that naturally occur in the Transformer architecture. A sub-optimal solution to this issue is to aggregate all frames’ visual tokens for the self-attention layer, which is inflexible as it increases the time complexity by a factor of (given frames per video clip) compared to our method.
Our All-in-one Transformer is pre-trained towards the objectives of video-text matching and masked language modeling, following the common practice of Image-Language models . The pre-trained model is then fine-tuned to perform downstream video-text tasks. To further reap the modality-agnostic benefit of our All-in-one Transformer, we claim that our pre-trained model can not only encode the joint representations of video-language multimodal inputs, but also embed the unimodal features by feeding only video or text data into the Transformer. By fine-tuning our pre-trained model with a contrastive loss between video and text features, our All-in-one Transformer can play the role of an ordinary dual-stream framework on the downstream text-video retrieval tasks, realizing fast retrieval.
Our contributions are summarized as three-fold. (1) We introduce the simplest, most lightweight, and most efficient video-language model for pre-training, namely All-in-one Transformer, which is the first to capture video-language representations from the raw visual and textual signals end-to-end in a unified backbone architecture. (2) We elucidate and tackle the difficulties of applying a unified and shared backbone for multimodal video and text data, that is, how to properly process the unique temporal information of videos. A novel temporal token rolling operation is proposed to capture the temporal representations of sparsely sampled frames without any extra parameters or increasing time complexity. (3) Comprehensive experiments on four downstream video-text tasks of nine datasets fully demonstrate the superiority of our pre-trained All-in-one Transformer on both effectiveness and efficiency compared to recent mainstream methods . Moreover, benefiting from the modality-agnostic characteristic of our model, our pre-trained Transformer can be treated as a dual-stream framework to encode separate video and text features for highly efficient retrieval.
Related Work
Video-Language Pre-training. Pre-training on large-scale video-text pairs and fine-tuning on specific downstream tasks gradually becomes the standard paradigm in the video-language domain. Pre-trained models show strong transfer ability in a series of popular downstream video-language tasks including Text-to-Video Retrieval , Video Question Answering , and Visual Storytelling . Previous approaches leverage offline video and text features extracted from off-the-shelf visual and language backbones. Some recent methods including ClipBERT and Frozen have attempted to train models in an end-to-end fashion but still rely on well-trained visual encoders for feature extraction. In addition, these works mainly pre-train models on image-text dataset, like Google Conceptual Captions and Visual genome , and finetune the pre-trained models for downstream video-language tasks. In this work, we try to challenge this paradigm and focus on exploring effective strategies for pre-training on pure large-scale video-text benchmarks with only one network, and adapt our approach to various video-language downstream tasks.
Temporal Modeling in Video Understanding. Temporal modeling is a fundamental yet challenging topic in video representation learning. Several classic ideas including sparse sampling , 3D-type operations are proposed for temporal modeling in both convolution and Transformer architectures. 3D-type temporal modeling like Timesformer is extremely time-consuming because of the increasing number of sampled frames, which can be disastrous for large-scale pre-training techniques. Sparse sampling along the temporal dimension, a type of data augmentation proposed in TSN , has been widely adopted to train video backbones. Based on this, more related works try to shift channels among different frames for temporal modeling in action recognition. Inspired by these works, we try to roll video tokens for better alignment between modalities. This work focuses on parameter-free temporal modeling based on sparsely sampled frames without heavy 3D-type operation.
Unified Architecture Design for Multimodal Data. Recently the unified model, which is capable of processing either unimodal inputs or multimodal inputs with a shared encoder, has attracted a lot of attention. VATT trains a shared transformer with unimodal inputs to process Video, Audio and Text via multimodal contrastive learning and improves the performance of action recognition. Omnivore converts image, video, and single-view 3D modalities into embeddings that are fed into a Transformer model and trains the model with multi-task learning, which focuses on image/video/scene classification. In image-text pre-training, the early work Unimo solves both understanding and generation tasks with cross-modal contrastive learning. More recently, UFO also uses contrastive learning and employs a momentum teacher to guide the pre-training of a image-text shared encoder, which incurs large computational costs. Based on cross-modal contrastive learning, our work can also process unimodal inputs and perform retrieval tasks in a dual stream manner, which is very efficient. To the best of our knowledge, All-in-one Transformer is the first unified network for video-language pre-training.
Method
We propose All-in-one Transformer, a generic framework that enables end-to-end learning on video and language data, by learning joint representations directly from raw video pixels and raw text tokens, instead of the deeper feature from separate deep embedder. All-in-one has a succinct architecture as a Video-Language Pre-training model with parameter-free temporal modeling layer. In model design we making the pipeline as simple as possible so that the model can be used almost out of the box.
Fig.2 gives an overview of All-in-one framework. It adopts a sparse sampling strategy using only segments (one frame in each segment) at each training step, instead of full-length videos. Formally, we denote a video-text pair as (for video) and (for text sequence), where is the number of channels, is the resolution of each raw frame, is the length of input sentence and is the length of the word dictionary.
These text tokens are connected in series with vision tokens of each frame and the joint input is recorded as . Then is fed into stacked blocks and each block consists of a temporal Token Rolling layer, a multi-head self-attention layer and a multilayer perceptron (MLP). We initialize the weights of both self-attention and MLP layers from pre-trained ViT or DeiT . The visual features of each sampled frame are independently encoded using a visual backbone model to extract the relationship between the frame and its associated textual representation. Formally,
where means multiheaded self-attention, is multilayer perceptron and is short for Temporal Token Rolling Module. Independent predictions from all the sampled frames are fused together to derive a consensus at the video level. Formally, . For pre-training, objectives are calculated based on this consensus to learn model parameters.
2 Temporal Token Rolling
Motivation. In VLP, the common usage for temporal modeling is to add additional time attention layers in vision encoder or use the feature from deep off-the-shelf video encoder . However, these techniques are particularly designed for video and thus can not be applied to process text signal, as well as bringing a large amount of parameters. For example, by simply adding a temporal attention layer to each block of the Transformer, the model becomes a normal Timesformer with parameters increased from 86M to 121.7M (an increase of 42%). Thus, these techniques cannot be used in our unified framework and we turn to find new ways to learn temporal information with modest parameters.
Approach. A straightforward approach, denoted as “Flatten”, is to concatenate video and text tokens together and flatten into one tensor which will be fed into the self-attention blocks. Given text token with length and video token with length , we show the flatten version in Fig.4. However, as the self-attention layer has quadratic complexity, the computation cost will be , about times more than 1-frame All-in-one The length of text tokens much smaller than video tokens in general.. To overcome such limitation, we try to exchange information for different time segments in the token level. The proposed Token Rolling module is described in Fig. 2 (b). The tokens at different time stamps are denoted as different colors in each row. Along the temporal dimension, we roll parts of the token by 1, leaving the rest unchanged. Then the self-attention is computed in each tokens and treat each token in the same way. In this way, we reduce the computation complex to , around of the Flatten version.
Taking advantage of Token Rolling, longer dependencies between texts and videos are gradually modeled in deeper layers, which helps to learn better video-text alignment. We try to visualize the cross-modality attention weight density among text and video tokens in Fig. 4. For each text token, we compute the similarity by dot product to reveal its corresponding high-weight video tokens (more details are given in appendix). The baseline is All-in-one without rolling layers. We observe a severe inductive bias in the baseline, i.e., text tokens pay more attention to the centric tokens. By introducing the Temporal Token Rolling, these rolled tokens contribute more to the cross-modality interaction.
3 Training Objectives
We train All-in-one with two objectives commonly used to train VLP models: video-text matching (VTM), and masked language modeling (MLM). In addition, in order to overcome the disadvantage of low retrieval efficiency of one-stream, we introduce the video-text contrastive loss (VTC).
Video-text Matching. Given a paired video-text input, we randomly replace the paired video with a different video with the probability of 0.5 and ask the model to distinguish them. For the cls token of the last block, a single linear layer VTM head projects tokens to logits over binary class. We compute negative log-likelihood loss as our VTM loss.
Masked Language Modeling. MLM aims to predict the ground truth labels of masked text tokens from the other text and video tokens. Following the common practices , we randomly mask text tokens with the probability of 0.15 and model it as a classification task.
Video-text Contrastive. Inspired by the recent success of contrastive learning in visual-language pre-training , we also introduce this loss to our unified frameworks when fine-tuning for downstream video-text retrieval task. Specifically, for video-text pairs, we input video and text independently to the shared encoder to obtain high-level features. We then feed these features into a modality-specific projection head to project them into the shared embedding space. Following common practice , we use a symmetric (both text-to-video and video-to-text) contrastive loss based on these features. When doing retrieval tasks, we only need to extract unimodal features once.
Experiments
To explore model scalability, we use large-scale Webvid-2.5M , HowTo100M and YT-Temporal 180M for Pre-training. We evaluate All-in-one on four popular downstream video-language tasks: text-to-video retrieval, video question answering, multiple-choice and visual commonsense reasoning across 9 different datasets. We also provide extensive ablation studies to analyze the key factors that contribute to All-in-one’s success, with insights and qualitative results. More tasks and datasets are reported in the appendix.
When considering the generality of All-in-one, we consider using three configurations based on ViT and DeiT, as summarized in Tab. 1. To simplify, we use brief notation to indicate the model size: for instance, All-in-one-B/16 means the “Base” variant with input patch size. Following ViLT , we use the bert-base-uncased tokenizer to tokenize text inputs. We random sample frames and resize each frame to . Then the patch projection of All-in-one yields patches for each frame.
1.2 Pre-training & Fine-tuning.
Considering YT-Temporal 180M partially overlaps with HowTo100M , we pretrain on WebVid2.5M+Howto100M as default. If the model is trained on all three datasets, we named it as All-in-one *. Due to the storage limitation, we use the first half of the YT-Temporal 180M. We train all models using AdamW optimizer with a base learning rate of and weight decay of . The learning rate was warmed up for 10% of the total training steps and was decayed linearly to zero for the rest of the training.
For pre-training, we train All-in-one-S and All-in-one-B for 200K steps on NVIDIA A100 GPUs with a batch size of 16 per GPU 32 GPUs take 7 days total, 128 GPUs take less than 2 days total.. For All-in-one-Ti, we pre-train for 100K steps with a batch size of 32 per GPU, as we found it converges very fast. We adopt mixed precision technique to speed up the training process. As the domain gap between pre-train dataset and downstream visual commonsense reasoning dataset is large, we use batch size 512 and train for 100 epochs for this task. For the other downstream tasks, we train for 20 epochs with a batch size of 256. Note that downstream performance may be further improved if we customize the hyperparameters to each task.
2 Downstream Tasks
Datasets: In this work, we explore TGIF-QA , MSRVTT-QA and MSVD-QA . We experiment with 3 TGIF-QA tasks: Repeating Action and State Transition for multiple-choice QA, and Frame QA for open-ended QA. Both MSRVTT-QA and MSVD-QA are open-ended VQA. Evaluation: VQA requires answering questions according to the context of the video. For open-ended VQA, the answers are originally in free-form natural language, but it is a common practice to convert the task to a classification task by representing the answer with a class label. Following this practice, we add a two-layer MLP with hidden size 768 on the cls token. For multiple choice VQA (both the questions and candidates are sentences), we concatenate the question and candidates together, with [SEP] to distinguish them. We select the candidate with maximum output logit of VTM head as prediction.
2.2 Text-video Retrieval.
Datasets: MSRVTT , DiDeMo , ActivityNet Captions are utilized for this task. Evaluation. We initialize the similarity score head from the pre-trained ITM head, particularly the part that computes the true-pair logits. We train this task with both VTC and VTM since we find these two objectives boost each other. During inference, we simply feed each modality input independently and match pairs according to the cosine similarity of the output feature. In this way, we take advantage of high efficiency of dual-stream frameworks in retrieval.
2.3 Multiple-choice.
Datasets: In this task, we adopt MSRVTT multiple-choice test set and LSMDC multiple-choice test set . Evaluation: Given a video query and 5 candidate captions, the task is to find the one that fits the query out of 5 possible candidates. The correct answer is the ground-truth (GT) caption, and four other negatives are chosen from other captions that have different activity-phrase labels from the correct answer. We initialize the VTM head from the pretrained model on-top of the CLS token. During the train, we simply concat each candidate with the given video together as input and only the correct answer is positive pair while the others negative pairs. We tune the model with cross-entropy loss to maximize the scores on positive pairs.
2.4 Visual Commonsense Reasoning.
VCR is a task and dataset where models must answer commonsense visual questions about images. This task test our model’s ability to transfer its video-level understanding to single image. To solve this challenge task, VCR provides additional information to models (in the form of bounding boxes around entities), and explicit groundings between those entities and references in questions. Following previous efforts , we incorporate the location and identity information by drawing mask around the referenced entity.
3 Comparing to State-of-the-art
In this experiment, we compare three variations of our All-in-one to state-of-the-art methods from the literature. For multiple-choice VQA, we evaluate our All-in-one on two sub splits of TGIF-QA and report the result in Tab. 2. We find All-in-one especially good at this type of VQA. With ouly 1 frame input, our All-in-one-B outperforms previous VIOLET about 5.8% on the Action subset. Interestingly, we find more frames not benefit to Action and Transition but FrameQA. We also report the result of All-in-one on the three open-ended datasets. Surprisingly, even though Just-Ask is specifically designed for VQA and pretrined on large scale HowToVQA69M, our method still achieves a similar even better result than Just-Ask on MSVD-QA.
3.2 Retrieval Tasks.
In this experiment, we fine-tune All-in-one on MSRVTT, ActivityNet Caption and DiDeMo datasets. Tab. 3 summarizes results on text-to-video retrieval. In Tab. 3(a), All-in-one achieves significant performance gain over existing methods on MSRVTT retrieval in both 9K and 7K train setting. Compre with these related works, we only use one Cross-modality Encoder and the parameter is half of the Frozen . All-in-one even leads to 2.1% relative improvement on R@1 when compare with OA-Trans , which use additional offline object feature and only focus on retrieval. When adopt on LSMDC and DiDeMo dataset, our method also show competitive result.
3.3 Multiple-choice.
Tab. 4 shows that All-in-one improves ClipBERT model by 3.2% on accuracy, on MSRVTT multiple choice test task. We also report the zero-shot performance for comparison, we find zero-shot accuracy already close to JSFusion in MSRVTT multiple choice with only three frames as input.
3.4 Visual Commonsense Reasoning.
After pre-training, we use a visual reasoning task to test the generality ability of our model. Our results on the VCR dataset, in comparison to other models at the same (“base”) scale, are given in Tab. 6. Moreover, to utilize identity information, we also mask the different identity with different color (as shown in Fig. 5). We observe our model outperforms MERLOT clearly in the same setting with different sources of data.
4 Analysis of Temporal Token Rolling
To better study Token Rolling, we also train our All-in-one in four different settings: Single Frame: Pre-training and inference with 1 frame. Time Average: Pre-training with 1 frame but inference with 3 frames. Time Average : Pre-training and inference with 3 frames. Channel Shift: we replace each Token Rolling layer with channel shift operation. Flatten: As presented in Sec. 3.2. We observe: i. Pre-training with multiple frames are essential for SSL tasks. e.g, from 35.16 to 47.33 on LSMDC. ii. The Token Rolling boosted an amazing 5.42% improvement over time average baseline. Compared with channel shift , the rolling on tokens also show superior performance. We guess VLP require learn alignment between patches and the channel operation will erase this boundary. iii. Even the computation complex of Flatten is three times as Token Rolling, the performance of Token Rolling is slightly better than Flatten in both setting. As we discussed in Sec. 3.2, the benefits might come from the Token Rolling is a natural extension of self-attention among patches.
4.2 Ablation on the Rolling Token Ratio.
In order to understand how many tokens are needed to roll during pre-training, we conduct an ablation study as indicated in Tab. 8 (b). We follow the All-in-one protocol in the pre-training setup except for a smaller 1024 batch size and 100K steps. Compared to the temporal average baseline (ratio equals to 0), we observe an amazing 5.69% improvement. The benefits come from more effective temporal modeling.
4.3 The Variations of Initialization.
To answer the question if initialization is crucial for large-scale VLP. We initialize All-in-one with three versions: Scratch, ImageNet and ImageNet-21K. We report the results in Tab. 6 with different train iterations and make following observations: i. Train from scratch convergence slower than train from ImageNet pretrained model. ii. The combination of Webvid2.5M and Howto100M is large enough to train the model from scratch. When training our All-in-one for 800K steps, we find the train from scratch is close to the ImageNet-21K initialized version in both Pre-training and downstream evaluation.
5 Objectives of Retrieval
To study the effect of objectives during fine-tuning, we experiment with three different combination of objectives. As shown in Tab. 7, the VTC loss have high R@1 and VTM is more effective in R@5 and R@10. With the combination of three objectives, All-in-one can achieve best performance in all measurement. Notice that these methods are trained in a dual-stream way.
6 Visualization
To better understand the pre-trained All-in-one, we analyze its internal representations. Specifically, given paired ground truth text and raw video, we mask some keywords (both verb and nouns) and ask the model to predict these masked words and further find out which video patch has strong correlations with the masked words. We use optimal transports to calculate the correlation between video and text. We only show the attention weight that is larger than the given threshold and give some examples of cross-modal alignment in Fig. 6. We find the model can predict correct nouns and verbs in most cases. Sometimes, it predicts the wrong word but with a similar meaning to the correct word. e.g. “guy” and “man”. Benefiting from temporal modeling, we also find that the model attends to the motion regions for verbs like “waving” and “walking”.
Conclusions
In this paper, we present the first unified end-to-end Video-Language Pre-training architecture with raw video and text as input, All-in-one Transformer. By learning only one fusion network, All-in-one is able to complete with a large number of counterparts equipped with additional heavy off-the-shelf video visual embedding networks and holds promise for the future. We hope that the VLP community will focus more on lightweight end-to-end modal interactions within Transformer modules, rather than only on heavier single-modality embedders or larger fusion models. While these initial results are encouraging, this new design of unified video-language interaction also brings new challenges, in particular fine-grained word region alignment. Furthermore, the temporal modeling is still not fully explored and we also encourage future work to use All-in-one for single-modality tasks.
Acknowledgement
This project is supported by the National Research Foundation, Singapore under its NRFF award NRF-NRFF13-2021-0008. We would like to thank David Junhao Zhang for his kindly help on Transformer training.
References
Appendix
In this appendix, we first evaluate the model scalability and provide more ablation studies about All-in-one. Then we transfer All-in-one to more downstream tasks and datasets. At last, we provide retrieval efficiency and more visualization analysis.
Appendix A Model Scalability
In this experiment, we evaluate the model from All-in-one-Ti to All-in-one-L using three assessment tasks: Fine-tune, Zero-shot and Linear Probe. Zero-shot means we directly test the pretrained model on downstream task without fine-tining, Linear Probe means we frozen the overall model and only the last linear layer is learned on downstream tasks. We varying the model size from 13M to 320M and do evaluation on 10 different datasets. For fair comparison, the pre-training and fine-tuning settings are consistent for models of different scales. The results are presented in Fig. 7 and we make the following observations:
i. For Zero-shot and Linear Probe task, we observe large model leads to better result in general. ii. However, we find sometimes All-in-one-L on the Fine-tune task leads to worse result than All-in-one-B in several benchmarks (circled). We show the train curve of MSRVTT-QA in the right of Fig. 7. Since MSRVTT-QA only contains 10K video-text pairs, we find that the model severely overfits the dataset with few iterations. For more large dataset like TVQA and TGIF-QA (ten times larger than MSRVTT-QA), All-in-one-L still lead to better results. We conclude that simply pursuing larger models is not suitable for all cases, especially for fine-tuning on small scale, and All-in-one-B is a better choice in most cases, as a trade-off between parameters and performance .
Appendix B Ablation Study
In this work Masked Language Modeling (MLM) and Video-text Matching (VTM) are modeled as binary and multiple classes classification tasks, correspondingly. So we also report top-1 classification accuracy (%) for these two pre-training objectives as reference in this section.
In addition to the parameter-free time token rolling operation proposed in this work, we also try different ways for temporal modeling: i. TimeSformer: For each self-attention block in All-in-one, we add a additional divided space attention and time attention before multi-head self-attention layer. As shown in the middle of Fig. 8. ii. Decouple Attn: For text modality and visual modality, we conduct self-attention independently first and then concatenate them again for cross-modality attention. As shown in the right of Fig. 8.
Pre-training on Webvid2.5M + HowTo100M, we report both the pretrain and downstream evaluation performance in Tab. 9. With more parameters in visual processing, we find that TimeSformer and Decouple Attn are particularly good at Video-text Matching, but not good at Masked Language Modeling. However, we find that these methods are difficult to train, cost about 2-3 times more expensive than All-in-one-B, and show worse results on downstream zero-shot tasks.
B.2 Do we need to sample more frames?
To further understand the relation between the number of frames and the quality of learned representation, we conduct pretrain and finetune experiments by varying the sampled frames for train. Considering the memory consumption, we vary the frames from 1 to 16. The difference between the default setting is that we pretrain on Webvid2.5M. We report both the pretrain VTM and MLM accuracy and downstream zero-shot multiple-choice accuracy on both LSMDC and MSRVTT in Fig. 9. We find that more frames leads to better results in general and All-in-one is already close to the best results when frames equals to 3. To balance the computation cost and performance, we use 3 frames as default.
B.3 Position Embedding and Modality Type Embedding
In this experiment, we explore the effect of Position Embedding and Modality Type Embedding. The results are given in Tab. 10, we observe spatio-temporal position embedding help the VTM and modality type embedding helps the MLM pretext. The combination of these two embedding leads to better results on both downstream zero-shot multiple choice result with limited parameters.
Appendix C In-depth Analysis of Token Rolling
In this experiment, we explore where to add our temporal rolling layer. Specifically, All-in-one-B contains 12 self-attention blocks and we add the Token Rolling Layers from the beginning, the 3rd block and the 6th block. The results are report in the left of Tab. 8. We observe more temporal token rolling layers leads to better representation.
C.2 The Sampling Strategy of Rolled Token
In this experiment, we explore the sampling strategy of sampling token. We try three versions: i. Random Selection: Random select 25% tokens. ii. Varying with layers: For the first layer, we roll the first top 25% and then the second top 25%. iii. Select a block that contains 25% tokens. As shown in the right of Tab. 11, selecting the 25% tokens leads to best result. We guess this is due to the multiple perceptron is position sensitive and the random selection or varying layers will lost this information. In this work, we adopt block selection as default.
Appendix D Transferability Evaluation
To evaluate the transfer ability of our model on single-modality task. We transfer the learned representation to downstream linear probe result on K400 and HMDB51 dataset. Specifically, we frozen the overall unified model and only learns linear layers based on the cls token of the last layer. By pre-training model on these two datasets, we compare the base model with Time Average and previous best method Frozen .
The linear probe results are given in Tab. 12. We observe the number of frames have large impact on this task. When adopt same 8 frames, our All-in-one-B clearly outperforms Frozen especially on large-scale K400 dataset.
D.2 Extension to Egocentric Video
Ego-4d is a egocentric dataset that has large domain gap with our third-view video from Youtube. We show some examples from Ego-4d in Fig. 10 and test multiple-choice (5 choices) task on this dataset. We report both the zero-shot result and fine-tune result in Tab. 13. Compared to other multiple-choice benchmarks such as LSMDC and MSR-VTT, zero-shot accuracy is lower, but our All-in-one still outperforms Frozen clearly by half the parameters in this challenge benchmark.
Appendix E Complexity Analysis of Retrieval
Due to the specifiy design for contrastive loss, Our All-in-one has a very fast inference running time even for retrieval on million-scale datasets. We use the popular similarity search/ranking library FAISS-GPU open-source library on a server with 8 A100 GPUs and 88 Kernel Intel(R) Xeon(R) Platinum 8255C CPU @ 2.50GHz. Given a new query, Table 14 below shows the time needed for visual encoding, textual encoding, and similarity ranking (1st row for thousand-scale and 2nd row for million-scale). Given a new query, the total search time on HowTo100M is (12.05 + 143.25 = 155.3) ms for text-to-video retrieval and (33.16 + 143.25 = 176.41) ms for video-to-text retrieval, which is acceptable in practice.
Appendix F Visualization (Cont’d)
In addition to the person-centric videos, we also visualize some samples about outdoor scene and objects in Fig. 11. For find our model can make correct prediction of masked words.