VindLU: A Recipe for Effective Video-and-Language Pretraining

Feng Cheng, Xizi Wang, Jie Lei, David Crandall, Mohit Bansal, Gedas Bertasius

Introduction

Fueled by the growing availability of video-and-text data and advances in the Transformer model design , the last few years have witnessed incredible progress in video-and-language (VidL) understanding . Since the initial transformer-based models for VidL, such as ClipBERT , the text-to-video retrieval accuracy has improved from 22.0%,22.4%22.0\%,22.4\%, and 21.3%21.3\% on MSR-VTT , DiDeMo , and ActivityNet to >45%>45\% R@1 accuracy on all three of these datasets, thus, marking an extraordinary relative improvement of more than 100%100\% in less than 22 years.

At the same time, the model architectures and pretraining/finetuning protocols used by modern VidL approaches have become significantly more complex and specialized over the last several years. As a result, it is increasingly difficult to reproduce, analyze and compare most recent VidL frameworks. For example, several recent approaches propose new architectures, new initialization strategies, pretraining objectives, pretraining datasets, and optimization protocols. Due to the large computational cost of ablating all these factors, it is difficult to understand which components are critical to the success of the proposed frameworks. Similarly, the key success factors of many other recent VidL approaches are also often obfuscated, which hinder future research.

In Table 1, we illustrate the complexity of modern VidL frameworks by dissecting them along multiple dimensions, including temporal modeling schemes, multimodal fusion modules, pretraining objectives, the source of the pretraining data, and the number of frames for pretraining, finetuning and inference. Based on this analysis, we observe that there exist significant differences among these VidL methods. Unfortunately, it’s not clear which differences are important for the overall VidL performance and which are not.

The recent METER work studies a subset of these components in the context of image-language modeling. However, their analysis is limited to images and, thus, ignores various aspects related to video modeling, such as spatiotemporal architecture design, video pretraining objectives, video pretraining data, and video-specific finetuning/evaluation protocols such as the number of frames. As we will show in our experimental section, many of the findings presented in the image-based studies do not hold for video. Beyond image-based analysis, we note that the concurrent work in conducts an empirical study of VidL transformers. However, unlike our work, which covers a broad range of VidL design factors, their analysis is focused predominantly on masked visual modeling objectives, which we also study in this work.

Our main objective in this work is to answer the question “What are the key steps needed to build a highly performant VidL framework?” To do this, we conduct a thorough empirical study that demystifies the importance of various VidL design choices and ultimately leads to a VidL framework that achieves state-of-the-art results on various VidL benchmarks. Using our empirical insights, we then develop a step-by-step recipe for effective VidL pretraining. Our recipe, dubbed VindLU (VIdeo aND Language Understanding), starts from a standard Vision Transformer (ViT) and uses a simple progressive expansion scheme where at each step, we investigate a particular aspect of VidL framework design (e.g., architecture, pretraining objective, pretraining data, etc.), and choose the best performing option. In particular, we study the following VidL design components: (i) the spatiotemporal architecture design, (ii) the multimodal fusion schemes, (iii) the pretraining objectives, (iv) the source of the pretraining data, (v) finetuning/inference protocols, and (vi) scaling of the data and model. We present our recipe in Fig. 1.

The key findings of our empirical study include:

Contrary to the conclusions of several prior works that a single frame is sufficient for VidL modeling, we discover that temporal modeling using multiple frames leads to a significant improvement over the spatial-only baselines (+6% averaged video retrieval accuracy on MSR-VTT, DiDeMo, and ActivityNet).

Multimodal fusion module that incorporates video features into text is critical for good VidL performance (+3.6%). Conversely, we find that adding text features to the video representation is not useful.

Masked language modeling objective significantly improves performance (+6.2%). However, to obtain such gains, a BERT-like language model pretrained on this objective is needed for initialization. Masked video modeling objective brings an additional +1% improvement.

Pretraining jointly on images and videos is beneficial (+2.7%). Also, contrary to prior methods , we find multi-stage training unnecessary.

Pretraining with a small number of frames (e.g., 4) is sufficient and it can significantly reduce the computational cost of large-scale pretraining. Pretraining with more frames does not lead to a substantial performance boost.

Compared to many recent CLIP-based VidL approaches , our recipe achieves comparable or even better performance with 20×\bf{20\times} less pretraining data.

Our final model, trained using our VindLU recipe, achieves state-of-the-art results on several VidL benchmarks. Specifically, on the video retrieval task, our method achieves 46.5%, 61.2%, 55.0% R@1 accuracy on MSR-VTT, DiDeMo, and ActivityNet, outperforming the state-of-the-art by 7.8% and 6.1% on the latter two datasets. Also, our approach obtains state-of-the-art video question-answering results on ActivityNet-QA, MSRVTT-QA, MSRVTT-MC and TVQA, where we achieve top-1 accuracy of 44.7%, 44.6%, 95.5%, and 79.0% respectively.

We want to make it clear that, in this paper, we do not claim technical novelty behind any of the individual design choices (i.e., different subsets of these design choices were already used by prior VidL methods as shown in Table 1). Instead, our main contribution, which we believe might be equally if not more important than proposing yet another specialized or obfuscated VidL model, is to investigate these components collectively and validate their importance. We also do not claim superiority over previous methods (despite better results). Due to the implementation complexities of each method, fair and complete comparisons are difficult and not our intent. Instead, we hope that our recipe for building an effective VidL framework will provide useful insights for future research on VidL understanding. To enable the VidL community to build on our work, we release our code and pretrained models.

Related Work

Image-and-Language Pretraining. Recent years have witnessed remarkable progress in image-and-language pretraining . However, most modern methods such as ViLBERT , UNITER , CoCa , LEMON , BEiT-3 typically employ complex transformer-based architectures and pretraining objectives. As a result, it is difficult to decipher which components are critical for good performance. A recent empirical study on image-language modeling METER studies a variety of components, including the choice of a vision encoder, multimodal fusion schemes, and pretraining objectives. However, since their analysis is done exclusively on images, it’s unclear whether these findings generalize to video. The analysis of METER also ignores many video-specific design choices such as temporal modeling schemes, video pretraining objectives and data, and video-specific finetuning/inference protocols. In comparison, our work thoroughly studies all of these components, the result of which is a detailed step-by-step recipe for effective video-language pretraining.

Video-and-Language Pretraining. In recent years, the large-scale VidL pretraining has shown strong transfer learning ability to downstream VidL tasks such as text-to-video retrieval , video question answering , video captioning , etc. Several methods achieve impressive results by building on the popular image-language pretrained model CLIP . Additionally, several recent approaches propose more sophisticated VidL frameworks to achieve comparable performance as CLIP-based methods without large-scale CLIP pretraining. However, with the impressive results, these methods also require more complex architectures and specialized video pretraining protocols (as shown in Table 1). The complexity of these frameworks and the large computational cost of VidL pretraining makes it challenging to decipher which VidL framework components are truly needed for good performance. Moreover, unlike in the image-language domain, there are few empirical studies investigating various VidL design components collectively. For instance, the concurrent work of Fu only studies masked video modeling pretraining objectives and is based on a slightly older VIOLET method. Furthermore, the recent works focus predominantly on spatial biases in modern VidL benchmarks. In contrast to these prior approaches, our work aims to investigate the importance of a broad range of factors in VidL framework design. We then use our empirical insights to provide a detailed step-by-step recipe for effective VidL pretraining.

A Recipe for Video-Language Pretraining

In this section, we describe our recipe for video-and-language (VidL) pretraining. We begin with a standard image transformer (e.g., ViT ) and progressively expand it to a model that achieves state-of-the-art results on various VidL datasets and tasks. At each step of our recipe, we study how various design choices affect VidL performance. In particular, we are interested in answering the following questions about the VidL pretraining design:

Does a VidL model benefit from a temporal modeling capability, especially considering that most VidL benchmarks are spatially biased as demonstrated by several prior methods ? If so, what is the best mechanism for temporal modeling?

What is the most effective way to do multimodal fusion? Some prior approaches use bidirectional whereas others employ unidirectional (e.g., text-to-video or video-to-text) multimodal fusion modules. Which of these fusion schemes works the best?

Which pretraining objectives are most useful for VidL representation learning? Previous methods adopt many pretraining objectives including video-text contrastive (VTC), video-text matching (VTM), masked-language-modeling (MLM), and masked-video-modeling (MVM). How important are each of these objectives? Are they complementary to each other?

What pretraining data is most useful for training VidL models? Should we train VidL models only on the video data or jointly on images and videos? If so, how do we do this effectively? Prior works propose a variety of different pretraining protocols (e.g., a single-frame training, curriculum learning, joint multi-frame pretraining, etc.). Which of these is the most effective?

How many frames are needed for pretraining, fine-tuning, and inference? Several recent approaches claimed that single frame pretraining is sufficient while others proposed to pretrain their models with 8 or even more frames. Furthermore, should we finetune and test the pretrained VidL models using the same number of frames as during pretraining? Is it helpful to use more frames during fine-tuning and inference?

Motivated by these questions, we next present our recipe while also studying these questions in more detail.

Image Transformer Baseline. We start with a standard ViT-B/16 transformer trained on single frames of the WebVid-2M dataset. For text encoder, we use BERT throughout all of our experiments. Formally, given the paired video and text input (v,t)(v,t), the image transformer randomly selects a single frame from the video as input to extract the video embeddings. A text encoder encodes the text tt to extract the text embeddings. We then use a video-text contrastive (VTC) loss to maximize the agreement between the paired video and text embeddings as in . Following , we use BEiT initialization for our image transformer, whereas the text encoder is initialized with BERTbase\text{BERT}_{base}.

Experimental Setup. As our initial pretraining data, we use WebVid-2M unless noted otherwise. Afterward, we finetune and evaluate our pretrained model on the three popular text-to-video retrieval datasets: MSR-VTT, DiDeMo, and ActivityNet-Captions, which include both short and long videos. As our evaluation metric, we report the averaged Top-1, Top-5, and Top-10 text-to-video retrieval accuracy across these three datasets. As shown in the Fig. 2, our Image Transformer baseline achieves an average accuracy of 50.4%.

Over the next several subsections, we progressively expand this baseline by adding more components of increasing complexity. In particular, we start by incorporating (i) temporal modeling blocks, (ii) a multimodal fusion encoder, and (iii) additional pretraining objectives. Afterward, we investigate the choice for the (iv) pretraining data, (v) finetuning and inference protocols, and (vi) dataset and model scaling schemes. We would like to note that due to the large computational cost, we cannot ablate the order of the steps in our recipe. Thus, the order of the steps is primarily determined by the computational cost (i.e., the steps that can be implemented most efficiently are studied first then, moving to the more computationally costly steps).

Step 1: Temporal Modeling

In the first step of our recipe, we extend our initial image transformer to video via a temporal modeling mechanism, which enables training our model on multiple frames. Such a temporal modeling mechanism would enable training our model on multiple frames for more robust VidL spatiotemporal representation learning. For compactness, in this part of our empirical study, we include the analysis of the four commonly used temporal modeling schemes. More temporal modeling baselines can be found in Appx. C.2.

Mean Pooling (MP). In this variant, the visual encoder processes input frames independently and averages their frame-wise scores for the video-level score as in .

Late Temporal Attention (L-TA). Following we use a late temporal modeling scheme by attaching 2 Transformer layers to an image encoder, which then aggregates temporal information across all input frames.

Temporal Convolution (TC). Many previous methods used 3D convolutions for temporal modeling. To validate its effectiveness, we inject a TC block, consisting of a linear down-projection layer with hidden size 384, a depth-wise 3×1×13\times 1\times 1 convolution as in , a ReLU activation, and a linear up-projection layer, before the spatial attention to each Transformer Layer.

Temporal Attention (TA). Inspired by TimeSformer , we experiment with divided space-time attention, which we insert before spatial attention as in .

As shown in the upper part of Fig. 2 and the Table below, the temporal modeling capability is critical for good VidL performance. This is indicated by a +6.3%\bf+6.3\% accuracy boost of our temporal attention variant (TA) over the spatial-only baseline. We also observe that late temporal modeling (L-TA) has nearly no effect. We conjecture that this is due to the limited temporal modeling capacity (i.e., only two layers) and the lack of temporal fusion in the early layers. Lastly, our results suggest that TA outperforms TC by 2.1%, which might indicate that long-range temporal attention is more useful than local 3D convolutions.

Interestingly, we note that our findings are contrary to the conclusions of several recent methods claiming that temporal modeling is not needed for many VidL tasks. Upon experimenting with the publicly released models of , we found that the temporal variants of their approach performed consistently better than the spatial-only variants, further strengthening our conclusions. We conjecture that even on the spatially-biased datasets, temporal modeling might be useful for resolving spatial ambiguities caused by appearance variations across different frames.

Takeaway #1: For all subsequent experiments, we adopt Temporal Attention (TA) as our temporal modeling mechanism and pretrain our model with 4-frame inputs unless otherwise noted.

Step 2: Multimodal Fusion Encoder

Building on the model from Step 1 (Fig. 3a), we next analyze the role of multimodal fusion modules. The purpose of the multimodal fusion encoder is to fuse multimodal cues from video and language for a more discriminative VidL feature representation. As shown in Fig. 3, we experiment with several variants of multi-modal fusion encoders:

Video-to-Text Multimodal Fusion (V2T-MF). As illustrated in Fig. 3b, V2T-MF injects relevant video cues into the textual features using Cross-Attention. For a fair comparison with previous baselines , we do not add any extra layers but instead re-purpose the last mm layer of our text encoder for V2T fusion. Specifically, a cross-attention operation is inserted into each of the mm last layers in the text encoder between Self-Attention and MLP. This scheme was also previously used by .

Text-to-Video Multimodal Fusion (T2V-MF). Similar to V2T-MF, we build T2V-MF (Fig. 3c) by re-purposing the last mm layers of the vision encoder and utilizing cross-attention to incorporate text cues into the video features.

Bidirectional Multimodal Fusion (B-MF). Instead of using unidirectional multimodal fusion modules, several prior approaches concatenate the visual features and text features and then feed them jointly to a subsequent mm-layer multimodal Transformer. However, this is often infeasible in the video domain due to a large number of input frames and hence, large computational cost. Thus, instead, we implement B-MF (Fig. 3d) by combining the T2V-MF and V2T-MF, which reduces the space and time complexity from O((L+S)2)O((L+S)^{2}) to O(L×S)O(L\times S) where L,SL,S is the number of video tokens and text tokens respectively.

To train each multimodal fusion encoder variant, we add the video-text matching (VTM) loss objective (described in the Sec. 3) as was done in several prior approaches . In the table below and Figure 2, we present our analysis. Based on these results, we report that the V2T-MF scheme performs the best (i.e., +3.6% improvement). Surprisingly, we observe that the reverse, T2V-MF scheme, leads to a substantially decreased performance (-1.3%). We conjecture that predicting the matching video-text pairs using a pretrained language rather than a visual representation is easier. Lastly, the bidirectional fusion scheme, B-MF, yields no improvement compared to V2T-MF. We conjecture that this happens because of the poor performance in the T2V-MF branch.

Takeaway #2: For our remaining experiments, we use V2T-MF as our multimodal fusion encoder.

Step 3: Pretraining Objectives

Building on the model from Step 2, we next study the following pretraining objectives:

Visual-Text Contrastive Learning (VTC). VTC aims to learn independent representations for video and text by maximizing the agreement between positive (visual, text) pairs while minimizing the agreement between negative pairs. Note that this objective is already used in previous steps, and thus, not included in Figure 2.

Visual-Text Matching (VTM). VTM objective is implemented as a standard cross-entropy loss that encourages a VidL model to produce binary predictions indicating whether a given video-text pair matches. Following , we attach this loss to our multimodal fusion encoder and use hard negative mining during training as in . The VTM objective is already used in Step 2 (i.e., the multimodal fusion step) and thus, not included in Figure 2.

Masked Language Modeling (MLM). MLM objective aims to predict the masked words by leveraging information from both visual and textual features. To implement this pretraining objective, We mask 50% text tokens using the same masking strategy as in BERT and attach a linear layer on top of our text-to-video multimodal fusion encoder (T2V-MF) to predict the masked words. See Appx. C.1 for masking ratio ablation.

Masked Video Modeling (MVM). Just like MLM, the MVM objective aims to recover the masked tokens but in the video modality. This pretraining objective has been recently adopted by many self-supervised learning video methods . To implement MVM, we apply a linear layer on top of the vision encoder and predict the masked tokens. Following , we randomly mask 75% tokens and predict the masked tokens quantized using discrete variational autoencoder (dVAE) .

Based on the results in the Table below and Fig. 2, we observe that the MLM pretraining objective leads to a substantial boost in performance (+6.2%). Furthermore, we note that adding MVM loss further improves the accuracy by 1%. Interestingly, our finding is contrary to the conclusions in the image-based analysis of METER , which finds that MVM objective applied to images substantially degrades the performance. We hypothesize that videos are more redundant compared to images, which might make the optimization easier, thus, leading to a performance boost. However, adding the MVM objective slows the training by about 40% (due to additional forward and backward passes). Thus, to speed up the training, we don’t use MVM loss in our remaining experiments.

Takeaway #3: For the remaining experiments, we use VTC, VTM, MLM as our pretraining objectives.

Step 4: Pretraining Data

In this section, we analyze the effect of (i) the pretraining data, and (ii) pretraining protocols.

Datasets. Several recent methods suggest that jointly pretraining on image and video data can lead to better performance. To investigate this, we consider an additional image-based CC3M consisting of 3M image-text pairs. Specifically, we experiment with pretraining our framework on the (i) image-only (CC3M), (ii) video-only (WebVid2M), and (iii) joint image and video (CC3M + WebVid2M) datasets. When pretraining on images, we replace our previously introduced temporal attention module with an identity connection. This enables our model to be easily applied to both images and videos.

As shown in the Table below and Fig. 2, training on videos is more beneficial than training on images (+2.7%), which makes sense as all of our downstream applications involve video. Furthermore, we observe that jointly pretraining on both images and videos leads to an additional 2.7% boost in performance. This suggests that the spatial and temporal cues are complementary and that a stronger spatial representation can boost VidL performance.

The Number of Input Frames for Pretraining. Prior approaches use a different number of input frames for pretraining (i.e., from 1 to 16). Thus, we next study how many frames are needed for effective VidL pretraining. The models are pretrained jointly on image and video (CC3M + WebVid2M) datasets. From the Table below and Fig. 2, we observe that multi-frame pretraining using 4 frames leads to 1.7% improvement compared to a single-frame pretraining. However, we also observe that the performance saturates with 4-frame inputs while the computational cost of pretraining with more frames increases significantly. In particular, we note that pre-training with 4-frame inputs leads to a speedup of 2.5×\times compared to pretraining with 16-frame inputs. Thus, our finding is useful as it can save lots of computing power and speed up the development of future research.

Multi-stage Curriculum Pretraining. Lastly, we also validate the necessity of multi-stage curriculum pretraining, which was used in several prior VidL approaches . Specifically, we experiment with two different pretraining protocols: (i) a two-stage pretraining that first trains a model for 10 epochs using single frames, and then uses a multi-frame training for 5 additional epochs using 4-frame inputs, and (ii) a three-stage pretraining that builds on (i) by adding a third stage where the model is trained for additional 3 epochs using 8-frame inputs. The model is pretrained jointly on image and video datasets. Our results in the Table below and Figure 2, indicate that multi-stage pretraining does not lead to any significant boost in performance, which is contrary to the findings of prior approaches . We conjecture that this might happen because prior approaches train their model for only several epochs at each stage whereas we train it until convergence (10 epochs for the first stage). We also note that compared to the 4-frame one-stage pretraining (described above), the two-stage 1→41\to 4 has a comparable pretraining cost as the latter model is trained for more epochs.

Takeaway #4: We adopt a single-stage pretraining on joint image and video datasets while using 4-frame inputs.

Step 5: Finetuning & Inference

Existing methods typically use the same number of frames either between pretraining and finetuning or between finetuning and inference . Here, we study whether we can use a different number of frames at different phases.

Finetuning. We experiment with finetuning our 4-frame pretrained model with K=1,4,8,12,24,32{K=1,4,8,12,24,32}-frame inputs while using MM frames during inference. We use M=12M=12 for all K≤12K\leq 12 and M=KM=K for K>12K>12 as we found inference with more frames leads to higher performance. Based on the results in the Table below, we observe that while finetuning with more frames leads to higher accuracy (70.5%) the performance saturates with about 12 frames. We also note that finetuning with a single-frame input is 22.4×\times faster than with 32-frame inputs but has a 5% lower accuracy. On the other hand, finetuning with 12-frame inputs yields only 0.3% lower accuracy but 2.6×\times speedup compared to finetuning with 32-frame inputs. Therefore, due to the favorable accuracy-cost tradeoff, we finetune most of our models with 12-frame inputs.

Inference. Next, we also experiment with using 12, 24, 32, 64 frames for testing our 4-frame pretrained and 12-frame finetuned model. In the table below, we report the averaged accuracies on the DiDeMo (D) / ActivityNet (A) datasets, which contain longer videos. Using more frames for inference is beneficial, but the performance also saturates quickly, and the inference speed slows down rapidly.

Takeaway #5: Considering the trade-off between computational cost and accuracy, we use 12 frames for finetuning and inference on all datasets except ActivityNet. On ActivityNet, we use 12 and 32 frames for finetuning and inference.

Step 6: Scaling Up

As our last step, we investigate scaling up the pretraining data and the model size.

Pretraining Data. For the pre-training data, we experiment with (a) adding 12M images from CC12M for a 17M Corpus, and (b) additional 10M videos from WebVid10M for a 25M Corpus. The results in the Table below and in Figure 2 indicate that scaling our corpus from 5M→17M5M\rightarrow 17M improves the downstream VidL performance by 2.2%. Furthermore, scaling the corpus from 17M→25M17M\rightarrow 25M leads to an additional boost of 1.2%.

Model Size. In the Table below, we also experiment with scaling the video encoder (ViTbase→ViTlarge\text{ViT}_{base}\to\text{ViT}_{large}) or text encoder (BERTbase→BERTlarge\text{BERT}_{base}\to\text{BERT}_{large}). Due to the large computational cost, we could only conduct these experiments on the 5M corpus. We report that scaling the vision encoder brings larger improvement ( +3.0%) than scaling the text encoder (+1.0%).

Final Takeaway: Our final scaled-up VindLU model improves the initial image transformer baseline by 23.2%.

Other Useful Empirical Tips

Isotropic vs Pyramid-based Vision Encoder. Pyramid-style ViTs that use downsampling along the spatial dimension (e.g., Swin, MViT) have shown stronger performance than isotropic ViTs (vanilla ViT) on many image/video classification tasks. Thus, several recent VidL approaches adopt pyramid ViTs as their vision encoders. However, in our study, we find that isotropic ViTs tend to have better performance. Specifically, in Tab. 9, we show that a ViT-based encoder outperforms VideoSwin by 1.6%. We hypothesize that this might happen because isotropic ViTs preserve more fine-grained spatial information needed for various VidL tasks.

A Linear Scaling Rule. Linear scaling strategy has been extensively used for large-scale pretraining on image/video classification tasks. However, in our setting, we observed that the linear scaling rule leads to similar or worse results (See Table 3). Therefore, for all of our experiments, we use a fixed learning rate (1e-4) for all batch sizes.

Initialization. We also found that the initialization of various modules in our model is critical for good VidL performance. In particular, we note that to make MLM and MVM pretraining objectives effective, we need to use text and video encoders pretrained with these objectives in a self-supervised manner (e.g., BERT and BEIT respectively). Otherwise, the performance will drop significantly (∼\sim5% averaged R@1,5,10 accuracy drop on MSR-VTT, DiDeMo, ActivityNet datasets).

Experimental Results

We validate our VindLU recipe on two mainstream VidL tasks. See implementation details in Appx. A and dataset descriptions in Appx. B.

Text-to-Video Retrieval. We compare our results with existing methods on three spatially-biased datasets MSR-VTT, DiDeMo, and ActivityNet and two temporally-heavy datasets, SSv2-label, and SSv2-template as shown in Tab. 4 and Tab. 5 respectively. Our method outperforms previous methods by a large margin on multiple datasets, achieving averaged accuracies of 79.3% (+5.6%), 75.4% (+4.7%), 84.6% (+4.6%) on DiDeMo, ActivityNet-Captions and SSv2 respectively. Our results on MSR-VTT are worse (66.5% vs. 68.6%) than OmniVL but our pretraining framework is significantly cheaper (i.e., 82 vs. 169 V100 GPU days). We also note that our method is significantly cheaper than other top-performing approaches including LAVENDER , All-in-one , and CLIP-ViP (82 vs. 640, 448, 984 V100 GPU days for pretraining respectively). Additionally, our cheapest VindLU variant requires only 1515 V100 GPU days for pre-training, which is the second cheapest model among all listed approaches, and it still achieves competitive results on all three benchmarks. Furthermore, compared to the other leading VidL approaches such as OmniVL and Singularity, which rely on a multi-stage curriculum pretraining, our framework is simpler since it can be trained in a single stage. Lastly, our results on the SSv2 dataset in Table 5 indicate that VindLU performs very well not only on spatially-biased datasets but also on temporally-heavy datasets, which require sophisticated temporal modeling capabilities. For fairer comparisons, we de-emphasize CLIP-based methods since they use a lot more pre-training data.

Video Question-Answering. In Table 6, we also present our results for the video question-answering task on ActivityNet-QA , MSRVTT-QA , MSRVTT-MC and TVQA . Our results indicate that compared to prior state-of-the-art approaches, VindLU achieves competitive results across all four of these datasets. In particular, our method outperforms existing approaches by 0.6% on ActivityNet-QA, 0.3% on MSRVTT-QA, 3.4% on MSRVTT-MC and 0.3% on TVQA. For fair comparison, we de-emphasize FrozenBiLM , since it is a lot larger than our model (1.2B vs. 201M parameters) and uses a lot more pretraining data (400M vs. 25M).

Action Recognition. We finetune our pretrained video encoder on Kinetics-400 directly using TimeSformer codebase with exactly the same hyperparameters as in . As shown in Table 7, our video encoder outperforms TimeSformer and OmniVL by 2.1% and 1.0% respectively with all models using exactly the same architecture . This indicates the usefulness of our VidL pretraining recipe for a pure video understanding task.

Conclusion

In this work, we demystify the importance of various components used in modern VidL framework design. Throughout our empirical study, we find that temporal modeling, multimodal fusion, masked modeling pretraining objectives, and joint training on images and videos are critical for good performance on the downstream VidL understanding tasks. Our empirical insights enable us to develop a step-by-step recipe for effective video-language (VidL) pretraining, which leads to a highly performant VidL model, dubbed VindLU. Compared to the existing VidL approaches, our method achieves competitive or even better results on 9 VidL benchmarks while also being simpler and more efficient. While our paper does not provide any novel individual contributions, we believe that our empirical insights and our VidL pretraining recipe will be useful and help advance further research in the VidL domain.

We thank Yan-Bo Lin, Md Mohaiminul Islam, Avinash Madasu and Maitrey Gramopadhye for helpful discussions. This work was supported by the Sony Faculty Innovation award, Lilly Endowment, Inc. via Indiana University Pervasive Technology Institute, Laboratory for Analytic Sciences via NC State University and NSF-AI Engage Institute DRL211263.

Appendix

Appendix A Implementation Details

Positional Embeddings. We use learnable absolute temporal positional embeddings as in and relative spatial positional embeddings as in . The temporal positional embeddings are applied after patchifying the tokens, while the relative spatial positional embeddings are applied at each Transformer layer. When adapting the pretrained model to downstream tasks with more frames, we use zero-padding for the temporal positional embeddings as in . When adapting to higher spatial resolutions, we linearly interpolate the spatial positional embeddings.

Video Retrieval. We finetune the pretrained model with VTC and VTM losses. During inference, we follow to first select top-KK (K=128K=128 in our experiments) candidates based on the video-text similarity scores of the unimodal encoders and then re-rank these candidates by calculating their pairwise VTM scores.

Open-ended Question-Answer. Following , we formulate this task as a text generation task. As shown in Fig. 4, we add a decoder that takes the multimodal encoder’s outputs as the cross attention key and value to generate the answers. The decoder starts with a [CLS] token and ends when a [SEP] token is generated. The decoder has the same architecture as the multimodal encoder and is initialized with the pretrained multimodal encoder’s weights. The model is optimized using the averaged cross-entropy loss of each token between the generated answer and the ground truth answer. For a fair comparison with prior works , we constrain the decoder to generate from the 3128 most common answers during inference.

Multiple-Choice Question-Answering. For Multiple-Choice QA, we follow and convert it to the text-to-video retrieval task. Specifically, for each question and mm candidate answers, we generate mm sentences by concatenating the question with each candidate’s answer. We then rank these sentences by ensembling the retrieval model’s video-text similarity and pairwise VTM scores. The ensembling weights are set to 0.3 for the similarity score and 0.7 for the VTM score.

Inference with More Frames. Following , we perform inference using more frames than our finetuned model. Specifically, we first linearly interpolate the temporal positional embeddings in the video encoder. Then all the visual tokens are concatenated and fed to the multimodal encoder.

Pretraining Datasets. As discussed in the main draft, in Steps 1-3 of our recipe, we pretrain our model on a 2M WebVid-2M corpus. For Steps 4-5, we use a joint image-video corpus consisting of 3M images from CC3M and 2M videos from WebVid-2M . Lastly, in Step 6, we scale our pretraining data from 5M→17M→25M5M\rightarrow 17M\rightarrow 25M.

Model Details. Our final VindLU uses a vision encoder based on ViT architecture initialized with BEITbase\text{BEIT}_{base} weights, pretrained on ImageNet-21k. The additional temporal attention modules are randomly initialized and added before spatial attention in each Transformer block as in . As our text encoder, we use the first 9 layers of BERTbase\text{BERT}_{base} . The multimodal fusion encoder is our previously described V2T-MF module built using the last 3 layers of the same BERTbase\text{BERT}_{base} model. Our final pretraining objective is the sum of VTC, VTM and MLM losses. The hyperparameters are shown in Table 8. When doing multi-stage pretraining in Step 4 in the main draft, we set the initial learning rate of 5e-5 for stage 2 and 1e-6 for stage 3. Our model is implemented using PyTorch with Mixed Precision Training and Gradient Checkpointing .

Training Time. We train 2M and 5M corpus on 8×8\times RTX A5000 GPUs, which takes about 1 day and 1.8 days, respectively. For 17M and 25M, we train our model using 32×32\times A5000 GPUs, which takes 1.3 days and 3 days, respectively. For downstream tasks, the finetuning time ranges from 2-40 hours depending on the dataset size. The speed of A5000 is 0.99×0.99\times as V100 and 0.5×0.5\times as the A100 according to Lambda’s benchmarkhttps://lambdalabs.com/gpu-benchmarks fp16, bert_base_squad.

Appendix B Dataset Descriptions

Pretraining. We pretrain our model on three corpora: C5M, C17M and C25M, which we describe below.

C5M (5M): WebVid-2M , and CC3M . It contains a total of 5.44M image/video and text pairs.

C17M (17M): C5M, COCO , Visual Genome , SBU Captions , and CC12M . It contains a total of 18.41M image/video and text pairs.

C25M (25M): C17M, and WebVid-10M (excluding 2M videos from WebVid-2M as WebVid-10M is a superset of WebVid-2M). It contains a total of 25.91M image/video and text pairs.

Text-to-Video Retrieval. We evaluate our model on 3 spatially biased datasets MSR-VTT , DiDeMo , ActivityNet- Captions and 2 temporally-heavy datasets SSv2-Template , SSv2-Label .

MSRVTT contains 10K YouTube videos with duration between 10-30 seconds and 200k captions. Following , we train on 9K videos and report results on 1K-A test set.

DiDeMo contains 10K Flicker videos with 41K captions. Following , we only keep the first 30 seconds of each video and evaluate paragraph-to-video retrieval, where all the descriptions for a video are concatenated to form a single query.

ActivityNet-Captions contains 20K YouTube videos with 100K captions. Following , we train on the train set with 10K videos and evaluate on the val set with 4.9K videos and evaluate paragraph-to-video retrieval.

SSv2-Template contains 169K videos for training and 2K videos for evaluation from dataset SSv2 . The queries are 174 template (e.g., “Holding [something] next to [something]”) in SSv2. In the 2K test set, each template has 12 videos.

SSv2-Label contains the same videos for train/test as in SSv2-Template except that the text queries are the annotated labels (e.g., “holding potato next to vicks vaporub bottle”) in SSv2.

Video Question Answering. We evaluate on two open-ended QA datasets ActivityNet-QA, MSRVTT-QA and two multiple-choice QA dataset MSRVTT-MC, TVQA.

ActivityNet-QA contains 58K open-ended questions on 5.8K sampled videos from ActivityNet .

MSRVTT-QA contains 244K open-ended questions on 10K MSRVTT videos.

MSRVTT-MC contains 3K sampled videos with one multiple choice question for each video with 5 candidates. We evaluate the performance using the retrieval model finetuned on MSRVTT 7K training set.

TVQA contains 22K video clips and 153K multiple-choice questions focused on popular TV shows. We use the official train/val/test splits and reports results on the test set.

Appendix C Additional Quantitative Results

In this section, we present additional quantitative results on temporal modeling.

MLM masking ratio. We found a larger masking ratio (50%) for the MLM objective is more helpful for VidL pretraining, compared to the 15% masking ratio used in BERT . We conjecture that we can use a higher mask ratio than text-only BERT because our model incorporates complementary video cues.

Analysis on More Tasks/Datasets. In Tab. 10, we further evaluate our recipe on VidQA on MSRVTT-QA and video retrieval on SSv2-Label , SSv2-Template . As our evaluation metrics, we report the averaged R@{1,5,10} on SSv2-* and R@1 on VidQA. Since VidQA needs a multimodal fusion (MF) encoder to generate the answers, we cannot report the results without the MF module (i.e., Columns 1,2 in Row 2 in Tab. 10). Our results indicate that our conclusions (i.e., the importance of temporal modeling, multimodal fusion, and joint image+video pre-training) also hold on these tasks/datasets.

C.2 Additional Temporal Modeling Baselines

As discussed in the main draft, our first step is to extend our initial image transformer to video via a temporal modeling mechanism. Such a temporal modeling mechanism would enable training our model on multiple frames for more robust VidL spatiotemporal representation learning. For this part of our empirical study, we experiment with the following temporal modeling schemes using 4-frame inputs and pretrained on WebVid-2M . Besides the four temporal modeling baselines (i.e., mean pooling (MP), late temporal attention (L-TA), temporal convolution (TC), and temporal attention (TA)) that we included in the main draft, we further study Temporal Attention via Prompts (TA-P) and Window Attention (WA). We describe each of these baselines in more detail below:

Temporal Attention via Prompts (TA-P). Following, several previous methods we implement a baseline that uses temporal attention via prompt tokens. As shown in Figure 5, we first add mm prompt tokens to each frame. Then, these prompt tokens attend to each other via temporal attention to exchange frame-level information. Finally, all frame-level image tokens and prompt tokens for that frame attend to each other via spatial attention. Our TA-P scheme follows the same implementation as in .

Window Attention (WA). Similar to Swin , the spatial-temporal tokens are divided into cuboids of size T×k×kT\times k\times k, where TT is the number of frames and kk is the window size. WA is performed inside each cuboid. Similar to Temporal Attention, the WA is inserted before the spatial attention as in . We experiment with k=2k=2 and k=7k=7. Larger kk leads to an out-of-memory error.

We also illustrate these attention mechanisms in Figure 5. Furthermore, for completeness, below, we also describe the four baselines included in the main draft of the paper.

Mean Pooling (MP). In this variant, the visual encoder processes input frames independently and averages their frame-wise scores for the video-level score as in .

Late Temporal Attention (L-TA). In this variant, we attach 2 Transformer layers to an image encoder, which then aggregates temporal information across all input frames.

Temporal Convolution (TC). We insert a TC block before the spatial attention in each ViT layer. The TC block consists of a linear down-projection layer with hidden size 384, a depth-wise 3×1×13\times 1\times 1 convolution as in , a ReLU activation, and a linear up-projection layer.

Temporal Attention (TA). We insert a TA before spatial attention in each layer as in TimeSformer .

As shown in Table 11, Temporal Attention outperforms Temporal Convolution and Temporal Attention via Prompts by 2.1% and 6.8% respectively on averaged top-{1,5,10} accuracy. Window Attention with window sizes of k=2k=2 and k=7k=7 outperforms Temporal Attention by 0.2% and 0.7% respectively. These results indicate that high temporal modeling capacity is important in VidL models. As Window Attention has k×k\times the computational and memory cost and limited performance improvement compared with Temporal Attention, we choose Temporal Attention as our final temporal modeling blocks.

References