LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling
Linjie Li, Zhe Gan, Kevin Lin, Chung-Ching Lin, Zicheng Liu, Ce Liu, Lijuan Wang
Introduction
Large-scale transformer-based pre-training is now the de facto practice for NLP and vision-language research . Together with the great success of image-text pre-training , video-language (VidL) pre-training has also received an increasing amount of attention. By pre-training an end-to-end multimodal transformer on a large number of video-text pairs, state-of-the-art performance has been achieved across a wide range of VidL tasks, including video question answering (QA) , text-to-video retrieval , and video captioning . These advances are encouraging; however, on the other hand, all existing VidL works require designing task-specific heads on top of the transformer encoder for each pre-training or downstream task. For example, during pre-training, separate Masked Language Modeling (MLM) and Video Text Matching (VTM) heads are used, while a new, separately parameterized head needs to be added for each downstream adaptation. Furthermore, due to the particular nature of different tasks, they are typically modeled using different training objectives. For example, multiple-choice video QA is formulated as a classification problem, while video captioning is inherently a generation task. A natural but challenging question arises: can we have a unified architecture that supports all the popular VidL tasks simultaneously without introducing task-specific heads?
To answer this, we present Lavender, a unified VidL framework where all pre-training and downstream tasks are formulated as simple MLM. As shown in Figure 1, we use two pre-training tasks: MLM and VTM. However, for VTM, instead of adding a head on top of the output of the commonly used [CLS] token, as used in all existing works, we propose to append the same [MASK] token that is used for MLM at the end of the video-text input, and use the same MLM head to predict whether the input video-text pair matches or not. Note that VTM is typically formulated as a binary classification problem; here, we simply treat the output of true or false from VTM as natural language tokens directly predicted from the whole vocabulary, so that the same set of parameters can be used for both MLM and VTM.
During downstream adaptation, instead of discarding the MLM head used in pre-training and adding new heads, which is the standard practice for all previous VidL works, we use the same MLM head used in pre-training for all downstream tasks. Specifically,
For text-to-video retrieval, we train the model in the same way as in the VTM pre-training task. During inference, for each text query, we concatenate it with each candidate video, and calculate the corresponding probability of the [MASK] token being predicted as true, and then rank all candidate videos based on that score.
For multiple-choice video QA, we concatenate the question and each answer candidate sequentially, and add a [MASK] token at the end of the sequence, and use the same MLM head to predict the answer as “n” (assuming the ground-truth choice is the n-th answer).
For open-ended video QA, since most of the ground-truth answers in our tested datasets only contain one word, we simply append a [MASK] token to the end of the video-question input, and let the model predict the answer from the whole vocabulary.
For video captioning, during training, we mask a certain percentage of the tokens, and then predict the masked tokens using a seq2seq attention mask . During inference, the full caption is auto-regressively predicted, by inserting [MASK] tokens one at a time.
Lavender is inspired by VL-T5 , UNICORN and OFA that aim to provide a unified pre-training framework for image-text tasks. However, ours is very different from theirs, as we use an encoder-only model and an additional lightweight MLM head on top of it, while a heavy transformer decoder is needed in . By unifying all the VidL tasks as MLM, Lavender can seamlessly adapt to different VidL tasks, and enables new capabilities over existing task-specific methods, such as () supporting different VidL tasks with a single set of parameter values when multi-task finetuned; () better generalizability to test data under few-shot finetuning; and () zero-shot inference on video question answering. Surprisingly, by using this simple generative approach, we outperform previously published state-of-the-arts on 12 out of 14 downstream tasks (Table 1), even when pre-trained with much fewer data (see Section 4.5 for detailed comparisons).
Related Work
Video-Language Pre-training. Branching out from large-scale image-text pre-training , researchers have been leveraging large-scale multimodal data to build pre-trained video-language (VidL) models for a wide range of generative and discriminative tasks. Prominent examples include VideoBERT , HERO , ActBERT , ClipBERT and MERLOT . Popular pre-training tasks include Masked Language Modeling (MLM) , Video Text Matching (VTM) , frame order modeling and masked visual modeling . For simplicity, we only use MLM and VTM losses for pre-training, as in our experiments, we observe that the other pre-training tasks are not that essential to improving the final performance on downstream tasks (results in Appendix A).
Although achieving strong performance, existing methods all require task-specific architectures or objectives for different downstream tasks. For example, text-to-video retrieval is modeled as binary classification or via contrastive learning ; video question answering is often formulated as multi-class classification with a set of pre-defined answer candidates ; and video captioning can be tackled via MLM with a multi-layer perceptron or prefix language modeling with a text decoder .
Unified Frameworks for Multimodal Understanding. There have been attempts in building an omnipotent model that can simultaneously handle different tasks with a unified architecture, which can be broadly categorized into two directions. The first is to insert expert-designed task-specific heads for each downstream task . These task-specific output layers not only require expert knowledge, but are also unlikely to generalize to new tasks. For example, when a new question answering task comes in, a new fully-connected layer with output dimension of answer vocabulary size is required. The second direction is to unify the input-output format of different downstream tasks . With a unified vocabulary, different downstream tasks (e.g., image question answering , image captioning and visual grounding ) can be formulated as the sequence-to-sequence generation with the shared encoder-decoder architecture .
Our work aims to provide a unified framework for VidL understanding, in contrast to task-specific architectures and objectives used in existing VidL models (Figure 2(c) vs. Figure 2(a)). Lavender differs from the previous unified image-text models in that all pre-training and downstream tasks are unified as MLM, and a simple encoder-only architecture with a lightweight MLM head is used, instead of sequence-to-sequence modeling as in that also requires a heavy transformer decoder (Figure 2(c) vs. Figure 2(b)).
Lavender
Given a pair of text sentence and a video , we first encode them separately via unimodal encoders (i.e., vision encoder and text encoder) to generate unimodal features. Here, is the number of tokens in a sentence and is the number of frames sampled from the input video. We follow previous works to only sparsely sample a few frames to ease the computational burden. A multimodal fusion encoder (dubbed as fusion encoder) projects textual features and visual features into a shared embedding space to learn cross-modal representations. As Lavender unifies both pre-training and downstream tasks as Masked Language Modeling (MLM), the same MLM head is used to generate the final outputs from the cross-modal representations, across different tasks. Next, we explain each component in detail.
Vision Encoder. Inspired by the success of vision transformers in modeling spatial details in images , different transformer architectures have been proposed to model the long-range temporal modeling in videos, achieving promising results on action recognition . Recent video-language works have started to embrace the success of video transformers, demonstrating stronger performance than encoding each video frame independently . In our work, we adopt Video Swin Transformer (VidSwin) as the vision encoder to encode the raw video frame inputs as a sequence of visual features. Given input video frames of size , we first split each frame into non-overlapping patches of size . VidSwin additionally enforces temporal downsampling of size 2 as a preprocessing step. To allow Lavender to have the flexibility of utilizing both video-text and image-text data for pre-training, we remove this temporal downsampling. As a result, we can extract a sequence of visual features of size from the last encoder block of VidSwin. Each feature is of size ( is the channel dimension), which is projected to the same dimensional space as text features via a fully-connected layer. We follow to add learnable positional embedding layers along both spatial and temporal dimensions. The resulting visual features are used as input to the fusion encoder to learn cross-modal representations.
Text Encoder. The input text sentence is first tokenized into the sequence of word tokens , following . Two special tokens [CLS] and [SEP] are inserted at the beginning and the end of the token sequence. We follow previous works to adopt a lightweight word embedding layer as the text encoder. The high-dimensional text embeddings are concatenated with visual features and then fed into the fusion encoder.
Multimodal Fusion Encoder. The fusion encoder is a 12-layer, 768-dimensional Transformer , mirroring the BERT-base architecture . To compute the cross-modal representations, the unimodal features from vision and text encoders are fused together via self-attention operations.
2 Our Unified Framework
Now, we introduce how Lavender can be trained in a unified way.
Video-language Pre-training. We adopt two objectives to pre-train Lavender. The first is Masked Language Modeling (MLM), which is directly adopted from language model pre-training . In MLM, we randomly replace 15% of word tokens with a [MASK] token, a random word, or the same word. The goal is to reconstruct the correct tokens based on the corresponding hidden representations from the output of the fusion encoder at the masked position. A multi-layer perceptron (MLP) with output dimension as vocab_size projects these hidden representations into the discrete word token space. Cross-entropy loss is used to supervise the model training. The second is Video Text Matching, but reformatted as MLM (VTM as MLM). Specifically, we append a [MASK] token to the textual sentence to mimic the masked textual inputs in MLM. At each training step, we randomly replace the corresponding text for a given video with a text description from a different video in the same batch. At the masked position, Lavender reuses the exact same MLP used in MLM to make a prediction. Although the ground-truth label is restricted to two tokens (i.e., true (false) for a positive (negative) video-text pair), but the model predictions are made across all vocabularies.
Downstream Adaptation. As shown in Figure 1, we can readily apply the pre-trained Lavender to 4 types of downstream tasks, including text-to-video retrieval, multiple-choice video question answering, open-ended video question answering and video captioning. For each task, we transform the text input by inserting or replacing existing tokens with [MASK] tokens, so that all tasks can be supervised with cross-entropy loss, and the final predictions are made based on the word token predicted at the masked position. Here, we explain in detail how to construct the masked textual inputs and generate model predictions for each downstream task.
For text-to-video retrieval, similar to VTM during pre-training, we insert a [MASK] token at the end of the text input. During training, we treat corresponding video-text pairs as positives (with ground-truth label true) and all other pairwise combinations constructed by replacing the ground-truth text with a randomly sampled one as negatives (with ground-truth label false). During inference, given a textual query, we rank the videos according to the model confidence of predicting true at the masked position. For multiple-choice video QA, we concatenate each answer choice (A) sequentially to the question (Q) with a [SEP] token in between. A [MASK] token is then added at the end, to allow the model to make a prediction of the correct index for the ground-truth answer choice. For example, for a question and 5 answer choices, we take Q+[SEP]+A0+[SEP]+...+A4+[MASK] as the text input. If A is the correct answer, the ground-truth label for the masked token is n. Through the MLM head, the model makes a prediction at the masked position over the whole vocabulary. During inference, to ensure a valid answer, we take the most probable predictions over all answer indices (e.g., {0,1,2,3,4}). For open-ended video QA, we similarly inject [MASK] tokens after the question. For simplicity, we only add one [MASK] token.Note that we can optionally insert multiple [MASK] tokens to allow answer predictions with variant lengths. However, as we will show in Appendix E, over 95% of the questions in the evaluation datasets considered can be answered with a single word. We then tokenize the ground-truth answers as the ground-truth label for masked prediction. If the tokenized answer is longer than 1 word, we simply ignore it during training, and regard it as a wrong prediction during inference. For video captioning, we use a causal self-attention mask where the caption token can only attend to the existing output tokens, which simulates a uni-directional seq2seq generation process, following . During training, we randomly “mask” some words with [MASK] token and apply the MLM objective. During inference, the caption is generated in an auto-regressive manner. At each generation step, the model sees the entire video input and previously generated tokens, plus a [MASK] token, at which the model makes a prediction for the current token.
Experiments
In this section, we first describe our experimental settings (4.1), and show the superiority of Lavender over comparable task-specific baselines under both single-task (4.2) and multi-task (4.3) finetuning settings. We then show that our model can better generalize to testing data under few-shot finetuning and has strong zero-shot capability on video question answering benchmarks (4.4). Lastly, we compare Lavender with prior arts and show we outperform published methods on 12 out of 14 benchmarks, even when pre-trained with much fewer data (4.5).
Pre-training Data. In our default setting, we follow to aggregate video-text pairs in WebVid2.5M and image-text pairs in CC3M to pre-train Lavender. As a scale-up recipe, we additionally crawl 11.9M video-text pairs from the web, following the same procedure in . We similarly scale up image-text pairs by assembling COCO , Visual Genome , SBU Captions , CC12M and CC3M. Combining these video-text and image-text datasets together results in 14M videos + 16M images. Unless otherwise specified, all results reported in this section as Lavender are pre-trained under the default setting with 2.5M videos + 3M images. In Section 4.5, we show that scaling up our pre-training data further improves model performance.
Downstream Tasks. We evaluate Lavender on 14 video-language benchmarks over popular VidL tasks, including text-to-video retrieval, video question answering (QA) in both multiple-choice (MC) and open-ended (OE) settings and video captioning. We briefly list the evaluation datasets for each task type below; detailed descriptions are included in Appendix E.
Text-to-video Retrieval: MSRVTT , DiDeMo , MSVD and LSMDC ;
MC Video QA: TGIF-Action, TGIF-Transition , MSRVTT-MC and LSMDC-MC ;
OE Video QA: TGIF-Frame , MSRVTT-QA, MSVD-QA and LSMDC-FiB ;
Implementation Details. We initialize our Vision Encoder with VideoSwin-Base , pre-trained on Kinetics-600 . Text Encoder and Multimodal Fusion Encoder are initialized from pre-trained BERT-Base . We train Lavender in an end-to-end manner for both pre-training and downstream finetuning. Our implementation of Lavender is based on PyTorch . We adopt AdamW as the optimizer with an initial learning rate of 2e-5, betas of (0.9, 0.98), and weight decay of 1e-3 for all pre-training and finetuning experiments. During pre-training, we sparsely sample =4 video frames and resize them into ==224 to split into patches with size ==32. For default setting with 2.5M videos + 3M images, we pre-train Lavender for 10 epochs on 32 NVIDIA V100 GPUs with a batch size of 28 per GPU, which takes about 2 days. The scale-up pre-training takes about 10 days on 64 NVIDIA V100 GPUs. For all downstream tasks, we adopt the same video frame size and patch size, but =5 video frames. More details can be found in Appendix C.
2 Comparison to Task-specific Baseline
To make a fair comparison to task-specific methods, we train a task-specific version of Lavender (denoted as Lavender-ts). We replace the shared Masked Language Modeling (MLM) head in Lavender with task-specific heads and adopt task-specific objectives. For text-to-video retrieval (and similarly for video text matching during pre-training), a multi-layer perceptron (MLP) with output dimension 1 is applied over the global video-language representation of the [CLS] token and binary cross-entropy loss is adopted to supervise the model training. For multiple-choice video question answering (QA), we concatenate questions with all answer candidates to form the input text, similar to what was described in Section 3.2, but without the added [MASK] token. A task-specific MLP with output dimension as the number of answer choices is applied over the representation of [CLS] token and cross-entropy loss is used to train a classifier over all answer indices (e.g., with 5 answer choices). For open-ended video QA, we follow the common practice to build a finite set of answer vocabularies covering the most common answers in the training split of each dataset. Similarly, a MLP with output dimension as the number of answers is added and cross-entropy loss is used to train a classifier over all answers. For video captioning, we simply adopt the same training strategy with MLM head as in our unified model.
Table 2 compares Lavender to the task-specific baseline Lavender-ts under different settings, on four representative benchmarks. For easier comparisons, we use Meta-Ave, the average across scores over all evaluation tasks, to measure the average model performance. Here, we focus our discussion on results under single-task finetuning, and delay the analysis on multi-task finetuning to Section 4.3. Without video-language (VidL) pre-training, task-specific baseline with different task heads (L3) outperforms our unified model (L1) on MSVD-QA and DiDeMo Retrieval. Captioning performance are similar as we adopt the same MLM head and finetuning strategy for both models. The only outlier is TGIF-Action, on which we empirically find the training of Lavender-ts struggles to converge, leading to a low Meta-Ave score.
We also compare the two models under VidL pre-training. We follow to pre-train task-specific Lavender-ts with MLM and the standard Video Text Matching (VTM) task, which is modeled as binary classification with an additional MLP layer. In comparison, the unified model Lavender is pre-trained with MLM and VTM as MLM, with the shared MLM head (Section 3.2). The unified VidL pre-training (L7) significantly enhances the performance of Lavender, with a gain of +23.4 on Meta-Ave over without pre-training (L1). Comparing both models under VidL pre-training, we also observe that Lavender (L7) outperforms Lavender-ts (L5) by a notable margin across all 4 tasks (+4.9 on Meta-Ave).
3 Multi-task Finetuning
In this section, we aim to answer the important question raised in Section 1: can we have a unified architecture that supports all downstream tasks simultaneously without introducing task-specific heads? We first compare Lavender with several multi-task learning baselines with different task-specific designs in Table 2 and report results under the extreme multi-task setting - a single model that can support all 14 datasets (all-in-one) in Table 3.
Comparison to task-specific baseline. We begin our comparison with the most common multi-task baseline in literature - adopting a task-specific head and objective for each task, while sharing the backbone. This is equivalent to finetuning the task-specific model Lavender-ts under multi-task setting. We compare the two models in Table 2 and summarize our findings below:
Single-task vs. Multi-task: Without video-language (VidL) pre-training (L1-4), we find that multi-task training greatly improves model performance in both cases, as it can take advantages of additional and diverse supervision. Combining multi-task finetuning with VidL pre-training renders slight performance drop (L5-8).
Single shared head vs. Task-specific heads: Results show that Lavender (single shared head) outperforms Lavender-ts (task-specific heads), with (L8 vs. L6, +5.9) or without (L2 vs. L4, +2.9) VidL pre-training. Lavender also saves more parameters (approximately 3H) from the additional 3 task-specific heads in Lavender-ts under multi-task finetuning. In addition, the performance drop of multi-task finetuning from single-task finetuning is less severe with the shared MLM head (-0.6 for Lavender vs. -1.6 for Lavender-ts) under VidL pre-training.
Multi-task Variants with Lavender. We also explore different multi-task variants with our unified framework Lavender in Table 2: () L8: the vanilla version without any task-specific design; () L9: with human-readable task-specific prompts (e.g., “is the video-text paired, true or false?" for video text matching), which has shown promising results for language understanding ; and () L10: with learnable task-specific tokens (e.g., a special token [VTM] for video text matching), which is in analogy to the task prefixes in . Different from observations in , both task-specific prompts and tokens do not show a clear advantage over the vanilla version. We conjecture the differences may be due to the weaker text encoder and the less diverse text prompts, which we leave as interesting directions for future study. Based on the above analyses, we simply extend the vanilla multi-task finetuning method from 4 to all 14 VidL benchmarks considered.
All-in-One. In Table 3, we finally attempt to answer the question with one model that can conquer all 14 downstream tasks simultaneously. We first establish the baseline performance by training single-task (ST) models with Lavender. We then report multi-task results with () a single set of parameters for all tasks (MT (all-in-one)); () the best-performing checkpoint for each task (MT (best)); and () with multi-task finetuning as 2nd stage pre-training and then finetune on each single task (MTST), from the learned weights with MT (all-in-one). As the results show, MTST achieves slightly better Meta-Ave across all settings. Surprisingly, the all-in-one model is very competitive, with only -0.5 performance drop on Meta-Ave, when compared with ST models.
4 Few-shot and Zero-shot Evaluation
Next, we showcase two new capabilities enabled by Lavender over task-specific baseline.
Few-Shot Generalizability. We first study how Lavender can generalize to testing data with limited training examples. Figure 3 presents the results of the proposed Lavender (red line), which is unified as Masked Language Modeling (MLM) for all downstream tasks, in comparison to the task-specific baseline Lavender-ts with different heads and finetuning objectives (blue line) on 4 representative benchmarks. For easier comparison, we further plot two dotted lines, which denotes 90% of model performance when trained with all the training data for each model. Lavender has shown a clear advantage, as it can easily achieve 90% of the full model performance, with much less training data. Specifically, approximately 40%, 10%, 10%, 6% of training data is needed for Lavender on TGIF-Action, MSVD-QA, DiDeMo-Retrieval and MSRVTT-Captioning, while Lavender-ts requires 60%, 60%, 25% and 10%, respectively.
Zero-Shot Evaluation on Video QA. Even without a heavy text decoder as in , a pre-trained Lavender can be directly evaluated on video question answering (QA) tasks without finetuning. Table 4 compares zero-shot (ZS) performance of Lavender with task-specific baseline Lavender-ts on 8 video QA benchmarks.
Since the model has neither learned to perform the multiple-choice QA task nor seen similar data during pre-training, we transform multiple-choice QA as Video Text Matching (VTM) for better ZS performance. Specifically, we let Lavender to predict true or false via MLM head, given a video-question-answer input, and we rank the probability of model prediction as true across all answer choices. Similarly, for Lavender-ts with binary classification head for VTM, we simply rank the probability of model prediction as “matched”. With the same pre-training data, Lavender evidently outperforms Lavender-ts on all multiple-choice QA benchmarks. On open-ended QA tasks, Lavender can be applied seamlessly, thanks to the shared MLM head. However, for Lavender-ts, the randomly initialized task-specific heads give meaningless ZS predictions.
We also compare Lavender against previous methods. Without the help of additional audio modality in or supervision signals in , Lavender achieves competitive ZS performance, even when pre-trained with much less data (5.5M vs. >69M). When scaling up the pre-training data by roughly 5 times, we observe notable performance improvements on most QA benchmarks. The performance drop on a few datasets may be due to the inclusion of more noisy data when scaling up.
5 Comparison to Prior Arts
In this section, we compare Lavender with prior arts, which are mostly designed to tackle a single type of video-language (VidL) task.
Table 5 summarizes results of Lavender on video question answering (QA) and video captioning. For video QA, Lavender achieves significant gains over existing VidL pre-trained models on 7 out 8 video QA benchmarks considered. On MSRVTT-QA, Lavender is only 1.8 points behind All-in-one pre-trained with 283M videos, which is 9 times more than ours (30M). It is worth mentioning that VIOLET adopts the same model architecture and finetuning objectives as our task-specific baseline Lavender-ts. Even when scaling up the VidL pre-training to 186M videos+images, the task-specific model VIOLET still underperforms Lavender, which further demonstrates the advantages of our unified framework. For video captioning, Lavender achieves the new state-of-the-arts on both datasets. Note that MV-GPT is pre-trained for multi-modal video captioning, where the auto-transcribed text from audio is used as additional input. With video-only inputs, Lavender is able to achieve comparable performance.
Table 6 presents the comparison on text-to-video retrieval. The most competitive methods on text-to-video retrieval are based on CLIP pre-trained on 400M images. However, with much fewer pre-training data, Lavender can still perform competitively on all 4 benchmarks, especially when compared to non-CLIP pre-trained methods . Notably, on DiDeMo and LSMDC, Lavender surpasses all baseline methods in Table 6. We hypothesize that the fusion encoder in Lavender is more effective in modeling interactions between video and long paragraph query (i.e., DiDeMo) or the contextualized queries collected from movie scripts (i.e., LSMDC), than the late dot-product fusion in .
Conclusion and Discussion of Broader Impact
We introduce Lavender, the first unified video-language (VidL) framework, that can tackle various VidL tasks with a unified Masked Language Modeling objective. Without any task-specific architectures, Lavender outperforms the prior state-of-the-art on 12 out of 14 benchmarks considered. Experiments show that Lavender is better suited for multi-task learning, few-shot generalization and zero-shot evaluation on video question answering tasks. There are several potential limitations of Lavender that would make for promising directions for future work, including: () extension to fine-grained VidL tasks (e.g., video corpus moment retrieval ); and () more effective in-context few-shot learning or prompt tuning. Like other data-driven systems, Lavender shares similar risks that may have negative societal impact, such as biases in training data and energy consumption with large-scale training. However, we believe that our unified framework combined with multi-task learning can most likely reduce both memory and energy costs, and potentially lead to more economical deployment in real-world applications.
References
Appendix A Additional Results
Multi-task Finetuning. For completeness, we include additional results under multi-task finetuning. Table 7 presents the results of single-task finetuning and multi-task variants from the scale-up pre-training on 14M videos + 16M images. For easier reference in future work, we report detailed retrieval results on R1/5/10 in Table 8.
Investigation on Other Pre-training Tasks. As mentioned in the main text, we only adopt Masked Language Modeling (MLM) and Video Text Matching (VTM) as pre-training tasks for both the proposed Lavender and the task-specific baseline Lavender-ts. Here we briefly discuss other popular pre-training objectives with Lavender-ts. The first is Frame Order Modeling , where the input video frames are randomly shuffled and the goal is to revert back its original order. Different from the video-ASR pairs utilized in these works, the paired text in our pre-training data is not temporally grounded. In most cases, the shuffled frame sequence will probably still be globally aligned with the textual description. Hence, such fine-grained temporal reasoning objective is not applicable in our case. The second is Masked Visual Modeling (MVM), where the model learns to reconstruct high-level semantics or low-level details for a certain percentage of “masked” visual inputs (i.e., features or patches). Different variants have been proposed and shown little-to-none effect in vision-language pre-training, such as predicting the object category of masked image regions and distilling region/frame features from well-supervised vision encoders . More recently, by taking advantage of pre-trained DALL-E , researchers have shown potentials in masked visual token modeling, which asks the model to recover the discrete latent codes of the masked image patches. explores image feature descriptors such as Histograms of Oriented Gradients (HOG) as the prediction target for self-supervised visual pre-training. In Table 9, we investigate three different MVM objectives on top of VTM + MLM pre-training for Lavender-ts: () VQ Token: to recover the discrete codes extracted from pre-trained DALL-E following ; () Pixel: to regress the RGB colors as in ; and () HOG: to regress the HOG values, following . Results show that only MVM with HOG achieves a marginal performance improvement of +0.3 on average. Therefore, we adopt a simple recipe for all other pre-training experiments in the paper, that is with only MLM and VTM.
Qualitative Comparisons to Task-specific Baseline. Figure 4 provides qualitative comparisons between Lavender and task-specific baseline Lavender-ts on video question answering (QA). The model predictions are sampled from MSVD-QA.
In Figure 4(a), the ground-truth answer “fold” is not in the top- ( for MSVD-QA, following ) most common answers in training split, hence excluded from the pre-defined answer vocabulary for training Lavender-ts. In Figure 4(b), the ground-truth answer “bowl” appears roughly 9 times more than “bag" in the training split. These visualization results on video QA suggest that () our Lavender can better fit the open-ended setting for QA tasks, as it does not restrict the predictions to be from a pre-defined answer vocabulary as in Lavender-ts (Figure 4(a)); and () the task-specific baseline is easier to fail on questions with out-of-distribution answers than Lavender (Figure 4(b)). Additionally, we show in Figure 4(c) when both models can provide reasonable answers to the question, which do not exactly match the ground-truth answer. This result reveals potential problems with the current evaluation metrics or existing datasets on video QA. Future work may consider collecting additional annotations to enrich the dataset and improve the evaluation metric to handle multiple ground-truth answers (e.g., similar to VQA scores ).
Appendix B Additional Comparison with Existing Work
Figure 2 in the main text shows the detailed comparison between Lavender and existing methods with the image/video question answering task as an example. In Figure 5, we illustrate the differences among these methods in pre-training. Unlike existing video-language models, which design task-specific heads and objectives for different pre-training tasks. Lavender unifies masked language modeling (MLM) and video text matching (VTM) as MLM. Compared with unified image-text models (e.g., VL-T5 ), which are typically pre-trained with a combination of complex pre-training tasks, such as visual question answering and grounded captioning. Although these pre-training tasks may enable the model with new abilities (e.g., generating region proposals as in ), the supervision often comes from human-annotated data. It remains unclear how to design and effectively pre-train a unified model with such capability but without dependency on human-labeling.
Appendix C Implementation Details
Task-specific Prompts and Tokens. As mentioned in Section 4.3 of the main text, we explore the vanilla multi-task finetuning without any task-specific designs and two additional variants with task-specific prompts and tokens for Lavender. Here, we describe the prompts and tokens used in these baselines.
For task-specific prompts, we insert a human-readable text prompt at the beginning of the text input.
For text-to-video retrieval and video-text matching during pre-training, the text prompt is “is the video-text paired, true or false”;
For multiple-choice video question answering (QA), the text prompt is “which answer choice is correct, choose from 0, 1, 2, 3, 4.”;
For open-ended video QA, the text prompt is “answer the question about the video.”;
For video captioning, the text prompt is “write a description about the video.”.
As discussed in the Experiments section of the main text, we only briefly investigate prompt tuning with Lavender. How to design more diverse prompts for more effective prompt tuning is an interesting direction for future work.
For task-specific tokens, we add several new tokens to the whole vocabulary, and learn these token embeddings from scratch during multi-task finetuning. For both training and inference, the task-specific token is inserted right after [CLS] token in the text input. Specifically, the new task-specific tokens are [VTM] for text-to-video retrieval and video-text matching, [MC] for multi-choice video QA, [OE] for open-ended video QA, [CAP] for video captioning.
Additional Training Details. We summarize the training configurations for downstream finetuning in Table 10. Due to various data scales and domains, we use task-specific batch size and training epochs based on the performance of the validation set for each downstream task (Table LABEL:subtab:task-specific). All other settings are shared across all datasets (Table LABEL:subtab:common). All experiments are conducted on Microsoft Azure , adopting mixed-precision training with DeepSpeed . All video data are pre-processed by evenly extracting 32 frames to avoid expensive decoding on-the-fly. During training, we randomly sample frames from 32 frames, resize the shorter side of all frames to 224 and random crop (224x224) at the same location for all the frames in a given video. During inference, we evenly sample frames from 32 frames and center crop (224x224) for all the frames.
For multi-task finetuning, since the same set of videos are shared among several downstream tasks, there might be overlaps between one’s training split and others’ validation or testing split (e.g., some video-text pairs in MSRVTT-Retrieval 9K-train is in the testing split of MSRVTT-Captioning). To avoid data contamination, we filter out validation and testing videos in all downstream datasets from the training splits, and use this cleaned version for multi-task finetuning. At each training step, we randomly sample one dataset from all 14 of them, and construct a batch of examples from that dataset. The training is conducted on 1680GB A100 for 20 epochs, and we adopt the same batch size for retrieval tasks as shown in Table LABEL:subtab:task-specific, and batch size 60 for all other tasks.
Appendix D Pre-training Data
Public Datasets We use the following publically available datasets to pretrain Lavender:
WebVid2.5M scrapes 2.5M video-text pairs from the web. The texts in this data are alt-text descriptions, which generally describe the global video semantics.
Conceptual Captions 3M (CC3M) consists of 3.3M image-text pairs, which are also harvested from the web. CC12M further enlarges CC3M by 4 times. Both have been used to pre-train large-scale image-text models.
SBU-Captions is another widely used dataset for image-text pre-training, web-crawled from Flickr. It contains 1M image-text pairs.
COCO and Visual Genome (VG) are two human-annotated image-text datasets. COCO contains 5 captions per image over 120K images. Unlike COCO captions that can describe the whole scene, VG collects 5M regional descriptions over more than 100K images.
Video-Text Data Collection For the scale-up pre-training, we additionally crawl 11.9M video-text pairs from the web, following the same procedure in . Here, we briefly describe how we collected the data.
WebVid2.5M has led to promising results in text-to-video retrieval tasks as shown in . This motivates us to further crawl more video-text pairs from the same source. We first use a search engine to identify the potential data sources based on sampled textual descriptions in WebVid2.5M, and then we scrape the video-text pairs from these data sources. Similarly, we follow to filter out offensive content and hide person and location names. In total, we have collected 11.9M videos, each accompanied with an alt-text description. The collected dataset shares similar characteristics as WebVid2.5M, with the average video duration as 20 seconds, and the average number of words in the textual description as 20. Note that at the time when we started this project, WebVid10M in has not been released yet. We later found our 11.9M data largely overlaps with WebVid10M. Hence, we refer future work to WebVid10M for scale-up pre-training.
Appendix E Downstream Datasets
In this section, we introduce all downstream datasets used for evaluating Lavender and discuss some dataset-specific training details below. Table 11 summarizes the number of examples in training/validation/testing split for each dataset.
Text-to-video Retrieval We evaluate Lavender on 4 popular text-to-video retrieval datasets, namely MSRVTT , DiDeMo , MSVD and LSMDC . MSRVTT contains 10K YouTube videos with 200K descriptions. We follow to train on 9K videos and evaluate on 1K-A testing split. DiDeMo consists of 10K Flickr videos, each annotated with 4 sentences. We concatenate all sentences from the same video into a paragraph and perform paragraph-to-video retrieval, following . Although this dataset comes with localisation annotations (ground-truth temporal proposals) for each sentence, we perform all experiments without leveraging this fine-grained information for both training and evaluation. Instead, we use the same procedure as described in Appendix C to sample frames from videos. MSVD is based on 2K YouTube videos and crowdsourced 40 textual descriptions per video. LSMDC is built upon 118K video clips from 202 movies. Each clip has a caption from movie scripts or descriptive video services. We use the standard splits for DiDeMo, MSVD and LSMDC, following . For paragraph-to-video retrieval on DiDeMo, we adopt the text augmentation technique proposed in Frozen , which is to randomly sample and concatenate a variable number of sentences as paragraph for each video.
Multiple-choice Video QA We evaluate Lavender on four multiple-choice QA datasets: TGIF-Action, TGIF-Transition , MSRVTT-MC and LSMDC-MC . Among them, TGIF-Action and TGIF-Transition aim to test the model’s ability to recognize repeating actions and state transitions in short GIFs. Each video-question pair is accompanied with 5 answer choices. We concatenate the 5 answer choices sequentially with the question, and the model is asked to predict the ground-truth answer index. MSRVTT-MC and LSMDC-MC are based on retrieval tasks, but reformulated as multiple-choice QA. A model needs to find the caption that describes the video out of 5 candidate captions. Due to its similarity to video-to-text retrieval, we formulate it as video-text matching, which is the same as zero-shot evaluation described in the Experiments section of the main text. Specifically, we let Lavender to predict true or false via MLM head, given a video-question-answer input, and we rank the probability of model prediction as true across all answer choices. As there is no training and validation data constructed in the same way for MSRVTT-MC, we follow to evaluate the retrieval model trained on MSRVTT to rank the 5 candidate answers.
Open-ended Video QA Four datasets are considered for open-ended video QA: TGIF-Frame , MSRVTT-QA, MSVD-QA and LSMDC-FiB . Among them, the question-answer pairs in all but TGIF-Frame are based on the linguistic transformation of captions for each video. Questions in TGIF-Frame is collected via crowd-sourcing, which are answerable with just a single frame in the video. MSRVTT-QA contains 243K open-ended questions over 10K videos and MSVD-QA consists of 47K questions over 2K videos. The Fil-in-the-blank (FiB) task of LSMDC-FiB is, given a video and a sentence with a blank in it, to predict a correct word for the blank. We replace the blank with a [MASK] token, and naturally it becomes a Masked Language Modeling (MLM) task.
As mentioned in the main text, Lavender answers the open-ended questions in these datasets with only one word, as there is only one [MASK] token appended to the text input. Table 12 summarizes the max answer length and the percentage of examples with answers longer than one word in all four datasets. As the statistics show, more than 92% of examples in these datasets are answerable with a single word.
Video Captioning MSRVTT and MSVD are used for captioning evaluation. As introduced before, MSRVTT consists of 10K videos with 20 captions per video, and MSVD contains 2K videos, with 40 captions per video. We follow the standard captioning splits in and for MSRVTT and MSVD, respectively. During training, we set the probability of random masking caption tokens to be 0.15, the same as what is used in MLM during pre-training. During inference, we perform generation until the model outputs a [SEP], which is defined as the sentence ending token or when it reaches the maximum generation step 50.
Appendix F Data Usage
We strictly comply with the data usage agreement for each public dataset used in this paper, which is to only use for non-commercial research purposes. To the best of our knowledge, there is no personal identification information or offensive content in any of the public datasets.