End-to-end Generative Pretraining for Multimodal Video Captioning

Paul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, Cordelia Schmid

Introduction

A long-standing goal of the AI community is the development of conversational multimodal systems that can both reliably perceive the world and effortlessly communicate with humans. An emerging benchmark of progress in this field is the task of multimodal video captioning - which tests both abilities; a successful model must not only accurately understand ‘multimodal’ streams of input video (including the speech and the video frames), but also generate coherent natural language descriptions of the content.

Unsurprisingly, a major challenge in the field of vision and language learning is the lack of large-scale, manually annotated data. Annotating captions for videos is time intensive, expensive and subjective (with low inter-annotator agreement ) – this is in contrast to fields such as image classification where fully annotated datasets are orders of magnitude larger . To overcome this limitation, there has been a flurry of recent works that pretrain their video-language models on instructional videos , a domain where the speech is particularly well aligned to visual content. Recently introduced datasets such as Cooking312K and HowTo100M leverage such instructional videos with associated captions from ASR (automatic speech recognition) to learn joint video-and-text embeddings or to train multimodal video encoders . However, the models in these works often do not contain a decoder, lacking the ability to generate sentences, and thus only the video encoder is transferred to the downstream tasks – indeed for the case of video captioning, the decoder is often learned from scratch . While one can still initialize the decoder using independently pretrained weights such as those from a GPT-2 model, we observe that this strategy is suboptimal and performance is significantly improved by optimizing the encoder and the decoder jointly.

For the task of multimodal video captioning, we require a model that can both encode multimodal videos (i.e. frames and textual inputs) and generate captions. Using multimodal information as input can greatly improve the quality of the generated captions (as illustrated in Figure 1(a)). However, learning such an encoder-decoder model jointly from unlabelled data is particularly challenging, as it requires two streams of textual data – naturally occurring transcribed speech accompanying the video for the encoder, and target sentences for the decoder – whereas unlabelled videos only come with a single stream of speech (Figure 1(b)). Recent works have attempted to solve this problem with a denoising autoencoder - wherein the input speech to the model is artificially ‘noised’, i.e. random words are masked out . The decoder is then tasked with simply reconstructing either the masked phrases or the original unmasked text, where the supervisory signals are provided only from the masked words. In these frameworks, additional losses are often required to strengthen the pretraining supervision, such as multimodal input alignment and segment ordering .

In our framework, we introduce a novel stronger loss. We leverage future utterances as another source of textual data and train a model to generate these entirely unseen sentences as depicted in Figure 1(b). To alleviate the problem that future utterances are not temporally aligned, we propose a backward generation objective where present aligned utterances are generated given future utterances. Experimental results show that a model pretrained with this bidirectional generation objective effectively transfers to multimodal video captioning and outperforms the state of the art by a margin.

We make the following contributions: (i) We propose a novel pretraining objective for multimodal video captioning that requires no manually annotated captions, and instead uses utterances sampled at different times in the same video. Our objective is bidirectional in time – i.e. we not only generate future utterances but also the present ones from the future; (ii) By using two sources of textual data, we are able to jointly train the entire encoder-decoder model. This is unlike previous works which pretrain only the (multimodal) encoder, thereby lacking the ability to generate captions ; (iii) Our encoder is trained from raw pixels and words directly, in contrast with existing methods that rely on pre-extracted visual features limiting transfer to new domains ; (iv) We achieve state-of-the-art results on four video captioning benchmarks – YouCook2, ViTT, MSR-VTT and ActivityNet-Captions – consistently outperforming existing methods by significant margins; and finally (v) Our pretraining objective yields strong multimodal video representations, which achieve state-of-the-art performance on other video understanding tasks such as VideoQA, video retrieval and action classification.

Related Work

Video captioning. Early works in video captioning consisted of rule-based methods , where subjects, verbs and objects (SVO-triplets) detected from the video were combined into sentence templates. Later work moved away from rule-based methods by framing captioning as a machine translation task , which developed the common encoder-decoder paradigm of today for the task – the encoder processes a set of video features and accumulates its hidden state, which is then passed to a decoder for producing a caption. Early works implemented the visual encoder as a 2D CNN (either frozen or finetuned) applied to video frames, which was then naturally extended to 3D CNNs , to better capture motion dynamics, with temporal aggregation over the entire video typically performed using attention strategies . Given the computational challenge of using expensive 3D CNNs applied to dense frame inputs (typically 30 fps), most of these works operated on pre-extracted features, only learning the fusion of features in the encoder. Unlike such works, we address this problem using a transformer-based encoder applied to raw pixels , sampled at a coarse rate to better capture long range context.

Pretraining with weakly paired data. Existing video captioning datasets are orders of magnitude smaller than video classification datasets . As a source of weakly paired video and language data, a number of works have used the visual frames and the Automatic Speed Recognition (ASR) transcripts of unlabelled videos to pretrain video representations . These approaches learn multimodal representations by formulating proxy tasks such as masked language/frame modeling , video-text matching or segment ordering . While these studies show improvements on visual representation or multimodal video representation learning, they are designed for discriminative tasks only, and lack the generation capability. Pretraining techniques for generative tasks such as ours, are fewer. While use multimodal translation as a generative objective, their encoder is limited to accept visual inputs only. Works that use multimodal inputs to the encoder, train with masking losses – wherein words or phrases are masked and the objective is to reconstruct the original sentences or the masked targets using an autoregressive generator. In contrast, we make use of utterances outside of the clip boundary, which are simply ignored in previous works. We leverage future utterances as a second source of textual data, and propose a bi-directional generation objective where the model generates the future utterance given the current utterance and vice versa. While we also use a masked language modelling loss, this is simply in addition to our primary generative bidirectional loss.

Method

Our objective is to pretrain a model that can effectively encode multimodal videos (visual frames and transcribed speech) as well as decode natural language sentences. This will allow us to use the model for multimodal captioning. In this section, we first describe the pretraining losses used to train the encoder and decoder jointly from unlabelled videos. We then describe our model, which consists of modality specific encoders, a multimodal encoder and a text decoder (Figure 1).

Our framework is designed to take advantage of unlabelled instructional video data, which consists of video frames and utterances often linked to the visual content . As mentioned earlier, our framework requires two textual streams – an input to the encoder and a captioning target for the decoder. Because unlabelled videos do not have captioning targets, we instead propose a simple objective – our model is trained to generate a future utterance in the video given the current video context and current utterances (forward generation). This gives us two sources of textual supervision, the current utterance allows us to learn how to optimally fuse modalities in the video encoder, while the decoder is tasked with predicting a new utterance it has never seen before. However, our goal is video captioning, and not ‘predicting the future’. To enable our model to generate text corresponding to the present video context, we also add in an additional backward generation loss – where the model must generate the current utterance given the current video frames and a future utterance (backward generation). This encourages generated sentences to be temporally aligned (and hence more tightly coupled) with the visual inputs.

Given a large set of unlabelled videos, we extract short clips consisting of visual frames F={f1,…,fNf}F=\{f_{1},\dots,f_{N_{f}}\} and transcribed speech utterances U={u1,…,uNu}U=\{u_{1},\dots,u_{N_{u}}\} aligned with FF. For each clip, we also consider the immediate future utterance W={w1,…,wNw}W=\{w_{1},\dots,w_{N_{w}}\} where uiu_{i} and wjw_{j} are tokenized words in the transcribed utterances. Note that we use the term ‘utterance’ to refer to a single sentence of transcribed speech.

Unlike UniVL where the MLM loss is applied to the outputs of the encoder, we apply it to the outputs of the decoder. This encourages the self attention layers in the decoder to focus on further multimodal contextualization of the textual tokens (since each masked token prediction requires knowledge of neighbouring context). As we show in the experiments, this leads to performance gains.

2 Model

Our model consists entirely of transformer blocks, and is trained end-to-end directly from pixels and word tokens.

Given a multimodal video input consisting of the visual frames F={f1,…,fNf}F=\{f_{1},\dots,f_{N_{f}}\} and text inputs X={x1,…,xNx}X=\{x_{1},\dots,x_{N_{x}}\}, we first extract features from the individual modalities independently. Note here that the textual input XX is the aligned utterance UU in general (for computing the forward generation loss and for downstream captioning tasks) but is set to WW when computing the backward generation loss.

Textual Encoder: We extract NxN_{x} contextualized textual embeddings E={ei}E=\{e_{i}\} from the input text using a BERT encoder.

Visual Encoder: Unlike previous approaches where visual features are pre-extracted by models pretrained on different datasets, we extract the visual features directly from pixels. We use the recent transformer-based video encoder ViViT , in particular, the tubelet embedding scheme and the factorized encoder architecture. For the tubelet embedding scheme we first extract spatio-temporal 3D tubes from the visual input volume resulting in S×TS\times T token embeddings where SS and TT correspond to the numbers of tokens in the spatial and temporal dimensions, respectively. Then, the spatial transformer first takes each group of SS embeddings from the same temporal index with a special CLS token embedding, and the temporal transformer models interactions between the output CLS embeddings of the individual spatial groups with another CLS embedding resulting in T+1T+1 visual features V={vj}V=\{v_{j}\} – see for further details.

Unlike 3D CNN visual encoders which operate on consecutive frames extracted at high frame rates (30 fps), our visual encoder can operate on coarsely sampled frames (1 fps), thus significantly reducing compute. This allows us to train the visual encoder end-to-end, and helps adapt our features across the domain gaps between pretraining and downstream datasets. It also allows the easy adoption of off-the-shelf video augmentation directly to RGB frames, which is useful for small-scale downstream benchmarks.

Once the two sets of textual features EE and visual features VV are extracted, our multimodal encoder fuses multimodal information using the co-attentional transformer used in . Each layer consists of two streams where each stream is a stack of two transformer blocks. In the textual stream, we first contextualize the features EE using a cross-attention transformer block attending to the visual features VV. Then, the output features are further contextualized by another transformer block with self-attention. The first transformer block performs inter-modality contextualization through a cross-attention process whereas the second transformer block carries out intra-modality contextualization through a self-attention process. In the same way, the visual stream VV attends to the textual stream. Our multimodal encoder repeats this process RR times resulting in the output multimodal features E^\hat{E} and V^\hat{V}.

Pretraining: Since our pretraining objective is bidirectional, each triplet (F,U,W)(F,U,W) consisting of the visual frames FF, the present utterances UU and the future utterance WW is processed by the network twice. For forward generation, the model takes FF and UU as inputs and generates WW, and it generates UU given FF and WW, in backward generation. To enable the model to recognize the different configurations, we attach distinct, special tokens CLS1 and CLS2 to the input text for the forward and backward generation losses respectively as illustrated in Figure 1. Similarly, we feed distinct BOS1 and BOS2 tokens to the decoder to initiate sentence generation.

Finetuning for captioning: In downstream video captioning datasets, video clips (consisting of frames FF and aligned utterances UU) are manually annotated with a natural language caption. During finetuning, we attach the CLS1 token to UU (as is done in forward generation), since UU is an aligned utterance, but for generation we feed in the BOS2 token (as is done in backward generation to predict the present utterance), so that we also generate a temporally aligned caption.

For our text encoder, we adopt the BERT-Base architecture with uncased wordpiece tokenization . Our visual encoder uses the corresponding ViViT-Base configuration with a 1-layer temporal transformer and a tubelet size of 16×16×416\times 16\times 4 . Our multimodal encoder consists of 2 layers following and finally, the decoder is based on the GPT-2 (117M parameters) architecture but we modify it to take both multimodal input context CC and a BOS token allowing conditional generation (the original GPT starts generation immediately by taking the first word as its input and only conditions on text). We initialize the text encoder and the decoder with the standard BERT and GPT-2 weights respectively pretrained on large-scale unlabelled corpora . Similarly, we initialize our visual encoder using the pretrained weights on Kinetics 400 in unless otherwise specified. Our model is pretrained end-to-end using the Adam optimizer for 1.5M iterations with the batch size of 2048. For more detailed hyperparameters and training strategies for pretraining and finetuning, please refer to the appendix.

Experiments

In this section, we first demonstrate our results on four different benchmarks for multimodal video captioning. We then also show that our pretrained model has the ability to generalise to other video understanding tasks such as video question answering (VideoQA), video retrieval and action classification.

We use HowTo100M as our pretraining dataset, and evaluate on four downstream captioning benchmarks. HowTo100M consists of 1.2M instructional videos from YouTube. Transcribed speech is obtained using the YouTube ASR API . Following , we extract 53M triplets of frames, current utterances and future utterances for pretraining.

YouCook2 is the most widely adopted benchmark for multimodal video captioning and contains 2,000 cooking videos for 89 different dishes with 14K video clips. Each video clip is annotated with a single captioning sentence.

Video Timeline Tags (ViTT) was created to better reflect the distribution of instructional videos in the wild. It consists of 8,169 videos, 5,840 of these videos for training and the remaining videos for validation and testing. Videos are divided into 7.1 segments on average, with each segment accompanied by a short timeline tag.

MSR-VTT is a standard benchmark with 10K open domain video clips for video captioning. The duration of each video clip is between 10 and 30 seconds, and 20 natural language descriptions are manually annotated per clip.

ActivityNet-Captions is a standard dense video captioning benchmark consisting of 100K temporally localized sentences for 20k videos. We follow the standard splits with 50/25/25% examples for training, validation and test sets. To evaluate our model’s ability to predict captions, we use ground truth temporal proposals following .

We pretrain a single model on HowTo100M, which is then transferred to all four captioning benchmarks through finetuning. We report results using the following established metrics: BLEU-4 (B-4) , CIDEr (C) , METEOR (M) and ROUGE-L (R-L) . For ViTT, we measure BLEU-1 (B-1) instead of BLEU-4 following .

In this section we ablate some key design choices, in particular the backbone and objective functions used in MV-GPT, and explore the impact of the end-to-end training. Finally, we compare our model to the state of the art.

Pretraining Losses: We implement a simple baseline, which consists of a masked language modelling loss given visual frames and ASR as input (Baseline PT). We also reimplement three state-of-the-art pretraining losses: (i) CoMVT , (ii) UniVL and (iii) M-MASS . For a fair comparison, we use our model architecture for all experiments, varying the loss function only. For the methods which pretrain the encoder only, we initialise the decoder with public GPT-2 weights . For ‘No PT’, the encoder is not pretrained either, but is initialized with public BERT and ViViT pretrained on ImageNet21k.

Table 1 compares these different losses. We can observe that pretraining the encoder only brings moderate gains over training from scratch, for all the losses investigated. This performance is greatly improved by pretraining both the encoder and decoder jointly. Finally, we observe that our approach MV-GPT outperforms existing joint pretraining losses.

Effect of each Loss Term in MV-GPT: Table 2 shows the effect of each term in our loss function. The forward generation (FG) loss already provides strong supervision. When applying the masked language modelling loss on the decoder outputs (MLM-D) instead of the encoder outputs (MLM-E), performance is slightly improved due to the additional input contextualization provided by the decoder. Adding the backward generation (BG) loss provides a boost across all metrics. Additionally, we observe that adding weight decay (WD) brings additional gains, and we report our scores in this full setting for the rest of the paper.

Visual Encoder and End-to-end Training: In Table 3, we first compare the ViViT encoder to commonly used S3D features . When both encoders are trained on Kinetics and fixed for multimodal pretraining and finetuning, they show comparable scores despite the large complexity of S3D due to the high frame rate required (30 fps vs. 1 fps for ViViT). Using HowTo100M to train a visual encoder, we observe large gains with both architectures as expected given the similarity in the domains – HowTo100M and YouCook2 are both instructional video datasets. However, we observe larger gains with ViViT where the visual encoder is optimized for generative losses within our framework and jointly trained with the other components thanks to the low complexity of the ViViT encoder. These results show the benefits of end-to-end pretraining.

We further investigate the effects of end-to-end training for finetuning. For YouCook2, we observe slight performance degradation when naively finetuning the network end-to-end from the beginning (row 4 to 5). This degradation is overcome by initially freezing the visual encoder and starting end-to-end training after convergence, which gives us a minor gain (row 6). These results indicate that our pretrained visual encoder already captures strong representations for inputs in a similar domain, and end-to-end finetuning is less critical in this case. However, we observe more significant gains on MSR-VTT since end-to-end finetuning becomes crucial given a larger domain gap (row 7 to 8).

Pretraining with Random Initialization: We also investigate the ability of the model to learn from scratch. We initialize the model either entirely randomly or using pretrained BERT, ViViT and GPT-2 weights. Table 4 shows that with random initialization, our method still performs very well (row 2), outperforming the model initalized with public BERT, GPT-2 and ViViT weights (row 3). Note that the pretrained ViViT weights were obtained from training on the fully supervised dataset Kinetics. Also, pretraining entirely from scratch even approaches the case where all parts of the model are intialized using public weights and pretrained (row 4).

Multimodal vs. Single Modality: In Table 5, we show results with text only and visual only inputs (we only feed the CLS token for the omitted modality). It is clear that both modalities are complementary and performance is best when combining both. Additionally, to assess the contribution of the visual modality, we test a model pretrained with text inputs only. Even when this pretrained model is finetuned with both modalities, the performance is significantly lower compared to a pretrained multimodal model (last row in Table 2): there is a 25% relative drop on all 4 metrics (e.g., 1.43 vs. 2.14 in CIDEr). When finetuned with text inputs only, the scores drop further (e.g., to 1.20 in CIDEr). These results confirm the importance of the visual inputs during pretraining.

Comparisons to the State of the Art: Finally, we compare MV-GPT to existing methods on all four datasets. Table 5 compares our method to the state of the art on YouCook2, where we outperform all prior work including works pretrained on HowTo100M. On ViTT (Table 6), the gap is even larger, with our model advancing the state-of-the-art by 15% (absolute) compared to M-MASS in B-1 and M scores.

Despite the domain gap between instructional videos in HowTo100M and general online videos in MSR-VTT, our model outperforms all existing work as shown in Table 7. Although UniVL also pretrains both the encoder and the decoder on HowTo100M, our method achieves relative improvements of over 31% thanks to our end-to-end training. Similarly, Table 8 shows that our pretraining method achieves state-of-the-art performance on ActivityNet-Captions despite the significant domain gap.

Qualitative Results: We show examples from YouCook2 and MSR-VTT in Figure 2. The first example illustrates that our model can use the visual modality to infer the term ‘sauce’ despite the ASR error ‘source’ and further recognizes its name ‘sriracha’. Similarly, the second example illustrates that our approach manages to take into account both modalities jointly. Finally, we show a failure case in the last row in which our model fails to capture the concept ‘ski lift’. A possible explanation is that the concept of a ski lift may be rarely seen in the pretraining dataset, a problem which may be alleviated by collecting more diverse pretraining videos, or incorporating external object knowledge through the use of pre-trained object detectors.

2 Non-generative Video Understanding Tasks

Although MV-GPT is a generative model and is particularly designed for multimodal video captioning, we also find that our pretraining technique learns a powerful multimodal video encoder that can be transferred easily to multiple video understanding tasks. In particular, we show results on VideoQA, video retrieval and action classification. For details on each task please refer to the appendix.

VideoQA: We use MV-GPT as an encoder (no BOS token is fed to the decoder so it only contextualizes the input tokens; see appendix for details) and the average pooled input embedding is fed to a two-layered MLP classifier to predict the answer. The question is simply concatenated to the ASR inputs. Following the standard protocols in , we measure the answer prediction accuracy on MSRVTT-QA and ActivityNet-QA .

Table 9 compares the accuracy of MV-GPT to existing methods that are pretrained on HowTo100M . Even though MV-GPT is not designed for this particular task, our model slightly outperforms the previous state-of-the-art VQA-T (which is specifically designed for VideoQA) on both datasets.

Video Retrieval: The common practice for retrieval is to train a video-text joint embedding using discriminative losses only, typically in the form of a standard NCE loss , where each video clip has a single corresponding textual caption. Here we investigate whether our generative pretraining loss can provide a boost to performance. Since each example forms two inputs-target triplets in our bidirectional framework, we apply NCE losses on both (Bi-NCE). We then add our generative pretraining loss to this framework and report results in Table 10. We evaluate our model with and without ASR to compare fairly to existing works. We report recall at k={1,5,10}k=\{1,5,10\} (R@kk) and median rank (MdR) on MSR-VTT following the standard 9K retrieval splits .

Our first observation is that our Bi-NCE serves as a strong baseline pretraining method for retrieval. We show that adding our generative losses further improves performance by a relative 6.3% in R@1, yielding state-of-the-art performance. Finally, adding ASR to our multimodal encoder further improves performance by a significant margin (+ 4%).

Action Classification: We test the visual encoder of MV-GPT on action classification following . We evaluate models using top-1 classification accuracy on Kinetics 400 and 600 . Note that we adopt the ViViT-Base architecture with factorized encoder following , however we use a tubelet size of 16×16×416\times 16\times 4 instead of 16×16×216\times 16\times 2 to reduce complexity. We compare our model with two different initializations for the visual encoder: random and pretrained weights on ImageNet21k. The baseline models are finetuned on the evaluation benchmarks immediately from these initializations whereas we first post-pretrain models in our MV-GPT framework and finetune for action classification.

Table 11 demonstrates that MV-GPT is an effective pretraining strategy for the visual encoder. High-capacity transformer models like ViViT are challenging to train from scratch, and overfit easily as shown in the first row. However, ViViT initialized from an MV-GPT visual encoder trained from scratch performs substantially better, obtaining absolute improvements of 24% on Kinetics-400 (a standard video classification benchmark). This number is close to the performance of ViViT initalized with ImageNet-21K pretraining, as done by the original authors (note that ImageNet-21K was created with high manual annotation cost, while we used no labels at all during pretraining). Finally, initialising the MV-GPT visual encoder with these same ImageNet-21K weights, and then pretraining the MV-GPT visual encoder weights on HowTo100M achieves the best results, improving upon the initialisation of by 1.5% and 1.8% on Kinetics-400 and Kinetics-600 respectively, which is the current state of the art on this dataset with this particular architecture.

Conclusion

We present a novel generative pretraining framework for multimodal video captioning. Our bi-directional generative objective jointly trains an encoder for multimodal inputs and a decoder to generate meaningful captions, by using utterances sampled at different times in unlabelled videos. The model is trained end-to-end both during pretraining and finetuning, and achieves state-of-the-art results on multiple video captioning benchmarks as well as on other video understanding tasks, namely VideoQA, video retrieval and action classification.

References

Appendix

In this appendix, we first provide additional experimental results and descriptions on the dataset configurations for our pretraining and downstreams tasks in Section A and B. Further implementation details for the downstream tasks are described in Section C. We then present more qualitative results in Section D. Finally, we discuss limitations and broader impacts of our method in Section E

We perform additional ablations on MSR-VTT (mirroring Table 1 in the main manuscript) and provide results in Table 12. We observe similar trends as on YouCook2 (Table 1), albeit with smaller gaps. We believe this is due to the larger domain gap between HowTo100M and MSR-VTT.

A.2 Impact of Pretraining Dataset Size

Figure 3 reports performance against dataset size. All four metrics show improvement (almost linear) when the dataset size is doubled. This signifies that our model could improve further by collecting more unlabelled videos for pretraining.

A.3 Open-ended Generative VideoQA

To further investigate our model’s decoding capability, we test our model on the open-ended long-form VideoQA (OL-VideoQA) benchmark . Note that the training set is smaller than the one reported in (26K vs. 53K examples) although we obtained the dataset directly from the authors. We test our model with/without pretraining to show its effectiveness and report the scores in B-1 and WUPS@α\alpha metrics where α\alpha is a threshold for word similarity (see for details). In Table 13, our model without pretraining (No PT) serves as a strong baseline outperforming almost all the scores of the existing methods despite the fewer training examples used. The pretrained model (MV-GPT) then boosts performances further in all the metrics. Note that the gaps in WUPS@0.0 are relatively small since all soft matches are equally weighted regardless of their semantic similarities.

A.4 Impact of Decoder as a Part of Encoder

As described in the main manuscript and depicted in Figure 4d, we use the pretrained decoder as a part of the encoder for the VideoQA model. To investigate the effectiveness of our decoder when used as a part of an encoder, we compare our model with and without the decoder for VideoQA, and observe a 1.0% and 0.8% gain in accuracy with the decoder on the MSRVTT-QA and ActivityNet-QA benchmarks respectively.

Appendix B Datasets

We prepare our pretraining dataset following and extract triplets (F,U,W)(F,U,W) of the video frames FF, the present utterance UU, and the future utterance WW, from the videos in HowTo100M . We obtained transcribed speech using the YouTube ASR API YouTube Data API. https://developers.google.com/youtube/v3/docs/caption, however these are noisy. To respect licensing terms, videos that have been removed from YouTube since the dataset was originally created are not used. We then divide these videos into shorter video clips. The duration of video clips is determined as follows: we start with a single ASR sentence and then iteratively expand the length of the video clip backwards by adding previous sentences until the segment is longer than 5 seconds. Each video clip therefore contains full sentences (no sentences are cut-off mid way). This process results in 53.5M training examples. Since we focus on the pretraining approach, we keep only 7.5K examples as a small validation split.

B.2 Datasets for Non-generative Tasks

In addition to the datasets used for multimodal video captioning, which are described in the main manuscript, we make use of the following datasets for the experiments on the non-generative video understanding tasks.

MSR-VTT is commonly adopted for video retrieval. We follow the standard splits for retrieval containing 9K and 1K examples in train and test sets, respectively.

MSRVTT-QA is a VideoQA benchmark derived from MSR-VTT, and contains 243K QA pairs. The dataset follows the standard splits released in MSR-VTT .

ActivityNet-QA contains 58K QA pairs for VideoQA where the train, val and test sets have 32K, 18K and 8K pairs, respectively.

Kinetics is the largest action classification benchmark. We evaluate on both Kinetics 400 and 600, containing approximately 267K clips from 400 classes and 446K clips from 600 classes, respectively.

Appendix C Implementation Details

As described in the main manuscript, we pretrain our model by the proposed bidirectional loss, which consists of the forward and backward generation loses. As described in the main manuscript, our framework pretrains a model consisting of a visual encoder (VE), a text encoder (TE), a multimodal encoder (MME) and a decoder (Figure 4a). After pretraining, a different subset of these components depending on the downstream task is transferred and finetuned, which is described in the following sections.

For pretraining, we initialize the text encoder and the decoder with the standard BERT and GPT-2 weights respectively pretrained on large-scale unlabelled corpora . Similarly, we initialize our visual encoder using the pretrained weights on Kinetics 400 in . Our entire model is pretrained end-to-end using the Adam optimizer for 1.5M iterations with the batch size of 2048. We adopt a weight decaying factor of 0.01, and use the cosine learning rate decay with a linear warm-up of 500 iterations.

C.2 Multimodal Video Captioning

Given an MV-GPT model pretrained on HowTo100M (Figure 4a), the entire pretrained MV-GPT is transferred for multimodal video captioning as our main target task as depicted in Figure 4b. The differences are the input and output configurations as described in the main manuscript; during pretraining, we feed present utterances (PU) as inputs predicting future utterances (FU) in forward generation and vice versa in backward generation whereas our model predicts captions given present utterances for captioning. Note also that we feed a special BOS token to initiate the sentence generation from the decoder.

We finetune the entire model end-to-end for 1K iterations with an initial learning rate of 0.0001 and a batch size of 512, and use the best validation checkpoint selected based on the Meteor score. For testing, we perform beam search with a beam size of 5 as in . Note that we initialize the decoder using the weights of GPT-2 when we test models trained by encoder-only pretraining methods.

C.3 Generative VideoQA

Generative VideoQA requires generating an open-ended answer given multimodal video and a question. While a question is given as an additional text input, we simply concatenate it to the present utterance; this allows us to use the original MV-GPT model for this task without any change as depicted in Figure 4c.

C.4 VideoQA

Following previous work , we formulate this task as a classification problem of predicting a predefined answer class. Note that we simply concatenate the input question to the utterances from the clip and feed the concatenation as a single textual input. Although we do not decode any textual outputs in this task, we still make use of the decoder as an additional multimodal encoder since our decoder is also trained to contextualize the input embeddings by applying the masked language modeling on the decoder outputs (see Section 3.1.2 in the main manuscript). Instead of feeding the BOS token and predicting next tokens, we first obtain the embeddings of the inputs from the decoder, average-pool these embeddings, and feed the pooled embedding to a two-layered MLP classifier to predict the answer (Figure 4d). Note that we use the entire pretrained model but append a randomly-initialized classifier.

For every experiment, we finetune the entire model end-to-end on the downstream benchmark for 20K iterations with a batch size of 512 and report the results using the checkpoint with the best answer accuracy.

C.5 Action Classification

Our goal with the experiments in action classification is to show the effectiveness of the pretrained visual encoder in MV-GPT, and therefore we simply discard all the other components and append a randomly initialized classification layer to the visual encoder as illustrated in Figure 4e. For finetuning, we follow all the exact evaluation protocols used in .

C.6 Video Retrieval

The common practice for retrieval is to train a video-text joint embedding using discriminative losses only, typically in the form of a standard NCE loss , where each video clip has a single corresponding textual caption. In the retrieval experiments, we investigate whether our generative pretraining loss can provide a boost to performance. Since each example forms two inputs-target triplets, i.e. (F,U,W)(F,U,W) and (F,W,U)(F,W,U), in our bidirectional frameworks, we apply NCE losses on both (Bi-NCE; Figure 5a). Note that we use an additional textual encoder to compute embeddings of the target texts. We then add our generative pretraining loss to this framework (Figure 5b). Finally for finetuning, we transfer the visual/textual/multimodal encoders and the additional text encoder of the pretrained model, and train the network using an NCE loss with the text query provided in the downstream benchmark.

For pretraining, we down-weight the bidirectional NCE losses with a factor of 0.001, and follow the same hyper-parameters used in the regular MV-GPT pretraining. For finetuning, we train the entire network end-to-end for 1K iterations with a batch size of 512 and we report the scores from the best checkpoints on the validation set.

Appendix D More Qualitative Examples

We show more qualitative examples on YouCook2 in Figure 6. These qualitative examples demonstrate that our MV-GPT model can capture both textual cues (e.g., the word ‘parsely’ in the first example) and visual cues (e.g., the action of ‘spreading sauce’ in the last example) whereas the model without pretraining is often unable to capture these.

Appendix E Limitations and Broader Impact

Limitations: Our approach is not always successful, in particular in the presence of a significant domain shift between the pretraining data and the downstream application. Future work will address this limitation, for example by collecting curated pretraining data.

Broader Impact: Large, uncurated datasets scraped from the web may contain unintended biases, and models pretrained on such datasets may inadvertently amplify these biases. Therefore, applications of our work beyond the academic setting presented here should first carefully examine and filter the pretraining dataset for potentially harmful biases in the data.