Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning

Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont-Tuset, Ivan Laptev, Josef Sivic, Cordelia Schmid

Introduction

Dense video captioning requires the temporal localization and captioning of all events in an untrimmed video . This differs from standard video captioning , where the goal is to produce a single caption for a given short video clip. Dense captioning is significantly more difficult, as it raises the additional complexity of localizing the events in minutes-long videos. However, it also benefits from long-range video information. This task is potentially highly useful in applications such as large-scale video search and indexing, where the video content is not segmented into clips.

Existing methods mostly resort to two-stage approaches , where events are first localized and then captioned. To further enhance the inter-task interaction between event localization and captioning, some approaches have introduced models that jointly solve the two tasks . However, often these approaches still require task-specific components such as event counters . Furthermore, they exclusively train on manually annotated datasets of limited size , which makes it difficult to effectively solve the task. To address these issues, we take inspiration from recent sequence-to-sequence models pretrained on Web data which have been successful on a wide range of vision and language tasks .

First, we propose a video language model, called Vid2Seq. We start from a language model trained on Web text and augment it with special time tokens that represent timestamps in the video. Given video frames and transcribed speech inputs, the resulting model jointly predicts all event captions and their corresponding temporal boundaries by generating a single sequence of discrete tokens, as illustrated in Figure 1 (right). Such a model therefore has the potential to learn multi-modal dependencies between the different events in the video via attention . However this requires large-scale training data, which is not available in current dense video captioning datasets . Moreover, collecting manual annotations of dense captions for videos is expensive and prohibitive at scale.

Hence we propose to pretrain Vid2Seq by leveraging unlabeled narrated videos which are readily-available at scale. To do this, we reformulate sentence boundaries of transcribed speech as pseudo event boundaries, and use the transcribed speech sentences as pseudo event captions. We then pretrain Vid2Seq with a generative objective, that requires predicting the transcribed speech given visual inputs, and a denoising objective, which masks spans of transcribed speech. Note that transcribed speech may not describe the video content faithfully, and is often temporally misaligned with the visual stream . For instance, from the example in Figure 1 (left), one can understand that the grey skier has descended a slope from the last speech sentence which is said after he actually descended the slope. Intuitively, Vid2Seq is particularly suited for learning from such noisy supervision as it jointly models all narrations and the corresponding timestamps in the video.

We demonstrate the effectiveness of our pretrained model through extensive experiments. We show the importance of pretraining on untrimmed narrated videos, the ability of Vid2Seq to use both the visual and speech modalities, the importance of the pretraining objectives, the benefit of joint caption generation and localization, as well as the importance of the language model size and the scale of the pretraining dataset. The pretrained Vid2Seq model achieves state-of-the-art performance on various dense video captioning benchmarks . Our model also excels at generating paragraphs of text describing the video: without using ground-truth event proposals at inference time, our model outperforms all prior approaches including those that rely on such proposals . Moreover, Vid2Seq generalizes well to the standard task of video clip captioning . Finally, we introduce a new few-shot dense video captioning setting in which we finetune our pretrained model on a small fraction of the downstream training dataset and show benefits of Vid2Seq in this setting.

In summary, we make the following contributions: (i) We introduce Vid2Seq for dense video captioning. Given multi-modal inputs (transcribed speech and video), Vid2Seq predicts a single sequence of discrete tokens that includes caption tokens interleaved with special time tokens that represent event timestamps. (ii) We show that transcribed speech and corresponding timestamps in unlabeled narrated videos can be effectively used as a source of weak supervision for dense video captioning. (iii) Finally, our pretrained Vid2Seq model improves the state of the art on three dense video captioning datasets (YouCook2, ViTT, ActivityNet Captions), two video paragraph captioning benchmarks (YouCook2, ActivityNet Captions) and two video clip captioning datasets (MSR-VTT, MSVD), and also generalizes well to few-shot settings.

Our code implemented in Jax and based on the Scenic library is publicly released at .

Related Work

Dense video captioning. Dense video captioning lies at the intersection of event localization and event captioning . The majority of existing methods for dense video captioning consist of a temporal localization stage followed by an event captioning stage. To enrich inter-task interactions, recent works jointly train the captioning and localization modules. In particular, Wang et al. propose to view dense video captioning as a set prediction task, and jointly perform event localization and captioning for each event in parallel. In contrast, our model generates event boundaries and captions conditioned on the previously generated events. Deng et al. propose to first generate a paragraph and then ground each sentence in the video. We also generate all captions as a single output sequence, however our output already includes event timestamps. Zhang et al. propose to generate event boundaries sequentially, but separately perform event localization and single event captioning, and only use visual input. Most related to our work, Zhu et al. also perform dense video captioning by generating a single output sequence. Their method, however, infers event locations directly from the timestamps of transcribed speech and, hence, can only detect events that closely follow the speech. In contrast, our model generates event timestamps as special tokens and can produce dense captions for videos with limited speech, as we demonstrate on the ActivityNet Captions dataset.

Video and language pretraining. Following the success of image-text pretraining , recent works have explored video-text pretraining . These methods show strong improvements on various tasks such as text-video retrieval , video question answering and video clip captioning . While these works mostly learn global video representations to tackle video-level prediction tasks, we here focus on learning detailed representations to address a dense prediction task requiring reasoning over multiple events in untrimmed videos. Several works have explored long-form video-text pretraining and video-text pretraining for temporal localization tasks . However these works focus on video understanding tasks while our pretraining approach is tailored for a generative task that not only requires the model to reason over multiple events in the video, but also to describe them by natural language.

A few works explore pretraining for dense video captioning. Zhang et al. pretrain on ActivityNet Captions to improve the downstream performance on the same dataset. In contrast, we propose a pretraining method that does not rely on any manual annotation, and show its benefits on multiple downstream datasets. Huang et al. explore pretraining on narrated instructional videos, but only consider event captioning using ground truth proposals as their model does not handle localization. Finally, explore pretraining on a domain specific text-only dataset . In contrast, we propose to pretrain on a generic video corpus and show benefits on various domains.

Unifying tasks as language modeling. Recent works have shown that it is possible to cast various computer vision problems as a language modeling task, addressing object detection , grounded image captioning or visual grounding . In this work we also cast visual localization as a language modeling task. However, unlike prior work focused on image-level spatial localization, we address the different problem of event localization in time, in untrimmed videos.

Method

The goal of dense video captioning is to temporally localize and describe with natural language all events in an untrimmed input video. Therefore a key challenge is to effectively model the relationships between the different events in the video, as for example, it is easier to predict that the dogs are pulling the sled if we know that the man has just fastened a dog (see Figure 1 (right)). Furthermore, due to the dense nature of the task, there can be many events in a long video and the requirement is to output a natural language caption for each event. Hence, another key challenge is that the manual collection of annotations for this task is particularly expensive. To tackle these challenges, we first develop a unified multi-modal model that jointly predicts event boundaries and captions as a single sequence of tokens, as explained in Section 3.1 and Figure 2. Second, we design a pretraining strategy that effectively leverages cross-modal supervision in the form of transcribed speech from unlabeled narrated videos by reformulating sentence boundaries as pseudo event boundaries, as presented in Section 3.2 and Figure 3.

We wish to design a model for dense video captioning that can capture relationships between events using visual and (transcribed) speech cues in order to effectively localize and describe these events in untrimmed minutes-long videos. To tackle this challenge, we cast dense video captioning as a sequence-to-sequence problem where the input and output sequences contain both the semantic information about the event in the form of natural language descriptions and the temporal localization of the events in the form of temporal timestamps. In addition, to best leverage both the visual and the language signal, we develop an appropriate multi-modal encoder-decoder architecture. As illustrated in Figure 2, our architecture takes as input video frames x={xi}i=1Fx=\{x_{i}\}_{i=1}^{F} together with the transcribed speech sequence y={yj}j=1Sy=\{y_{j}\}_{j=1}^{S}. The output of our model is an event sequence z={zk}k=1Lz=\{z_{k}\}_{k=1}^{L}, where each event contains both its textual description and timestamps corresponding to the temporal event locations in the video. Below we explain the structure of the transcribed speech and event sequences constructed for our model as well as details of our model architecture.

To model inter-event relationships in dense event captioning annotations (or the readily-available transcribed narration, see Section 3.2), we cast dense video captioning as predicting a single output sequence of tokens zz. This output event sequence is constructed by leveraging a text tokenizer augmented with special time tokens. Furthermore, we enable our architecture to jointly reason about the semantic and temporal information provided in the transcript of the input narration by constructing the input transcript sequence yy in a similar manner as the event sequence zz. Details are given next.

Time tokenization. We start from a text tokenizer with a vocabulary size VV, and augment it with NN additional time tokens, resulting in a tokenizer with V+NV+N tokens. The time tokens represent relative timestamps in a video, as we quantize a video of duration TT into NN equally-spaced timestamps. In detail, we use the SentencePiece tokenizer with vocabulary size V=32,128V=32,128 and N=100N=100.

Event sequence. Our introduced tokenizer enables us to construct sequences that contain both video timestamps and text video descriptions. We next explain how we construct the output event sequence zz. Note that videos have a variable number of events in standard dense video captioning datasets . Each event kk is characterized by a text segment, a start time and an end time. We first construct for each event kk a sequence by concatenating its start time token tstartkt_{start_{k}}, its end time token tendkt_{end_{k}} and its text tokens [zk1,...,zklk][z_{k_{1}},...,z_{k_{l_{k}}}]. Then we order all these sequences in increasing order of their start times and concatenate them. In practice, each text segment ends with a dot symbol indicating the separation between different events. Finally, the event sequence is obtained by prepending and appending a BOS and an EOS tokens to indicate the start and the end of sequence, respectively, i.e. z=[BOS,tstart1,tend1,z11,...,z1l1,tstart2,...,EOS]z=[BOS,t_{start_{1}},t_{end_{1}},z_{1_{1}},...,z_{1_{l_{1}}},t_{start_{2}},...,EOS].

Transcribed speech sequence. To enable the model to use both the transcribed speech and its corresponding timestamps, we convert the speech transcript into a speech sequence yy similarly as the input training dense event captions zz. This is done by segmenting the raw speech transcript into sentences with the Google Cloud API1, and using each transcribed speech sentence with its corresponding timestamps analogously as an event in the previously explained process.

We wish to design an architecture that can effectively model relationships between different events in untrimmed minutes-long videos. To tackle this challenge, we propose a multi-modal encoder-decoder architecture, illustrated in Figure 2, that gradually refines and outputs the event sequence described above. In detail, given an untrimmed minutes-long video, the visual encoder ff embeds its frames while the text encoder gg embeds transcribed speech and the corresponding timestamps. Then a text decoder hh predicts event boundaries and text captions using the visual and transcribed speech embeddings. The individual modules are described next.

2 Training

In this Section, we describe how we leverage a large amount of unlabeled narrated videos to train the previously described dense event captioning model. We first present the pretraining method used to effectively train Vid2Seq using cross-modal supervision in readily-available narrated videos in Section 3.2.1 and Figure 3. Then we explain how we finetune our architecture for various downstream tasks including dense event captioning in Section 3.2.2.

We wish to leverage narrated videos for pretraining as they are easily available at scale . However these videos do not contain dense event captioning annotations. Therefore we use as supervisory signal the transcribed speech sentences and their corresponding timestamps. As speech transcripts are not always visually grounded and often temporally misaligned , we note that they only provide weak supervision. Furthermore, speech transcripts drastically differ from dense event captioning annotations. For instance, in the YT-Temporal-1B dataset , a video contains 120 speech sentences on average which is an order of magnitude more than the number of events in standard dense video captioning datasets . Our Vid2Seq model is particularly suitable for using such weak supervision as it constructs the speech sequence similarly as a manually annotated event sequence, and jointly contextualizes the speech boundaries and semantic information on the level of potentially minutes-long videos (see Section 3.1) rather than at a shorter clip-level, enabling our model to learn long-term relationships between the different speech segments: in experiments we show that pretraining on entire minutes-long videos is highly beneficial.

We next describe the two proposed training objectives, which are both based on a maximum likelihood objective. Formally, given visual inputs xx, encoder text sequence yy and a decoder target text sequence zz, both objectives are based on minimizing the following loss:

where LL is the length of the decoder target sequence, wkw_{k} is the weight for k-th token in the sequence, which we set to wk=1w_{k}=1 ∀k\forall k in practice, θ\theta denotes the trainable parameters in the model and pθp_{\theta} is the output probability distribution over the vocabulary of text and time tokens.

Generative objective. This objective uses the transcribed speech as a (pseudo-)supervisory signal to teach the decoder to predict a sequence of events given visual inputs. Given video frames xx, which are fed to the encoder, the decoder has to predict the transcribed speech sequence yy (see Figure 3), which serves as a proxy dense event captioning annotation. Note that no text input is given to the encoder for this task as using transcribed speech both as input and target would lead the model to learn text-only shortcuts.

2.2 Downstream task adaptation

Our architecture and task formulation enables us to tackle dense video captioning with a generic language modeling training objective and inference procedure. Note that as a by-product of our generic architecture, our model can also be used to generate paragraphs about entire videos by simply removing the time tokens from the output sequence, and can also be easily adapted to video clip captioning with the same finetuning and inference recipe.

Finetuning. To finetune our model for dense video captioning, we use a maximum likelihood objective based on the event sequence (see Equation 1). Given video frames xx and speech transcripts yy, the decoder has to predict the event sequence zz.

Inference. The text decoder autoregressively generates the event sequence by sampling from the model likelihood. In practice, we use beam search as we find that it improves the captioning quality compared with argmax sampling or nucleus sampling. Finally, the event sequence is converted into a set of event predictions by simply reversing the sequence construction process.

Experiments

This section demonstrates the effectiveness of our pretrained Vid2Seq model and compares our method to the state of the art. We first outline our experimental setup in Section 4.1. We then present ablation studies in Section 4.2. The comparison to the state of the art in dense video captioning, video paragraph captioning and video clip captioning is presented in Section 4.3. Next, we present results in a new few-shot dense video captioning setting in Section 4.4. Finally, we show qualitative results in Section 4.5.

Datasets. For pretraining, following prior work showing the benefits of pretraining on a diverse and large dataset , we use the YT-Temporal-1B dataset , which includes 18 million narrated videos collected from YouTube. We evaluate Vid2Seq on three downstream dense video captioning datasets: YouCook2 , ViTT and ActivityNet Captions . YouCook2 has 2K untrimmed videos of cooking procedures. On average, each video lasts 320s and is annotated with 7.7 temporally-localized sentences. ViTT consists of 8K untrimmed instructional videos. On average, each video lasts 250s and is annotated with 7.1 temporally-localized short tags. ActivityNet Captions contains 20k untrimmed videos of various human activities. On average, each video lasts 120s and is annotated with 3.7 temporally-localized sentences. For video clip captioning, we use two standard benchmarks, MSR-VTT and MSVD . For all datasets, we follow the standard splits for training, validation and testing. Note that we only use videos available on YouTube at the time of the work, resulting in 10 to 20% less videos than in the original datasets.

Implementation details. We extract video frames at 1FPS, and subsample or pad the sequence of frames to FF frames where we set F=100F=100. The text encoder and decoder sequence are truncated or padded to L=S=1000L=S=1000 tokens. Our model has 314M trainable parameters. We use the Adam optimizer . We pretrain our model for 200,000 iterations with a batch size of 512 videos split on 64 TPU v4 chips, which lasts a day. We sum both pretraining objectives with equal weighting to get our final pretraining loss. More details are included in Appendix Section B.

Evaluation metrics. For captioning, we use CIDEr (C) and METEOR (M). For dense video captioning, we follow the commonly used evaluation tool which calculates matched pairs between generated events and the ground truth across IoU thresholds of {0.3, 0.5, 0.7, 0.9}, and compute captioning metrics over the matched pairs. However, these metrics do not take into account the story of the video. Therefore we also use SODA_c (S) for an overall dense video captioning evaluation. To further isolate the evaluation of event localization, we report the average precision and average recall across IoU thresholds of {0.3, 0.5, 0.7, 0.9} and their harmonic mean, the F1 Score.

2 Ablation studies

The default Vid2Seq model predicts both text and time tokens, uses both visual frames and transcribed speech as input, builds on the T5-Base language model, and is pretrained on untrimmed videos from YT-Temporal-1B with both the generative and denoising losses. Below we ablate the importance of each of these factors on the downstream dense video captioning performance by reporting results on YouCook2 and ActivityNet Captions validation sets.

Pretraining on untrimmed narrated videos by exploiting transcribed speech sentence boundaries. In Table 1, we evaluate the effectiveness of our pretraining task formulation that uses untrimmed videos and integrates sentence boundaries of transcribed speech via time tokens. In contrast, most video clip captioning pretraining methods use short, trimmed, video-speech segments for pretraining. We adapt this strategy in our model and find that it indeed yields significant performance improvements over the baseline that uses no video-text pretraining (row 2 vs row 1). However, larger improvements are obtained by using untrimmed video-speech inputs (row 3 vs row 2). Moreover, using time tokens to integrate time information from transcribed speech drastically improves performance (row 4 vs row 3). This shows the benefits of exploiting sentence boundaries of transcribed speech via time tokens and of using untrimmed videos during pretraining. In Appendix Section C.2, we show additional ablations to quantify how the performance improves by pretraining on longer narrated videos that contain more speech sentences.

Input modalities and pretraining objectives. In Table 2, we analyze the importance of input modalities and pretraining tasks on the downstream dense video captioning performance. The model with visual inputs only (no transcribed speech as input) benefits significantly from pretraining with the generative objective (row 3 vs row 1). This shows the effectiveness of using the transcribed speech as a proxy annotation for dense video captioning pretraining. However, this model is pretrained with visual inputs only and its performance largely drops when it is finetuned with both visual and transcribed speech inputs (row 4 vs row 3). With both modalities, adding the denoising loss strongly benefits our model (row 5 vs rows 4 and 2). We conclude that the denoising objective benefits multi-modal reasoning.

Effect of captioning on localization. In Table 3, we compare the event localization performance of our model with a localization-only variant that only predicts event boundaries. We find that the model that jointly predicts event boundaries and captions localizes better and benefits more from pretraining than the localization-only baseline (row 4 vs row 3), which demonstrates the importance of contextualizing the noisy timestamps of the transcribed speech with the speech semantic content during pretraining.

Model size and pretraining data. In Table 4, we show that the language model size has a great importance on the performance, as the model with T5-Base outperforms its variant with T5-Small (row 7 vs row 1). We also evaluate the importance of the size of the pretraining dataset of narrated videos by constructing subsets such that larger subsets include the smaller ones. We find that scaling up the size of the pretraining dataset is beneficial, and that our pretraining method yields important benefits when only using 150K narrated videos for pretraining (row 4). We further show that our pretraining method generalizes well to the HowTo100M dataset . The model pretrained on HowTo100M (row 6) actually achieves best results on YouCook2, as these datasets are from a similar domain. Finally, we ablate the importance of the size and pretraining of the visual backbone in Appendix Section C.2.

3 Comparison to the state of the art

Dense video captioning. In Table 5, we compare our approach to state-of-the-art dense video captioning methods using cross-entropy training We do not include methods directly optimizing the test metric . on the YouCook2, ViTT and ActivityNet Captions datasets. Vid2Seq sets new state of the art on all three datasets. In particular, our method improves the CIDEr metric by 18.2 and 0.8 points on YouCook2 and ActivityNet Captions over PDVC. Our method also outperforms E2ESG which uses in-domain text-only pretraining on Wikihow. These results demonstrate the strong dense event captioning ability of our pretrained Vid2Seq model.

Event localization. In Table 6, we evaluate the event localization performance of our dense video captioning model in comparison with prior work. On both YouCook2 and ViTT, Vid2Seq outperforms prior work tackling dense video captioning as a single sequence generation task. However, our model underperforms compared to PDVC and UEDVC on ActivityNet Captions. We emphasize that our approach integrates less prior knowledge about temporal localization than both these approaches, which include task specific components such as event counters or separately train a model for the localization subtask .

Video paragraph captioning. In Table 7, we compare our approach to state-of-the-art video paragraph captioning methods on the YouCook2 and ActivityNet Captions datasets. Vid2Seq outperforms all prior methods on both datasets, including the ones using ground-truth event boundary proposals at inference time , showing strong video paragraph captioning ability.

Video clip captioning. In Table 8, we compare our approach to state-of-the-art video clip captioning methods on the MSR-VTT and MSVD datasets. Vid2Seq improves over prior methods in their respective pretraining data setting while using a comparable number of trained parameters. We conclude that our pretrained Vid2Seq model generalizes well to the standard video clip captioning setting.

4 Few-shot dense video captioning

To further evaluate the generalization capabilities of our pretrained Vid2Seq model, we propose a new few-shot dense video captioning setting where we finetune Vid2Seq using only a fraction of the downstream training dataset. From Table 9, we observe important improvements when using 10% compared to 1% of training data (row 2 vs 1). In Appendix Section C.1 we further show that pretraining is essential in this few-shot setting.

5 Qualitative examples

In Figure 4, we show an example of dense event captioning predictions from Vid2Seq. This example shows that our model can predict meaningful event boundaries and captions, and that the predicted captions and boundaries differ considerably from the transcribed speech input (showing the importance of the visual tokens in the input). More examples are provided in Appendix Section A.

Conclusion

We introduced Vid2Seq, a visual language model that performs dense video captioning by generating a single sequence of tokens including both text and time tokens given multi-modal inputs. We showed that Vid2Seq benefits from large-scale pretraining on unlabeled untrimmed narrated videos by leveraging transcribed speech sentences and corresponding temporal boundaries. Vid2Seq achieves state-of-the-art results on various dense event captioning datasets, as well as multiple video paragraph captioning and standard video clip captioning benchmarks. Finally, we believe the sequence-to-sequence design of Vid2Seq has the potential to be extended to a wide range of other video tasks such as temporally-grounded video question answering or temporal action localization .

Acknowledgements. The work was partially funded by a Google gift, the French government under management of Agence Nationale de la Recherche as part of the ”Investissements d’avenir” program, reference ANR-19-P3IA-0001 (PRAIRIE 3IA Institute), the Louis Vuitton ENS Chair on Artificial Intelligence, the European Regional Development Fund under project IMPACT (reg. no. CZ.02.1.01/0.0/0.0/15 003/0000468). We thank Anurag Arnab, Minsu Cho, Anja Hauth, Ashish Thapliyal, Bo Pang, Bryan Seybold and the entire Ganesha team for helpful discussions.

References

Appendix

In this Appendix, we present the following additional material:

Additional qualitative examples of dense video captioning predictions (Section A).

Additional information about our experimental setup (Section B);

Additional experimental results (Section C), including an ablation on the importance of pretraining for few-shot dense video captioning (Section C.1) and additional ablation studies in the standard fully-supervised dense video captioning setting (Section C.2).

Appendix A Qualitative examples of dense video captioning predictions

In Figure 4, we show qualitative results of dense event captioning by our Vid2Seq model. Here in Figures 5 and 6 , we show additional results on examples from the YouCook2 and ActivityNet Captions datasets. These results show that Vid2Seq can predict meaningful dense captions and event boundaries in diverse scenarios, with or without transcribed speech input, e.g. series of instructions in cooking recipes (Figure 5) or actions in human sports or leisure activities (first three examples in Figure 6). The last example in Figure 6 illustrates a failure case where the model hallucinates events that are not visually grounded such as ‘one man hats off to the camera‘.

Appendix B Experimental setup

In this section, we complement the information provided in Section 4.1 about the datasets we use (Section B.1). We also give additional implementation details (Section B.2).

YT-Temporal-1B consists of 18.821M unlabeled narrated videos covering about 150 years of video content for pretraining. Compared with HowTo100M , this dataset was created to cover a wider range of domains and not only instructional videos.

HowTo100M consists of 1.221M unlabeled narrated instructional videos covering about 15 years of video content for pretraining.

YouCook2 has 1,790 untrimmed videos of cooking procedures. On average, each video lasts 320s and is annotated with 7.7 temporally-localized imperative sentences. The dataset is split into 1,333 videos for training and 457 videos for validation.

ViTT consists of 7,672 untrimmed instructional videos from the YouTube-8M dataset . Compared to YouCook2, ViTT was created to better reflect the distribution of instructional videos in the wild. On average, each video lasts 250s and is annotated with 7.1 temporally-localized short tags. The dataset is split into 5,476, 1,102 and 1,094 videos for training, validation and testing, respectively. Videos in the validation and test sets are provided with multiple sets of dense event captioning annotations. Following , we treat each set of annotations as a single example during evaluation and discard videos with more than 3 sets of annotations.

ActivityNet-Captions contains 14,934 untrimmed videos of various human activities. Different from YouCook2 and ViTT where most videos contain transcribed speech content, we find that 68% of videos in ActivityNet Captions do not have transcribed narration. On average, each video lasts 120s and is annotated with 3.7 temporally-localized sentences. The dataset is split into 10,009 and 4,925 videos for training and validation, respectively. Videos in the validation set are provided with two sets of dense video captioning annotations. Following prior work , we use both sets of annotations for evaluation, by computing the average of the scores over each set for SODA_c and by using the standard evaluation tool for all other dense event captioning metrics. For video paragraph captioning, we follow and report results on the ’val-ae’ split that includes 2,460 videos .

MSR-VTT consists of 10,000 open domain video clips. The duration of each video clip is between 10 and 30 seconds. 20 natural language descriptions are manually annotated for each clip. The dataset is split into 6,513, 497 and 2,990 videos for training, validation and testing, respectively.

MSVD consists of 1,970 open domain video clips. The duration of each video clip is between 10 and 30 seconds. Each video clip has roughly 40 manually annotated captions. The dataset is split into 1,200, 100 and 670 videos for training, validation and testing, respectively.

B.2 Implementation details

The visual temporal transformer encoder ftf^{t}, the text encoder gtg^{t} and the text decoder hth^{t} all have 12 layers, 12 heads, embedding dimension 768, and MLP hidden dimension of 2048. The text encoder and decoder sequences are truncated or padded to L=S=1000L=S=1000 tokens during pretraining, and S=1000S=1000 and L=256L=256 tokens during finetuning. At inference, we use beam search decoding where we track the top 4 sequences and apply a length normalization of 0.6.

We use the Adam optimizer with β=(0.9,0.999)\beta=(0.9,0.999) and no weight decay. During pretraining, we use a learning rate of 1e−41e^{-4}, warming it up linearly (from 0) for the first 1000 iterations, and keeping it constant for the remaining iterations. During finetuning, we use a learning rate of 3e−43e{-4}, warming it up linearly (from 0) for the first 10% of iterations, followed by a cosine decay (down to 0) for the remaining 90%. During finetuning, we use a batch size of 32 videos split on 16 TPU v4 chips. We finetune for 40 epochs on YouCook2, 20 epochs on ActivityNet Captions and ViTT, 5 epochs on MSR-VTT and 10 epochs on MSVD. We clip the maximum norm of the gradient to 0.1 during pretraining, and 1 during finetuning. For data augmentation, we use random temporal cropping. For regularization, we use label smoothing with value 0.1 and dropout with probability 0.1.

Appendix C Experiments

In this section, we provide additional experiments that complement the results presented in Section 4. We first show the importance of pretraining in our proposed few-shot setting in Section C.1. Then we provide additional ablation studies in the standard fully-supervised setting in Section C.2, where we ablate various factors including pretraining on long narrated videos, the pretraining dataset and the size of the visual backbone, the time tokenization process and the number of time tokens, the sequence construction process, the temporal positional embeddings and the initialization of the language model.

In Section 4.2, we show the benefits of our pretraining method in the fully-supervised setting, i.e. when using 100% of the downstream training dataset. In Table 10, we further show that our pretraining method has a considerable importance in the few-shot setting defined in Section 4.4, i.e. when using a smaller fraction of the downstream training dataset. In particular, our pretraining method enables our Vid2Seq model to have a non zero performance when using only 1% of the downstream training dataset (rows 1 and 2).

C.2 Additional ablation studies

We here complement ablation studies reported in Section 4.2, using the same default settings, evaluation metrics and downstream datasets.

In Table 1, we show the benefits of pretraining on untrimmed videos in comparison with the standard practice of pretraining on short, trimmed, video-speech segments . In Table 11, we further evaluate the importance of sampling long narrated videos during pretraining. By default, at each training iteration, we randomly temporally crop each narrated video without constraints, resulting in a video that can span over hundreds of transcribed speech sentences. We here evaluate a baseline that constrains this cropping process such that the cropped video only spans over a given maximum number of narration sentences. Even with a maximum of 10 narration sentences, this baseline significantly underperforms our model trained in default settings where we sample longer untrimmed narrated videos (rows 1, 2 and 3). This demonstrates that our model benefits from pretraining on long narrated videos.

In Table 4, we show the benefits of scaling up the size of the pretraining dataset of narrated videos and the size of the language model. In Table 12, we further analyze the importance of the pretraining dataset and size of the visual backbone fsf^{s}. We find that CLIP pretraining considerably improves over ImageNet pretraining with the same ViT-B/16 visual backbone model (row 2 vs 1). Furthermore, scaling up the visual backbone size from ViT-B/16 to ViT-L/14 brings additional improvements (row 3 vs 2).

In Table 13, we further ablate the time tokenization process presented in Section 3.1. Our default time tokens represent relative timestamps in a video, as we quantize a video of duration TT into NN equally-spaced timestamps. Another possibility is to use time tokens that represent absolute timestamps in the video, i.e. the k-th token represents the k-th second in the video. For both these variants, we vary the number of time tokens NN. For the relative time tokens, increasing NN makes the quantization more fine-grained but also spreads the data into more time tokens. On the other hand, for the absolute time tokens, increasing NN increases the video duration that the time tokens can cover. We find that the best dense video captioning results are obtained with the relative time tokens and N=100N=100 time tokens (row 5).

In Table 14, we further ablate the sequence construction process presented in Section 3.1. Our default sequence inserts the start and end time tokens of each segment before its corresponding text sentence. Another possibility is to insert time tokens after each corresponding text sentence. We find that both variants achieve similar results (rows 2 and 4), with the default sequence (row 4) resulting in slightly higher event localization performance (F1 Score) but slightly lower dense captioning results overall. Furthermore, we observe that the dot symbols indicating the separation between different events have low importance (rows 1 and 2, rows 3 and 4).

In Table 1, we show that time tokens in the speech sequence provide temporal information about the speech transcript to our model. In Table 15, we also evaluate the importance of the temporal positional embeddings which communicate temporal information from the visual stream to our model. We find that these temporal embeddings are beneficial (row 2 vs 1).

In Table 4, we show the benefits of using T5-Base instead of T5-Small. In Table 16, we further investigate the importance of initializing the language model from weights pretrained on Web text. Without pretraining on narrated videos, we find that text-only initialization is helpful (rows 1 and 2). Interestingly, after pretraining on narrated videos, we find that text-only initialization has little importance (rows 3 and 4), as it slightly improves the performance on ActivityNet Captions while resulting in a slight drop of performance on YouCook2. We believe that this may be because of the domain gap between Web text and the imperative-style dense captions in YouCook2, which are more similar to transcribed speech in YT-Temporal-1B.