Multimodal Pretraining for Dense Video Captioning
Gabriel Huang, Bo Pang, Zhenhai Zhu, Clara Rivera, Radu Soricut
Introduction
YouTube recently reported that a billion hours of videos were being watched on the platform every day (YouTubeBlog, 2017). In addition, the amount of time people spent watching online videos was estimated to grow at an average rate of 32% a year between 2013 and 2018, with an average person forecasted to watch 100 minutes of online videos per day in 2021 (ZenithMedia, 2019).
An important reason for this fast-growing video consumption is information-seeking. For instance, people turn to YouTube “hungry for how-to and learning content” O’Neil-Hart (2018). Indeed, compared to traditional content format such as text, video carries richer information to satisfy such needs. But as a content media, videos are also inherently more difficult to skim through, making it harder to quickly target the relevant part(s) of a video. Recognizing this difficulty, search engines started showing links to “key moments” within videos in search results, based on timestamps and short descriptions provided by the content creators themselves.https://www.blog.google/products/ search/key-moments-video-search/ This enables users to get a quick sense of what the video covers, and also to jump to a particular time in the video if so desired. This effort echoes prior work in the literature showing how users of instructional videos can benefit from human-curated meta-data, such as a timeline pointing to the successive steps of a tutorial Kim et al. (2014); Margulieux et al. (2012); Weir et al. (2015). Producing such meta-data in an automatic way would greatly scale up the efforts of providing easier information access to videos. This task is closely related to the dense video captioning task considered in prior work Zhou et al. (2018a, c); Krishna et al. (2017), where an instructional video is first segmented into its main steps, followed by segment-level caption generation.
To date, the YouCook2 data set Zhou et al. (2018a) is the largest annotated data set for dense video captioning. It contains annotations for 2,000 cooking videos covering 89 recipes, with per-recipe training / validation split. Restricting to a small number of recipes is helpful for early exploratory work, but such restrictions impose barriers to model generalization and adoption that are hard to overcome. We directly address this problem by constructing a larger and broader-coverage annotated dataset that covers a wide range of instructional topics (cooking, repairs, maintenance, etc.) We make the results of our annotation efforts publicly available as Video Timeline Tags (ViTT)Available at https://github.com/google-research-datasets/Video-Timeline-Tags-ViTT, consisting of around 8,000 videos annotated with timelines (on average 7.1 segments per video, each segment with a short free-text description).
Using YouCook2 and the new ViTT dataset as benchmarks for testing model performance and generalization, we further focus on the sub-problem of video-segment–level caption generation, assuming segment boundaries are given Hessel et al. (2019); Sun et al. (2019b); Luo et al. (2020). Motivated by the high cost of collecting human annotations, we investigate pretraining a video segment captioning model using unsupervised signals – ASR (Automatic Speech Recognition) tokens and visual features from instructional videos, and unpaired instruction steps extracted from independent sources: Recipe1M Marin et al. (2019) and WikiHow Koupaee and Wang (2018). In contrast to prior work that focused on BERT-style pretraining of encoder networks (Sun et al., 2019b, a), our approach entails jointly pretraining both multimodal encoder and text-based decoder models via MASS-style pretraining Song et al. (2019). Our experiments show that pretraining with either text-only or multi-modal data provides significant gains over no pretraining, on both the established YouCook2 benchmark and the new ViTT benchmark. The results we obtain establish state-of-the-art performance on YouCook2, and present strong performance numbers on the ViTT benchmark. These findings help us conclude that the resulting models generalize well and are quite robust over a wide variety of instructional videos.
Related Work
Language pretraining models based on the Transformer neural network architecture Vaswani et al. (2017a) such as BERT (Devlin et al., 2018), GPT Radford et al. (2018), RoBERTa Liu et al. (2019), MASS Song et al. (2019) and ALBERT Lan et al. (2020) have achieved state-of-the-art results on many NLP tasks. MASS (Song et al., 2019) has been recently proposed as a joint encoder-decoder pretraining strategy. For sequence-to-sequence tasks, this strategy is shown to outperform approaches that separately pretrain the encoder (using a BERT-style objective) and the decoder (using a language modeling objective). UniLM Dong et al. (2019), BART Lewis et al. (2019), and T5 Raffel et al. (2019) propose unified pretraining approaches for both understanding and generation tasks.
Multimodal Pretraining.
VideoBERT Sun et al. (2019b), CBT Sun et al. (2019a) and ActBERT Zhu and Yang (2020) use a BERT-style objective to train both video and ASR text encoders. Alayrac et al. (2016) and Miech et al. (2020) use margin-based loss functions to learn joint representations for video and ASR, and evaluate them on downstream tasks such as video captioning, action segmentation and anticipation, and action localization. An independent and concurrent work (UniViLM) by Luo et al. (2020) is closely related to ours in that we share some similar pretraining objectives, some of the pretraining setup – HowTo100M Alayrac et al. (2016), and the down-stream video captioning benchmark using YouCook2 Zhou et al. (2018a). The main difference is that they use BERT-style pretraining for encoder and language-modeling style pretraining for decoder, whereas we use MASS-style pre-training to pretrain encoder and decoder jointly.
Other approaches such as ViLBERT (Lu et al., 2019), LXMERT (Tan and Bansal, 2019), Unicoder-VL (Li et al., 2019), VL-BERT (Su et al., 2019), and UNITER (Chen et al., 2019) focus on pretraining joint representations for text and image, evaluating them on downstream tasks such as visual question answering, image-text retrieval and referring expressions.
Dense Video Captioning.
In this paper, we focus on generating captions at the segment-level, which is a sub-task of the so-called dense video captioning task Krishna et al. (2017), where fine-grained captions are generated for video segments, conditioned on an input video with pre-defined event segments. This is different from the video captioning models that generate a single summary for the entire video Wang et al. (2019).
Hessel et al. (2019) make use of ASR and video for segment-level captioning on YouCook2 and show that most of the performance comes from ASR. Shi et al. (2019); Luo et al. (2020) train their dense video captioning models on both video frames and ASR text and demonstrate the benefits of adding ASR as an input to the model. There are also a number of video captioning approaches that do not use ASR directly Zhou et al. (2018c); Pan et al. (2020); Zheng et al. (2020); Zhang et al. (2020); Lei et al. (2020).
Instructional video captioning data sets.
In addition to YouCook2 Zhou et al. (2018a), there are two other smaller data sets in the instructional video captioning category. Epic Kitchen Damen et al. (2018) features 55 hours of video consisting of 11.5M frames, which were densely labeled for a total of 39.6K action segments and 454.3K object bounding boxes. How2 Sanabria et al. (2018) consists of instructional videos with video-level (as opposed to segment-level) descriptions, authored by the video creators themselves.
Data
We present the datasets used for pretraining, finetuning, and evaluation in Table 1. We also describe in detail the newly introduced dense video captioning dataset, Video Timeline Tags (ViTT).
Our goal is to generate captions (CAP) for video segments. We consider two datasets with segment-level captions for fine-tuning and evaluating ASR+VideoCAP models.
Up to this point, YouCook2 (Zhou et al., 2018a) has been the largest human-annotated dense-captioning dataset of instructional videos publicly available. It originally contained 2,000 cooking videos from YouTube. Starting from 110 recipe types (e.g., “shrimp tempura”), 25 unique videos per recipe type were collected; the recipe types that did not gather enough videos were dropped, resulting in a total of 89 recipe types in the final dataset. In addition, Zhou et al. (2018b) “randomly split the videos belonging to each recipe into 67%:23%:10% as training, validation and test setsNote that no annotations are provided for the test split; we conducted our own training/dev/test split over available videos.,” which effectively guarantees that videos in the validation and test sets are never about unseen recipes. Annotators were then asked to construct recipe steps for each video — that is, identify the start and end times for each step, and provide a recipe-like description of each step. Overall, they reported an average of 7.7 segments per video, and 8.8 words per description. After removing videos that had been deleted by users, we obtained a total of 11,549 segments.
ViTT.
One limitation of the YouCook2 dataset is the artificially imposed (almost) uniform distribution of videos over 89 recipes. While this may help making the task more tractable, it is difficult to judge whether performance on its validation / test sets can be generalized to unseen topics.
The design of our ViTT dataset annotation process is aimed at fixing some of these drawbacks. We started by collecting a large dataset of videos containing a broader variety of topics to better reflect topic distribution in the wild. Specifically, we randomly sampled instructional videos from the YouTube-8M dataset (Abu-El-Haija et al., 2016), a large-scale collection of YouTube videos that also contain topical labels. Since much of prior work in this area revolved around cooking videos, we aimed at sampling a significant proportion of our data from videos with cooking labels (specifically, “Cooking” and “Recipe”). Aside from the intentional bias regarding cooking videos, the rest of the videos were selected by randomly sampling non-cooking videos, including only those that were considered to be instructional videos by our human annotators.
Once candidate videos were identified, timeline annotations and descriptive tags were collected. Our motivation was to enable downstream applications to allow navigating to specific content sections. Therefore, annotators were asked to identify the main steps in a video and mark their start time. They were also asked to produce a descriptive-yet-concise, free-text tag for each step (e.g., “shaping the cookies”, “removing any leftover glass”). A subset of the videos has received more than one complete annotation (main steps plus tags).
The resulting ViTT dataset consists of a total of 8,169 videos, of which 3,381 are cooking-related. A total of 5,840 videos have received only one annotation, and have been designated as the training split. Videos with more than one annotation have been designated as validation / test data. Overall, there are 7.1 segments per video on average (max: 19). Given the dataset design, descriptions are much shorter in length compared to YouCook2: on average there are 2.97 words per tag (max: 16) — 20% of the captions are single-word, 22% are two-words, and 25% are three words. Note that the average caption length is significantly shorter than for YouCook2, which is not surprising given our motivation of providing short and concise timeline tags for video navigation. We standardized the paraphrases among the top-20 most frequent captions. For instance, {“intro”, “introduction”} “intro”. Otherwise, we have preserved the original tags as-is, even though additional paraphrasing most definitely exists. Annotators were instructed to start and end the video with an opening and closing segment as possible. As a result, the most frequent tag (post-standardization) in the dataset is “intro”, which accounts for roughly 11% of the 88,455 segments. More details on the data collection process and additional analysis can be found in the Supplementary Material (Section A.1).
Overall, this results in 56,027 unique tags, with a vocabulary size of 12,509 token types over 88,455 segments. In this paper, we consider two variants: the full dataset (ViTT-All), and the cooking subset (ViTT-Cooking).
2 Pretraining Datasets: ASR+Video
We consider two large-scale unannotated video datasets for pretraining, as described below. Time-stamped ASR tokens were obtained via YouTube Data API,https://developers.google.com/youtube/v3/docs/captions and split into ASR segments if the timestamps of two consecutive words are more than 2 seconds apart, or if a segment is longer than a pre-specified max length (in our case, 320 words). They were paired with concurrent video frames in the same segment.
We extract the cooking subset of YouTube-8M (Abu-El-Haija et al., 2016) by taking, from its training split, videos with “Cooking” or “Recipe” labels, and retain those with English ASR, subject to YouTube policies. After preprocessing, we obtain 186K ASR+video segments with an average length of 64 words (24 seconds) per segment.
HowTo100M.
This is based on the 1.2M YouTube instructional videos released by Miech et al. (2019), covering a broad range of topics. After preprocessing, we obtain 7.99M ASR+video segments with an average of 78 words (28.7 seconds) per segment.
3 Pretraining Datasets: CAP-style
We also consider two text-only datasets for pretraining, containing unpaired instruction steps similar in style to the target captions.
is a collection of 1M recipes scraped from a number of popular cooking websites (Marin et al., 2019). We use the sequence of instructions extracted for each recipe in this dataset, and treat each recipe step as a separate example during pretraining. This results in 10,767,594 CAP-style segments, with 12.8 words per segment.
WikiHow
is a collection of 230,843 articles extracted from the WikiHow knowledge base (Koupaee and Wang, 2018). Each article comes with a title starting with “How to”. Each associated step starts with a step summary (in bold) followed by a detailed explanation. We extract the all step summaries, resulting in 1,360,145 CAP-style segments, with 8.2 words per segment. Again, each instruction step is considered as a separate example during pretraining.
4 Differences between Pretraining and Finetuning Datasets
First, note that video segments are defined differently for pretraining and finetuning datasets, and may not match exactly. For ASR+Video pretraining datasets, which are unsupervised, the segments are divided following a simple heuristic (e.g., two consecutive words more than 2 seconds apart), whereas for finetuning ASR+VideoCAP datasets, which are supervised, the segments are defined by human annotators to correspond to instruction steps. Otherwise, the ASR data are relatively similar between pretraining and finetuning datasets, as both come from instructional videos and are in the style of spoken language.
Second, compared to the target captions in finetuning datasets, the CAP-like pretraining datasets are similar in spirit — they all represent summaries of steps, but they may differ in length, style and granularity. In particular, the CAP-like pretraining datasets are closer in style to captions in YouCook2, where annotators were instructed to produce a recipe-like description for each step. This is reflected in their similar average length (YouCook2: 8.8 words, Recipe1M: 12.8 words, WikiHow: 8.2 words); whereas captions in ViTT are significantly shorter (2.97 words on average).
Despite these differences — some are inevitable due to the unsupervised nature of pretraining datasets — the pretraining data is very helpful for our task as shown in the experimental results.
Method
To model segment-level caption generation, we adopt MASS-style pretraining Song et al. (2019) with Transformer Vaswani et al. (2017b) as the backbone architecture. For both pre-training and fine-tuning objectives, we have considered two variants: text-only and multi-modal. They are summarized in Table 2 and more details are given below.
Both ASR tokens and video segment features are given as input in the multimodal variants. We consider an architecture with a separate transformer for each modality (text or video), see Figure 2 for details. When available, the text and video encoders attend to each other at every layer using cross-modal attention, as in ViLBERT Lu et al. (2019). The text decoder attends over the final-layer output of both encoders. We discuss in more detail the differences between using a separate-modality architecture vs. a vanilla-Transformer approach for all modalities in Appendix A.2.
The inputs to the text encoder is the sum of three components: text token embeddings, positional embeddings and the corresponding style embeddings,This is similar to the way language-ID embeddings are used in machine translation. depending on the style of the text (ASR or Caption-like). The inputs to the video encoder could be either precomputed frame-level 2D CNN features or 3D CNN features, pretrained on the Kinetics Carreira and Zisserman (2017); Kay et al. (2017) data set. The visual features are projected with fully-connected layers to the same dimension as the text embeddings.
The main architecture we consider is a 2-layer encoder (E2), 6-layer decoder (D6) Transformer. We use E2D6 to refer to the text-only version, and E2vidD6 to refer to the multimodal version with an active video encoder. We also experiment with E2D2 and E2vidD2 (2-layer decoder).We found in a preliminary study that using 6-layer encoders did not improve performance for our application.
2 Pretraining with Text-only MASS
Text-only pretraining is essentially the unsupervised learning of the style transfer between ASR-style and caption-style texts using unpaired data sources: ASR strings from video segments in YT8M-cook or HowTo100M; and CAP-style instruction steps found in Recipe1M or HowTo100M. Just like using MASS for unsupervised machine translation involves pretraining the model on unpaired monolingual datasets, we alternate between asrasr and capcap MASS steps during our pretraining stage, which does not require the “source” (ASR) and “target” (CAP-style) data to be aligned.
In an asrasr step, we mask a random subsequence of the ASR and feed the masked ASR to the text encoder. The text decoder must reconstruct the hidden subsequence while attending to the encoder output. A capcap step works similarly by trying to reconstruct a masked sequence of a CAP-style text. The encoder and decoder are trained jointly using teacher-forcing on the decoder. We denote this text-only strategy as MASS in the experiments.
3 Pretraining with Multimodal MASS
During multimodal pretraining, we alternate between text-only capcap MASS steps and multimodal MASS steps. During each multimodal MASS step asr+videoasr, we feed a masked ASR to the text-encoder and the co-occurring video features to the video-encoder. The text decoder must reconstruct the masked ASR subsequence. We denote this pretraining strategy as MASSvid in the experiments. This trains cross-modal attention between the text-encoder and video-encoder at every layer, jointly with the text decoder that attends to the output layer of both the text and video encoders. In preliminary experiments, we had attempted to directly adapt the MASS objective (Song et al., 2019) to video reconstruction — by masking a subsequence of the input video and making the video decoder reconstruct the input using the Noise Constrastive Estimator Loss (Sun et al., 2019a). Due to limited success, we did not further pursue this approach.
To force more cross-modal attention between encoder and decoder, we also investigate a strategy of hiding the text-encoder output from the decoder for some fraction of training examples. We refer to this strategy as MASSdrop in the experiments.
4 Pretraining with Alignment and Ordering Tasks
We also explore encoder-only multimodal pretraining strategies. We take the last-layer representation for the CLS (beginning of sentence) token from the encoder, and add a multi-layer perceptron on top of it for binary predictions (Figure 2). Given a pair of ASR and video segment, we train the encoder to predict the following objectives:
Segment-Level Alignment. An (ASR, video) pair is aligned if they occur in the same pretraining segment; negative examples are constructed by sampling pairs from the same video but at least 2 segments away.
Segment-Level Ordering. We sample (ASR, video) pairs that are at least 2 segments away, and train the model to predict whether the ASR occurs before or after the video clip.
During this MASSalign pretraining stage, we alternate between two text-only MASS steps (capcap, asrasr) and the two binary predictions (Alignment and Ordering) described above.
5 Finetuning on Video Captioning
For text-only finetuning, we feed ASR to the text encoder and the decoder has to predict the corresponding CAP (asrcap). For multimodal finetuning, we also feed additional video representations to the video encoder (asr+videocap). When finetuning a multimodal model from text-only pretraining, everything related to video (weights in the video encoder and any cross-modal attention modules) will be initialized randomly. In addition to these uni-directional (UniD) finetuning, we also experiment with several variants of bidirectional (BiD) finetuning (Table 2). For instance, adding capasr (predicting ASR from CAP) to text-only finetuning. In the experiments, we find some variants of bidirectional finetuning beneficial whether training from scratch or finetuning from a pretrained model.
Experiments
We tokenize ASR and CAP inputs using byte-pair–encoding subwords (Sennrich et al., 2015), and truncate them to 240 subwords. We truncate video sequences to 40 frames (40 seconds of video), compute the 128-dim features proposed by Wang et al. (2014) (which we will refer to as Compact 2D features), and project them to the embedding space using a two-layer perceptron with layer normalization and GeLU activations.
We instantiate the E2xDx models defined in Section 4.1 with 128-dimensional embeddings and 8 heads respectively for self-attention, encoder-decoder, and cross-modal attention modules. We define each epoch to be 3,125 iterations, where each iteration contains one repetition of each training step as represented in Table 2. We pretrain for 200 epochs and finetune for 30 epochs.
For evaluation, we consider Bleu-4 (Papineni et al., 2002), Meteor (Denkowski and Lavie, 2014), Rouge-L (Lin and Och, 2004) and CIDEr (Vedantam et al., 2015) metrics.
Please refer to Appendix A.3 for full implementation details, hyperparameters and computation cost.
With the exception of Rouge-L, all other metrics are sensitive to short groundtruth. 67% of the groundtruth tags in ViTT have less than 4 words, where a perfect prediction will not yield a full score in, say, Bleu-4. Thus, we focus mainly on Rouge-L, report Bleu-1 instead of Bleu-4 for ViTT, and provide the other two metrics only as reference points.
We had originally decided to use videos with multiple annotations as validation and test data, so that we could explore evaluation with multiple reference groundtruth captions. But as annotators do not always yield the same set of segment boundaries, this became tricky. Instead, we simply treat each segment as a separate instance with one single reference caption. Note that all segments annotated for the same video will be in either validation or test to ensure no content overlap.
2 Main Results
We run several variants of our method on YouCook2, ViTT-All and ViTT-Cooking, using different architectures, modalities, pretraining datasets, pretraining and finetuning strategies.
For YouCook2, we report our method alongside several methods from the literature (Hessel et al., 2019; Sun et al., 2019b; Zhou et al., 2018c; Lei et al., 2020), as well as state-of-the-art concurrent work (Luo et al., 2020). The related work is provided for reference and to give a ballpark estimate of the relative performance of each method, but results are not always strictly and directly comparable. Beyond the usual sources of discrepancy in data processing, tokenization, or even different splits, an additional source of complication comes from the fact that videos are regularly deleted by content creators, causing video datasets to shrink over time. Additionally, when comparing to other work incorporating pretraining, we could differ in (videos available in) pretraining datasets, segmentation strategies, etc. To this end, we perform an extensive ablation study, which at least helps us to understand the effectiveness of different components in our approach.
Effect of pretraining
The main experimental results for the three datasets we consider are summarized in Table 3 (YouCook2) and Table 4 (ViTT-All and ViTT-Cooking). Across all three datasets, the best performance is achieved by finetuning a multimodal captioning model under the Multimodal Pretraining condition. For instance, on YouCook2, E2vidD6-MASSvid-BiD improves over the no-pretraining model E2vidD6-BiD by 4.37 Rouge-L, a larger improvement than UniViLM with pretraining (#5) vs without (#2) Luo et al. (2020). This improvement also holds in ViTT-Cooking (+4.22 in Rouge-L) and ViTT-All (+2.97 in Rouge-L). We do not observe consistent and significant trends among the different multimodal pretraining strategies: MASS pretraining with video (MASSvid), with video and droptext (MASSdrop), or with alignment tasks (MASSalign).Limited improvement with MASSalign suggests that such alignment tasks are better suited for retrieval (Luo et al., 2020). Furthermore, we observe that most of the pretraining improvement is achievable via text-only MASS pretraining. Across all three datasets, while Multimodal Pretraining (E2vidD6-MASSvid-BiD) is consistently better than Text Pretraining (E2vidD6-MASS-BiD), the differences are quite small (under one Rouge-L point).
It’s worthy noting that for MASSalign, the best validation accuracies for the pretraining tasks are reasonably high: for YT8M-cook, we observed 90% accuracy for the alignment task, and 80% for the ordering task (for HowTo100M: 87% and 71.4%, respectively), where random guess would yield 50%. This suggests that our video features are reasonably strong, and the findings above are not due to weak visual representations.
Effect of other modeling choices
We experiment with 2-layer decoder (D2) vs 6-layer decoder (D6), combined with either unidirectional fine-tuning (UniD) or bidirectional fine-tuning (BiD). Table 5 shows ablation results of the four possible combinations when finetuning a multimodal model using text-only pretraining on YouCook2 (a more complete list of results can be found in Appendix A.5, showing similar trends). The D6xBiD combination tends to yield the best performance, with the differences among the four configurations being relatively small (under one Rouge-L point). For visual features, we also explored using 3D features (Xie et al., 2018) instead of 2D features during finetuning (with no pretraining or text-only pretraining), and do not find much difference in model performance on YouCook2. As a result, we use the simpler 2D features in our multimodal pretraining. We leave more extensive experiments with visual features as future work.
Generalization implications
An important motivation for constructing the ViTT dataset and evaluating our models on it has been related to generalization. Since the YouCook2 benchmark is restricted to a small number of cooking recipes, there is little to be understood about how well models trained and evaluated on it generalize. In contrast, the ViTT benchmark has a much wider coverage (for both cooking-related videos and general instructional videos), and no imposed topic overlap between train/dev/test. As such, there are two findings here that are relevant with respect to generalization: (a) the absolute performance of the models on the ViTT benchmark is quite high (ROUGE-L scores above 0.30 are usually indicative of decent performance), and (b) the performance on ViTT vs. YouCook2 is clearly lower (31.5 ROUGE-L vs. 39.0 ROUGE-L, reflecting the increased difficulty of the new benchmark), but it is maximized under similar pretraining and finetuning conditions, which allows us to claim that the resulting models generalize well and are quite robust over a wide variety of instructional videos.
Conclusions
Motivated to improve information-seeking capabilities for videos, we have collected and annotated a new dense video captioning dataset, ViTT, which is larger with higher diversity compared to YouCook2. We investigated several multimodal pretraining strategies for segment-level video captioning, and conducted extensive ablation studies. We concluded that MASS-style pretraining is the most decisive factor in improving the performance on all the benchmarks used. Even more to the point, our results indicate that most of the performance can be attributed to leveraging the ASR signal. We achieve new state-of-the-art results on the YouCook2 benchmark, and establish strong performance baselines for the new ViTT benchmark, which can be used as starting points for driving more progress in this direction.
Acknowledgements
We send warm thanks to Ashish Thapliyal for helping the first author debug his code and navigate the computing infrastructure, and to Sebastian Goodman for his technical help (and lightning fast responses!). We also thank the anonymous reviewers for their comments and suggestions.
References
Appendix A Appendix
Supplementary Material for “Multimodal Pretraining for Dense Video Captioning”.
The goal of the ViTT dataset design is to mirror topic distribution in the “wild”. Therefore, instead of starting from specific how-to instructions and searching for corresponding videos, we sampled videos from the validation set of the YouTube-8M dataset (Abu-El-Haija et al., 2016), a large-scale collection of YouTube videos with topical labels, subject to YouTube policies.
Exclusion criteria were lack of English ASR and the topic label “Game”. The latter was motivated by the fact that in this type of videos, the visual information predominantly features video games, while the ViTT dataset was intended to contain only videos with real-world human actions. Cooking videos can be easily identified by sampling videos that came with “Cooking” or “Recipe” topic labels. Given the convenience and the fact that much of prior work in this area had focused on cooking videos, approximately half of the dataset was designed to include cooking videos only, while the remaining videos would be randomly sampled non-cooking videos, as long as they were verified as instructional by human annotators.
Annotation process
Annotators were presented with a video alongside its timestamped, automatic transcription shown in sentence-length paragraphs. They were asked to watch the video and first judge whether the video was instructional. For the purpose of our dataset, we determine that a video is instructional if it focuses on real-world human actions that are accompanied by procedural language explaining what is happening on screen, in reasonable details. Also for our purposes, instructional videos need to be grounded in real life, with a real person in the video exemplifying the action being verbally described.
For videos judged to be instructional, annotators were then asked to:
Determine their start time if different from the automatically suggested start time (explained below).
Provide a label summarizing or explaining the segment.
Annotation guidelines
Annotators were instructed to identify video segments with two potential purposes:
Allow viewers to jump straight to the start of a segment for rewatch.
Present viewers with an index to decide whether to watch the video in full or directly skip to the segment of interest.
Our guidelines suggested a range of five to ten segments as long as the the structure and content of the video permitted. For short videos, the direction was to prioritize quality over quantity and to only define those segments that formed the narrative structure of the video, even if the resulting number of segments was below 5.
To help annotators determine segment start times, transcriptions were shown in “sentences” — we expected that sentence start times might be good candidates for segment start times. We obtained sentence boundaries automatically as follows. Given the stream of timestamped ASR tokens for a video, we first separated them into blocks by breaking two consecutive tokens whenever they were more than 2 seconds apart. We then used a punctuation prediction model to identify sentence boundaries in each resulting block. Each sentence was shown with the timestamp corresponding to its first token. Annotators were advised that transcriptions had been automatically divided into paragraphs that may or may not correspond to a video segment — if they decided that a segment started from a particular sentence, they could choose to use the start time of the sentence as the start time for the segment, or, if needed, they could put in an adjusted start time instead.
Once the start time had been identified, annotators were asked to provide a free-text label to summarize each segment. We instructed the annotators to use nouns or present participles (-ing form of verbs) to write the labels for the video segments, whenever possible. Additionally, we asked that the labels be succinct while descriptive, using as few words as possible to convey as much information as possible.
Data statistics and post-processing
The resulting dataset consists of 8,169 instructional videos that received segment-level annotations, of which 3,381 are cooking-related. Overall there are an average of 7.1 segments per video (max: 19). Given our instructions, the descriptions are much shorter in lengths compared to a typical captioning dataset: on average there are 2.97 words per description (max: 16); 20% of the captions are single-word, 22% are two-words, and 25% are three words. We refer to these descriptions as “tags” given how short they are.
When possible, annotators were also asked to start and end the video with an opening and closing segment. As a result, most annotations start with an introduction segment: this accounts for roughly 11% of the 88455 segments in the dataset (“intro”: 8%, “introduction”: 2.3%). Note that while “intro” and “introduction” are clearly paraphrases of each other, an automatic metric will penalize a model predicting “intro” when the groundtruth is “introduction”. Similarly, the ending segment was described in several varieties: “outro”: 3.4%, “closing”: 1%, “closure”, “conclusion”, “ending”, “‘end of video”: each under 1%. Penalizing paraphrases of the ground truth is an inherent weakness of automatic metrics. To mitigate this, we decided to reduce the chance of this happening for the most frequent tags in the dataset. That is, in our experiments, we identified three groups of tags among the top-20 most frequent tags, and standardized them as follows.
Note that this does not mean we can solve this problem as a classification task like in visual question answering (VQA): overall, there are 56,027 unique tags with a vocabulary size of 12,509 for the 88,455 segments; 51,474 tags appeared only once in the dataset, making it infeasible to reduce the segment-level captioning problem into a pure classification task. Table 7 shows the top 10 most frequent tags after standardization.
Estimate of human performance.
A subset of the candidate videos were given to three annotatorsA small set were unintentionally given to six annotators., to help us understand variations in human annotations. 5,840 videos received dense captioning from exactly one annotator and were used as training data. Videos with more than one annotation were used as validation / test data. Note that not all the videos with multiple timeline annotations have exactly three sets of them — in fact, 1368 videos received 3-way segment-level annotations. This is because not all annotators agreed on whether a video was instructional. Computing annotator agreement for the annotated timelines is non-trivial. Here we focus on an estimate of tagging agreement when a pair of annotators agreed over the segment start time. Specifically, we go through each video that received multiple segment-level annotations. For each segment where two annotators chose the same ASR sentence as its starting point, we take the tags they produced for this segment and consider one of them as groundtruth, the other as prediction, and add that into our pool of (groundtruth, prediction) pairs. We can then compute standard automatic evaluations metrics over this pool. The results are as follows.
Note that METEOR, and CIDEr scores are both penalized by the lack of n-grams for higher n. That is, when both groundtruth and prediction are single-word, say, “intro”, this pair will not receive a full score from any of these metrics. But the Rouge-L score is in the same ballpark as estimate of human performance in prior work Hessel et al. (2019). One might note that perhaps this pool of label pairs contains a higher share of “intro”, since annotators might be more likely to agree over where an opening segment starts. Indeed, 20% of the time, one of the tags is “intro”. Interestingly, in spite of standardization of top tags, 14% of the time one tag is “intro”, the other tag is not “intro”: they can be less frequent paraphrases (e.g., “welcoming”, “greeting”, “opening and welcoming”) or something semantically different (e.g., “using dremel tool”).
A.2 Separated vs. Concatenated-Modality Architecture
Prior work has explored both concatenating different modalities and feeding them into the same multimodal Transformer encoder (Sun et al., 2019b; Hessel et al., 2019), as well as separating them into unimodal transformers (Sun et al., 2019a; Lu et al., 2019). We opt for the separated architecture because it offers more flexibility. First, the concatenated architecture requires embedding the text and video features into the same space. When the video features are projected using a simple network, there is no guarantee that we can meaningfully project them into the text embedding space. VideoBERT (Sun et al., 2019b) gives more flexibility to the video embeddings by quantizing video features and learning an embedding for each codeword. However, the quantization step has subsequently been claimed to be detrimental (Sun et al., 2019a). Moreover, the concatenated architecture uses the same sets of forward and attention weights to process text and video, and performs layer normalization jointly between the two modalities, which is not necessarily meaningful. Finally, the separated architecture makes it easy to switch between variable length text-only, video-only, or text+video modalities, whereas concatenated architectures might rely on separating tokens, modalities embeddings, and using fixed sequence lengths (Luo et al., 2020).
A.3 Additional Implementation Details
We optimize all models on a nVidia v100 GPU using the Adam optimizer with inverse square root schedule, batch size 32, warm-up period of 4,000 iterations, and maximum learning rate of , following MASS (Song et al., 2019). The positional embeddings are initialized randomly. We use dropout and attention dropout with probabilities . With E2vidD6, pretraining takes 3-6 days depending on the objective and bidirectional finetuning takes up to 1.5 days, however those times could be improved by optimizing the data pipeline.
A.4 Example Predictions
We show examples of good and bad predictions on YouCook2 (Figure 5 and ViTT-All (Figure 4 and 5). The captions are generated by E2vidD6-BiD (no pretraining) and E2vidD6-MASS-BiD (text-only MASS pretraining).
A.5 Full result tables
We present here tables with all the ablation results that we run. There are two main takeaway messages from the results involving the pretraining approach: (a) the accuracy improvements, as measured across all the metrics we use, indicate the value of using a pretraining approach to this problem, specifically one that is capable of leveraging the ASR signals at both pretraining and finetuning stages, and (b) the training speedup achieved from pretraining is impressive, as a pretrained model converges much faster than training from scratch. This is especially visible on ViTT-All where finetuning after MASS pretraining reaches best Rouge-L score at epoch 2, whereas it takes around 11 epochs to converge when training from scratch.