MERLOT: Multimodal Neural Script Knowledge Models

Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, Yejin Choi

Introduction

The human capacity for commonsense reasoning is shaped by how we experience causes and effects over time. Consider the still image of people dining at a restaurant in the bottom right of Figure 1: while a literal, concrete description like “people sitting at a table eating" might be technically correct for the static scene, it doesn’t capture the richer temporal, commonsense inferences that are nonetheless obvious: before sitting down, the people had to meet up, agree where to go, and enter the restaurant; at present, the man is pointing because the server just came to the table, and she might want to know whose food is whose; and after, it is likely the server will return to the kitchen to help another table.

Teaching machines this type of script knowledge is a significant challenge in no small part because enumerating all facts, inferences, and counterfactuals is prohibitive. As a result, the highest performing models on vision-and-language tasks, including Visual Commonsense Reasoning (VCR) (where Figure 1’s scene originates from), learn about the visual world exclusively through static images paired with literal captions . Though some captions might hint at the past and future, it is not obvious that even training on, e.g., 400M literal image/text pairs will result in models capable of temporal reasoning.

In this paper, we introduce MERLOT, short for Multimodal Event Representation Learning Over Time. MERLOT is a model that learns commonsense representations of multimodal events by self-supervised pretraining over 6M unlabelled YouTube videos. With the goal of learning multimodal reasoning capacity beyond static images/literal captions, we train MERLOT to a) match individual video frames with contextualized representations of the associated transcripts, and to b), contextualize those frame-level representations over time by “unmasking" distant word-level corruptions and reordering scrambled video frames.

We validate our model on a diverse suite of video tasks, requiring both recognition- and cognition-level reasoning across long and short timescales; when finetuned, MERLOT achieves a new state-of-the-art on 12 such tasks. Additionally, we show that our script-knowledge representations transfer to the single image domain. On Visual Commonsense Reasoning (VCR; ), our model achieves particularly strong performance, outperforming models that require heavy visual supervision (in the form of object detection bounding boxes, or images paired with pristine captions).

Beyond finetuning, we show both quantitatively and qualitatively that MERLOT has a strong out-of-the-box understanding of everyday events and situations. Given a scrambled visual story, , MERLOT can sort image sequences to match captions which tell a globally coherent narrative. Despite considerable domain shift from videos to static images, MERLOT outperforms strong baselines like CLIP and UNITER , which independently match images to text and thus cannot reason over long-term contexts as effectively. This capacity for temporal coherence emerges during pretraining: analysis of MERLOT’s attention patterns (Figure 15) show that regions attend to captions that are distant in time (and vice versa), allowing it perform cross-modal coreference to piece together a holistic view of situations.

Finally, ablations of MERLOT show that 1) pretraining works better when we train on videos rather than still images, aided crucially by our strategy of corrupting highly visual words in the masked language modeling task, 2) using a diverse set of videos covering many aspects of everyday situations improves downstream performance compared to curated instructional video corpora which both cover a smaller slice of the visual world (confirming hypotheses from past work ); and 3) MERLOT’s performance does not saturate even after many epochs of training on the pretraining corpus we curated, YT-Temporal-180M, as it continues to improve performance simply with more pretraining. The combination of these results suggests that learning full-stack visual reasoning and multimodal world knowledge from video data is a promising path forward for future research.

MERLOT a performant end-to-end vision and language model, that learns powerful multimodal world representations from videos and their transcripts – using no labeled data.

YT-Temporal-180M, a diverse corpus of frames/ASR derived from a filtered set of 6M diverse YouTube videos, which we show greatly aids performance, and

A set of experiments/ablations demonstrating the strong performance of MERLOT on a set of 14 tasks, spanning finetuning and zero-shot transfer, and images and videos.

At rowanzellers.com/merlot, we have released code, data, and models for public research use.

Related Work

There is a long history of work on learning joint text-image representations . Recently, several papers have proposed “Visual BERT” models , trained on image captioning datasets such as MSCOCO . In general, features are extracted using Anderson et al. 2018’s frozen object detector, which was originally trained on Visual Genome . Some exceptions are Zhang et al. 2021, who use an even larger object detector trained on more labeled data; Kim et al. 2021, who use an ImageNet-pretrained backbone , and Shen et al. 2021, who study a CLIP backbone pretrained on web image-caption pairs.

Overall, these approaches all learn visual representations of static images, and rely on significant human annotation in doing so (e.g. through literal image descriptions). Instead, our approach learns dynamic visual representations purely from videos – their frames, and a transcript of what is said – thus using no human annotation.

2 Learning from videos, with automatic speech recognition (ASR) transcripts

Prior works have used web videos with ASR to build weakly-supervised object detectors , action detectors/classifiers , instruction aligners , video captioners , and visual reference resolvers . Of late, works have sought to learn multimodal representations transferable to many tasks from uncurated sets of (usually how-to) videos ; generally these are applied to video understanding tasks like activity recognition. One challenge is designing an appropriate objective for learning video-level representations. Lei et al. 2021’s ClipBERT model learns vision-language representations from image captions, which more literally describe image content versus the longer ASR transcripts we consider. Tang et al. 2021 use a pretrained dense image captioner to provide auxiliary labels for web how-to videos. Both approaches use (supervised) ResNets pretrained on ImageNet as their visual backbones. MERLOT is trained using a combination of objectives requiring no manual supervision; it nonetheless outperforms both prior approaches on downstream tasks.

3 Temporal ordering and forecasting

There has been a large body of work on analyzing ‘what happens next’ in videos . Some modeling choices include using pixels , graphs , euclidean distance using sensors , or studying cycle consistency across time . In addition to extrapolation, past work has studied deshuffling objectives in videos , though this has mostly been limited to the visual modality. In contrast to these papers, our goal is learning multimodal script knowledge representations: using both language and vision as complementary views into the world, instead of just tracking what changes on-screen.

MERLOT: Multimodal Event Representation Learning Over Time

We now present our unified model for learning script knowledge through web videos; including our pretraining dataset, architecture, and objectives.

We collect YT-Temporal-180M, a dataset for learning multimodal script knowledge, derived from 6 million public YouTube videos. Our YT-Temporal-180M intentionally spans many domains, datasets, and topics. We began with 27 million candidate video IDs (which we then filtered), including instructional videos from HowTo100M , lifestyle vlogs of everyday events from the VLOG dataset , and YouTube’s auto-suggested videos for popular topics like ‘science’ or ‘home improvement.’ Our intent (in making the corpus as diverse as possible) was to encourage the model to learn about a broad range of objects, actions, and scenes : we will later show through an ablation that limiting our pretraining to only instructional videos indeed hurts performance downstream.

We filtered videos using the YouTube API, which provides access to videos themselves, their ASR track (automatically transcribed speech tokens), and other metadata. We discard videos 1) without an English ASR track; 2) that are over 20 minutes long; 3) that belong to visually “ungrounded" categories like video game commentaries; and 4) that have thumbnails unlikely to contain objects, according to a lightweight image classifier. We add punctuation to the ASR by applying a sequence-to-sequence model trained to add punctuation to sentences/paragraphs from news articles. Full details of the scraping and filtering are in Appendix A.

Each video V\mathcal{V} might contain thousands of frames. In this work, we represent a video V\mathcal{V} as a sequence of consecutive video segments {st}\{\bm{s}_{t}\}. Each segment st\bm{s}_{t} consists of:

an image frame It\bm{I}_{t}, extracted from the middle timestep of the segment,

the words wt\bm{w}_{t} spoken during the segment, with a total length of LL tokens.

To split the videos into segments, we byte-pair-encode (BPE; ) each video transcript and align tokens with YouTube’s word-level timestamps. This enables us to split the videos into segments of L=32L{=}32 BPE tokens each (Appendix A.4); our final dataset has 180 million segments of this form.

2 MERLOT Architecture

A diagram of MERLOT is given in Figure 2. MERLOT takes a sequence of video frames {st}\{\bm{s}_{t}\} as input. We encode each frame It\bm{I}_{t} using an image encoder, embed the words wt\bm{w}_{t} using a learned embedding, and jointly encode both using a Transformer . After pretraining, the architecture can be applied to a variety of vision-and-language tasks with minimal modification. For video QA, for example, we pass several video frames to the image encoder, the question to the text encoder, and extract a single vector representation from the CLS token position. For each task, we learn a lightweight classification head mapping from this hidden state to the task’s label space; specific modeling/optimization details are given in Appendix E.2.

Image encoder. We train our image encoder end-to-end, alongside the rest of the model, from random initialization (thus without learning from supervised data). While most performant vision-and-language models pre-extract features from a (supervised) object detector , for the sake of pre-training efficiency we use a grid-based hybrid ResNet/Vision Transformer. Standard object detectors have expensive operations for proposing regions, and extracting features from those regions (RoI-pooling); our grid approach avoids these. Recent work has proposed using ‘grid features’ broadly , yet on tasks like VCR these approaches have so far underperformed the more expensive object detector backbones ; our results suggest that ‘grid features’ can perform well broadly.

Specifically: our encoder uses a ResNet-50 backbone, followed by a 12-layer, 768-dimensional Vision Transformer . We made additional modifications that improve efficiency, including: 1) we trained on smaller, widescreen images of size 192x352 (because most YouTube videos are widescreen) using a patch size of 16x16 pixels; 2) we mirror ’s alterations of removing the C5 block in ResNet-50; and 3) we save compute further by average-pooling the final-layer region cells using a kernel size of 2×22\times 2. With these modifications, our image encoder requires 40 gigaFLOPs for a forward pass, which is 2% of the 2 teraFLOPs required for the Faster-RCNN.

In summary: given an image of size W×HW\times H, the image encoder will output a W/32×H/32W/32{\times}H/32 feature map, along with two CLS hidden states: one for pooling a global representation of the image, and another for pretraining (Task 1.).

Joint Vision-Language Encoder. The joint encoder is a 12-layer, 768-dimensional Transformer , mirroring the RoBERTa base architecture ; we initialize it with pretrained RoBERTa weights. To compute joint representations, we first embed the tokens {wt}\{\bm{w}_{t}\} via lookup, and then add position embeddings to both language and vision components (i.e., {It}\{\bm{I}_{t}\}). The position embeddings differ between different segments, so as to distinguish between images and captions at different timesteps. Finally, we pass the independent visual and textual feature maps to our joint encoder.

The tokens wt\bm{w}_{t} in each segment begin with a CLS token; recall that the feature maps for each frame It\bm{I}_{t} start with one as well. At those positions, we will later pool final-layer hidden-state representations, for use in pretraining along with downstream tasks.

3 Pretraining Tasks and Objectives

We use the following three objectives to pretrain MERLOT, that cover ‘full-stack’ visual reasoning – from recognition subtasks (like object detection) that operate at the frame level, to more ‘cognitive’ tasks that operate at the video level.

Contrastive frame-transcript matching . We want to ensure that the underlying image encoder produces helpful image representations. Thus, we use the video transcript to compute a ‘language-only’ representation of each video segment; and use a contrastive loss to maximize its similarity to corresponding representations from the image encoder. To save memory, our ‘language-only encoder’ for this subtask shares parameters with the joint vision-and-language encoder.

Unlike what is the case for many image captions, the words wt\bm{w}_{t} in each segment are often not sufficient to describe the gist of It\bm{I}_{t}, or even what the key objects might be – for that, video-level contextualization is often required. We thus pass the entire transcript into the language-only encoder, which then extracts hidden states for each segment at the segment-level CLS tokens.

Given matching representations for each frame It\bm{I}_{t} and caption wt\bm{w}_{t} as positive examples, the negative examples come from all other frame-caption pairs in the batch – whether or not they come from the same video. We project both of these representations into a size-768 hidden state which is then unit-L2-normalized, and compute an all-pairs dot-product between all image and text representations. We divide these logits by a temperature of τ=0.05\tau=0.05, and then apply a pairwise cross entropy loss to encourage matching captions and frames.

(Attention) Masked Language Modeling When providing words into the joint vision-and-language encoder, we randomly replace 20% with a MASK token, a random word, or the same word; MERLOT must then reconstruct the correct word with a cross-entropy loss, following .

This approach is commonly used by ‘visual BERT’ models in the image captioning domain, where captions are concise, and thus the identity of masked concrete words is difficult for models to recover given language context alone. However, we observed qualitatively that videos break these assumptions: people tend to ramble, and often mention key objects multiple times. Thus, applying vanilla BERT-style masking often causes ungrounded fillers like ‘umm’ or ‘yeah’ to get masked, while the (repeated) names of important objects are often partially masked, penalizing the learning of multimodal representations.

We introduce a simple solution to this problem, that we call attention masking: we use attention weights from a language-only transformer (introduced in the previous objective) as a heuristic for which words are grounded. 50% of the time, we mask out a random token; the other 50% of the time, we mask out one of the top 20% most-attended-to-tokens. We then apply SpanBERT masking , randomly corrupting the following or preceding tokens with an average length of 0.5 tokens in each direction; this makes it harder for models to over-rely on BPE artifacts. We show in ablations that this improves performance.

Temporal Reordering. We have the model order the image frames in a video, forcing it to explicitly learn temporal reasoning and giving it an interface to measure such temporal reasoning. Here, 40% of the time, we randomly pick an integer ii between 22 and NN (the number of segments provided to the joint encoder). Then we randomly scramble ii video frames chosen at random, by replacing the segment-level position embeddings (e.g. [image_t]) for that frame with a random and unique position embedding, e.g. [image_unk_0]). These random position embeddings are learned, and separate from the ‘unshuffled’ position embeddings. This allows the model to order each ‘shuffled’ frame conditioned on frames provided in the correct order (if any).

To compute the reordering loss, we extract hidden states from each frame at the CLS token position. For each pair of frames, we concatenate their hidden states htih_{t_{i}} and htjh_{t_{j}} and pass the result through a two-layer MLP, predicting if ti<tjt_{i}<t_{j} or ti>tjt_{i}>t_{j}. We optimize this using a cross-entropy loss.

4 Pretraining MERLOT

We pretrain our model for 40 epochs over our video dataset. We preprocess the dataset into examples with sequences of N=16N{=}16 video segments each, each containing up to L=32L{=}32 BPE tokens. To train the model on as much data as possible, we merged together the segments of short videos, and split up longer videos, such that all preprocessed examples in our dataset have exactly N=16N{=}16 video segments. The language-only encoder computes contrastive representations given this entire sequence, its total length is thus 512 tokens. To save memory, we provide the joint vision-language encoder 4 groups of N=4N=4 segments each. At an image training resolution of 192×352192\times 352, the joint model’s sequence length is 396 tokens. To combine the losses, we multiply the contrastive loss by a coefficient of 0.250.25, which we found scaled its gradient magnitudes to roughly the same magnitude as the Mask LM loss.

We train the model using a v3-1024 TPU pod, at a batch size of 1024 sequences (or 16k segments) in total. This pretraining process on this hardware takes 30 hours. We provide additional information about hyperparameters and experimental setup in Appendix E.1.

Experiments: Transferring MERLOT to Downstream Tasks

In this section, we explore MERLOT on 14 different tasks, covering vision-language reasoning on static images as well as videos; we present analysis and ablations to dig deeper into our performance.

VCR. We consider VCR , a task and dataset where models must answer commonsense visual questions about images. These questions, about e.g. ‘what might happen next’ or ‘what are people’s intentions,’ force MERLOT to transfer video-level understanding to the world of single images.

VCR provides additional ‘referring expression’ information to models in the form of bounding boxes around named entities. For example, if Person1 is referenced in the question, the location of Person1 is also given in the image. We provide this information to models by drawing (in pixel space) a colored highlight around the referenced entity (Appendix E.3.1), this differs from prior works (that integrate these entities into detection architectures).

Our results on the three VCR settings, in comparison to other models at the same (‘base’) scale, are given in Table 1. Our model outperforms these other models, that all learn from exclusively static images (paired with captions and supervised object detections).

Unsupervised ordering of Visual Stories. To probe our model’s ability to do out-of-the-box commonsense reasoning over events in images, we next consider the Visual Storytelling dataset . Each story in this dataset contains five images and captions in a certain order; the order tells a joint narrative between the captions and the images. Past work has considered unshuffling image-caption pairs , but we take a slightly different approach in this work to avoid language-only biases, which can rely on discursive clues to order text . In our formulation, models are given the captions in sorted order, and must match frames to the captions. Our formulation disarms language-only baselines, while still allowing us to quantify MERLOT’s capacity for commonsense temporal reasoning.

We compare MERLOT with two strong out-of-the-box baselines for text-image matching: CLIP , which encodes each caption and image separately and computes similarity through a dot product, and UNITER which jointly represents each image/caption pair, and is trained in part using a ‘text-image matching’ objective. We use our temporal reordering loss to find the most probable ordering of the video frames (Appendix E.1.1); for CLIP and UNITER we compute a maximum-weight bipartite matching over the pairwise image-text similarity scores.

Results over 5K stories are given in Table 2. MERLOT’s performance in comparison to the algorithms trained from image-literal caption pairs suggests that, with no fine-tuning, our model has strong capability to reason about past and future events expressed in collections of temporal visual stories.

2 Video Reasoning

We report results on 12 video reasoning tasks: TVQA , TVQA(+) , VLEP , MSRVTT-QA , MSRVTT-Multichoice , LSMDC-Multichoice, LSMDC fill-in-the-blank QA , ActivityNetQA , TGIFQA , and DramaQA . We apply MERLOT to these tasks in the same way. We sample a sequence of 5 to 7 still frames from each video clip, initialize new parameters only to map the model’s pooled CLS hidden state into the output labels, and finetune MERLOT with a softmax cross entropy loss; see Appendix E.2 for details.

As shown in Table 3, for all these datasets MERLOT sets a new state-of-the-art. Given the diversity of tasks and the strengths of the comparison models, these results provide strong evidence that MERLOT learned strong multimodal and temporal representations.

3 Ablations

We present ablations over VCR and TVQA+ to study the effect of several modeling decisions.

Context size. Table 4a shows the effect of varying the number of segments NN given to the joint vision-and-language encoder during pretraining. In the first two rows, we provide only a single video segment (N=1N{=}1) to the model. We keep the effective batch size the same, so that we use 4×4\times the number of sequences at 14\frac{1}{4}th the length. In this limited regime, we find that our ‘attention masking’ approach (preferential masking of tokens that were highly attended-to by the contrastive language-only encoder) does not outperform a strong baseline of masking spans randomly . Yet, when we expand the sequence length to N=4N{=}4 segments/128 tokens, our masking becomes more effective, improving by 1 point over the baseline. This supports our hypothesis (Section 3.3.2.) that text-only shortcuts become increasingly viable with length, and that our attention-masking approach counteracts them. Additional qualitative analyses of the attention patterns produced by the language-only encoder are in Appendix C.1; we find that highly attended-to tokens are typically more ‘visual’, and, thus, masking them may make the Masked LM objective require more cross-modal reasoning.

Losses. In Table 4b, we ablate the losses. We find that the contrastive frame-transcript matching loss is crucial to performance, suggesting that an explicit objective is critical for the (randomly initialized) image backbone to learn visual representations. The temporal ordering loss appears less critical for downstream tasks; it helps for TVQA but performance drops slightly for VCR. Thus, we find that it helps primarily as an interface by which we can query the model about temporal events (i.e. for the story ordering experiments); the model might be learning this information from other objectives.

Drawing bounding boxes. Table 4c shows the effects of providing grounding information to VCR models by drawing boxes. Performance drops 5% when they are removed, suggesting that they help.

Dataset source. In Table 4d, we investigate pretraining MERLOT on two datasets beyond YT-Temporal-180M. First, we train on 3 million static image-caption pairs from Conceptual Captions combined with MSCOCO ; for fair comparison, we train for the same number of steps as 5 epochs on our dataset. The resulting model achieves 58.9% accuracy on VCR. We suspect this might be due to 1) a smaller context window (Table 4a), and 2) overfitting (5 epochs on YT-Temporal-180M corresponds to 300 epochs on the caption data). Because our vision pipeline is trained from scratch, the scale of the curated/supervised image pairing corpora is a concern.

We next investigate the impact of video selection, comparing YT-Temporal-180M with HowTo100M . To control for number of videos, we train for an equivalent amount of steps: 5 epochs on our dataset, 30 epochs on HowTo100M, and likewise 30 epochs on a ‘HowTo100M-sized YT-Temporal-180M’. Using diverse YT-Temporal-180M data vs. only instructional videos improves VCR performance by 6.5 points. This suggests that the how-to domain is limited in terms of visual phenomena covered, and that other domains (like web dramas and VLOGs) provide helpful signal for tasks like VCR . Using all the data gives an additional 2.4-point performance boost.

Last, we investigate our choice to preprocess the YouTube ASR text with a language model (adding punctuation, etc); using ‘raw ASR’ instead of this preprocessing reduces performance by 2.4 points.

Pretraining longer. Last, in Table 4e, we investigate the effect of pretraining MERLOT for longer. The performance increases monotonically and doesn’t begin to plateau, which suggests that had we pretrained MERLOT for even longer, its performance could improve even further.

4 Qualitative examples

In Figure 3, we show two qualitative examples of MERLOT’s zero-shot story ordering capability. More examples (and a comparison with the best-scoring baseline, CLIP ) are in Appendix C.2. The examples here show that MERLOT has a strong understanding of events, transcending individual frames. In the first row, it orders the story correctly, performing vision-and-language coreference across several frames (e.g. frames and captions 2 and 3 use ‘he’ to refer to ‘the old man’ only mentioned in the first caption). Without resolving this coreference (establishing the subject as an elderly family member), it seems unlikely that anyone would describe the adults in frame (3) as ‘kids.’ Investigating the attention patterns of MERLOT (Appendix C.3) backs up this claim; they show that MERLOT frequently addresses video tasks by merging attention across (distant) video segments.

MERLOT gets the second row ‘wrong’, but for an interesting reason. It reverses the order of frames (3) and (4), which groups the merry-go-round pictures together – even though caption (3) mentions a barn. This seems to capture the temporal commonsense intuition that people might ride a merry-go-round for a while, i.e., it is not an atomic event .

Conclusion, Limitations, and Broader Impacts

We introduced Multimodal Event Representation Learning Over Time (MERLOT). We trained the model through a combination of self-supervised objectives on 6M YouTube videos, in service of learning powerful multimodal representations that go beyond single frames. The model achieves strong performance on tasks requiring event-level reasoning over videos and static images. We hope that MERLOT can inspire future work for learning vision+language representations in a more human-like fashion compared to learning from literal captions and their corresponding images.

There are several potential limitations of MERLOT that would make for promising avenues of future work, including: 1) exploring finer-grained temporal reasoning pretraining objectives vs. frame ordering e.g., a temporal frame localization within transcripts; and 2) learning multilingually from non-English videos and communities on YouTube.

Like other pretraining work, MERLOT risks some potential negative impacts. We discuss these in more detail below, in addition to the steps we took to reduce these harms.

As with other corpora gathered from the web used for pretraining data, YT-Temporal-180M contains publicly available content posted by users. We thus shaped our data gathering and release strategy to minimize inherent privacy and consent harms (Appendix A.5). Perhaps most importantly, we plan to only share video IDs for download, following a release strategy from prior work and giving users the right to opt out of not just YouTube, but our dataset as well.

2 Social biases.

The curation choices we made in this work could cause the model to exhibit undesirable social biases – for this reason, along with others, we do not advocate for deployed use-cases. For example, 30% of the data selected for by our filtering pipeline was local broadcast news (uploaded to YouTube). Including these news videos seems to perform better than filtering them out and only using how-to videos (Table 4b), however, there are risks when training on them. Local broadcast news (at least in the US) dedicates significant time to covering crime, sometimes in a racist and sensationalized manner . Indeed, running a topic model over our data identifies several ‘crime’ categories (Appendix B). Past work has shown correlation between watching local news and having more explicit racialized beliefs about crime ; it seems likely therefore that training models on this data could teach them learn the same racist patterns.

Additionally, there are inherent social biases on YouTube – and treating these videos as equivalent to ‘the world’ can embed hegemonic perspectives . Most popular YouTubers are men and video practices emerging on YouTube are often gendered . YouTube also has problems with hate, including radical alt-right and ‘alt-lite’ content . These problems – as with other problems in representation and power – are themselves amplified by the ‘YouTube algorithm’ that recommends content to users. Though we downloaded videos independently of YouTube’s recommender system, by filtering based on what content has views, we are implicitly filtering based on this algorithm. The dynamics of YouTube (i.e., which videos get popular/monetized) influence the style and content of videos that get made and uploaded to the platform; this in turn shapes and is shaped by culture more broadly .

3 Dual use.

The video QA tasks that we studied carry risk of dual use, through possible downstream applications like surveillance . It seems unlikely that purely technological fixes and defenses – which themselves can be problematic – could resolve these dynamics. Studying how well video-level pretraining enables surveillance applications might be an important avenue for future work, if only to inform stakeholders and policymakers about these risks.

4 Energy consumption.

The pretraining that we used in this work was expensive upfront . Our results suggest that scaling up the amount of data and compute that we used might yield additional performance gains – but at increased environmental cost. To pretrain more efficiently, we used a much more lightweight architecture (in terms of FLOPs) than is standard for today’s vision and language models. We hope that our public release of the model (for research use) can further amortize this cost.

5 Synthesizing these risks.

With these issues in mind, we release MERLOT and YT-Temporal-180M for researchers. We view our work, and our research artifacts, to be part of a larger conversation on the limits of pretrained ‘foundation models’ . These models have broad impact to real-world areas like healthcare, law, and education. At the same time, these models have significant risks, including the harms that we outlined. We believe that further academic research into this video-and-language pretraining paradigm is important – especially to probe its limits and possible harms. We hope that our paper, code, and data release can contribute to this direction.

Acknowledgements and Funding Transparency Statement

We thank the anonymous reviewers for their helpful feedback that improved this work, along with Oren Etzioni and Gabriel Ilharco. Thanks also to Zak Stone and the Google Cloud TPU team for providing access to the TPU machines used for conducting experiments, and for help with the computing infrastructure. Last, but not least, thanks to all the YouTubers who share interesting videos with the world. This work was funded by DARPA MCS program through NIWC Pacific (N66001-19-2-4031), and the Allen Institute for AI.

References

Supplemental Material

We present the following items in the supplemental:

An exploration of the data in our corpus (Section B)

Qualitative analysis of model representations (Section C)

An exploration of the intermediate visual representations (Section D)

Hyperparameters and experimental setup used for all experiments (Section E)

A Datasheet for our YT-Temporal-180M dataset (Section F)

Appendix A Collecting Videos and Transcripts from YouTube

We adopt the following high-level process to collect YouTube videos and their accompanying transcripts:

Collect channel pages that are likely to cover visually-textually grounded events (A.1),

Download videos from each channel, while filtering out videos without English ASR captions, or unlikely to have (changing) real-world scenes and objects (A.2),

‘Denoise’ the transcripts – using a language model to rewrite transcripts in a style more similar to written English, as opposed to spoken English (A.3),

Last, align words in the transcript to video frames, and extract the segments for pretraining (A.4).

As we will discuss in more detail in the following subsections, we designed our strategy to preserve user privacy as much as possible – an imperative when constructing a corpus on public-facing multimodal data. We conclude with a high-level summary of these privacy-preserving decisions, as well as about our release strategy (A.5).

The first stage in our pipeline was to collect YouTube video IDs that could potentially be relevant for learning visual-textual relationships. We opted to search for interesting channels rather than search for videos directly, as we found the API limits for searching for videos somewhat restrictive. Once a channel was downloaded, we could then download its videos.

We found channels using YouTube’s auto-generated ‘topic’ pages, corresponding to entries in FreeBase like ‘Science’ or ‘Home Improvement.’ We identified 18 of these topics, and retrieved the IDs for all channels that were linked to by each topic page. We also used YouTube channels that appeared in the VLOG dataset , as well as a selection of viral ‘How-To’ and ‘Cooking’ channels. Last, we searched YouTube for concrete nouns, using the object list from MSCOCO (‘baseball’, ‘snowboard’, etc.) as a starting point; we retrieved channel IDs for each video that appeared.

Channels on YouTube often feature other (often similar) channels; so we downloaded more channel IDs by performing a graph breadth-first search over the initial set of channels. We identified 50k channels total and filtered out any more ‘personal’ channels (with fewer than 10k views between all videos). Last, we gathered all video IDs that came from our list of channels, which left us with 27 million video IDs, which formed our final candidate list.

Privacy implications. Our high-level goal was to preserve user privacy by mainly using popular (and more monetized) YouTube videos and channels in our dataset, as opposed to personal ones. The YouTube search algorithm helped us do that, by ordering results (in part) by the popularity of a video / channel. Downloading all videos from a channel, and filtering out channels with fewer than 10k views, favors popular content (like for celebrities, professional YouTubers, and cable news stations). Our analysis in Appendix B shows this strategy was largely successful.

Connection with HowTo100M. As discussed in the paper, we used both a diverse selection of YouTube videos (coming from this process), as well as the video list from HowTo100M . We simply concatenated the video IDs from HowTo100M with the video IDs from this searching step. This means first, that the HowTo100M videos were also filtered by the next steps (and thus our copy of HowTo100M is slightly smaller than the original), though we found that the filtering step had minimal impact on those videos (that were already filtered by ). Second, it means that the HowTo100M videos do contain some instructional videos from less-popular channels. Our intuition here is that this might be okay from a privacy standpoint: few of these people are discussing personal topics; a typical example might be a grainy video of somebody baking cookies. Nonetheless, given the scale that we operated at ourselves, we tried to be more cautious with the filtering.

A.2 Filtering out videos

After retrieving a set of video IDs, our next step was to download ones likely to be appropriate for pre-training MERLOT. Not all videos would are likely to work well: many videos have no spoken words, are not in English, or otherwise do not have automatically-generated (ASR) captions. Likewise, many videos are not grounded: some just have still images (like podcasts), some are of people talking to each other or to the camera, and many are of people playing video games. Our intention was to filter out these videos, ideally without having to download them (so as to conserve bandwidth).

For each video ID, we perform the following steps:

Downloading info: YouTube allows us to download the video metadata separately from each video. We do this first as the video info file is much smaller than the video itself. We thus first (try to) download this file. We exit here if one of the following conditions are met:

the video is categorized as a ‘Gaming’ video,

the video does not contain any English ASR captions,

the video is over 20 minutes long (and thus might be overly expensive to download).

Inspecting thumbnails: the YouTube API has a hidden feature that allows us to download four thumbnails ; in terms of bandwidth usage, this is often much cheaper than downloading the whole video. We use these thumbnails as a proxy as to whether the entire video is likely suitable for pretraining. Note that YouTube thumbnails are also (algorithmically) curated: when thumbnails aren’t hand-selected by the uploader, YouTube’s thumbnail selection algorithm selects high quality, clear frames. https://ai.googleblog.com/2015/10/improving-youtube-video-thumbnails-with.html We trained a lightweight MobileNet-V2 CNN to score whether a COCO object class is present in an image or not, using a sigmoid cross entropy loss. We exit here if one of the following conditions are met:

the CNN classifies fewer than four COCO objects as being ‘present’ over the four frames, using a minimum threshold of 30% probability for an object to be counted as being ‘present.’ This is mainly to recognize scenes with people, as opposed to animations, landscape footage, or blank/placeholder slides.

The average cosine similarity between all feature representations (computed by the classifier) is over 0.9; this allows us to skip videos that have no visual variance (like a person sitting in front of a camera for the whole video, or an album cover while a song is playing).

Downloading the video: if we have not exited yet, we download the video.

A.3 Denoising ASR Captions

One concern with pretraining on ASR is that written text may differ from spoken text: thus, when transferring to downstream tasks based on written corpora, models pretrained on spoken transcriptions may not transfer well. Also, ASR generated by YouTube does not include punctuation or capitalization. Furthermore, ASR transcripts can contain errors, e.g., by mistranscribing rare words/proper nouns and instead predicting incorrect, but similarly pronounced, words. And finally, YouTube’s ASR system sometimes attempts to translate text from a different language to English, which is sometimes successful, but other times produces nonsense.

We aim to sidestep these issues by using a language model to ‘denoise’ ASR text, as well to filter out excessively noisy transcripts. We use a GROVER-Large language model to do this , as it was exclusively pretrained on written text from news articles. Then, we finetuned it in a sequence-to-sequence setting to ‘denoise’ ASR.

We created data for our ‘denoising’ task using the following procedure. Given an article from RealNews , we would trim it to 600 BPE tokens, and perform the following corruptions:

We lowercase all text, and remove all punctuation.

For each word (splitting by whitespace), we replace it with a random word 1% of the time. Within this 1%, 25% of the time, we use the CMU Pronouncing Dictionary http://www.speech.cs.cmu.edu/cgi-bin/cmudict to swap-in a word with identical pronunciation (to simulate mistranscriptions), and 75% of the time we use a random sequence of BPE tokens of the same length as the actual word.

For each word, 1% of the time we insert a ‘filler word’ before it, such as ‘umm,’ ‘hmm,’ or ‘yeah.’

The model was trained to generate the ‘noisy’ news article, followed by a ‘START’ token, then the original ‘clean’ news article, and then an ‘END’ token; all using a standard cross-entropy loss. We prioritize learning the ‘clean’ text by multiplying the loss on the initial ‘noisy’ tokens by 0.010.01. We trained this model using a batch size of 256 sequences of maximum sequence length 1536, a learning rate of 1e-5, and 80k steps.

The result is a model that not only attempts to fix mistranscriptions and corruptions, but also adds punctuation and capitalization. The model also produces an estimated likelihood of the ASR caption track, which we later use to filter out videos with very low quality ASR transcripts, e.g., poorly translated transcripts.

We apply the model to each video’s transcript that survived the described filtration, breaking up long transcripts into groups of 512 tokens. These groups are handed as input to the model, and Nucleus Sampling (with pp=0.9) is used to generate a cleaned transcript for the group. We exit, filtering out the entire video, if any group has a perplexity of over 200. Finally, we concatenated all the groups together to form a ‘clean’ transcript.

A.4 Putting everything together: aligning videos and cleaned transcripts to frames

To recap, at this stage in the pipeline, for each video, we have the video file, along with the original ASR transcript (with words, as well as timestamps for each word), and the cleaned ASR caption (without timing info). To estimate timing info for the clean transcript, we align the noisy and cleaned transcripts on a word-by-word level using Dynamic Time Warping ; word-word distance is computed using Levenstein distance. The timing estimate for a cleaned token was computed as the average of the noisy tokens assigned to it in this alignment.

Finally, given a video and its cleaned, per-word timed transcript, we sought to extract corresponding video frames – the data format we rely on for pretraining. We start with (empty) buffers of at most L=32L=32 tokens for both the original, and noisy transcripts. We loop through the (aligned) clean and noisy transcripts, and add the tokens to their respective buffers. If adding the next word would cause the buffer to exceed L=32L=32 tokens in length, we commit the segment – returning the noisy ASR text, along with the clean text, and timing information. We then extract a frame from the video corresponding to the middle of that segment. We do this until the end of the video. We use the GPT2 BPE encoder for this , as was also widely adopted in later work (e.g. RoBERTa ).

Not all videos fit neatly into 16 segments, which was the format we used for training. Thus, we merged segments from videos shorter than 16 segments, and for longer videos, we split them into multiple examples. We didn’t use any video sequence-level padding: all of our dataset examples have 16 valid frames, even though we did include padding at the token level (so many segments had fewer than L=32L=32 tokens).

A.5 Summary - scraping while preserving privacy

As we discussed in the sections above, we tailored our scraping process to protect user privacy. It should be mentioned here that we focused on public videos. Possibly due to cues of engagement like view/subscriber counts, users on YouTube appear to understand the privacy implications of uploading a ‘public’ video , differentiating YouTube from more private venues, like email and social media. Under Marwick and boyd 2014’s framework of networked privacy, when web users (particularly those with less viewership) upload public videos, they are often ‘being in public without being public.’ The idea behind this distinction is that web users, understanding that their content might be visible to others, tend to avoid sharing overly private data (like their phone number or date of birth); the information that they do share is often encoded (i.e., referring to a friend by their first name, not their full name). Finally, we took extra steps to filter out more ‘personal’ videos (without many views); our analysis in Appendix B shows this strategy was largely successful.

An additional aspect of our approach, as it relates to privacy, was our decision to use a diverse selection of channels. We did this to minimize risks of models ‘overfitting’ to specific individuals – a risk evidenced by a large GPT2 model memorizing users’ phone numbers . We believe that training a base-sized model in a large- and diverse-data regime minimizes many of the harms in this case; that said, the risk in the multimodal (video) space is unclear as of yet, and more research is needed.

Finally, we do not plan on releasing videos for download, only their IDs, following a strategy from prior work . This gives users an explicit ‘right to be forgotten’ not just from YouTube, but our data as well. We understand that this might make exact reproducibility difficult; we address this by releasing code for our filtering process. Thus, if in the future, if NN videos get deleted from YT-Temporal-180M, a practitioner can download NN new YouTube videos that pass through the same filters that we used.

Appendix B Data Exploration

Curating large pretraining corpora necessitates some ad-hoc decisions, e.g., what data to search for, what data to keep/discard, etc., and our work is no exception. The described data extraction pipeline contains several heuristics that we developed based on our subjective experiences (and per-step, heuristic validations) curating the corpus. While it isn’t computationally feasible ablate each stage of this pipeline (and examine each decision’s effect on downstream performance), we seek to quantify some basic of the properties of the corpus.

We randomly sampled 100 videos from the corpus, and answered the following basic questions for each of the videos: Q1: Does the video contain language utterances? Q2: If so, is the language primarily English? Q3: Is the video an instructional video, i.e., is it an attempt to teach the viewer how to undertake a task? A similar definition was proposed in . Q4: What type of entity created the video: a small youtuber (<10K subscribers); a medium youtuber (<100K, >10K subscribers); or a large youtuber (>100K subscribers); a news station; or a media company. Q5: Is the video a music video? Q6: Is the video a video game commentary?

Of the 100 examined videos, none were music videos or video game commentaries (Q5/Q6). The videos were mostly not instructional (84%) (Q3) and mostly in English (86%) (Q2); non-English videos nonetheless can have an English ASR track provided by the YouTube API if the spoken language is transcribed by YouTube via its auto-translate feature. And while all contained language utterances (Q1), at least one translated transcript had a very low quality transcription, which was only loosely semantically related to the underlying content. Finally, the most common video creators were news studios (29%; e.g., local news channels); big YouTubers (26%; e.g., popular vloggers), and media companies (24%; e.g., Major League Baseball). Also included, but in lesser proportion, were small YouTubers (8%), and TV studios (1%; e.g., official movie trailers).

What topics are covered by the corpus? We randomly sampled 55K video transcripts, and ran an LDA topic model implemented in MALLET with 100 topics. We used a vocab size of 25K word types that appear in at least 25 transcripts, but in no more than 10% of transcripts. The topics suggest diverse coverage, e.g., topics about specific sports (boxing, soccer), US and world politics, fashion, construction, fantasy settings, nail painting, etc. We use TSNE to visualize the per-document topic distributions, and color a sample of documents according to their top topic in Figure 4 (topic details in Table 5).

Overall, the topical coverage of YT-Temporal-180M, at least according to a topic model trained on the transcripts of a sample of videos, is broader than comparable-in-size video corpora like HowTo100M . And, experiments in the main paper demonstrate that this diversity is apparently helpful for a number of downstream tasks.

Appendix C Qualitative Analysis of Model Representations

In this section, we provide more qualitative analysis about the representations learned by MERLOT.

Early on in this project, when inspecting qualitative examples, we observed that using BERT-style masked language modeling – choosing 15% randomly selected BPE tokens as the prediction targets, and replacing them with MASK 80% of the time, or a random token 10% of the time – produced overly easy examples.

This has been observed by other work in the text-only setting: when long words get partially masked, it is often easy to recover the missing BPE token from the context, which motivated Joshi et al. 2020’s choice to mask out entire spans instead. However, our goal in multimodal pretraining is different. We want the model to learn grounded representations of events, such that even when we scale up the number of segments given to the model, the model has to construct a multimodal representation of what happened. Thus, in our setup, we wanted to encourage masking out highly visual words, to learn cross-modal representations.

Instead of masking randomly, recall that we used the attention weights produced by the language-only encoder (trained to match a sequence of captions to individual frames) to inform which tokens to mask. While we do not claim that these attention weights provide a full explanation of the model behavior , they do play some role in the model’s decision , and we find that our masking strategy improves performance on downstream tasks by around 1% (Table 4), versus a SpanBERT baseline .

We show qualitative examples that seem to back up our hypothesis in Figures 7 and 10. In Figure 7, for instance, the video shows a VLOG of an adult playing with children and talking the camera. Tokens flagged by our approach as having high attention weights (being in the top 20% of all tokens in the sequence, in terms of other positions attending to that token) include concrete words like ‘scissors’ and ‘toys.’ Even though scissors are not shown in the selected frames, that word might be a good prediction target, insofar as it might complete a picture of what is going on in the first few frames: somehow, the adult is able to open the package with the child’s toy, which could require scissors. Additionally, in Figure 10, showing an instructional video for both making iced tea and putting it in a sealed mason jar, concrete nouns such as ‘o-rings’ get masked out.

Nevertheless, there are still several cases where the model seems to assign attention weights to apparently non-visual tokens. The model places a lot of attention on the START token, a pattern noticed by prior work as well , perhaps because we pool representations from those positions (for matching with the video frames). However, we never select the START token for masking in our work, so this might not highly affect the learning signal. Perhaps more strangely, language-only encoder seems to attend highly to the final token in contractions (like ’t and ’s). It is not clear to us whether these represent something important visually, or noise; we leave a more in-depth investigation of this phenomenon to future work.

C.2 More qualitative examples for zero-shot story ordering

In this section, we show more examples of MERLOT unshuffling visual stories in SIND . We compare our model’s zero-shot results (using the logits from its temporal-ordering objective) to CLIP’s independent matching of each caption with each image (using the Hungarian algorithm to find the best-scoring assignment ).

In Figures 11 and 12, we show expanded versions of Figure 3, comparing to CLIP. The examples show that MERLOT has a strong understanding of events that transcends individual frames. Unlike MERLOT, CLIP can only match captions independently to images, so in the first row it struggles to connect ‘his kids’ with the middle-aged children of ‘the old man’ In the second row, it matches the barn image with the caption ‘they also had a barn’, while it is unable to keep all the merry-go-round images together (as MERLOT does).

We show additional examples in Figures 13 and 14. Our model provides a reasonable ordering to the ‘kayaking’ example (Figure 13), which is evident of multimodal script knowledge: first, people have to get ready to go kayaking (which they do on land!) and then they go out onto the water, and finally come back. The ordering of the tennis match (Figure ) seems reasonable as well. Unlike CLIP, MERLOT groups together frames (3) and (4) – the players first serving the tennis ball, and then awaiting the return.

C.3 Attention patterns

Finally, we show examples of the attention patterns produced by MERLOT, when it reasons over both vision-and-language content at a video level. Plots are shown in Figure 15. Overall, the model frequently links together visual regions with similar concepts in text, even when they get mentioned far away in time.

Though these attention patterns should be taken with a grain of salt, as they are not necessarily explanatory of the model’s decision , we find it promising that the model attends globally over all frames and captions – rather than ignoring one modality or ignoring the temporal dimension. We leave further investigation of the model’s attention patterns and behavior to future work.

Appendix D Linear Probe of Intermediate Visual Representations

Our goal with MERLOT was to learn about situations expressed through videos and language. However, as it includes a vision encoder that we trained from scratch, a reasonable question is how this visual encoder compares to other encoders (e.g., that were trained through image captions). To this end, we performed linear probing experiments over two activity recognition datasets: HMDB-51 and UCF-101 . These tasks are 51 and 101 class classification tasks, respectively: they challenge algorithms to predict which human activity is present in a video clip. Following prior work, for both datasets, we average over the three standard train/test splits. We evaluate in the linear probe setup, where models represent video clips as a single fixed vector, and a linear maximum entropy classifier is trained on top, freezing the rest of the model’s parameters.

In addition to a random prediction baseline, we compare against ’s RSPNet reported results (they use a 3DResNet-18 backbone pretrained on Kinetics400), and CLIP ViT-B/16 . For MERLOT and CLIP, we extract a single central frame from each video, and extract a feature vector from it. For MERLOT, we represent the frame as the concatenation of the two [CLS] tokens (one was for the image-transcript alignment task, the other was for passing to the joint encoder).

The results, shown in Table 6, show that CLIP performs best in this setup – though MERLOT does outperform an RSPNet baseline. At first, this might appear surprising, as MERLOT was trained on web videos, which might be closer to activity recognition datasets (as opposed to image captions). However, common benchmarks for activity recognition tend to have strong object and background bias – for example, to recognize the UCF action ‘playing guitar,’ it is sufficient to detect a guitar in an image (as guitars are unlikely to show up for the other activities like ‘playing basketball’) . Temporal self-supervised learning from transcripts may not lead to as powerful zero-shot object detectors because speakers in videos may be less likely to state the obvious , e.g., in this case, a speaker is probably unlikely to say ‘I will now play a guitar while sitting in a chair.’

Appendix E Experimental setup and hyperparameters

We used AdamW with a learning rate of 3e−43e-4, weight decay with value 0.10.1, and set β2=0.98\beta_{2}{=}0.98. We used minimal data augmentation on the image frames. We randomly scale them between 1.125 and 1.5 times what would fit in our 192×352192\times 352 resolution, and take a random crop. We use a random resize algorithm when doing this scaling, to make the model robust to different ways of preprocessing images . Last, for 80% of images, we randomly jittered either their brightness or contrast to between 0.70.7 and 1.31.3 their original values, which we suspect did not play a major role in performance.

On the text side, we note that we have both the original copies of each transcript – what was retrieved from YouTube – and versions “cleaned up” by our denoisifier. We can use both kinds of transcript as additional data augmentation. However, although the words are time aligned, there might be inconsistencies if alternating between cleaned and noisy versions inside of a single video. Thus, for each iteration, we randomly choose either the ‘clean’ or ‘noisy’ ASR transcript and use that one.

To slightly speed up convergence, we initialize the joint vision-and-language model, and the word embeddings, with parameters from RoBERTa . However, we suspect that due to the scale of our dataset and pretraining time, this might not have been required.

For the unsupervised scrambling of visual stories task, we did not do any finetuning on the SIND dataset . However, there is a slight mismatch between the model that we pretrained initially, and the format of the task – the visual stories in the SIND dataset have 5 images and captions each, whereas we initially pretrained with at most 4 segments. We handled this discrepancy by pretraining MERLOT for 10 more epochs, using a peak learning rate of 2e-5, and a new resolution of 384 x 384. This slightly bigger size was to account for the (not necessarily) widescreen images in SortStory, as opposed to the (mostly) widescreen videos on YouTube.

Recall that MERLOT’s pairwise loss is defined over pairs of segments. However, how to best combine these into a unified score for story ordering is an open question. To briefly explore this, during this additional pretraining of MERLOT, we applied three variants of our temporal loss: one over caption-caption pairs, one over caption-frame pairs, and one over frame-frame pairs. We also experimented with randomly shuffling the captions as well, in the same way as the frames, we found however that this did not boost downstream task performance (perhaps because using shuffled captions as input incentivizes models to learn exclusively language-language interactions). The loss is computed the exact same way everywhere; the only differences is that for caption-frame pairs, we have four options:

the caption (at tit_{i}) and frame (at tjt_{j}) are of the same segment, so ti=tjt_{i}=t_{j},

the caption precedes the frame, so ti<tjt_{i}<t_{j},

the caption comes after the frame, so ti>tjt_{i}>t_{j},

the caption comes from a different video as the frame, so comparing tit_{i} and tjt_{j} is undefined.

The model learns to distinguish between those four options with a cross-entropy loss. We found that using this version of the temporal loss over vision-language pairs produced slightly better results on story ordering (as judged on the validation set) compared with the loss applied over the frames. We hypothesize that this might be due to the additional ‘ti=tjt_{i}=t_{j}’ option allowing models to assign a probability to a frame-caption match, but are not sure. With this approach, to produce a unified score for (length-NN) permutations σL\sigma_{L} over the captions, and σV\sigma_{V} over frames, we then sum over pairwise log-probabilities:

For story ordering, the order of the captions is always fixed: σL=(1,2,3,4,5)\sigma_{L}=(1,2,3,4,5) and N=5N=5; we thus feed MERLOT captions with the correct order. However, the model should have no information about the order of the frames. Embarassingly, we found a slight leakage of this in the V1 of this arxiv paper which inflated the story ordering performance by a few percentage points (of pairwise accuracy), which we have corrected in this version. Recall that we handle this through position embeddings (3.3); e.g. one possible ordering might be

and those position embeddings would get added to each frame, respectively. This allows the network to disambiguate between distinct frames even though no order is revealed. However, we found that the model was sometimes sensitive to the exact order of these position embedding tokens, and so for each example we randomly sampled two orderings and averaged the model’s pairwise probabilities. We found no difference in performance when using more than two orderings. We hypothesize that this could be an issue with how (absolute) position embeddings are handled by Transformers, but are not fully confident; we leave a more thorough investigation for future work.

E.2 Per-downstream fine-tuning details.

In this section, we discuss implementation details for finetuning MERLOT on downstream tasks. For each downstream task, given images I1:N\bm{I}_{1:N} and language context w\bm{w}, we first encode I1:N\bm{I}_{1:N} via the image encoder. We concatenate this with word embeddings of w\bm{w}, apply position embeddings, and feed the result into the joint vision-language encoder to extract joint representation. The input images I1:N\bm{I}_{1:N} are either provided by the task or extracted from given video, where we uniformly select NN frames from the video clips (spaced evenly, so with an equal amount of time between sequential frames). For supervised tasks, we use as the ‘head’ a two-layer MLP from random initialization on top of the CLS token of the language context together with the rest of MERLOT.

For downstream tasks, we note that we found it effective to finetune on different resolutions than what we used during pretrianing. Our default image resolution here was 384×\times 704. To do this, we note that all parameters in the model remain the same, except for position embeddings on the image patches. We expanded the size of the position embedding matrix by initializing the upper-left-side 192x352 region from the pretrained model, and used random initialization for new position embeddings.

For all downstream tasks, we followed the standard training, validation, and test splits of the original datasets. We used the AdamW optimizer, with β2=0.98\beta_{2}=0.98, and warmed up the learning rate linearly for the first 10% of iterations, followed by a linear decay of the learning rate (down to 0) for the remaining 90%. For regularization, we used L2 weight decay with a value of 0.01, and a dropout rate of 10%. For tuning other hyperparameters, we first did a larger random hyperparameter search over VCR, and used those hyperparameters as defaults for the other tasks. We used a batch size of 64, and searched over learning rates in the range [1e-5, 2e-4] on VCR, we found that 1.2e-5 worked well, so we used it as the default for other tasks. We also trained with early stopping, validating every epoch and returning the best-performing model across epochs. Due to our choice of early stopping, we trained for a slightly larger-than-typical number of epochs (18 by default for every tasks, as we found training longer did not help on VCR).

We follow the standard evaluation metrics for these tasks, which is usually accuracy for QA-style configurations. Alongside brief descriptions of each downstream task, we provide hyperparameter and training details in the following section.

E.3 Static Image Reasoning Tasks

VCR contains two different subtasks: question answering (Q→\rightarrowA) and answer justification (QA→\rightarrowR), both of which are multiple choice questions over a given image. These subtasks are combined in the joint Q→\rightarrowAR metric, which requires a model to both pick the right answer and the right rationale for the model to get a question ‘right.’ VCR has 290k questions over 110k movie scenes.

As mentioned in the main text, VCR provides bounding boxes around entities, with explicit groundings between those entities and references in questions. We draw colored highlights around the referenced entity directly in the image, with consistent mapping between color code and entity name (e.g. person1 with red box, person2 with green box, etc). Though no text is written on the image, because we always associate each string (e.g. person1) with a deterministic color, the model can learn through finetuning to associate that color with the entity. Figure 16 illustrates one such example.

We jointly finetune MERLOT on Q→\rightarrowA and QA→\rightarrowR, with two separate MLP heads. We concatenate the question (the question and the ground truth answer) and each answer (rationale) choice from the four possible answer (rationale) candidates. On-top of the CLS token of the question, we train the classifier to predict the confidence for each candidate to be correct with cross-entropy loss, and take softmax over four possible candidates for each question. We used a widescreen resolution of 384×\times704 set the batch size as 64, and train for 60k training steps, which is roughly 18 epochs. We started with this and then tuned the learning rate (from candidates chosen randomly); here, we found that a learning rate of 1.2e-5 worked well. We then used this learning rate as a default for the other tasks.

Note that our pretraining setup is different from other work. Previous works conduct what they call ‘second-stage pretraining’ with VCR training data. Here, they use a masked language model objective over the VCR dataset (instead of answering the question correctly). In particular, UNITER reports 2.8 % point performance boost due to the second-stage pretraining. We suspect that this might be because the caption data (that models like UNITER rely on) are quite different from VCR. We tried performing secondary pretraining and found it did not help. One possible reason might be that our large-scale pretraining corpus covers diverse and complex event space thus we don’t need additional data domain adaptation.

E.4 Video Reasoning Tasks

MSRVTT-QA is a question-answering task with 244K questions posed over 10K videos. For each video clip, we uniformly selected 5 image frames (spaced evenly through the video). We follow the protocols of the original work and use an answer vocabulary containing the most common 1K answers in the training set as answer candidates. The questions with out-of-vocabulary answer will automatically get wrong. We encode the answers in a one-hot fashion, and train 2-layer MLP classifier over all answer candidates with a binary cross-entropy loss on-top of the CLS token of the question. We train for 60k training steps with batch size 16. A few additional fine-tuning runs were conducted to examine the effect of changing the resolution from 384×\times704 to 704×\times704, a batch size of 16 vs. 32, and and using 1.5K answers instead of 1K, but none had much impact on validation accuracy. We undertook a light hyperparameter optimization over the validation set, wherein we considered 3 possible learning rates (1.2e-5, 6e-5, 2.4e-6), but the default worked best. MSRVTT-QA splits questions by type, and we report our per-type test set results in comparison to in Table 7.

TVQA is a multiple choice task with 152K questions posed over 21K video clips. For each clip, we uniformly select 6 image frames. We concatenate the question and each answer choice from the five possible answer candidates. On-top of the CLS token of the question, we train 2-layer MLP classifier to predict the confidence for each candidate to be correct with cross-entropy loss, and take softmax over five possible candidates for each question. We set the batch size as 64, and train for 35k training steps (roughly 18 epochs over the corpus). We used the default learning rate of 1.2e-5, and a resolution of 384×\times704.

TVQA+ is a subset of TVQA, where bounding boxes are provided in video clips, linking depicted objects to visual concepts in questions and answers. TVQA+ contains 29.4K questions posed over 4.2K video clips. We uniformly select 6 image frames per video, and draw bounding boxes on each frame following the same manner with VCR. We train the classifier in the same way with TVQA. We trained with the same hyperparameters as TVQA, but for 16k steps (18 epochs still).

VLEP is a binary choice task to infer which of the two events is more likely to happen next following the given video. VLEP contains 28.7K questions posed over 10K video clips. For each clip, we uniformly select 6 image frames. On-top of the CLS token of the event, we train 2-layer MLP classifier to predict the confidence for each event to happen next with cross-entropy loss, and take softmax over two possible events for each instance. We trained the model for 8k steps (18 epochs over the dataset), and with otherwise default hyperparameters.

DramaQA is a multiple choice task with 17.9K questions posed over 23.9K video clips. For each clip, we uniformly select 6 image frames. We concatenate the question and each answer choice from the five possible answer candidates. On-top of the CLS token of the question, we train 2-layer MLP classifier to predict the confidence for each candidate to be correct with cross-entropy loss, and take softmax over five possible candidates for each question. We trained for 3.5k steps (18 epochs) with otherwise default hyperparameters. A few additional fine-tuning runs were conducted to examine the effect of changing the resolution between 384×\times704, 512×\times512 and 704×\times704, and we found 512×\times512 works the best for this task.

TGIF-QA is web GIF VQA, which requires spatio-temporal reasoning from visual frames to answer questions correctly. We finetuned MERLOT on three tasks in TGIF-QA benchmark,

Action is defined as a multiple choice question about identifying an action that has been repeated in a video.

Transition is asking about transitions of certain states. The benchmark provides a multiple choice question about identifying the state before or after another state.

FrameQA is asking open-ended questions about the given video. The model selects answer from a dictionary of words, given a question in a complete sentence.

For each video clip, we uniformly select 5 image frames. We serialized 5 candidate answers and a question, where we put a special token QSEP between the candidate answers and question to concatenate them into one question. On-top of the CLS token of the question, we trained 2-layer MLP to predict the confidence of the five candidates with cross-entropy loss. We set the batch size as 16, and train for 70k training steps (Action : 56 epoch, Transition : 22 epoch, FrameQA : 28 epoch) for each task with 1.2e-5 learning rate. We used a longer training duration for each task as we found that performance increased when we did so (and we used the same number of training steps for each TGIF-QA task). All other hyperparameters were default.

ActivityNetQA is a question-answering with 58K questions posed over 5.8K videos. For each video clip, we uniformly select 5 image frames. We use an answer vocabulary containing the most common 1K answers in the training set as answer candidates. The questions with out-of-vocabulary answer will automatically get wrong. We encode the answers in a one-hot fashion, and train 2-layer MLP classifier over all answer candidates with a binary cross-entropy loss on-top of the CLS token of the question. We set the batch size as 16, and train for 34K training steps for each task. We undertook a light hyperparameter optimization over the validation set, wherein we considered 3 possible learning rates (1.2e-5, 6e-5, 2.4e-6), but the default worked best. A few additional fine-tuning runs were conducted to examine the effect of changing the resolution from 384×\times704 to 704×\times704, a batch size of 16 vs. 32, and using 1.5K answers instead of 1K, but none had much impact on validation accuracy. ActivityNetQA splits questions by type, and we report our per-type test set results in comparison to in Table 9.

The Fill-in-the-blank (FiTB) task is, given a video clip and a sentence with a blank in it, to predict a single correct word for the blank. The test set includes 30,000 examples from 10,000 clips (i.e. 3 blanks for each description). For each clip, we uniformly select 5 image frames. We constructed answer vocabulary containing the most common word for blank in the training set as answer candidates. We replace the blank in the sentence with BLANK token, so the question query should be a blanked sentence with the special token. On-top of the CLS token of the blanked sentence query, we trained 2-layer MLP classifier to predict the word for the blank over answer vocabulary. We set the batch size as 16, and train for 150k training steps (8 epoch) with 1.2e-5 learning rate.

Given a video query and 5 candidate captions, the task is to find the one that fits the query out of 5 possible candidates. The correct answer is the ground-truth (GT) caption, and four other negatives are chosen from other captions that have different activity-phrase labels from the correct answer. We randomly created 100,000 video and candidates pairs for training. For each video clip, we uniformly select 5 image frames. We put a special token QSEP between the candidate captions to concatenate 5 candidates into one question. At the end of the 5 captions, we put CLS token as an end of the question. On-top of the CLS token, we trained 2-layer MLP to predict the confidence of the five candidates with cross-entropy loss. We set the batch size as 16, and train for 80k training steps (12 epoch) with 1.2e-5 learning rate.

The task objective for the MSRVTT Multichoice benchmark is identical to those of corresponding tasks in the LSMDC benchmark . The benchmark has 2,990 questions in total for the multiple choice test, using all the test video clips of MSR-VTT. For each test video. We finetuned our model on MSR-VTT train split, and evaluated on the evaluation set. We trained the same model specification as the LSMDC Multichoice task. For training, we set the batch size as 16, and train for 80k training steps (12 epoch) with 1.2e-5 learning rate.

Appendix F Datasheet for YT-Temporal-180M

In this section, we present a DataSheet for YT-Temporal-180M, synthesizing many of the other analyses we performed in this paper.

Why was the dataset created? In order to investigate learning events from videos – involving a collection of frames and captions over time, that together form a view about the world.

What (other) tasks could the dataset be used for? Possibly other types of representation learning, with or without ASR captions.

Who funded dataset creation? This work was funded by DARPA MCS program through NIWC Pacific (N66001-19-2-4031), and the Allen Institute for AI.

What are the instances? The instances that we consider in this work are videos, paired with ASR transcripts aligned over time.

How many instances are there? We include 6 million videos. The total length of all the ASR transcripts is 5 billion BPE tokens. Altogether, we extracted 180 million image frames from this data.

What data does each instance consist of? The instances have ‘raw’ video frames and text, which we preprocess through BPE tokenization and extracting frames for every 32 BPE tokens.

Is there a label or target associated with each instance? We only use the ASR captions as labels in this work, though it might be also possible to use auxiliary information (like tags or video titles).

Is any information missing from individual instances? No.

Are relationships between individual instances made explicit? Not applicable – we do not study relations between different videos (e.g. made by the same creator), though this is a possibility for future work

Does the dataset contain all possible instances or is it a sample? Just a sample.

Are there recommended data splits (e.g., training, development/validation, testing)? We do not provide recommended data splits at this time, as this data was built only for pretraining rather than evaluation. We suspect that the data is large enough that overfitting is not a major concern.

Are there any errors, sources of noise, or redundancies in the dataset? If so, please provide a description. Yes. YouTube ASR is often noisy, and though we presented a pipeline to correct some of these errors, there are many that we cannot fix.

Is the dataset self-contained, or does it link to or otherwise rely on external resources (e.g., websites, tweets, other datasets)? The dataset is self-contained. However, we plan to only release the video URLs, rather than the videos themselves, so as to protect user privacy (allowing users to delete videos).

What mechanisms or procedures were used to collect the data? We used the YouTube API and the youtube-dl library.

How was the data associated with each instance acquired? Was the data directly observable (e.g., raw text, movie ratings), reported by subjects (e.g., survey responses), or indirectly inferred/derived from other data? The data was directly observable (from YouTube).

If the dataset is a sample from a larger set, what was the sampling strategy (e.g., deterministic, probabilistic with specific sampling probabilities)? We used a probabilistic strategy with many heuristics, more details in Appendix A.

Who was involved in the data collection process (e.g., students, crowdworkers, contractors) and how were they compensated (e.g., how much were crowdworkers paid)? Data collection was primarily done by the first authors of this paper.

Over what timeframe was the data collected? Does this timeframe match the creation timeframe of the data associated with the instances (e.g., recent crawl of old news articles)? If not, please describe the timeframe in which the data associated with the instances was created. The data was collected from November 2020 to April 2021, even though the YouTube videos are often much older (dating back to when the platform was first created).

Was any preprocessing/cleaning/labeling of the data done (e.g., discretization or bucketing, tokenization, part-of-speech tagging, SIFT feature extraction, removal of instances, processing of missing values)? Yes, we discuss this in Appendix A: of note, we use a sequence-to-sequence model to ‘denoise’ ASR transcripts (Appendix A.3), BPE-tokenize text, turn everything into segments, and extract the middle image frame for each video segment.

Was the “raw” data saved in addition to the preprocessed/cleaned/labeled data (e.g., to support unanticipated future uses)? If so, please provide a link or other access point to the ‘raw’ data. The raw data was saved, but at this time we do not plan to release it directly due to copyright and privacy concerns.

Is the software used to preprocess/clean/label the instances available? If so, please provide a link or other access point. We will make our code public to support future research.

Does this dataset collection/processing procedure achieve the motivation for creating the dataset stated in the first section of this datasheet? If not, what are the limitations? We believe our dataset does allow for study of our goal – indeed, it covers grounded temporal situations from a variety of domains – but with significant limitations. Some of the key ones we are aware of involve various biases on YouTube, which we discuss in Section 5.

How will the dataset be distributed? At this time, we plan to distribute all the metadata (transcripts, etc) that we used, as well as links to the YouTube videos that we used. We will do this on our website.

When will the dataset be released/first distributed? What license (if any) is it distributed under? We will release it as soon as possible, using a permissible license for research-based use.

Are there any copyrights on the data? We believe our use is ‘fair use,’ however, due to an abundance of caution, we will not be releasing any of the videos themselves.

Are there any fees or access restrictions? No.

Who is supporting/hosting/maintaining the dataset? The first authors of this work.

Will the dataset be updated? If so, how often and by whom? We do not plan to update it at this time.

Is there a repository to link to any/all papers/systems that use this dataset? Not right now, but we encourage anyone who uses the dataset to cite our paper so it can be easily found.

If others want to extend/augment/build on this dataset, is there a mechanism for them to do so? Not at this time.

Were any ethical review processes conducted (e.g., by an institutional review board)? No official processes were done, as our research is not on human subjects, but we had significant internal deliberation when choosing the scraping strategy.

Does the dataset contain data that might be considered confidential? No, we only use public videos.

Does the dataset contain data that, if viewed directly, might be offensive, insulting, threatening, or might otherwise cause anxiety? If so, please describe why Yes – many of these videos exist on YouTube; we discuss this more in Section 5.

Does the dataset identify any subpopulations (e.g., by age, gender)? Not explicitly (e.g. through labels)

Is it possible to identify individuals (i.e., one or more natural persons), either directly or indirectly (i.e., in combination with other data) from the dataset? Yes, our data includes celebrities, or other YouTube-famous people. All of the videos that we use are of publicly available data, following the Terms of Service that users agreed to when uploading to YouTube.