MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound

Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, Yejin Choi

Introduction

The world around us is dynamic. We experience and learn from it using all of our senses, reasoning over them temporally through multimodal script knowledge . Consider Figure 1, which depicts someone cooking popcorn. From the images and dialogue alone, we might be able to imagine what sounds of the scene are: the process might begin with raw kernels scattering in an empty, metallic pot, and end with the dynamic ‘pops’ of popcorn expanding, along with the jiggling of a metal around the stove.

Predicting this sound is an instance of learning from reentry: where time-locked correlations enable one modality to educate others. Reentry has been hypothesized by developmental psychologists to be crucial for how we as humans learn visual and world knowledge, much of it without need for an explicit teacher . Yet, we ask – can we build machines that likewise learn vision, language, and sound together? And can this paradigm enable learning neural script knowledge, that transfers to language-and-vision tasks, even those without sound?

In this work, we study these questions, and find that the answers are ‘yes.’ We introduce a new model that learns self-supervised representations of videos, through all their modalities (audio, subtitles, vision). We dub our model MERLOT Reserve Short for Multimodal Event Representation Learning Over Time, with Re-entrant Supervision of Events., henceforth Reserve for short. Our model differs from past work that learns from audio-image pairs , from subtitled videos , or from static images with literal descriptions . Instead, we learn joint representations from all modalities of a video, using each modality to teach others. We do this at scale, training on over 20 million YouTube videos.

We introduce a new contrastive masked span learning objective to learn script knowledge across modalities. It generalizes and outperforms a variety of previously proposed approaches (e.g. ), while enabling audio to be used as signal. The idea is outlined in Figure 1: the model must figure out which span of text (or audio) was MASKed out of a video sequence. We combine our objective with a second contrastive learning approach, tailored to learning visual recognition from scratch: the model must also match each video frame to a contextualized representation of the video’s transcript . Through ablations, we show that our framework enables rapid pretraining of a model and readily scales to ‘large’ transformer sizes (of 644M parameters).

Experimental results show that Reserve learns powerful representations, useful even for tasks posed over only a few of the studied modalities. For example, when finetuned on Visual Commonsense Reasoning (a vision+language task with no audio), it sets a new state-of-the-art, outperforming models trained on supervised image-caption pairs by over 5%. It does even better on video tasks: fine-tuning without audio, it outperforms prior work on TVQA by a margin of over 7% (and given TVQA audio, performance increases even further). Finally, audio enables 91.1% accuracy on Kinetics-600 . These performance improvements do not come at the expense of efficiency: our largest model uses one-fifths the FLOPs of a VisualBERT.

Reserve also performs well in zero-shot settings. We evaluate on four diverse benchmarks: Situated Reasoning (STAR) , EPIC-Kitchens , LSMDC-FiB , and MSR-VTT QA . These benchmarks require visual reasoning with respective emphasis on temporality, future prediction, and both social and physical understanding. With no fine-tuning or supervision, our model obtains competitive performance on each. Of note, it nearly doubles ’s SoTA zero-shot accuracy on MSR-VTT QA, and it outperforms supervised approaches (like ClipBERT ) on STAR.

Finally, we investigate why, and on which training instances audio-powered multimodal pretraining particularly helps. For instance, predicting audio rewards models for recognizing dynamic state changes (like cooked popcorn) and human communication dynamics (what are people’s emotions and towards whom). Our model progressively learns these phenomena as pretraining progresses. These signals are often orthogonal to what snippets of text provide, which motivates learning from both modalities.

In summary, our key contributions are the following:

Reserve, a model for multimodal script knowledge, fusing vision, audio, and text.

A new contrastive span matching objective, enabling our model to learn from text and audio self-supervision.

Experiments, ablations, and analysis, that demonstrate strong multimodal video representations.

Overall, the results suggest that learning representations from all modalities – in a time-locked, reentrant manner – is a promising direction, and one that has significant space for future work. We release code and model checkpoints at rowanzellers.com/merlotreserve.

Related Work

Our work brings together two active lines of research.

Joint representations of multiple modalities. Many language-and-vision tasks benefit from early fusion of the modalities . A family of ‘VisualBERT’ models have been proposed for this: typically, these use a supervised object detector image encoder backbone, and pretrain on images paired with literal captions . Cross-modal interactions are learned in part through a masked language modeling (mask LM) objective , where subwords are replaced with ‘MASK’, and models independently predict each subword conditioned on both images and unmasked tokens.Recent papers propose extensions, like generating masked-out spans or text , but it is unclear whether they can outperform the VisualBerts on vision-language tasks like VCR . Another extension involves learning from text-to-speech audio in a captioning setting – yet this lacks key supervision for environmental sounds and emotive speech.

Perhaps closest to our work is MERLOT , which learns a joint vision-text model from web videos with automatic speech recognition (ASR). Through a combination of objectives (including a variant of mask LM), MERLOT established strong results on a variety of video QA benchmarks when finetuned. However, it lacks audio: it is limited to representing (and learning from) video frames paired with subtitles. Our proposed Reserve, which represents and learns from audio, outperforms MERLOT.

Co-supervision between modalities. A common pitfall when training a joint multimodal model is that complex inter-modal interactions can be ignored during learning, in favor of simpler intra-modal interactions . For example, when using the aforementioned mask LM objective, models can ignore visual input completely in favor of text-text interactions ; this issue is magnified when training on videos with noisy ASR text .

A line of recent work thus learns independent modality-specific encoders, using objectives that cannot be shortcutted with simple intra-modal patterns. Models like CLIP learn image classification by matching images with their captions, contrastively . Recent work has explored this paradigm for matching video frames with their transcripts , with their audio signal , or both ; these works likewise perform well on single-modality tasks like audio classification and activity recognition. These independent encoders can be combined through late fusion , yet late fusion is strictly less expressive than our proposed joint encoding (early fusion) approach.

Our work combines both lines of research. We learn a model for jointly representing videos, through all their modalities, and train it using a new learning objective that enables co-supervision between modalities.

Model: Reserve

In this section, we present Reserve, including: our model architecture (3.1), new pretraining objectives (3.2), and pretraining video dataset (3.3). At a high level, Reserve represents a video by fusing its constituent modalities (vision, audio, and text from transcribed speech) together, and over time. These representations enable both finetuned and zero-shot downstream applications.

More formally, we split a video V\mathcal{V} into a sequence of non-overlapping segments in time {st}\{\boldsymbol{s}_{t}\}. Each segment has:

A frame vt\boldsymbol{v}_{t}, from the middle of the segment,

The ASR tokens wt\boldsymbol{w}_{t} spoken during the segment,

The audio at\boldsymbol{a}_{t} of the segment.

Segments default to 5 seconds in length; we discuss details of how we split videos into segments in Appendix C.

As the text wt\boldsymbol{w}_{t} was automatically transcribed by a model given audio at\boldsymbol{a}_{t}, it is reasonable to assume that it contains strictly less information content.Despite being derived from the audio, pretraining with text is still paramount: 1) in §3.2 we discuss how jointly modeling audio+text prevents models from shortcutting pretraining objectives via surface correlations; 2) in §4.2 we show that incorporating both transcripts and audio during fine-tuning improves performance; and 3) a textual interface to the model is required for downstream vision+language with textual inputs. Thus, for each segment st\boldsymbol{s}_{t}, we provide models with exactly one of text or audio. We will further mask out portions of the text and audio during pretraining, to challenge models to recover what is missing.

An overview of Reserve is shown in Figure 2. We first pre-encode each modality independently (using a Transformer or images/audio; a BPE embedding table for text). We then learn a joint encoder to fuse all representations, together and over time.

Image encoder. We use a Vision Transformer (ViT; ) to encode each frame independently. We use a patch size of 16 and apply a 2x2 query-key-value attention pool after the Transformer, converting an image of size H×WH{\times}W into a H/32×W/32H/32{\times}W/32 feature map of dimension dhd_{h}.

Audio encoder. We split the audio in each segment at\boldsymbol{a}_{t} into three equal-sized subsegments, for compatibility with the lengths at which we mask text (Appendix C). We use an Audio Spectrogram Transformer to encode each subsegment independently . The three feature maps are concatenated; the result is of size 18×dh18{\times}d_{h} for every 5 seconds of audio.

Joint encoder. Finally, we jointly encode all modalities (over all input video segments) using a bidirectional Transformer. We use a linear projection of the final layer’s hidden states for all objectives (e.g. w^t\widehat{\mathbf{w}}_{t} and a^t\widehat{\mathbf{a}}_{t}).

Independently-encoded targets. We will supervise the joint encoder by simultaneously learning independently-encoded ‘target’ representations for each modality. Doing this is straightforward for the image and audio encoders: we add a CLS to their respective inputs, and extract the final hidden state vt\mathbf{v}_{t} or at\mathbf{a}_{t} at that position. For text, we learn a separate bidirectional Transformer span encoder, which computes targets wt\mathbf{w}_{t} from a CLS and embedded tokens of a candidate text span. This enables zero-shot prediction (4.4).

Architecture sizes. We consider two model sizes in this work, which we pretrain from random initialization:

Reserve-B, with a hidden size of 768, a 12-layer ViT-B/16 image encoder, and a 12-layer joint encoder.

Reserve-L, with a hidden size of 1024, a 24-layer ViT-L/16 image encoder, and a 24-layer joint encoder.

We always use a 12-layer audio encoder, and a 4-layer text span encoder. Details are in Appendix B.

2 Contrastive Span Training

We introduce contrastive span training, which enables learning across and between the three modalities. As shown in Figure 3, the model is given a sequence of video segments. For each one, we include the video frame, and then three ‘subsegments’ that are each either text or audio. The subdivided audio segments are encoded independently by the Audio Encoder, before being fused by the Joint Encoder. We train by replacing 25% of these text and audio subsegments with a special MASK token. The model must match the representation atop the MASK only with an independent encoding of its span.

Our approach combines past success at matching images to their captions along with ‘VisualBERT’-style prediction of independent tokens – though, crucially, we predict representations at a higher-level semantic unit than individual tokens. Our approach also enables the model to learn from both audio and text, while discouraging memorization of raw perceptual input, or tokens – which can harm representation quality .

Formally, we minimize the cross entropy between the MASKed prediction w^t\hat{\mathbf{w}}_{t} and its corresponding phrase representation wt\mathbf{w}_{t}, versus others in the batch W\mathcal{W}:

We first L2-normalize w\mathbf{w} and w^\hat{\mathbf{w}}, and scale their dot product with a parameter σ\sigma .Following past work, we optimize σ\sigma and clip it at 100, which enables the model to ‘warm-up’ its emphasis placed on hard negatives . We then add this to its transposed version Ltext→mask\mathcal{L}_{\textrm{text}\rightarrow\textrm{mask}}, giving us our text-based loss Ltext\mathcal{L}_{\textrm{text}}. Analogously, we define Laudio\mathcal{L}_{\textrm{audio}} for audio, between the MASKed prediction a^t\hat{\mathbf{a}}_{t} and its target at\mathbf{a}_{t}, versus others a\mathbf{a} in the batch.

In addition to these masked text and audio objectives, we simultaneously train the model to match video frames with a contextualized encoding of the transcript.In MERLOT , this objective was found to be critical for learning visual recognition from self-supervised videos. Here, the joint encoder encodes the entire video’s transcript at once, extracting a single hidden representation per segment v^t\hat{\mathbf{v}}_{t}. We use the same contrastive setup as Equation 1 to maximize the similarity of these vectors with the corresponding vt\mathbf{v}_{t} vectors from the frames, giving us a symmetric frame-based loss Lframe\mathcal{L}_{\textrm{frame}}. The final loss is the sum of the component losses:

Avoiding shortcut learning. Early on, we observed that training a model to predict a perceptual modality (like audio or vision) given input from the same modality, led to shortcut learning – a low training loss, but poor representations. We hypothesize that this setup encourages models to learn imperceptible features, like the exact model of the microphone, or the chromatic aberration of the camera lens . We avoid this, while still using audio as a target, by simultaneously training on two kinds of masked videos:

Audio only as target. We provide only video frames and subtitles. The model produces representations of both audio and text that fill in MASKed blanks.

Audio as input. We provide the model video frames, and subtitles or audio at each segment. Because the model is given audio as an input somewhere, the model only produces representations for MASKed text.

Another issue is that YouTube’s captions are not perfectly time-aligned with the underlying audio. During our initial exploration, models took ready advantage of this shortcut: for instance, predicting an audio span based on what adjacent (overlapping) words sound like. We introduce a masking algorithm to resolve this; details in Appendix C.

Pretraining setup. We train on TPU v3-512 accelerators; training takes 5 days for Reserve-B, and 16 days for Reserve-L. We made pretraining more efficient through several algorithmic and implementation improvements. Of note, we simultaneously train on written (web) text, which enables more text candidates to be used. We use a batch size of 1024 videos, each with N=16N{=}16 segments (split into two groups of 8 segments each). We use AdamW to minimize Equation 2. More details and hyperparameters are in Appendix B.

3 Pretraining Dataset

Recent prior work on static images that demonstrates empirical improvements by increasing dataset size – all the way up to JFT-3B . The same pattern emerges in videos: prior work that has shown promising empirical improvements not only by scaling to 6 million videos/180M frames , but also by collecting a diverse set (i.e., going beyond instructional videos ).

To this end, we introduce a new training dataset of 20 million English-subtitled YouTube videos, and 1 billion frames, called YT-Temporal-1B. At the same time, we take steps to protect user privacy, directing scraping towards public, large, and monetized channels. We detail our collection, preprocessing, and release strategy in Appendix E.

Experiments

In this section, we present model ablations (4.1.1), and show that a finetuned Reserve obtains state-of-the-art results on VCR (4.1.2), TVQA (4.2), and Kinetics-600 (4.3). We then show that our model has strong zero-shot capability, over four challenging zero-shot tasks (4.2).

We evaluate Reserve first through finetuning on VCR . Most competitive models for VCR are pretrained exclusively on images paired with captions, often with supervised visual representations (e.g. from an object detector). To the best of our knowledge, the only exception is MERLOT , which uses YouTube video frames and text as part of pretraining; no VCR model to date was pretrained on audio.

VCR Task. A model is given an image from a movie, and a question. The model must choose the correct answer given four multiple choice options (Q→AQ{\rightarrow}A); it then is given four rationales justifying the answer, and it must choose the correct one (QA→RQA{\rightarrow}R). The results are combined with a Q→ARQ{\rightarrow}AR metric, where a model must choose the right answer and then the right rationale, to get the question ‘correct.’

Finetuning approach. We follow ’s approach: ‘drawing on’ VCR’s detection tags onto the image, and jointly finetuning on Q→AQ{\rightarrow}A and QA→RQA{\rightarrow}R. For both subproblems, we learn by scoring each Q→AQ{\rightarrow}A (or QA→RQA{\rightarrow}R) option independently. We pool a hidden representation from a MASK inserted after the text, and pass this through a newly-initialized linear layer to extract a logit, which we optimize through cross-entropy (details in Appendix D.1.1.)

While we present our final, state-of-the-art VCR performance in 4.1.2, we first use the corpus for an ablation study. We use the same architecture and data throughout, allowing apples-to-apples comparison between modeling decisions. We start with a similar configuration to MERLOT and show that contrastive span training improves further, particularly when we add audio.

Contrastive Span helps for Vision+Text modeling. We start by comparing pretraining objectives for learning from YouTube ASR and video alone:

Mask LM. This objective trains a bidirectional model by having it independently predict masked-out tokens. We make this baseline as strong as possible by using SpanBERT-style masking , where text spans are masked out (identical to our contrastive spans). Each span w\boldsymbol{w} is replaced by a MASK token, and we predict each of its subwords wiw_{i} independently.Like , we concatenate the MASK’s hidden state with a position embedding for index ii, pass the result through a two-layer MLP, and use tied embedding weights to predict wiw_{i}.

VirTex . In this objective, we likewise mask text subsegments and extract their hidden states. The difference is that we sequentially predict tokens wi∈ww_{i}\in\boldsymbol{w}, using a left-to-right language model (LM) with the same architecture details as our proposed span encoder.

Results are in Table 1. Versus these approaches, our contrastive span objective boosts performance by over 2%, after one epoch of pretraining only on vision and text. We hypothesize that its faster learning is caused by encouraging models to learn concept-level span representations; this might not happen when predicting tokens individually .

Audio pretraining helps, even for the audio-less VCR:

Audio as target. Here, the model is only given video frames and ASR text as input. In addition to performing contrastive-span pretraining over the missing text spans, it does the same over the (held-out) audio span (Equation 2. This boosts VCR accuracy by 0.7%.

Audio as input and target. The model does the above (for video+text input sequences), and simultaneously is given video+text+audio sequences, wherein it must predict missing text. This boosts accuracy by 1% in total.

Sans strict localization. We evaluate the importance of our strict localization in time. Here, in addition to correct subsegments at the true position tt as a correct match, we count adjacent MASKed out regions as well. An extreme version of this was proposed by , where a positive match can be of any two frames in a video. Yet even in our conservative implementation, performance drops slightly, suggesting localization helps.

Putting these all together, we find that contrastive span pretraining outperforms mask LM, with improved performance when audio is used both as input and target. For our flagship model, we report results in Table 1 on simultaneously training on web-text sequences as well (Appendix C.4), this improves performance by an additional 1%.

1.2 VCR Results

Encouraged by these results, we train our models for 10 epochs on YT-Temporal-1B. Figure 6 demonstrates that finetuned VCR performance tracks with the number of pretraining epochs, as well as the validation loss.The plot suggests that if we pretrained longer, VCR performance might continue to increase, though a confounding factor might be the learning-rate schedule. With access to compute beyond our current capacity, future work would be well-suited to consider this and other pre-training modifications.

Finally, in Table 6, we compare Reserve against the largest published models from the VCR leaderboard. Of note, Reserve-L outperforms all prior work, by over 5% on Q→\rightarrowAR metric. It outperforms even large ensembles (e.g. 15 ERNIE-Large’s) submitted by industry , though we do not show these on this table to focus on only single models.

Efficiency. The accuracy increase of Reserve is not simply due to compute.Here, we use FLOPs as our key efficiency metric, as they are a critical bottleneck in model scaling . On the other hand, we argue that parameter count can be misleading – for instance, many Transformer parameters can be tied together with minimal performance loss . In fact, our Reserve-L requires one-fifth the FLOPs of detector-based systems, like UNITER-Large (Appendix B.3). Moreover, because Reserve-L uses a pure ViT backbone versus MERLOT’s ViT-ResNet hybrid, it uses fewer FLOPs than MERLOT, while scoring 7% higher. Meanwhile, Reserve-B outperforms ‘base’ detector-based models, while using less than one-tenth their FLOPs.

In terms of parameter count, Reserve-B is comparable to prior work. On VCR, including the vision stack, Reserve-B has 200M finetunable parameters and performs similarly to the 378M parameter UNITER-Large. Reserve-L has 644M parameters.

2 Finetuning on TVQA

Next, we use TVQA to evaluate our model’s capacity to transfer to multimodal video understanding tasks. In TVQA, models are given a video, a question, and five answer choices. The scenes come from American TV shows, and depict characters interacting with each other through dialogue – which past work represents through subtitles.

Audio-Subtitle Finetuning. To evaluate how much audio can help for TVQA, we finetune Reserve jointly between the ‘Subtitles’ and ‘Audio’ settings. Like on VCR, we consider one sequence per candidate: each contains video frame features, the question, the answer candidate, and a MASK token (from where we pool a hidden representation). During training, each sequence is duplicated: we provide one sequence with subtitles from the video, and for the other, we use audio. This lets us train a single model, and then test how it will do given subtitles, given audio, or given both (by averaging the two softmax predictions).

Results. We show TVQA results in Table 6. With subtitles and video frames alone, our Reserve-B outperforms all prior work by over 3%. Combining subtitle-only and audio-only predictions performs even better, improving over 4% versus the prior state-of-the-art, MERLOT (and in turn over other models). The same pattern holds (with additional performance gains) as model size increases: Reserve-L improves over prior work by 7.6%.

3 Finetuning on Kinetics-600 Activity Recognition

Next, we use Kinetics-600 to compare our model’s (finetuned) activity understanding versus prior work, including many top-scoring models that do not integrate audio. The task is to classify a 10-second video clip as one of 600 categories. We finetune Reserve jointly over two settings: vision only, and vision+audio.

Results. We show Kinetics-600 results on the validation set, in Table 8. Reserve improves by 1.7% when it can jointly represent the video’s frames with its sound. This enables it to outperform other large models, including VATT which learns to represent audio independently from vision (and so cannot early-fuse them), along with the larger MTV-Huge model by 1.5%.

4 Zero-Shot Experiments

Next, we show that our model exhibits strong zero-shot performance for a variety of downstream tasks. Our zero-shot interface is enabled by our contrastive span objective. For QA tasks that require predicting an option from a label space of short phrases, we encode this label space as vectors, and predict the closest phrase to a MASKed input. We consider:

Situated Reasoning (STAR) . This task requires the model to reason over short situations in videos, covering four axes: interaction, sequence, prediction, and feasibility. The model is given a video, a templated question, and 4 answer choices. We convert templated questions into literal statements (which are more similar to YouTube dialogue); the label space is the set of four options.

Action Anticipation in Epic Kitchens . Here, the goal is to predict future actions given a video clip, which requires reasoning temporally over an actor’s motivations and intentions. The dataset has a long tail of rare action combinations, making zero-shot inference challenging (since we do not assume access to this prior). As such, prior work trains on the provided in-domain training set. To adapt Reserve to this task, we provide it a single MASK token as text input, and use as our label space of all combinations of verbs and nouns in the vocabulary (e.g. ‘cook apple, cook avocado’, etc.).

LSMDC . Models are given a video clip, along with a video description (with a MASK to be filled in). We compare it with the vocabulary used in prior work .

MSR-VTT QA . This is an open-ended video QA task about what is literally happening in a web video. We use GPT3 , prompted with a dozen (unlabelled) questions, to reword the questions into statements with MASKs. This introduces some errors, but minimizes domain shift. We use a label space of the top 1k options.

For these tasks, we use N=8N{=}8 video segments (dilating time when appropriate), and provide audio input when possible. Details and prompts are in Appendix D. We compare against both finetuned and zeroshot models, including running CLIP on all tasks. CLIP is a strong model for zero-shot classification, particularly when encyclopedic knowledge about images is helpful; our comparisons showcase where multimodal script knowledge helps.

Results. Table 8 shows our model performs competitively:

On STAR, it obtains state-of-the-art results, with performance gain when audio is included. Interestingly, Reserve-B outperforms its larger variant; we hypothesize that this is due to limited prompt searching around question templates. We qualitatively observed that Reserve-L sometimes excludes topically correct options if they sound grammatically strange (to it).

On EPIC-Kitchens, our model obtains strong results at correctly anticipating the verb and noun - despite the heavy-tailed nature of both distributions. It is worse on getting both right (‘action’), we suspect that this might be due to priors (motifs) between noun and verb . These are easy to learn given access to training data, but we exclude these as we consider the zero-shot task.

On LSMDC, our model obtains strong results at filling-in-the-blank, likewise despite a heavy (unseen) frequency bias. Notably, it outperforms CLIP significantly, with CLIP often preferring templates that use visually-relevant words, even if they don’t make sense as a whole. For instance, given a clip of a mailman, CLIP chooses ‘the mailman smiles off,’ versus ‘the mailman takes off.’

Finally, our model performs well on MSR-VTT QA, outperforming past work that directly rewords subtitled instructional videos into video QA instances .

Qualitative Analysis: Why does audio help?

What can Reserve learn from both text and audio? Three validation set examples are shown in Figure 9. The model is given the displayed text and video frames, and must match the MASK to the correct missing text and audio span (out of 48k total in the batch). The plots show Reserve-B’s probability of correctly identifying the correct audio or text span, as it progresses through 10 epochs of pretraining.

Audio’s supervisory signal. In the first two rows of Figure 9, audio provides orthogonal supervision to text:

In the first row, the MASKed audio contains the sound of popcorn pops slowing. By the final epoch, Reserve-B selects this specific auditory cue with 60% probability, over others (including from adjacent segments, at different stages of popping). Here, sound provides signal for joint vision-text understanding of the situation, as evidenced by its greater match probability.

The second row contains only the text ‘why,’ with the audio providing greatly more information — a female-presenting speaker (shown in the next frame) laughs, astonished that the child (in the frame afterwards) might want a better relationship with their parents.

In the third row, matching performance is similar between modalities, possibly as the yogi is narrating over a (muted) video recording, and not adding much information.

Role of text. Text is still a crucial complement to audio, in terms of the supervision it provides. Consider the second row: Reserve-B learns to match the audio almost perfectly (perhaps reasoning that the speaker is shown in the next frame, and is laughing). In later epochs, its text-match probability increases: knowing that a ‘why’ question is likely to be asked is a valid social inference to make about this (tense) situation.

Learning through multimodal reentry. Developmental psychologists have hypothesized that human children learn by reentry: learning connections between all senses as they interact with the world . Using a held-out modality (like audio) might support learning a better world representation (from e.g. vision and text), by forcing models to abstract away from raw perceptual input. Our work suggests that reentry has potential for machines as well.

Conclusion, Limitations, Broader Impact

We introduced Reserve, which learns jointly through sound, language, and vision, guided through a new pretraining objective. Our model performs well in both finetuned and zero-shot settings, yet it has limitations. Our model only learns from 40-second long videos; relies on ASR models for subtitles, and can only match (not generate) text and audio.

Still, we foresee broad possible societal impact of this line of work. Video-pretrained models might someday assist low vision or d/Deaf users . Yet, the same technology can have impacts that we authors consider to be negative, including surveillance, or applications that hegemonize social biases. We discuss these further in Appendix A: key dimensions include respecting user privacy during dataset collection, exploring biases in YouTube data, dual use, and energy consumption. We discuss our plan to release our model and data for research use so others can critically study this approach to learning script knowledge.

Acknowledgements

We thank the anonymous reviewers, as well as Jae Sung Park, Oren Etzioni, Gabriel Ilharco, and Mitchell Wortsman for feedback on this work. Thanks also to Zak Stone and the Google Cloud TPU team for providing access to the TPU machines used for conducting experiments. Thanks to James Bradbury and Skye Wanderman-Milne for help with JAX on TPUs. Thanks to the AI2 ReVIZ team, including Jon Borchardt and M Kusold, for help with the demo. This work was funded by DARPA MCS program through NIWC Pacific (N66001-19-2-4031), and the Allen Institute for AI. Last, but not least, thanks to the YouTubers whose work and creativity helps machines to learn about the multimodal world.

References

Appendix A Broader Impact Statement

In this paper, we have presented a model for learning multimodal neural script knowledge, through incorporation of audio as a first-class citizen alongside text and video frames. We argue that academic study of this learning paradigm is important, in part because it relates to how we as humans understand the world. We as humans process situations by perceiving through multiple modalities and interpreting the result holistically.

At the same time, the work and methodology that we outlined risks dual use. Like other large machine learning systems pretrained on web data, our system may reproduce harmful social biases present in its training data. While a variety of past work has studied risks of language-only pretraining , the video-centric pretraining that we explore in our work might have different benefits and risks. We discuss these below, along with how we worked to mitigate them through our work.

A significant risk with training on data at YouTube scale is protecting user privacy. We took several proactive steps to ensure this, that in turn build off prior work and community norms :

We release only the video IDs for download, following prior work . Thus, if a user deletes a video off of YouTube, it becomes removed from YT-Temporal-1B as well, giving content creators a right to opt out of all uses of their videos.

Building off of past work , we directed our data collection towards public and monetized channels. These channels are identifiable insofar as they contain more subscribers, and more videos. They include companies that have official accounts, including journalism outlets like the New York Times and Vox. They also include individuals for whom making public YouTube videos is their full time job. In either case, our use videos in question for research purposes can be seen as fair use.

Framing of privacy. Privacy is a nuanced topic with many societally, culturally, and generationally-specific interpretations. We took inspiration from Marwick and Boyd ’s framework of networked privacy, which posits that users posting public videos might encode private information – enough so that their intended viewership (friends, possibly) can catch the gist, but not enough so as to leak private details like phone numbers to the world.

Through the lens of networked privacy, we see key differences between studying videos on a moderated platform, versus NLP work that trains models from the open web (e.g. ). When YouTube users upload videos, they tend to understand details of its privacy policy, beyond consenting to it . Likewise, YouTubers typically upload their own videos ; the platform deters users from re-posting other users’ content. These factors differ from text on the open web. Today, ‘data brokers’ post private details (like phone numbers) to the web for profit ; concerningly, a study on language models suggests that models are vulnerable at memorizing this private information .

It is worth examining our research through other framings of privacy as well. For example, internet platforms profit off of user data, whereas users do not share equally in these profits . For this, and for the other reasons mentioned, we aim to release our model only for research-based use.

Inspired by work studying language model memorization of private information , we wish to empirically probe Reserve’s ability to recognize individuals. Our goal during model development was not to optimize for this ability. Instead, our goal was to study models for multimodal script knowledge (what people might be doing in a situation over time, and why) instead of long-tailed visual recognition (including who those individuals are). These goals might trade off – for instance, our training data only has individuals’ names when they are mentioned in ASR subtitles, a pairing that might be significantly noisier than images and text on the open web.

We study this capacity on the VoxCeleb2 and VGGFace2 datasets , where we created a test set of 120 celebrities, with 100 samples of each. We study these datasets not to promote them, but to establish a conservative upper-bound for the capacity of a model to recognize non-celebrities. We hypothesize that if Reserve struggles to select the right celebrity out of 120 predefined options, it would struggle much more at identifying random people (where the set of candidate names is much greater). We test this hypothesis over three zero-shot settings:

Voice to name. Given an audio clip sampled for a celebrity, we encode it with our model’s audio encoder. We provide our model’s joint encoder the text ‘the sound of MASK’, followed by the encoded audio. A blank image is provided. We extract the representation on top of the MASK, and choose the most similar celebrity name.

Image+voice to name. Here, we adopt the same format as ‘Audio to name,’ except we additionally encode an image of the celebrity’s face in question.

Image to name. Here, Reserve encodes an image of the celebrity in question, and we provide it with text ‘A picture of MASK.’ No audio is provided. Using our model’s joint encoder, we select the closest encoded celebrity name, out of all options.

We use this format to compare to a CLIP model, which was trained on web images with captions . For the CLIP comparison, we use it to encode each image, and for all considered celebrity names, the sentence ‘A picture of ${name}’. We choose the closest encoded sentence to the encoded image.

We show our results in Table 2. In all modes, our model is less than 11% accurate at recognizing celebrities. Curiously, the accuracy drops given both the image and the voice, suggesting that the way we fused a celebrity’s image and voice together might be outside the model’s training distribution. These results are significantly lower than CLIP’s 86% accuracy at classifying a person from their image.

In Figure 10, we investigate more into which celebrities our model is best at recognizing. Only a few celebrities are reliably classified; these tend to be very famous celebrities like Oprah Winfrey and Justin Bieber. Several sports players are recognized well (including Lebron James and Roger Federer), which could imply that our model learned their identities from watching sports replays or commentary. Most other celebrities are hardly recognized, whereas CLIP does well across the board.

Results summary. Together, these results show that while models like CLIP focus on encyclopedic knowledge that results in strong zero-shot person recongition accuracy, Reserve is not as effective as other models in memorizing particular celebrities– and, thus, perhaps not as effective as memorizing particular non-celebrities. These results suggest that Reserve’s objectives and data might make it less of a concern to release privacy-wise, versus models trained on web images with captions.

As the rest of the paper emphasizes however, Reserve performs well on tasks with temporal understanding and commonsense reasoning as the primary goal. On a broader level, these results suggest that it is possible to learn strong models about temporal reasoning without person-level memorization, though more work is needed.

A.2 Biases in (pre)training data.

The ‘knowledge’ that our model learns should be viewed as situated within YouTube , which has numerous biases (that we will discuss next). Past work has made similar observations for language model pretraining on the open web. One of the root causes of such bias is learning objectives that encourage memorization of surface level cooccurences, rather than truly causal factors . Though it is possible that in the very long term, a paradigm of grounded learning might resolve some of these issues, the objectives in this work still likely reify biases that exist in the YouTube data.

Platform biases. Unlike many other pretraining efforts, that scrape data from the open internet (e.g. ) which directly leads to toxic biases (e.g. ); we trained our model on YouTube, which is a moderated platform . Though the content moderation might perhaps reduce overtly ‘toxic’ content, social media platforms like YouTube still contain harmful microagressions , and alt-lite to alt-right content . Additionally, it should be mentioned that the content moderation on YouTube disproportionately filters out minoritized voices . Thus, despite us not using any word-based ‘blocklist,’ our model’s pretraining data is still biased . Even without videos being explicitly removed, the ‘YouTube algorithm’ incentivizes the production of certain types of content over others ; e.g. people’s roles in YouTube videos tend to be highly gendered , which might bias situation understanding .

Bias amplification. In this work, we pretrained a model primarily on ASR text, which is itself produced by another model. The automatic captions in YouTube are known to suffer from gender bias , which our model (like neural models generally) might in turn amplify . The transcriptions on YouTube are also likely poor at handling important identity markers, like pronouns. Already, text-only models like BERT struggle with pronouns like they/them and zi/zir; our reliance on ASR text makes our corpus likely worse in this regard . While past work, namely MERLOT , ‘cleaned’ this ASR text – through another large language model – we opted not to do so for this work due to computational expense. Though in that work, the ASR-denoisification was found to boost performance in VCR, it seems unlikely that it would solve this core issue of model bias.

A.3 Dual use.

Learning connections between video, audio, and text – though an important area of study as we have argued – can be used for undesirable applications, beyond what we have outlined under ‘biases.’ We outline and discuss a few below.

Generating fake content. A concern for pretrained models is that they can generate fake content, that could be used by ‘bad’ actors for their ends . It should be noted that our model cannot explicitly ‘generate’ text, audio, or vision in a direct sense. Nonetheless, however, it is possible that a finetuned or expanded version of this model could be used for that purpose – and that our model would be more helpful to such an actor versus them training their own (perhaps larger) model from scratch.

Surveillance. Our model might contain representations that enable it to be used in surveillance applications. As we note in Appendix A.1.1, our model’s low performance on person recognition suggests that it might perform poorly recognition-focused applications. Still, one possibility is that a neural script knowledge could ‘summarize’ surveillance videos in some form (like identifying an activity of interest), without identifying the person(s).

We suspect (but cannot definitively prove) that the reporting bias of the YouTube data that it was trained on might make it poor for such a surveillance-focused task . Namely, most surveillance videos are sparse in nature – finding an activity of interest is like finding a ‘needle in a haystack’ . Though, some surveillance videos are inevitably posted on YouTube and then captioned, these disproportionately contain interesting events (like somebody’s car crashing into a house). It is not clear whether our system could be easily adapted to such a sparse problem; the amount of work required suggests that it might be out-of-scope at least for low-skill actors. On the other hand, this broad research agenda, and perhaps all of computer vision for that matter, might enable large actors to do just that ; which might not be addressable through purely technical solutions .

Harmful outcomes if deployed. Beyond the biases that our system possesses, some applications of our system – if deployed in production – could cause harm, particularly to groups already harmed by AI systems. Of note, linking someone’s voice with their appearance is not always a good thing . Likely some of the key features that our model learns – though we did not teach it this explicitly – involve recognizing gender, and this is harmful especially to transgender individuals .

A.4 Energy consumption.

Our model cost a lot amount of energy to pretrain ; roughly 3 weeks of time on a TPU v3-512. The total carbon footprint of our work was a net 8.23 tons of CO2 equivalent, which is roughly 4.5% of the emissions of a jet plane flying round-trip from San Francisco to New York.CO2 Calculation. It is also important to consider the location where these TPUs are located, as the renewables portion at each datacenter is not equal . Our TPUs were in the ‘europe-west4’ region, which uses on average 60% carbon-free energy, and a Grid Carbon intensity of 0.410 kgCO2eq / kWh. A single TPU v3 processor (with 8 cores over 2 chips) has a power average of 283 W, so after performing the math from , our training cost 20,000 kWh. This gives us a net 8.23 tons of CO2 equivalent. It should be mentioned that this figure only covers the electricity usage given the chips (and the datacenter), not the raw materials involved in making these chips (which is significant ).

At the same time, it is possible that our model could save energy overall, when shared with researchers who build off of our system. Indeed, Reserve-B uses less energy than MERLOT (due to a smaller vision backbone, and smaller image sizes), MERLOT in turn is more efficient than past work which used expensive detector-based backbones (e.g. ), that are made more expensive because some of their computational primitives (like non-maximum suppression) are difficult to make efficient on-device.

A.5 Synthesis.

With these risks in mind, we release our video IDs, as well as Reserve’s checkpoints, exclusively for research use. We believe that at this point in time, we as a field lack full knowledge of the privacy, bias, and dual-use risks of video-based models – though, we hope that our analysis in this section provides a good starting point. For instance, while the objectives that we have studied were designed to promote learning general neural script knowledge above encyclopedic memorization, they have not yet been tested in all possible cases. By opening our models to the research community, we hope to promote fundamental work in uncovering both promising aspects of these systems, alongside examining their risks. We hope to contribute to these lines of research as well.

Appendix B Model implementation details

In this section, we discuss at a more in-depth, technical level, how we implement certain aspects of Reserve, and other details (like its runtime in FLOPs). We discuss our use of rotary position encodings (B.1), how we set the sequence lengths for the model (B.2), measure the model’s computational footprint (B.3), list hyperparameters (B.4 ), and discuss several training strategies (B.5.

We use a rotary position encoding to model the relative location of input sequences . We chose this primarily because we did not want to use absolute (additive) position embeddings, which would have to be added to the inputs of each encoder, and possibly at multiple levels in the hierarchy (e.g. for the joint encoder, the video segment index tt would be needed as well).

The rotary encoding uses no parameters, and instead uses a kernel trick to allow the model to recover relative distances between key and query elements in a Transformer’s attention head. This can be seen as ‘rotating’ pairs of elements; we apply the rotation to only the first half of each 64-dimensional head, and the second half is kept as is.

Video Frame Encoder (ViT): just the h,wh,w coordinates of the image; so (h,w,0,0)(h,w,0,0).

B.2 Sequence lengths

We briefly remark on the sequence lengths used by parts of the model.

Video Frame Encoder (ViT): Most YouTube videos are widescreen (16x9). We thus used a widescreen resolution for our video frame encoder. It takes in patches of size 16x16, and we used a layout of 12 patches (in height) by 20 patches (in width). This corresponds to 192x320. Among other factors that are important are ensuring that TPUs do not execessively pad the sequence length . The sequence length is 241 in this case, as there is a CLS token, and it gets padded to 256.

Attention pooling. As we note in the main text, afterwards we apply attention pooling in a 2x2 grid (ignoring the CLS token here). Similar to Transformer-style query,key,value attention , the query is the average of the vectors in the 2x2 grid; the keys and values are learned projections of the vectors. This gives us a H/32H/32 by W/32W/32 grid for the joint encoder (6 x 10).

Audio Encoder. Our model independently encodes each 1.6 second of audio (a segment has three such ‘subsegments’). We do this through spectrograms. Each window involves 1536 samples at a sample rate of 22500 Hz, and there are 588 samples ‘hops’ between windows. We chose these hyperparameters largely around efficiency. We found that the Discrete Fourier Transform is fastest if the window size is close to a multiple of 2. We used a small number of mel spectrogram bins (64) because we found that at that threshold, we could reconstruct the original sequence at an acceptable level using the Griffin-Lim algorithm, which itself might be a lower bound on quality as neural methods trained for this purpose have been shown to do better .

In our implementation, we compute the spectrogram for an entire video segment (5 seconds) at once; this is of size 6464 mel bins by 192 windows. During pretraining, we perform what is effectively a ‘random crop’ over the spectrogram: we extract three sequential 64x6064x60 sub-spectrograms, for each audio subsegment. We constrain them to not overlap, which means that 12 (random) windows are held out.

We note that our Audio Encoder AST is quite different from the one proposed by . Though it operates over spectrograms, we opted for a linear ‘1-dimensional’ layout rather than a two-dimensional (image-like) one. We also did not pretrain our audio encoder on any supervised data (they used ImageNet and found, perhaps surprisingly, that it helped initialize the model). We used a patch size of 6464 mel bins by 22 windows; the resulting (1D) sequence is of size 30. After adding a CLS token, the result is a sequence of length 31.

As we note in the main text, we apply attention pooling afterwards (for all elements except the CLS token), pooling by a factor of five to resize the length-30 sequence to a length of 6 ‘audio tokens.’

Text Span Encoder: We operate on spans that are at most of length 15, with an additional CLS token. Its length is thus 16.

Joint encoder. Let LL be the number of text or pooled audio tokens given to the model per segment, on average; we set L=20L{=}20. Let TT be the number of video segments. Then, the joint model’s sequence length is T×(L+W/32×H/32)T\times(L+W/32{\times}H/32). We set T=8T{=}8 (8 video segments given to the model at a time) and used a H=192H{=}192 by W=320W{=}320 resolution. Our total sequence length was thus 640.

To better adapt our model to downstream tasks – particularly single-image tasks like VCR , where past work tends to use a resolution much higher than 192x320, after pretraining, we performed FixRes pretraining (for one epoch on Reserve-B, and one half epoch on Reserve-L .We had intended to do a full epoch for Reserve-L, but our job got preempted, and the loss seemed to have already converged. Here, we trained the model on larger images – simultaneously on 288x512 widescreen images (18 patches by 32 patches), and on 384x384 square images (24 patches on each side). The joint encoder, correspondingly, uses a sequence length of 1312.

During 10 epochs of pretraining, we used a cosine decay of the learning rate down to 0.02 its maximum. During FixRes pretraining afterwards, we warmed up the learning rate to 0.02x its peak, over the first 1/5th of an epoch, and afterwards used a cosine schedule to anneal it towards 0.

B.3 Efficiency metrics of our model

In Table 3, we report efficiency metrics of Reserve, versus others. We calculate these metrics in the context of scoring a single VCR question and answer candidate. This requires encoding one image, and using 128 tokens for each question and answer combined (for all models). We compare against a UNITER , which is a representative VisualBERT style model, along with MERLOT . Our models are far more efficient in terms of FLOPs, with Reserve-L being roughly on par with MERLOT, yet outperforming it by 6% in terms of VCR accuracy. We discuss key differences below:

UNITER. We note that UNITER, like other VisualBERT models, uses a supervised object detection backbone . This processes images using a ResNet 101 model , at a resolution of 600x800; the final ResNet ‘C4’ block is applied densely over the entire image to obtain object-detection potentials everywhere in the image. Both factors greatly increase the FLOPs count.

When computing UNITER’s FLOPs count, we exclude operations like non-max suppression, which is an operation that is difficult to implement (and thus whose FLOP count might vary significantly depending on implementation). Our FLOPs count is thus a lower-bound. 36 detection regions are extracted, which is why the ‘joint encoder’ for UNITER is smaller than the equivalents for MERLOT and Reserve.

MERLOT. This model has two key differences versus our Reserve. First, it uses a larger image resolution for VCR: 384x704, versus our 288x512. Second, it uses a hybrid ViT-ResNet50 backbone for encoding images. The backbone here is lighter weight than the object detection backbone of UNITER (in particular, the final ‘C4’ block is removed), and thus, as shown in Table 3, though it uses more FLOPs than does our Reserve-L, it uses far fewer FLOPs than UNITER.

We choose flops as our primary comparison metric as past work shows that it is one of the key factors in model scaling . Parameters are arguably more fungible. For instance, in text-only representation learning, ALBERT demonstrates that it is possible to tie parameters together at all layers of a BERT-like transformer, reducing parameters by an order of magnitude (while not modifying compute), with a minimal performance drop. We did not do this for this work, as we wanted to use a more ‘vanilla’ Transformer architecture; however, it suggests that representation learning models with hundreds of millions of parameters might be FLOPs bound as opposed to parameter-bound.

Nonetheless, UNITER-Base has 154 million parameters, though some are frozen (86 million from their Transformer, 23 million from the word embedding layer, and then 44 million from their object detector ). UNITER-Large has 378 million parameters (303 from their Transformer, 31 million from word embeddings, and 44 million from the same object detector. Meanwhile, MERLOT has 223M parameters. Versus our Reserve-B, 14 million extra parameters are due to a larger vocabulary, and 10 million parameters are due to a ResNet50 encoder – but these parameters have a disproportionate impact in FLOPs count.

B.4 Full model hyperparameters

In Table 4, we present full hyperparameters for our model. Among other details, we used AdamW as our optimizer, with β2=0.98\beta_{2}=0.98 and ϵ=1e−6\epsilon=1e-6. We increased the learning rate linearly to its peak value (4e-4 for Reserve-B, 3e-4 for Reserve-L) over 3750 steps (120\frac{1}{20}th of an epoch). Our number of warmup steps is lower than many other pretraining work; we note that all of our contrastive objectives involve learning a σ\sigma parameter, which functions as a secondary ‘warmup.’

We did not use gradient clipping. We trained and evaluated in 16-bit bfloat16 precision wherever we could – casting all gradients to that precision as well, and saving the AdamW running mean and variance to be 16-bit as well. A few times during pretraining Reserve-L, we found that some values in gradients would be NaN. We addressed this by always setting NaN values to be 0. This seemed to address the symptoms of training instability – though sometimes the training loss would spike to roughly around the same loss as random initialization, it always converged back to slightly better than it was before the spike. We are not currently sure why this happens.

B.5 Speed improvements during pretraining

We made several high-level algorithmic and engineering implementations to our implementation, which made pretraining run faster, and that we discuss here.

Duplicated video copies. As mentioned in the main text, we create two copies per each video – allowing us to learn separately how to handle audio as an input as well as how to learn from audio. We chose this in part because copying a video does not increase the total compute requried by a factor of two. Instead:

We use the image and audio encoders, to encode the underlying video frames and audio clips only once (for the two video copies), and then duplicate the encodings; this is far more efficient than encoding them both separately from scratch.

For the two video copies, we sampled two disjoint sets of masks (for which audio and text subsegments are replaced with MASK) at a 25% rate. This increases the pool of negative samples for contrastive learning, again increasing training efficiency.

Reducing memory usage. The memory usage of our Transformer implementation scales quadratically with sequence length, which could pose a problem since we operate on sequences of videos. We split the video into two groups of 8 segments, and encode each group separately by the joint encoder.

Vectorization. We vectorize all joint transformer inputs together into a single call. During this vectorization, we also encode the transcript (for the transcript-frame matching objective).

We note that this vectorization is incompatible with the Mask LM variant proposed by MERLOT . In this variant, which the authors called ‘attention masking,’ two transformer calls must happen sequentially – first, a language only encoder must encode the inputs and mark down (what is presumably) visually-grounded tokens; second, these tokens are masked for the joint encoder. We found that such an objective was unnecessary when pretraining under our contrastive span approach, which in turn enabled more efficient pretraining.

We discuss the exact pretraining data formatting technique that we used in the next section.

Appendix C Pretraining Data Formatting: alignment and masking

In this section, we discuss how we turn a video V\mathcal{V} into a (masked) list of segments {st}\{\boldsymbol{s}_{t}\} for pretraining.

Recall that each segment contains a video frame vt\boldsymbol{v}_{t}, ASR tokens wt\boldsymbol{w}_{t}, and audio at\boldsymbol{a}_{t}. We generate the list of segments by iterating through the video with a 5-second sliding window.Sometimes there are long ‘pauses’ in videos where nothing gets said. When this happens – if two segments in a row have fewer than 8 BPE tokens – we merge them 90% of the time, in effect ‘fast-forwarding’ the audio and still extracting a frame from the middle. We do this at most twice, so the total length is at most 15 seconds here (in effect, a ‘playback’ rate of 1x, 2x, or 3x). In roughly 90% of cases, the segments are 5 seconds of length.

Audio and text subsegments for masking. We want audio to be used in part as a target for contrastive prediction. However, during early exploration we found that 5 seconds of audio could correspond to many BPE tokens; roughly 15 on average. We use past work in language modeling as a guide and wanted an average span length of around 5 tokens. To get this, we split each audio segment into three equal subsegments, each with a duration of 1.66 seconds. We can then perform masked language modeling at the aligned subsegment level, where we mask out the text corresponding to an audio subsegment, and have the model (contrastively) predict the masked-out span of text, as well as the corresponding span of audio. We use a masking rate of 25%, which means that a quarter of the subsegments will be corrupted and replaced by a MASK token.

In theory, splitting the videos into (masked) segments ought to be straightforward. However, the key challenge that we ran into is that the YouTube caption timing information is unreliable. Problems might arise when we perform pretraining with both audio and text, on misaligned data. Suppose the model is given audio in segment st−1\boldsymbol{s}_{t-1} that ends with somebody saying the word ‘pasta.’ If the alignment between audio and text is off, the model might be able to cheat the desired task by simply predicting the word ‘pasta’ for segment st\boldsymbol{s}_{t} – thereby turning the challenging masked-prediction task into an easier speech recognition task; we discuss this in more detail in Appendix C.1.

One way of addressing the timing issue would be to run our own ASR model over all videos, but we chose not to do this due to computational expense. Instead, we adopted two complementary strategies. First, we trained a lighweight regressor to refine the timing information (C.2); second, we mask audio and text conservatively, to minimize alignment errors (C.3). Finally, we discuss how we combine everything efficiently (in a vectorized way) in C.4.

YouTube provides automatically generated captions for accessibility purposes, which include timing information on each word. In the subtitle encoding that we used (vtt), each word ww contains a single timestamp tt which corresponds to when the word should flash on-screen. The timings are mostly accurate, but we found two key issues:

First, they show up on average roughly 0.1 seconds before each word is spoken, which we suspect might be for usability purposes (perhaps so that while the viewer is reading the caption, they hear the word).

Second, with a single timestamp tt for each word, it is difficult to infer about pauses. For each word ww, we can use its timestamp tt, and the timestamps of adjacent words, to loosely infer an interval [ts′,te′][t_{s}^{\prime},t_{e}^{\prime}] around when the word is said. However, the interval is not tight. We can only infer that the word is being actively spoken for some subinterval [ts,te][t_{s},t_{e}] such that ts′≤ts≤te≤te′t_{s}^{\prime}\leq t_{s}\leq t_{e}\leq t_{e}^{\prime}.Note that this is compounded with the first problem, the ground truth interval [ts,te][t_{s},t_{e}] might not be fully contained in the provided interval [ts′,te′][t_{s}^{\prime},t_{e}^{\prime}] due to ‘captions being shown before audio’, the error here is typically small though (0.1 seconds).

This can lead to high absolute error (in terms of a difference between timesteps), when pauses occur. For example, suppose a speaker says a word, and then pauses. The interval given by the subtitles, [ts′,te′][t_{s}^{\prime},t_{e}^{\prime}], might be rather large (possibly a few seconds), even though the actual word was spoken for a fraction of that time.

C.2 Refining timing information

We trained a simple multilayer perceptron regressor to correct the timing information of YouTube transcripts. For data, we used 2000 videos with transcripts from YT-Temporal-180M, and also used Google Cloud’s (highest quality, paid) ASR service to transcribe them. After aligning the words for these transcripts, this gave us tuples of the YouTube ASR word w\boldsymbol{w}, its provided interval [ts′,te′][t_{s}^{\prime},t_{e}^{\prime}], and the ‘ground truth’ interval [ts,te][t_{s},t_{e}].When matching YouTube ASR to Google Cloud’s ASR, we skipped words without an ’exact-match’ alignment, as well as words that were over 0.25 seconds apart (i.e., where either δs>0.25\delta_{s}>0.25 or δe>0.25\delta_{e}>0.25 Our modeling objective was then to predict the desired offsets with respect to the provided interval: δs=ts−ts′\delta_{s}=t_{s}-t_{s}^{\prime} and δe=te−te′\delta_{e}=t_{e}-t_{e}^{\prime}. We took a feature based approach.

For each input (w,ts′,te′)(w,t_{s}^{\prime},t_{e}^{\prime}), we used as features:

the number of punctuation characters in ww,

the value of te′−ts′t_{e}^{\prime}-t_{s}^{\prime}.

We provided these features as input to the model, as well as the corresponding features for the next word, and the previous word. We z-normalized all features and used a two-layer multilayer perceptron, with a hidden size of 32 and RELU activations. We used a tanh activation at the end to bound the regression. The final predictions for δs\delta_{s} (analogously for δe\delta_{e}) were then given by the following equation:

where h is the hidden state, and with learnable parameters cc, w\mathbf{w}, b1b_{1}, and b2b_{2}. The learned bounds mean that, no matter what the input, the model will never predict an offset of above c+b2c+b_{2} (of which it learned for both parameters c≈0.2c\approx 0.2 and b2≈0.11b_{2}\approx 0.11, so the offsets can never be above 0.3 seconds). We trained our lightweight regression model using an L1 loss, and used it to correct the timing on all of the transcripts.

C.3 Handling worst-case scenarios in masking, when alignment isn’t perfect

The regressor that we described reduces the average timing error of a transcript, as a preprocessing step, but it is not perfect. Thankfully, however, we find that most of the remaining alignment errors are single words that are slightly misaligned. For instance, for three words wt,wt+1,wt+2w_{t},w_{t+1},w_{t+2}, the audio corresponding to the time interval around wtw_{t} might contain sound from wt+1w_{t+1} being spoken, but rarely wt+2w_{t+2}. We suspect this is primarily due to the difficulty inferring pauses: by definition, no other word can be said in a pause, so the errors are local.

We present a high level approach for masking audio and text, that in turn addresses these alignment issues (making it difficult for models to cheat). A diagram is in Figure 11.

Recall that in our framework, we only either go from ‘vision and text →\rightarrow text and audio’ (VT→\rightarrowTA), or, ‘vision, text, and audio →\rightarrow text’ (VTA→\rightarrowT). One of the reasons we did this is to avoid allowing a model to cheat by performing speaker identification (or even ‘microphone identification’), which might be feasible if audio was given to the joint model as input. We can handle the two cases separately:

Vision and text →\rightarrow text and audio (VT→\rightarrowTA). Here, the text as input (to the joint encoder) might overlap with the audio we are trying to predict. Our solution here is thus to donate nearby tokens from the predicted span, to the input. Let the span that we are trying to predict (and that we will ‘mask out’) have a start time of tst_{s} and an ending time of tet_{e}. If the final token in the previous text span, if any, has a timestamp of greater than ts−0.125t_{s}{-}0.125, we move it to the predicted span; likewise, if the first token in the next text span has a timestamp of less than te+0.125t_{e}{+}0.125, we move it to the predicted span as well.

Vision, text, and audio →\rightarrow text (VTA→\rightarrowT). In this prediction task, models are given information from all modalities as input, and must predict masked-out text spans. Note that models are only given a single ‘speech’ modality – either text, or audio – at each timestep. What this means is that we can carefully choose which input subsegments to turn into ‘audio subsegments,’ and which to turn into ‘text subsegments.’ Our strategy is, given a masked out subsegment, to turn 80% of adjacent subsegments into ‘text subsegments.’

We give an illustration of this in Figure 11, part 2. Here the word ‘focus’ is part of a4,1\boldsymbol{a_{4,1}} but also w3,3)\boldsymbol{w_{3,3}}). This might make w3,3)\boldsymbol{w_{3,3}}) overly easy to predict, if we gave the model a4,1\boldsymbol{a_{4,1}} as input. Our solution is thus to give the model text from w3,2)\boldsymbol{w_{3,2}}) and from w4,1)\boldsymbol{w_{4,1}}) as input; we are guaranteed that there is no misalignment overlap here between input and prediction spans. All of the other subsegments (not adjacent to one of the 25% that we mask out) will be provided as audio.

C.4 Putting it all together, along with web text

Finally, we discuss how we combine the various masking approaches into the prediction tasks outlined in the main text.

Each video has N=16N=16 video segments, and three subsegments of audio or text spans per segment. We consider two sub-problems for this video sequence:

in VT→\rightarrowTA, vision and text are provided as input, and the model must predict masked-out text and audio. These are done on top of separately-encoded MASK tokens and MASKAUDIO tokens, to enable the model to learn different predictions for each modality over two separate transformer ‘columns.’

In VTA→\rightarrowT, vision, text and audio are provided as input, and models must predict masked-out text. Here, we use the term ‘predict’ as a shorthand for our contrastive objective – in which a model must match a context (a jointly-encoded MASK) to the exact missing span in question, where many negative contexts and spans are provided.

We use a masking rate of 25% for audio and text subsegments, and there are 3 subsegments per segment. This means that a single video instance gives us 48×0.25=1248\times 0.25{=}12 masked-out spans of text, for each of VT→\rightarrowTA and VTA→\rightarrowT, so 24 in total (as we use disjoint masked-out subsegments). Likewise, it gives us 12 masked-out spans of audio. If we scaled these to the whole batch of 1024 videos, we would have 12k audio span options and 24k text span options. This might suffice, but scaling up the pool of candidates boosts performance in a contrastive setting, as suggested from prior work (e.g. ), and as our ablations (Table 1) support as well. Thus, we do the following:

Text candidates. We scale up the text candidates by simultaneously training the model on web text, from The Pile . The joint encoder – which can handle pooled video, pooled audio, and BPE-encoded text – is simultaneously given a sequence of web text, for each video that we have. By performing the span-contrastive objective with this piece of web text as well, we can not only teach the model about written (as opposed to spoken) language, but we can scale up the set of text candidates as well.

Next, we mask the fake subsegments. During pretraining, we use text sequences of length L=800L=800, but a model sequence length of only 640. Because we are masking spans and not individual tokens, the text sequences ‘shrink’ when we mask them. We extract exactly 38 masked-out spans, which corresponds to around 25% of total text.

Finally, we combine the target spans that we took from the webtext sequence, with the target spans from the video. We note that sometimes – especially in a video – text spans might be empty. Not every 1.6 second slice of a video has someone speaking. We thus try to not use these empty spans in our contrastive objective. For each video (which is paired with text for implementation reasons) we select the ‘best’ 48 text spans out of the (38+24) options – penalizing empty spans, and choosing spans from videos 4x as often.

These ‘best 48’ text spans, as well as the pooled contexts that they were paired with, will be used in the contrastive objective. Aggregating over the entire batch of 1024 videos (and 1024 web text sequences), this gives us 49152 text spans as candidates, for the all-pairs symmetric softmax between text spans and contexts.

Audio candidates. For each video, we note that we have exactly 12 pooled MASKAUDIO tokens, where the model is trying to predict the corresponding audio span. One option would be be to just use those 12 corresponding audio spans as the targets, aggregate these over the batch, and do a symmetric-cross-entropy loss.

However, we can do even better for free. Note that for the VTA→\rightarrowT direction, we might have to encode many of the audio spans anyways, using the lower level audio encoder (which simultaneously extracts a CLS representation and a sequence-level pooled representation). To simplify implementation, we encode all 48 audio spans per video. We can use these audio spans as candidates.

Thus, we do the following when computing the loss over audio prediction. We aggregate all 12288 contexts from the MASKAUDIO tokens in the batch, and we aggregate all 49152 candidate audio spans. We perform an all-pairs dot product between these two sets, and use it to compute a symmetric cross-entropy loss over both directions. We did not encounter any trouble using the same temperature for both directions (even though for one direction, there are 12288 options, and for the other, there are 49152).

The combination of these design decisions provide more ‘hard negatives’ for the model during training. We also found that they worked well to reduce wasted computation on a TPU. For each video, the joint transformer uses one L=640L=640 length sequence for transcript-frame matching, two length-LL sequences for the VT→\rightarrowTA direction (as we break it up into two groups of 8 frames each), two length LL sequences for the VTA→\rightarrowT direction, and finally one length-LL sequence of text. These sequences can all be vectorized together, and the total batch size is 6×6\times the number of videos. This is helpful because using an even-numbered batch size reduces wasted computation on a TPU.

Appendix D Downstream Task Implementation Details

In this section, we present information for how we adapted Reserve on downstream tasks.

For adapting Reserve in a finetuned setting, we take the following approach. We use a linear warmup of the learning rate over the first half of the first epoch, with a linear decay thereafter to 0. To find the learning rate, we did a small grid search generally centered around 1e-5. Our full hyperparameters are shown in Table 4.

When finetuning (and pretraining), we did not use any dropout to make implementation simpler. Instead, as a way to apply regularization, we used the same L2L_{2} penalty as in pretraining (a weight decay of 0.1), but with respect to the pretrained weights. This idea was used in among other works, and although it often tends to underperform dropout , it is simple to implement.

As mentioned in the main text, VCR considers two subtasks: Q→AQ{\rightarrow}A, where models are given a question and must choose the right answer given four options; and QA→RQA{\rightarrow}R, where models are given a question (and the right answer) and must select the right rationale.

In our setup for this task, we treat it as a four-way classification problem, extracting a single score from each answer or rationale candidate. An example Q→AQ{\rightarrow}A is:

What is going to happen next? answer: person2 is going to say how cute person4’s children are. MASK

What is going to happen next? person2 is going to say how cute person4’s children are. rationale: It looks like person4 is showing the photo to person2, and person2 will want to be polite. MASK

We extract representations from the MASK position (which are of dimension dhd_{h}), score them with a newly-initialized dh×1d_{h}\times 1 weight matrix, and optimize scores with softmax-cross entropy.

Both VCR subtasks use only a single image. We also followed past work in ‘drawing on’ the provided detection tags to the image . These are unambiguous references to entities that are then referred to in the question, answer, and rationale. For example, text might reference a ‘person1’, which corresponds to an image region. When drawing on these detection tags, we do so in a deterministic way – for example, ‘person1’ always gets the same box color. We determine the box color by hashing the object’s ID (in this case, ‘person1’) and using that to determine the hue. The model learns the connection between boxes with different hues, and the names, during finetuning.

We randomly flip images left or right, so long as there is no instance of the word ‘left’ or ‘right’ in the question, answer, or rationale candidates. We did no other data augmentation (other than randomly resizing images to between 100% to 110% of the network’s size).

D.1.2 TVQA

TVQA provides models with a video, a question, and five answer candidates; we represent this as five distinct sequences for the model to score (one per candidate). The version of TVQA that we used also gives models annotations for the time region in the video that is being referenced. It is not clear that only using this region would provide enough context to be able to understand what is going on – enough to answer correctly. Thus, for each question, we extract 35 seconds of video around the provided time region. We then provided the model with two numbers corresponding to the time region, relative to the cropped time interval. For example, if the provided timestamp annotation is [t0,t1][t_{0},t_{1}], we use the following region:

The location of [t0,t1][t_{0},t_{1}] in relative coordinates is then:

We provide models with t0rt_{0}^{r} and t1rt_{1}^{r}, multiplied by 100 and casted to an integer. Thus, an example TVQA instance might look like:

1 to 28 What is Janice Holding on to after Chandler sends Joey to his room? Chandler’s tie. MASK[subtitles or audio]

This text input corresponds to the first ‘segment’ of a video; to it we append subtitles (or audio representations) from seven segments from the provided TVQA video (with accompanying frames).

D.1.3 Kinetics-600

We evaluate Reserve on Activity Recognition over the Kinetics-600 dataset . Here, the model has to classify a short 10-second video clip into a mutually-exclusive set of 600 categories, like ‘assembling bicycle’ or ‘alligator wrestling’. We consider performing this task in a finetuned setting, so as to better compare to prior work. We format each example by extracting 4 video frames from the clip (sampled uniformly), and extracting 6 audio subsegments (totalling 10 seconds of audio). The model processes these inputs along with a MASK token, where we extract a vector representation. We initialize the 600-way classification layer with the activations of our Text Span Encoder, over the names of the 600 categories.

We finetune the model jointly over two settings: a setting where audio is provided, and a setting where no audio is provided, to allow us to investigate both settings. We tried to closely follow VATT’s finetuning approach , including their exact data augmentation settings. We used a batch size of 64 videos (that we process simultaneously ‘with audio’ and ‘without audio’). We used the same image augmentation code as VATT , and finetuned for 15 epochs. We used a learning rate of 5e-6 for Reserve-L and 1e-5 for Reserve-B.

D.2 Setup and prompting for Zero-shot tasks

Here, we discuss how we set up various tasks for Reserve in a fully zero-shot setting. In addition to evaluating Reserve, we also evaluate CLIP in the same zero-shot setting. CLIP is not pretrained on videos, and it cannot jointly encode text. For each task, we construct CLIP’s label space by taking our prompt and substituting in each possible answer option. We average together the logits over all frames, and take a softmax, giving us a distribution over the task-specific label space.

We study the task of action anticipation from the EPIC-Kitchens dataset , a large egocentric video dataset with 700 unscripted and untrimmed videos of cooking activities. In action anticipation, a model must predict a future action that comes τa\tau_{a} seconds after a given video clip. The observed segments are of arbitrary length; we follow prior work and set τa=1\tau_{a}=1.

The model tries to choose the correct noun and verb that happens next, given a list of predefined options for each. We report results on each category using the class-mean top-5 recall.

Zero-shot inference approach. We directly evaluate the pretrained Reserve on action anticipation to verify the knowledge learned during pre-training. All prior work reported on the official leaderboard use supervision from the in-domain training set, which we do not use at all .

For each action segment, we sample at most N=8N=8 image frames and their associated audio, with fixed time interval t=2.0t=2.0 preceding it and ending τa\tau_{a} seconds before the start of the action. We append a MASK token as the sole text input (at the last frame, after audio is optionally included).We were unable to find a better text based prompt than this, as we found that they often biased the model towards linguistically relevant words; however, we suspect that such a prompt does exist. We create short phrases out of all candidate nouns and verbs, and use that as our label space to simultaneously predict them both. We compute the score for each verb and noun independently by averaging their scores, over all labels for which they appear.

Results. We show the full zero-shot action anticipation results in Table 6. We also show our results on the test set here for our best performing model (Reserve-L, with audio provided). It gets competitive results on verb and noun prediction – with only 1.6% and 3.3% lower compared to the challenge winner method AVT+ , which is fully supervised and use additional object-level annotations. On Unseen Kitchen and Tail Classes, our model outperforms AVT+ on noun and verb. Overall, audio significantly improves the results – Reserve-L (+audio) outperforms Reserve-L with an average 3.0%, which suggests that it is useful for this task.

D.2.2 Zero-shot Situated Reasoning

Next, we evaluate on situated reasoning (STAR) which requires the model to capture the knowledge from surrounding situations and perform reasoning accordingly. STAR dataset includes four types of questions, including interaction, sequence, prediction, and feasibility. A model is given a video clip, a templated question, and 4 answer choices.

Zero-shot inference approach. For each video clip, we sample N=8N=8 image frames uniformly from the video, we also optionally include the video’s sound.

To reduce domain shift between YouTube data – where people don’t typically ask visual questions, and where ASR typically does not insert question marks – we convert the question-answer pair into a statement. We did so using the question-answer templates provided by the author, with the answer replaced by a MASK. For example, “Q: What did the person do with the bottle? – A: Put down.” will be converted to “The person MASK the bottle.”.

We put the converted statement into the first frame and use the four candidate answers as a unique label space (that differs from example to example). Like with EPIC-Kitchens, we also evaluate how much audio can help by masking the audio inputs.

Results. We show our zero-shot STAR results in Table 8 in the main text. Our base model outperforms all supervised prior work by 3.7%. The model with audio performs better, with average 1.1% improvement. Interestingly, Reserve-L is worse than Reserve-B, we suspect the reason is Reserve-L is sensitive to grammar details. Given the previous example, we note that while ‘Put down’ is a valid answer that might make sense both semantically and syntactically, a different answer ‘pick up’ might be flagged by some English speakers as being ungrammatical: the instantiated template would then be ‘the person pick up the bottle.’ We noticed instances of the larger model paying greater attention to these syntax-level details, even though they were not the focus of the task. It does suggest, however, that additional prompting (or label space augmentation) could resolve these issues and increase performance even further.

D.2.3 Zero-shot LSMDC

We evaluate our model on Movie Fill-in-the-Blank task, which based on descriptive audio description for the visually impaired. Given a movie clip and an aligned description with a blank in it, the task is to fill in the blank with the correct word. Following , we report prediction accuracy in test set of 30,354 examples from 10K movie clips.

Zero-shot Inference approach. We sample N=8N=8 video segments uniformly over the movie clip, and extract the audio and middle frame of each segment. We replace the ‘blank’ token in each description with a MASK token, and provide it (as text-based input) to the model at its final segment. For the other segments, we optionally provide the model with audio; for all segments, we provide the associated image frame. We use the vocabulary set in the LSMDC dataset as our label space (for what the ‘missing word’ might be).

Results. Our results are shown in Table 8 in the main text. Our model obtains 31% when audio is included, which outperforms human text-only performance (30.2 %) , predicted by human annotators. A supervised LSTM obtains 34.4% in this text-only setting which suggests that there is a certain textual bias in this task, which our model cannot learn (as it is zero-shot). This also suggests that state-of-the-art supervised models exploit patterns in this vocabulary distribution.

Without such an advantage, our model performs well, outperforming CLIP (2%) by a large margin. This suggests that jointly reasoning over both the visual situation, and the linguistic context of the provided sentence, is helpful for zero-shot performance on LSMDC fill-in-the-blank.

D.2.4 Zero-shot MSRVTTQA

Finally, we evaluate our model on MSR VTT-QA, a question-answering task over videos . We provide a model with N=8N=8 video segments sampled uniformly from the video clip, and extract an image from each one. For the first seven segments, we optionally include audio extracted from that point; at the last segment, we insert a converted version of the question, along with a MASK. We compare the similarity of that hidden state to the top 2000 most common answers, similar to past work .

Similar to STAR, we convert the questions into statements to minimize drift away from the pretraining distribution. We use GPT3 prompted with several examples for this. Our exact prompt is the following:

Input: what is a car being driven through? Output: a car is being driven through _. Input: who are running across screen? Output: _ are running across screen. Input: when is a girl performing? Output: a girl is performing at _. Input: what is a cartoon doing? Output: a cartoon is _. Input: how many women talk in a bedroom? Output: _ women talk in a bedroom. Input: what a man playing while dancing with others? Output: a man is playing _ while dancing with others. Input: where is a flag hoisted? Output: a flag is hoisted in _. Input: who talks to another man on the couch? Output: _ talks to another man on the couch. Input: what does a teenage girl try to get at a public restroom? Output: a teenage girl tries to get _ at a public restroom. Input: when do the models walk as the audience watches? Output: the models walk as the audience watches at _. Input: what shows a person killing animals in a green forest? Output: _ shows a person killing animals in a green forest. Input: who does a man ask to go on a date? Output: a man asks _ to go on a date. Input: what are three people sitting on? Output: three people are sitting on _. Input: ${question} Output:

Then, given a new question ${question}, GPT3 generates a converted output, wherein we can replace it’s underscore with a MASK. GPT3 works well at this conversion, though sometimes it generates a sentence where inserting the ‘correct answer’ feels gramatically strange. For example, the question ‘how many women talk in a bedroom?’ suggests any integer might be a reasonable answer. On the other hand, ‘_ women talk in a bedroom’ implies that ‘one’ is not a valid answer (since ‘women’ is plural). We note that the errors caused by this conversion technique are specific to English grammar, and so if such a question-conversion approach was done in other languages, there could be more (or less) errors that directly result.

Our results are shown in Table 8. Of note, our model through automatic question-conversion outperforms Just Ask , which performs an analogous (supervised-guided) question conversion on all its YouTube transcripts, before pretraining. Our model also outperforms CLIP, which cannot naturally handle dynamic situations.

Appendix E Dataset Collection

In this section, we discussed how we curated data for YT-Temporal-1B. We had several goals in mind. We wanted to use only public-facing data, which motivated our choice of YouTube as it is a public platform that users understand is public . We wanted to use this platform to examine to what extent we can learn multimodal neural script knowledge from web data alone.

Our data collection strategy in this work was informed by past work, notably MERLOT . That paper found that increasing the diversity and scale of a video corpus both allowed for better learned representations. At the same time, the data collected by MERLOT (YT-Temporal-180M) has issues. Of note, the authors’ scraping strategies – to prioritize monetized content – also led to a lot of U.S. local news being in that corpus (roughly 30% of all data). Local news might be problematic to learn from, particularly in that quantity, due to its numerous biases (e.g. racist coverage on ‘crime’ ). Our goal was to expand the dataset in both diversity and size to 20 million videos, while having less local news and without scraping private content.

High level approach. We adopt a similar dataset collection strategy as in MERLOT . In the first phase, we identify a candidate set of videos ID to download. In the second phase, we open each video ID in YouTube and apply several filtering steps that go from inexpensive to expensive. The filtering steps allow us to exit early and possibly avoid downloading the video if the video seems unsuitable for our purpose from the title, description, and captions alone.

For a Datasheet , please see the MERLOT paper .

For MERLOT’s YT-Temporal-180M, the bulk of the video IDs were identified by applying breadth-first-search on YouTube channels from HowTo100M and VLOG . Each channel often links to other channels, and given a channel it is inexpensive to obtain a list of all its videos using the youtube-dl Python package.

In this paper, we considered numerous approaches to search for diverse, visually grounded videos. We ended up using an approach where we used YouTube’s recommended videos algorithm to suggest similar videos to YT-Temporal-180M. We went through all non-news and non-sports videos YT-Temporal-180M, and opened each video up in YouTube. For each other video that YouTube recommended, we retrieved its channel ID – giving us access to not just that video, but all other videos. This approach yielded 2 million channels, with 200 million videos among them.

E.2 Filtering video IDs by channel

Given this (large) list of channels, each with many videos, we took steps to filter it further. We used the python cld3 library to remove channels whose titles might not be in English. We then finetuned, and used, a language model to identify channels likely to have visually grounded videos, which we describe next.

In more detail, we selected 2000 videos, and asked workers on Mechanical Turk to rate their level of groundedness, their genre, and whether they had explicit content or not. The questions we asked are shown in Figure 12. We annotated 2k videos under this schema, and trained a model to predict the annotations given video metadata.

For model training, we used a slightly different setting to what we gave the crowdworkers. We trained a model to predict the labels, given a formatted list of 5 video titles from the same channel. During training, we made the weak-supervision assumption that all videos from a channel have exactly the same rating (as the video we annotated). This enabled us to collect 84k examples from our 2k annotations. The model we chose was T5-base model , which generates the labels left-to-right in text form (and which we converted automatically to a structured representation).

We then used this model to identify channels that seem especially promising. For each channel with at least 5 videos, we randomly sampled 8 sets of length-5 videos, and used the finetuned T5 model to classify them. We filtered out any channel that had at least 25% of likely non-English or irrelevant-English videos, any channel that had at least 25% of slideshows, and any channel that likely had racist or sexist content.

One side benefit of this model is that it allowed us to estimate our videos’ genre breakdown before downloading them. We found 1% Gaming videos, 11% News videos, 20% How-To videos, 20% ‘chatting’ videos, 5% sports videos, 5% Music videos, 3% Movies/Drama videos, 4% Documentary videos, and 31% Miscellaneous. The Gaming videos were then filtered out.

We used the classification model to create a budget for how many videos to download from each channel; with the aim to download more videos from likely more-grounded channels. Using the answers to Q3 (from Figure 12), we gave each channel 1 point for likely having ‘a variety of objects’, 2 points for ‘a variety of actions’, and 0.5 points for ‘a variety of scenes.’ We subtracted 3 points if it was likely to be a slideshow. (Likely-racist or sexist channels were already filtered out.) We then z-normalized and softmaxed the channel scores, and used the result as the channel-level budgets. Any channel with an aggregate ‘interestingness’ score of 1 standard deviation above the mean would then have a budget of 8x larger than the mean. We clipped the channel-level budgets to include at most 500 videos per channel.

This process (finally!) gave us 30 million YouTube video IDs that were likely to be high-quality.

E.3 Filtering videos from their metadata

Last, we filtered and downloaded these videos using a filtering approach similar to . We first retrieved the video metadata and used it to filter out ‘gaming’ videos. We then retrieved the video’s transcript, and filtered out any video without a ‘dense’ span of spoken words – defined as an interval of 30 seconds where at least 50 words are spoken. Additionally, we used the Python package cld3 to filter out any transcript with a probability of less than 80% of being English. Last, we used a hidden feature in the YouTube API to download four thumbnails of the video. Using the image classification model from , we filtered out videos whose four thumbnails had an average cosine similarity of above 85%, or that contained fewer than 1 object from COCO.

Unlike , we did not use a sequence-to-sequence model to ‘translate’ spoken text to text that appears more stylistically like written English (i.e., by adding capitalization and punctuation, and removing filler words).

Appendix F Additional Experiments and Exploration

In this section, we briefly include additional experiments, showcasing our model’s performance on specific tasks that do not necessarily require multimodal script knowledge.

We evaluate Reserve on the task of zero-shot audio classification, to study to what extent its learned audio representations can directly predict text-based labels. We conduct this evaluation on environmental sounds from ESC50 , urban sounds from US8K , and (as part of the privacy-minded exploration in Appendix A) celebrity voices from VoxCeleb2 .

We consider the format where we encode an audio input into a CLS level representation, and retrieve the most-similar label given a set of encoded options. We encode the audio input with our encoder, which takes in as input audio clips of length at most 1.6 seconds. For shorter audio clips (like many sounds in ESC50), we repeat them in time until their length is at least 1.6 seconds. For longer audio clips, we encode multiple CLS representations and then average the resulting vectors.

We consider the following ways to encode the labels:

Text-only. Inspired by the prompt ‘a photo of’, which is used in CLIP’s zero-shot image classification task , we give Reserve’s joint encoder a blank image, with associated tokens the sound of ${label}. We do this once for each label, giving us a single ‘target’ vector for each possible label in the dataset.

Image-only. Inspired by YouTube videos of sound effectsFor instance, youtu.be/VmgKryu4__k., we created image-only prompts that suggest a sound (of the target class) is playing in the background. An example is shown in Figure 13. We encode each image with our joint encoder, and do this once for each label.

We note that for VoxCeleb2, we use face images of celebrities rather than this image-based prompt, due to our interest in exploring whether models can perform person-level recognition due to the privacy issue (Appendix A.1.1).

Image and text. Here, we combine both of the above options: encoding one input for each label, using both the image and text prompt.

For each prompt, we append the token ‘MASKAUDIO’ and extract the hidden state from there, as our final representation for that label.

We present our results in Table 7. The results show, possibly surprisingly, that Reserve can perform optical character recognition over image prompts like Figure 13 – given just the image, its accuracy on ESC50 is higher than given just text. Its accuracy on ESC50 and US8K improves further when given both an image and text.

These results are slightly different for VoxCeleb2, which emphasizes long-tail recognition of people – something that might be more encyclopedic than semantic, and that we did not wish to optimize in this work. There, when given an image of a celebrity’s face, it demonstrates some capacity at linking it with one of their audio clips – a capacity that decreases if prompted with additional text. We suspect that this is due to interpreting the given text as spoken, for example, Justin Bieber himself saying ‘the sound of Justin Bieber.’ On all celebrities, Reserve struggles versus recognition-focused models like CLIP (Appendix A.1.1).

Overall, our model displays strong audio understanding ability. In comparison, AudioCLIP (which is supervised on human-annotated labels from AudioSet ), performs 16% higher on ESC50, and 6.4% higher on US8K.

F.2 Additional Qualitative Analysis

In Figure 14, we include an additional figure of examples, of the same format as Figure 9. The examples are chosen randomly – not by how much Reserve improved at retrieving their audio or text spans over the course of training.