Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval

Max Bain, Arsha Nagrani, Gül Varol, Andrew Zisserman

Introduction

Joint visual-text models have become increasingly popular as they enable a wide suite of downstream tasks, including text-to-visual retrieval , visual captioning , and visual question and answering . Their rapid development is due to the usual improvements on three fronts: new neural network architectures (e.g. transformers for both text and visual inputs); new large-scale datasets; and new loss functions that are, for example, able to handle label noise . However, their development mostly proceeds on two independent tracks: one for images, with its own architectures, training datasets and benchmarks ; and the other for videos with a similar separation of training datasets and benchmarks . The only common link between the two is that often video networks are initialized by pre-training image networks on image datasets . This separation of effort is suboptimal given the overlap in information that images and video convey over multiple tasks. For example, although classifying some human actions requires the temporal ordering of video frames, many actions can be classified from just their distribution over frames or even from a single frame .

In this paper we take a step towards unifying these two tracks, by proposing a dual encoder architecture which utilises the flexibility of a transformer visual encoder to train from images-with-captions, from video clips-with-captions, or from both (Fig. 1). We do this by treating images as a special case of videos that are ‘frozen in time’. Using a transformer-based architecture allows us to train with variable-length sequences, treating an image as if it was a single frame video, unlike in standard 3D CNNs where to train on images jointly with videos one must incur the cost of actually generating a static video. Furthermore, unlike many recent methods for video-text dual encoding, we do not use a set of ‘expert networks’ that are pre-trained on external image datasets and then fixed, but instead train the model end-to-end.

This end-to-end training is facilitated by scraping the web for a new large-scale video-text captioning dataset of over two million video alt-text pairs (WebVid-2M). We also take advantage of large-scale image captioning datasets such as Conceptual Captions .

We make the following contributions: (i) we propose a new end-to-end model for video retrieval that does not rely on ‘expert’ features, but instead, inspired by employs a transformer architecture with a modified divided space-time attention applied directly to pixels; (ii) because our architecture can gracefully handle inputs of different lengths, it is versatile and can be flexibly trained on both video and image datasets (by treating images as a single-frame video). We build on this flexibility by designing a curriculum learning schedule that begins with images and then gradually learns to attend to increasing temporal context when trained on video datasets through temporal embedding interpolation. We show that this increases efficiency, allowing us to train models with far less GPU time; (iii) we introduce a new dataset called WebVid-2M, consisting of 2.5M video-text pairs scraped from the web; and finally (iv) we achieve state-of-the-art performance by only using the video modality on MSR-VTT , MSVD , DiDeMo and LSMDC – outperforming works that use pre-extracted experts from multiple modalities, as well as those that are pretrained on the noisy HowTo100M, which is 20x larger than our dataset in the number of video-text pairs.

Related Works

Pretraining for video-text retrieval. Given that most video-text retrieval datasets tend to be small-scale, the dominant paradigm for video retrieval has been to use a combination of pre-extracted features from ‘expert’ models, including models trained for various diverse tasks and on multiple modalities such as face, scene and object recognition, action classification and sound classification. MoEE , CE , MMT and concurrent work HiT all follow this paradigm, with the overall similarity for a video-text pair obtained as a weighted sum of each expert’s similarity with the text.

However, since the release of the HowTo100M dataset , a large-scale instructional video dataset, there has been a flurry of works leveraging large-scale pretraining to improve video-text representations for tasks such as video question-answering , text-video retrieval and video captioning . Although semantically rich and diverse, text supervision from instructional videos is extremely noisy, and hence incurs a large computational cost, as scale is required for competitive results. A few approaches have been proposed to combat the noise – e.g. using loss functions such as MIL-NCE or using the raw audio directly to increase robustness. Given the large size of existing image-captioning datasets, some have naturally tried to overcome the lack of video-caption training data with joint image-text pretraining (such as in MoEE and ClipBERT ). MoEE trains on images jointly by feeding in zeros to all expert streams that require videos, such as the motion and audio features, while ClipBERT restricts their feature extractors to 2D CNNs. Instead we propose an elegant transformer-based encoder that works well with either images or videos and can be trained effectively on both.

Similar to our work, although only suitable for images is CLIP , which learns an effective joint image-text representation from millions of text-image pairs scraped from the internet using contrastive loss.

End-to-end video representation learning. A large number of architectural developments have been driven by action recognition on datasets such as Kinetics where manual labelling has been relatively easier than obtaining textual descriptions for datasets. For a long time this space was dominated by spatio-temporal CNNs such as I3D , 3D ResNets , S3D or ‘R(2+1)D’ CNNs . Here, images are used simply to initialise video models, through inflation . Multigrid scheduling has been proposed for efficient training .

Transformers for vision. A number of works use self-attention for images, either in combination with convolutions or even replacing them entirely.

Works that use only self-attention blocks tend to apply them at an individual pixel level , often requiring tricks to ensure computational tractability, including restricting the scope of self-attention to a local neighbourhood , adding global self-attention on heavily downsized versions, or sparse key-value sampling . To increase efficiency, ViT decompose images into a sequence of patches and then feeds linear embeddings of these patches as inputs to a transformer, effectively adding a single convolutional layer to the image at the start. This idea has been extended in DeiT . For video, previous works also employ self-attention blocks together with CNN layers, for action recognition and video classification .

In contrast, our architecture consists entirely of self-attention units and is heavily inspired by ViT and particularly the Timesformer , which uses divided space and time attention. Unlike these works, we use expandable temporal embeddings to allow flexible training of variable-length videos and images both jointly and separately. We are unaware of any previous works that use self-attention to train on both images and videos in the same model.

Method

In this section, we describe our transformer-based spatio-temporal model architecture (Section 3.1), and our training strategy (Section 3.2).

Spatio-temporal patches. Following the protocol in ViT and Timesformer , the input video clip is divided into M×NM\times N non-overlapping spatio-temporal patches of size P×PP\times P, where N=HW/P2N=HW/P^{2}.

such that all patches within a given frame mm (but different spatial locations) are given the same temporal positional embedding EmtE^{t}_{m}, and all patches in the same spatial location (but different frames) are given the same spatial positional embedding EpsE^{s}_{p}. Thus enabling the model to ascertain the temporal and spatial position of patches.

In addition, a learned [CLS] token is concatenated to the beginning of the sequence, which is used to produce the final visual embedding output embedding of the transformer.

Space-time self-attention blocks. The video sequence is fed into a stack of space-time transformer blocks. We make a minor modification to the Divided Space-Time attention introduced by , by replacing the residual connection between the block input and the temporal attention output with a residual connection between the block input and the spatial attention output (see Section C.3 of the Appendix for details). Each block sequentially performs temporal self-attention and then spatial self-attention on the output of previous block. The video clip embedding is obtained from the [CLS] token of the final block.

Text encoding. The text encoder architecture is a multi-layer bidirectional transformer encoder, which has shown great success in natural language processing tasks . For the final text encoding, we use the [CLS] token output of the final layer.

Projection to common text-video space. Both text and video encodings are projected to a common dimension via single linear layers. We compute the simliarity between text and video by performing the dot product between the two projected embeddings.

Efficiency. Our model has independent dual encoder pathways (such as in MIL-NCE and MMV networks ), requiring only the dot product between the video and text embeddings. This ensures retrieval inference is of trivial cost since it is indexable, i.e. it allows application of fast approximate nearest neighbour search, and is scalable to very large scale retrieval at inference time. Given tt text queries and vv videos in a target gallery, our retrieval complexity is O(t+v)O(t+v). In contrast, ClipBERT which inputs both text and video as input to a single encoder, has retrieval complexity O(tv)O(tv) since every text-video combination must be inputted to the model. Other expert-based retrieval methods such as MoEE , CE and MMT also contain a dual encoder pathway, however they still require query-conditioned weights to compute the similarity scores for each expert, while our model does not.

2 Training Strategy

Loss. We employ in a retrieval setting, where matching text-video pairs in the batch are treated as positives, and all other pairwise combinations in the batch are treated as negatives. We minimise the sum of two losses, video-to-text and text-to-video:

where xix_{i} and yjy_{j} are the normalized embeddings of ii-th video and the jj-th text respectively in a batch of size BB and σ\sigma is the temperature.

Joint image-video training. In this work, we train jointly on both image-text pairs as well as video-text pairs, taking advantage of both for larger-scale pretraining. Our joint training strategy involves alternating batches between the image and video datasets. Since the attention mechanism scales with the square of input frames O(M2)O(M^{2}), the alternate batch training allows the image batches (M=1M=1) to be far greater in size.

Weight initialisation and pretraining. Following , we initialise the spatial attention weights in the space-time transformer model with ViT weights trained on ImageNet-21k, and initialise the temporal attention weights to zero. The residual connections mean that under these initialisation settings, the model is at first equivalent to ViT over each input frame – thereby allowing the model to learn to attend to time gradually as training progresses. Since transformer architectures have demonstrated most of their success from large-scale pretraining, we utilise two large-scale text-image/video datasets with a joint training strategy, resulting in large improvements in performance.

Temporal curriculum learning. The space-time transformer architecture allows a variable length input sequence and therefore a variable number of input video frames. If the model has only trained on videos up to length mm however, then the temporal positional embedding Et\boldsymbol{E}^{t} will only be learned up to E:mt\boldsymbol{E}^{t}_{:m}. Therefore, applying the model to input video of sequences up to length MM will result the addition of Em:Mt\boldsymbol{E}^{t}_{m:M}, which would not yet be learned.

Two temporal expansion methods are investigated: interpolation and zero-padding. Zeros can be filled in, 0→Em:Mt\boldsymbol{0}\rightarrow\boldsymbol{E}^{t}_{m:M}, allowing the model to learn the additional temporal positions from scratch during training. Alternatively, interpolation could be used to upsample the temporal embeddings in the temporal dimension, E:mt→E:Mt\boldsymbol{E}^{t}_{:m}\rightarrow\boldsymbol{E}^{t}_{:M}. We investigate two methods of interpolation: nearest neighbour and bilinear. The effects of these different initialisations can be found in the Appendix, Section C.4.

We employ this expansion strategy in order to perform curriculum learning in the number of input frames. Initially training on fewer frames has drastic savings in computation, whilst having comparable or even better performance (see Section 4.5).

Frame sampling. Given a video containing LL frames, we subdivide it into MM equal segments where MM is the desired number of frames for the video encoder. During training, we sample a single frame uniformly from each segment (in a similar manner to TSN and GST ). At test time, we sample the ithi^{th} frame in every segment, to get a video embedding viv_{i}. The values for ii are determine using a stride SS, resulting in an array of video embeddings v=[v0,vS,v2S,vM]\boldsymbol{v}=[v_{0},v_{S},v_{2S},v_{M}]. The mean of these video embeddings is used as the final embedding for the video.

Experiments

We first describe the pretraining datasets including our WebVid-2M video-text dataset (Section 4.1), followed by the downstream datasets used for the evaluations in our experiments (Section 4.2). We then describe implementation details of our model (Section 4.3). Next, we ablate various training components on the MSR-VTT dataset, in particular the effects of pretraining and our space-time attention modification (Section 4.4), and our proposed curriculum strategy (Section 4.5). Then, we compare to the state of the art on four benchmarks: MSR-VTT, MSVD, DiDeMo and LSMDC (Section 4.6).

We jointly pretrain our model on image and video data.

Video pretraining: The WebVid-2M Dataset. We scrape the web for a new dataset of videos with textual description annotations, called WebVid-2M. Our dataset consists of 2.5M video-text pairs, which is an order of magnitude larger than existing video captioning datasets (see Table 1).

The data was scraped from the web following a similar procedure to Google Conceptual Captions (CC3M). We note that more than 10% of CC3M images are in fact thumbnails from videos, which motivates us to use such video sources to scrape a total of 2.5M text-video pairs. The use of data collected for this study is authorised via the Intellectual Property Office’s Exceptions to Copyright for Non-Commercial Research and Private Studywww.gov.uk/guidance/exceptions-to-copyright/. We are currently performing further analysis of the dataset on its diversity and fairness.

Figure 2 provides sample video-caption pairs. There are a variety of different styles used in caption creation, as can be seen from Figure 2 (left to right) where the first video has a longer, poetic description compared to the succinct description for the second video. The third video caption has a less defined sentence structure, with keywords appended to the end, while the fourth video mentions a specific place (maldives). Time-specific information is important for the second and third example, where details such as “talking on walkie-talkie” or “playing billiards” would be missed when looking at certain frames independently.

We note that our video dataset is 10x smaller than HowTo100M in video duration and over 20x smaller in the number of paired clip-captions (Table 1). Our dataset consists of manually generated captions, that are for the most part well formed sentences. In contrast, HowTo100M is generated from continuous narration with incomplete sentences that lack punctuation. The clip-text pairs are obtained from subtitles and may not be temporally aligned with the video they refer to, or indeed may not refer to the video at all . Our captions, on the other hand, are aligned with the video and describe visual content.

Moreover, there is no noise from imperfect ASR transcription and grammatical errors as is the case for HowTo100M. Our dataset also has longer captions on average (12 vs 4 words for HowTo) which are more diverse (Measure of Textual Lexical Diversity, MTLD = 203 vs 13.5).

Image pretraining: Google Conceptual Captions . This dataset consists of about 3.3M image and description pairs. Unlike the curated style of COCO images, Conceptual Captions (CC3M) images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. The raw descriptions are harvested from the Alt-text HTML attribute associated with web images.

2 Downstream Datasets

We now describe the downstream text-video datasets that our model is evaluated on. MSR-VTT contains 10K YouTube videos with 200K descriptions. Following other works , we train on 9K train+val videos and report results on the 1K-A test set. MSVD consists of 80K English descriptions for 1,970 videos from YouTube, with each video containing 40 sentences each. We use the standard split of 1200, 100, and 670 videos for training, validation, and testing .

DiDeMo contains 10K Flickr videos annotated with 40K sentences. Following , we evaluate paragraph-to-video retrieval, where all sentence descriptions for a video are concatenated into a single query. Since this dataset comes with localisation annotations (ground truth proposals), we report results with ground truth proposals (where only the localised moments in the video are concatenated and used in the retrieval set as done by ) as well as without (as done by ).

LSMDC consists of 118,081 video clips sourced from 202 movies. The validation set contains 7,408 clips and evaluation is done on a test set of 1,000 videos from movies disjoint from the train and val sets. This follows the protocol outlined in .

Flickr30K . We also evaluate on a text-to-image retrieval benchmark to demonstrate the versatility of our model in that it can be used to achieve competitive performance in image settings as well as state-of-the art in video retrieval. The Flickr30K dataset contains 31,783 images with 5 captions per image. We follow the standard protocol of 1,000 images for validation, 1,000 images for testing and the remaining for training.

For downstream datasets with separate val and test splits, we train all models for 75 epochs and use the epoch with the lowest validation loss for reporting test results. For downstream datasets without a val set we report results at 50 epochs.

3 Implementation Details

All experiments are conducted with PyTorch . Optimization is performed with Adam, using a learning rate of 1×10−51\times 10^{-5}, we use batch sizes of 16, 24, and 96 for 8, 4, and 1-frame inputs respectively. The temperature hyperparameter σ\sigma for the loss defined in Eq. 2 & 3 is set to 0.05. The default pretraining is WebVid-2M and CC3M.

The text encoder of all models, unless specified otherwise, is instantiated as DistilBERT base-uncased pretrained on English Wikipedia and Toronto Book Corpus. The dimensionality of the common text-video space is set to 256. For visual augmentation, we randomly crop and horizontally flip during training, and center crop the maximal square crop at test time. All videos are resized to 224×224224\times 224 as input. At test-time we compute clip-embeddings for the video with a stride of 2 seconds. For paragraph-retrieval settings, we employ text augmentation during training by randomly sampling and concatenating a variable number of corresponding captions per video.

Finetuning time. A large motivation for using pre-extracted expert models for video retrieval is to save computational cost. Finetuning our 4-frame model for 50 epochs on MSR-VTT takes 10 hours on 2 Quadro RTX 6000k GPUs (with 24GB RAM each), which is similar to other works using pre-extracted expert features . This shows that our model is lightweight and can be finetuned end-to-end on the downstream video datasets quickly with sufficient pretraining (which is of one-time cost).

4 Ablation Study

In this section we study the effect of different pretraining strategies. In the Section C of the Appendix, we provide architectural ablations on different temporal expansion methods, different visual backbones, different text backbones and the improvement when using our modified space-time attention block.

Effect of pretraining. We compare performance on MSR-VTT with our model (i) trained from scratch, (ii) initialised with ImageNet weights and then finetuned, as well as (iii) initalised with ImageNet, and then pretrained on a number of different visual-text datasets before finetuning. For the video data, 4 frames are sampled at both pretraining and finetuning. Results on the MSR-VTT 1KA test set are shown in Table 2. For HowTo100M, we pretrain on a random 17M subset due to computational constraints (the largest subset we could obtain at the time of writing) totalling 19K hours. To generate text-video pairs, we sample 5 contiguous speech-video pairs and concatenate them to form a longer video. This allows for robustness to the noisy alignment of speech and vision. We find that training on CC3M alone does reasonably well, outperforming the HowTo-17M subset. This demonstrates the benefit of our flexible encoder that can be cheaply trained on images and easily applied to videos. Training on WebVid2M also outperforms training on the HowTo17M subset, despite being much smaller, confirming that the HowTo100M dataset is noisy. The best performance is achieved by jointly training on both CC3M and WebVid2M, effectively exploiting image and video data.

5 Curriculum strategy

Next, we evaluate the ability of our curriculum schedule to gradually learn the temporal dimension of videos by increasing the input number of frames. Table 3 summarises the results. Here, we show performance when pretraining on WebVid2M and finetuning on MSR-VTT. We explore two types of expansion in time: at pretraining and at finetuning stages. First, we observe that a single frame is not sufficient to capture the video content (18.8 R@1). Performing the temporal expansion at pretraining stage is better than doing so at finetuning (26.0 vs 24.9 R@1 with 4 frames). Finally, we obtain similar performance (slightly better at R@5) at half the computational cost in GPU hours by employing a curriculum strategy at pretraining (26.6 R@1). For 8 frames, the curriculum is even more useful, as we start training on 1 frame and then move to 4 before finally moving to 8 frames. Here, we obtain similar or better performance than training on 8 frames from the start, with almost a third of the computational cost. This is to be expected, as fewer frames significantly reduces forward pass times and enables larger batch sizes. Note that for a fair comparison, we allow the same number of training iterations for each row in the table.

We further analyse our proposed temporal curriculum strategy and its effects on training time and accuracy. Figure 3 shows the zero-shot results on MSR-VTT for various checkpoints with and without curriculum. It shows that our curriculum method yields a significant training speedup with a gain in accuracy. Shorter frame models are able to pass through more of the dataset in a shorter amount of time, which can lead to significant performance benefits in a constrained setting.

Expansion of temporal embeddings. We experiment with both zero padding and interpolation, and find that our model is robust to the type of temporal expansion strategy. More detailed results are provided in the Appendix, Section C.4.

6 Comparison to the State of the Art

Results on MSR-VTT can be seen in Table 4. We outperform all previous works, including many that pretrain on HowTo100M which is an order of magnitude larger than our pretraining dataset both in the number of hours (135K vs 13K) and in the number of caption-clip pairs (136M vs 5.5M). We also note that we outperform works that extract expert features (CE uses 9 experts, MMT uses 7) including object, motion, face, scene, sound and speech embeddings. We even outperform/perform on par with Support Set , which uses expert features from a 34-layer, R(2+1)-D model pretrained on IG65M, concatenated with ImageNet ResNet152 features, after which they add a transformer network and train end-to-end on HowTo100M.

We also report zero-shot results (Table 4) with no finetuning on MSR-VTT, outperforming both MIL-NCE and Support Set that trains on HowTo100M. This shows that our model is more generalisable, and can be used out of the box, and also perhaps that the domain of WebVid-2M is closer to that of MSR-VTT than HowTo100M. We will release the weights of our models publicly.

For both the zero-shot and finetuned setting we show that the addition of the COCO Captions image dataset further boosts our state-of-the-art MSR-VTT performance, indicating that the model is not yet saturated and additional pretraining dataset will lead to even better downstream performance.

For MSVD , we outperform all previous methods (Table 5). In particular, we outperform Support Set even though they train on an order of magnitude more data.

Results on DiDeMo can be found in Table 6. Note that on this dataset, our zero-shot performance is equivalent to CLIPBERT’s results with finetuning, and after we finetune our model on the DiDeMo training set we get an additional 14.2% boost in R@1.

We demonstrate further state-of-the-art results on LSMDC text-to-video retrieval. We outperform all previous methods, except for MMT in Median Rank, which pretrains on HowTo100M, a dataset consisting of over 100M clip-text pairs and contains multiple experts as well as audio modalities. Our model uses visual information alone.

To demonstrate the effectiveness of our model for downstream video and image tasks, we additionally report results on Flickr30K the image retrieval dataset in Table 8. Unlike other works which utilise high resolution regions extracted using a Faster-RCNN detector, our model is single stage and does not require any object detections. We compare to works with a similar number of training image-text pairs, and find that our model is comparable. We also note that training on WebVid2M provides a sizeable boost (5% improvement in R@1). Note that there are other recent text-image works such as UNITER and OSCAR , however these are trained on almost twice the number of samples. Recent works scale this up even further to billions of samples (ALIGN ).

Extension: Scaling up Further

To investigate the effects of downstream performance on additional pretraining datasets and increased scale, we train models on the following datasets:

WebVid-10M: An extension to our WebVid-2M dataset, we increase the size of the dataset fourfold to 10 million text-video pairs, following the same data collection protocol. The captions and video url’s can also be found at https://m-bain.github.io/webvid-dataset/.

Conceptual-Captions 12M : A dataset comprising of 12 million captioned images, intended for large-scale vision language pre-training. It is larger and more diverse than the Conceptual Captions (CC3M), albeit with noisier captions.

COCO Captions : A smaller dataset of 113.3k images with five captions per image, resulting in a total of 567k image-text pairs.

Downstream performance of these additional datasets can be found in Table 9. We find that restricting the model’s pretraining to only a small number of text-image pairs (COCO Captions) expectedly performs worse on downstream data, but still achieves competitive results. Thereby demonstrating the strength of our proposed method and that reasonable performance can be achieved on downstream video data with image pretraining alone.

Increasing the number of pretraining pairs consistently improves downstream performance, albeit with diminishing returns. It appears to be more efficient to add smaller datasets from diverse sources rather than add an increasingly larger dataset from a single source, shown by the boost of adding COCO captions (567k pairs) to the CC3M+WV2M pretraining compared to adding an extra 7.5 million pairs (WebVid10M) of that same source of data.

Conclusion

To conclude, we introduce a dual encoder model for end-to-end training of text-video retrieval, designed to take advantage of both large-scale image and video captioning datasets. Our model achieves state-of-the-art performance on a number of downstream benchmarks, however we note that the performance of our model is not saturated yet, and performance could be further improved by training on the full HowTo100M dataset, larger weakly paired image datasets such as Google3BN , as well as multi-dataset combinations thereof.

Acknowledgements. The authors would like to thank Samuel Albanie for his useful feedback. We are grateful for funding from a Royal Society Research Professorship, EPSRC Programme Grant VisualAI EP/T028572/1, and a Google PhD Fellowship.

Appendix

ActivityNet Captions contains 20K YouTube videos focused on actions, annotated with 100K sentences. The training set consists of 10K videos, and we use the ‘val1’ set of 4.9K videos to report results. At test time we use paragraph-to-video retrieval as is standard protocol set by other works, where the segment descriptions are concatenated to give a video-level description. We compare to prior work in Table 10 and achieve comparable results to the state of the art by using much less training data.

Appendix B Architectural Details

The patch embedding layer is implemented as a 2D convolutional layer with a kernel and stride size equivalent to the target patch size P=16P=16, and d=768d=768 output channels (the chosen embedding dimensionality of the video encoder).

The positional space and time embeddings are instantiated with shape M×dM\times d and N×dN\times d respectively, where MM is the maximum number of input video frames and NN is the maximum number of non-overlapping patches of size PP within a frame (196 for a video resolution of 224×224224\times 224). The [CLS] embedding is instantiated with shape 1×d1\times d.

Each space-time attention block consists of norm layers, temporal and spatial self-attention layers, and an MLP. The order and connections of these layers is shown in Figure 4.

B.2 Text Encoder

Our text encoder is instantiated as distilbert-base-uncased . Distilbert follows the same general architecture as BERT , but with the number of layers reduced by a factor of 2 and the token-type embeddings and the pooler removed. We use the HuggingFacehttps://huggingface.co/ transformers library implementation.

Appendix C Architectural Ablations

We investigate the effects of using different video backbone architectures (Table 11) and find that the space-time transformer encoder leads to large improvements in performance on MSR-VTT when compared to ResNets and 3D variants thereof.

During testing, all frame-variants see an equal number of frames, since the video embeddings are averaged over multiple strides.

For the video backbone ablation, we fix the text backbone to distilbert-base-uncased. For the text backbone ablation, we fix the video backone to the base space-time transformer with an input resolution of 224 and a patch size P=16P=16.

C.2 Text Backbone

The choice of text backbone has a significant impact on downstream performance (Table 12), with the t5 models performing significantly worse with more or similar numbers of parameters. DistilBERT and normal BERT achieve similar performance, with DistilBERT having far fewer parameters, therefore we chose to use DistilBERT in our work for efficiency.

C.3 Space-Time Attention

Space-time attention. Our modified space-time attention block, shown in Fig. 5, improves retrieval performance, as show in Table 13. We compare both variants during pretraining on WebVid-2M by reporting zero-shot results on MSR-VTT. We find once again that our modification leads to modest performance gains.

C.4 Temporal Expansion

We explore 3 different methods for expanding temporal positional embeddings (zero-padding and two interpolation methods), and observe robustness to all 3 (see Table 14).

Appendix D WebVid-2M Dataset Details

In this section, we show further details of the new WebVid-2M dataset. More qualitative examples of video-text pairs can be found in Figure 6 and histograms of caption lengths and video durations can be found in Figure 7. Note that 275,000 videos are longer than 30 seconds, providing many examples of videos which can be used for training long-range video models.

Appendix E WebVid-10M Extension

To facilitate further text-video pre-training, we extend the WebVid dataset fourfold to 10 million text-video pairs, following the same data collection protocol.

References