Use What You Have: Video Retrieval Using Representations From Collaborative Experts

Yang Liu, Samuel Albanie, Arsha Nagrani, Andrew Zisserman

Introduction

Videos capture the world in two important ways beyond a simple image: first, video contains temporal information – semantic concepts, actions and interactions evolve over time; Second, video may also contain information from multiple modalities, such as an accompanying audio track. This makes videos both richer and more informative, but also more challenging to represent. Our goal in this paper is to embed the information from multiple modalities and multiple time steps of a video segment into a compact fixed-length representation. Such a compact representation can then be used for a number of video understanding tasks, such as video retrieval, clustering and summarisation. In particular, we focus on retrieval; our objective is to be able to retrieve video clips using a free form text query that may contain both general and specific information.

Learning a robust and compact representation tabula rasa for this task is made extremely challenging by the high dimensionality of the sensory data contained in videos—to do so with discriminative training would require prohibitively expensive textual annotation of a vast number of videos. The primary hypothesis underpinning our approach is the following: the discriminative content of the multi-modal video embedding can be well approximated by the set of semantic representations of the video data learnt by individual experts (in audio, scenes, actions, etc). In essence, this approximation enables us to exploit knowledge from existing individual sources where the cost of annotation is significantly reduced (e.g. classification labels for objects and scenes in images, labels for actions in videos etc.) and where consequently, there exist very large-scale labelled datasets. These large-scale datasets can then be used to train independent “experts” for different perception tasks, which in turn provide a robust, low-dimensional basis for the discriminative query-content approximation described above.

The two key aspects of this idea that we explore in this paper are: (i) General and specific features: in addition to using generic video descriptors (e.g. objects and actions) we investigate encodings of quite specific information from the clip, for example, text from overlaid captions and text from speech to provide effective coverage of the “queryable content” of the video (Fig. 1, left). While such features may be highly discriminative for humans, they may not always be available (Fig. 1, right) and as we show through experiments (Sec. 4.3), making good use of these cues is challenging. We therefore also propose (ii) Collaborative experts: a framework that seeks to make effective use of embeddings from different ‘experts’ (e.g. objects, actions, speech) by learning their combination in order to render them more discriminative. Each expert is filtered via a simple dynamic attention mechanism that considers its relation to all other experts to enable their collaboration. This pairwise approach enables, for instance, the sound of a dog barking to inform the modulation of the RGB features, selecting the features that have encoded the concept of the dog. As we demonstrate in the sequel, this idea yields improvements in the retrieval performance.

Concretely, we make the following three contributions: (i) We propose the Collaborative Experts framework for learning a joint embedding of video and text by combining a collection of pretrained embeddings into a single, compact video representation. Our joint video embeddings are independent of the retrieval text-query and can be pre-computed offline and indexed for efficient retrieval; (ii) We explore the use of both general video features such as motion, image classification and audio features, and specific video features such as text embedded on screen and speech obtained using OCR and ASR respectively. We find that strong generic features deliver good performance, but that specific, rarely available features remain challenging to use for retrieval Note that this finding differs from the previous version of this paper (see appendix A.1).. (iii) We assess the performance of the representation produced by combining all available cues on a number of retrieval benchmarks, in several cases achieving an advance over prior work.

Related Work

Cross-Modal Embeddings: A range of prior work has proposed to jointly embed images and text into the same space , enabling cross-modal retrieval. More recently, several works have also focused on audio-visual cross-modal embeddings , as well as audio-text embeddings . Our goal in this work, however, is to embed videos and natural language sentences (sometimes multiple sentences) into the same semantic space, which is made more challenging by the high dimensional content of videos. Video-Text Embeddings: While a large number of works have focused on learning visual semantic embeddings for video and language, many of these existing approaches are based on image-text embedding methods by design and typically focus on single visual frames. Mithun el al. observe that a simple adaptation of a state-of-the-art image-text embedding method by mean-pooling features from video frames provides a better result than many prior video-text retrieval approaches . However, such methods do not take advantage of the rich and varied additional information present in videos, including motion dynamics, speech and other background sounds, which may influence the concepts in human captions to a considerable extent. Consequently, there has been a growing interest in fusing information from other modalities— utilise the audio stream (but do not exploit speech content) and use models pretrained for action recognition to extract motion features. These methods do not make use of speech-to-text or OCR for additional cues, which have nevertheless been used successfully to understand videos in other domains, particularly lecture retrieval (where the videos consist of slide shows) and news broadcast retrieval, where a large fraction of the content is displayed on screen in the form of text. Our approach draws particular inspiration from the powerful joint embedding proposed by (which in turn, builds on the classical Mixtures-of-Experts model ) and extends it to investigate additional cues (such as speech and text) and make more effective use of pretrained features via the robust collaborative gating mechanism described in Sec. 3. Annotation scarcity: A key challenge for video-retrieval is the small size of existing training datasets, due to the high cost of annotating videos with natural language. We therefore propose to use the knowledge from existing embeddings pretrained on a wide variety of other tasks. This idea is not new: semantic projections of visual inputs in the form of ‘experts’ was used by for the task of image retrieval and has also been central to modern video retrieval methods such as (Miech et al. 2018; Mithun et al. 2018). More recently, alternative approaches to addressing the issue of annotation scarcity have been explored, which include self-supervised Sun et al. 2019 and weakly-supervised Zhukov et al. 2019 video-text models.

Collaborative Experts

Given a set of videos with corresponding text captions, we would like to create a pair of functions ϕv\phi_{v} and ϕt\phi_{t} that map sensory video data and text into a joint embedding space that respects this correspondence—embeddings for paired text and video should lie close together, while embeddings for text and video that do not match should lie far apart. We would also like ϕv\phi_{v} and ϕt\phi_{t} to be independent of each other to enable efficient retrieval: the process of querying then reduces to a distance comparison between the embedding of the query and the embeddings of the collection to be searched (which can be pre-computed offline). The proposed Collaborative Experts framework for learning these functions is illustrated in Fig. 2. In this work, we pay particular attention to the design of the video encoder ϕv\phi_{v} and the process of combining information from different video modalities (Sec. 3.1). To complete the framework, we then discuss how the query text is encoded and the ranking loss used to learn the joint embedding space (Sec. 3.2).

Collaborative Gating: The collaborative gating module comprises two operations: (1) Prediction of attention vectors for every expert projection T={T(1)(v),…,T(n)(v)}T=\{T^{(1)}(\mathbf{v}),\dots,T^{(n)}(\mathbf{v})\}; and (2) modulation of expert responses. Inspired by the relational reasoning module proposed by Santoro et al. 2017 for visual question answering, we define the attention vector of the ithi^{th} expert projection TiT_{i} as follows:

where functions hϕh_{\phi} and gθg_{\theta} are used to model the pairwise relationship between projection Ψ(i)\Psi^{(i)} and projection Ψ(j)\Psi^{(j)}. Of these, gθg_{\theta} is used to infer pairwise task relationships, while hϕh_{\phi} maps the sum of all pairwise relationships into a single attention vector. In this work, we instantiate both hϕh_{\phi} and gθg_{\theta} as multi-layer perceptrons (MLPs). Note that the functional form of Equation (1) dictates that the attention vector of any expert projection should consider the potential relationships between all pairs associated with this expert. That is to say, the quality of each expert Ψ(j)\Psi^{(j)} should contribute in determining and selecting the information content from Ψ(i)\Psi^{(i)} in the final decision. It is also worth noting that the collaborative gating module uses the same functions gθg_{\theta} and hϕh_{\phi} (shared weights) to compute all pairwise relationships. This mode of operation encourages greater generalisation, since gθg_{\theta} and hϕh_{\phi} are encouraged not to over-fit to features of any particular pair of tasks. After the attention vectors T={T(1)(v),…,T(n)(v)}T=\{T^{(1)}(\mathbf{v}),\dots,T^{(n)}(\mathbf{v})\} have been computed, each expert projection is modulated follows:

where σ\sigma is an element-wise sigmoid activation and ∘\circ is the element-wise multiplication (Hadamard product). This gating function re-calibrates the strength of different activations of Ψ(i)(v)\Psi^{(i)}(\mathbf{v}) and selects which information is highlighted or suppressed, providing the model with a powerful mechanism for dynamically filtering content from different experts. A diagram of the mechanism is shown in Fig. 2 (right). The final video embedding is then obtained by passing the modulated responses of each expert through a Gated Embedding Module (GEM) Miech et al. 2018 (note that this operation produces l2-normalized outputs) before concatenating the outputs together into a single fixed-length vector.

2 Text Query Encoder and Training Loss

To construct the text embeddings, query sentences are first mapped to a sequence of feature vectors with pretrained contextual word-level embeddings (see Sec. 4.1 for details)—as with the video experts, the parameters of this first stage are frozen. These are then aggregated, again using NetVLAD Arandjelovic et al. 2016. Following aggregation, we follow the text encoding architecture proposed by Miech et al. 2018, which projects the aggregated features to separate subspaces for each expert using GEMs (as with the video encoder, producing l2-normalized outputs). Each projection is then scaled by a mixture weight (one scalar weight per expert projection), which is computed by applying a single linear layer to the aggregated text-features, and passing the result through a softmax to ensure that the mixture weights sum to one (see Miech et al. 2018 for further details). Finally, the scaled outputs are concatenated, producing a vector of dimensionality that matches that of the video embedding.

With the video encoder ϕv\phi_{v} and text encoder ϕt\phi_{t} as described, the similarity sijs_{i}^{j} of the ithi^{th} video, vi\mathbf{v}_{i}, and the and jthj^{th} caption, tj\mathbf{t}_{j}, can then be directly computed as the cosine of the angle between their respective embeddings ϕv(vi)Tϕt(tj)\phi_{v}(\mathbf{v}_{i})^{T}\phi_{t}(\mathbf{t}_{j}). During optimisation, the parameters of the video encoder (including the collaborative gating module) and text query encoder (the coloured regions of Fig. 2) are learned jointly. Training proceeds by sampling a sequence of minibatches of corresponding video-text pairs {vi,  ti}i=1NB\{\mathbf{v}_{i},\;\mathbf{t}_{i}\}_{i=1}^{N_{B}} and minimising a Bidirectional Max-margin Ranking Loss Socher et al. 2014:

where NBN_{B} is the batch size, and mm is a fixed constant which is set as a hyperparameter. When assessing retrieval performance, at test time the embedding distances are simply computed via their inner product, as described above.

When a set of expert features are missing, such as when there is no speech in the audio track, we simply zero-pad the missing experts when estimating the similarity score. To compensate for the implicit scaling introduced by missing experts (the similarity is effectively computed between shorter embeddings), we follow the elegant approach proposed by Miech et al. 2018 and simply remove the mixture weights for missing experts, then renormalise the remaining weights such that they sum to one.

Experiments

In this section, we evaluate our model on five benchmarks for video retrieval tasks. The description of datasets, implementation details and evaluation metric are provided in Sec. 4.1. A comprehensive comparison on general video retrieval benchmarks is reported in Sec. 4.2. We present an ablation study in Sec. 4.3 to explore how the performance of the proposed method is affected by different model configurations, including the aggregation methods, importance of different experts and number of captions in training.

Datasets: We perform experiments on five video datasets: MSR-VTT Xu et al. 2016, LSMDC Rohrbach et al. 2015, MSVD Chen and Dolan 2011, DiDeMo Anne Hendricks et al. 2017 and ActivityNet-captions Krishna et al. 2017, covering a challenging set of domains which include videos from YouTube, personal collections and movies. Expert Features: In order to capture the rich content of a video, we draw on existing powerful representations for a number of different semantic tasks. These are first extracted at a frame-level, then aggregated to produce a single feature vector per modality per video. RGB “object” frame-level embeddings of the visual data are generated with two models: an SENet-154 model Hu et al. 2019 (pretrained on ImageNet for the task of image classification), and a ResNext-101 Xie et al. 2017 pretrained on Instagram hashtags Mahajan et al. 2018. Motion embeddings are generated using the I3D inception model Carreira and Zisserman 2017 and a 34-layer R(2+1)D model Tran et al. 2018 trained on IG-65m Ghadiyaram et al. 2019. Face embeddings are extracted in two stages: (1) Each frame is passed through an SSD face detector Liu et al. 2016; Bradski 2000 to extract bounding boxes; (2) The image region of each box is passed through a ResNet50 He et al. 2016 that has been trained for the task of face classification on the VGGFace2 dataset Cao et al. 2018. Audio embeddings are obtained with a VGGish model, trained for audio classification on the YouTube-8m dataset Hershey et al. 2017. Speech-to-Text features are extracted using the Google Cloud speech API, to extract word tokens from the audio stream, which are then encoded via pretrained word2vec embeddings Mikolov et al. 2013. Optical Character Recognition is done in two stages: (1) Each frame is passed through the Pixel Link Deng et al. 2018 text detection model to extract bounding boxes for text; (2) The image region of each box is passed through a model Liu et al. 2018 that has been trained for scene text recognition on the Synth90K datasetJaderberg et al. 2014. The text is then encoded via a pretrained word2vec embedding model Mikolov et al. 2013. Temporal Aggregation: We adopt a simple approach to aggregating the features described above. For appearance, motion, scene and face embeddings, we average frame-level features along the temporal dimension to produce a single feature vector per video (we found max-pooling to perform similarly). For speech, audio and OCR features, we adopt the NetVLAD mechanism proposed by Arandjelovic et al. 2016, which has proven effective in the retrieval setting Miech et al. 2017. As noted in Sec. 3.1, all aggregated features are projected to a common size (768 dimensions). Text: Each word is encoded using pretrained word2vec word embeddings Mikolov et al. 2013 and then passed through a pretrained OpenAI-GPT model Radford et al. 2018 to extract contextual word embeddings. Finally, the word embeddings in each sentence are aggregated using NetVLAD. Dataset-specific details: Except where noted otherwise for ablation purposes, we use each of the embeddings described above for the MSR-VTT, ActivityNet and DiDeMo datasets. For MSVD, we extract the subset of features which do not require an audio stream (since no audio is available with the dataset). For LSMDC, we re-use the existing face, text and audio features made available by Miech et al. 2018, and combine them with the remaining features described above. Training Details: The CE framework is implemented with PyTorch Paszke et al. 2017. Optimisation is performed with the Lookahead solver Kingma and Ba 2014 in combination with RAdam Liu et al. 2019 (implementation by Wright 2019). Optimisation settings and the hyperparameter selection procedure is described in the appendix. Evaluation Metrics: We follow prior work (e.g. Dong et al. 2016; Zhang et al. 2018; Mithun et al. 2018; Yu et al. 2018; Miech et al. 2018) and report standard retrieval metrics (where existing work enables comparison) including median rank (lower is better), mean rank (lower is better) and R@K (recall at rank K—higher is better). When computing video-to-sentence metrics for datasets with multiple independent sentences per video (MSR-VTT and MSVD), we follow the evaluation protocol used in prior work Mithun et al. 2018; Dong et al. 2018; Dong et al. 2019 which corresponds to reporting the minimum rank among all valid text descriptions for a given video query. For each benchmark, we report the mean and standard deviation of three randomly seeded runs.

2 Comparison to Prior State-of-the-Art

We first compare the proposed method with the existing state-of-the-art on the MSR-VTT benchmark for the tasks of sentence-to-video and video-to-sentence retrieval Tab. 1. Driven by strong expert features, we observe that Collaborative Experts (CE) consistently improves retrieval performance for both sentence and video queries. We next evaluate the performance of the CE framework on the LSMDC benchmark for sentence-to-video retrieval (Tab. 2, left) and observe that CE matches or outperforms all prior work, including the prior state-of-the-art method Miech et al. 2018 which incorporates additional training images and captions from the COCO benchmark during training, but uses fewer experts. We observe similar trends in the results for the MSVD retrieval benchmark (Tab. 2, right). In Tab. 4, we compare with prior work on the ActivityNet paragraph-video retrieval benchmark (note that we compare to methods which use the same level of annotation as our approach i.e. video-level annotation), and see that CE is competitive. Finally, in Tab. 3 we provide a comparison with previously reported numbers on the DiDeMo benchmark and see that CE again outperforms prior work.

3 Ablation Studies

In this section, we provide ablation studies to empirically assess: (1) the effectiveness of the proposed collaborative experts framework vs other aggregation strategies; (2) the importance of using of a diverse range of experts with differing levels of specificity; (3) the relative value of using experts in comparison to simply having additional annotated training data.

Aggregation method: We compare the use of collaborative experts with several other baselines (with access to the same experts) for embedding aggregation including: (1) simple expert concatenation; (2) CE without projecting to a common dimension, without mixture weights and without the collaborative gating module described in Sec. 3.1; (3) the state of the art MoEE Miech et al. 2018 method (equivalent to CE without the common projection and collaborative gating) and (4) CE without collaborative gating. The results, presented in Tab. 6 (left), demonstrate the contribution of collaborative gating which improves performance and leads to a more efficient parameterisation than the prior state of the art.

Importance of different experts: The value of different experts is assessed in Tab. 5 (note that since several experts are not present in all videos, we combine them with features produced by a “scene” expert pretrained on Places365 Zhou et al. 2017—the expert with the lowest performance that is consistently available as a baseline to enable a more meaningful comparison). There is considerable variance in the effect produced by different choices of expert. Using stronger features within a given modality (pretraining on Instagram Mahajan et al. 2018 rather than Kinetics Carreira and Zisserman 2017 (resp. ImageNet) Deng et al. 2009 for actions (resp. object) experts can yield a significant boost in performance). The cues from scarce features (such as speech, face and OCR) which are often missing from videos (see Fig. 1, right) provide significantly weaker cues and bring a limited improvement to performance when used in combination.

Number of Captions in training: An emerging idea in our community is that many machine perception tasks might be solved through the combination of simple models and large-scale training sets, reminiscent of the “big-data” hypothesis Halevy et al. 2009. In this section, we perform an ablation study to assess the relative importance of access to pretrained experts and additional video description annotations. To do so, we measure the performance of the CE model as we vary (1) the number of descriptions available per-video during training and (2) the number of experts it has access to. The results are shown in Tab. 6 (right). We observe that increasing the number of training captions per-video from 1 to 20 brings an improvement in performance, approximately comparable to adding the full collection of experts, suggesting that indeed, adding experts can help to compensate for a paucity of labelled data. When multiple captions and multiple experts are both available, they naturally lead to the most robust embedding. Some qualitative examples of videos retrieved by the multiple-expert, multiple-caption system are provided in Fig. 3.

Conclusion

In this work, we introduced collaborative experts, a framework for learning a joint video-text embedding for efficient retrieval. We have shown that using a range of pretrained features and combining them through an appropriate gating mechanism can boost retrieval performance. In future work, we plan to explore the use of collaborative experts for other video understanding tasks such as clustering and summarisation.

Funding for this research is provided by the EPSRC Programme Grant Seebibyte EP/M013774/1 and EPSRC grant EP/R03298X/1. A.N. is supported by a Google PhD Fellowship. We would like to thank Antoine Miech, YoungJae Yu and Bowen Zhang for their assistance with experiment details. We would like to particularly thank Valentin Gabeur for identifying a bug in the software implementation that was responsible for the inaccurate results reported in the initial version of the paper. We would also like to thank Zak Stone and Susie Lim for their help with cloud computing.

References

Appendix A Supplementary Material

Following the release of the initial version of this paper (which can be viewed for reference at https://arxiv.org/abs/1907.13487v1), a bug was discovered in our open-source software implementation which resulted in: (i) an overestimate of model performance; (ii) inaccurate conclusions about the relative importance of different experts on retrieval performance.

This correction to the paper contains repeats of each of the experiments reported in the initial paper, with the following changes: (1) the removal of the bug which affected previous results; (2) a systematic approach to hyperparameter selection (discussed in more detail below); (3) the inclusion of additional “expert” pretrained features (described in Sec. A.5) to assess the influence of feature strength within a modality. In addition to results, the written analysis has also been updated to reflect the corresponding changes in results. The authors would like to express their gratitude to Valentin Gabeur who identified the bug in the software implementation and enabled this correction.

Bug details: The bug caused information about feature availability in the ground truth target video to become available to the query encoder during both training and testing when computing embedding distances. The leak occurred through incorrect weighting of the embedding distances due to: (1) a leaking broadcasting operation in an existing open-source library Miech et al. 2018 that was imported into our codebase; (2) incorrect NaN handling (introduced in our codebase), producing the same effect. The bug has now been patched in each of the open-source codebases that were known to have used this implementation.

A.2 Detailed Description of Datasets

MSR-VTT Xu et al. 2016: This large-scale dataset comprises approximately 200K unique video-caption pairs (10K YouTube video clips, each accompanied by 20 different captions). The dataset is particularly useful because it contains a good degree of video diversity, but we noted a reasonably high degree of label noise (there are a number of duplicate annotations in the provided captions). The dataset allocates 6513, 497 and 2990 videos for training, validation and testing, respectively. To enable a comparison with as many methods as possible, we also report results across other train/test splits used in prior work Yu et al. 2018; Miech et al. 2018. In particular, when comparing with Miech et al. 2018 (on splits which do not provide a validation set), we follow their evaluation protocol, measuring performance after training has occurred for a fixed number of epochs (100 in total). MSVD Chen and Dolan 2011: The MSVD dataset contains 80K English descriptions for 1,970 videos sourced from YouTube with a large number of captions per video (around 40 sentences each). We use the standard split of 1,200, 100, and 670 videos for training, validation, and testing Venugopalan et al. 2015; Xu et al. 2015 Note: referred to by Mithun et al. 2018 as the JMET-JMDV split. Differently from the other datasets, the MSVD videos do not have audio streams. LSMDC Rohrbach et al. 2015: This dataset contains 118,081 short video clips extracted from 202 movies. Each video has a caption, either extracted from the movie script or from transcribed DVS (descriptive video services) for the visually impaired. The validation set contains 7408 clips and evaluation is performed on a test set of 1000 videos from movies disjoint from the training and val sets, as outlined by the Large Scale Movie Description Challenge (LSMDC). https://sites.google.com/site/describingmovies/lsmdc-2017 ActivityNet-captions Krishna et al. 2017: ActivityNet Captions consists of 20K videos from YouTube, coupled with approximately 100K descriptive sentences. We follow the paragraph-video retrieval protocols described in Zhang et al. 2018 training up to 200 epochs and reporting performance on val1 (this train/test split allocates 10,009 videos for training and 4,917 videos for testing). DiDeMo Anne Hendricks et al. 2017: DiDeMo contains 10,464 unedited, personal videos in diverse visual settings with roughly 3-5 pairs of descriptions and distinct moments per video. The videos are collected in an open-world setting and include diverse content such as pets, concerts, and sports games. The total number of sentences is 40,543. While the moments are localised with time-stamp annotations, we do not use time stamps in this work.

A.3 Optimisation details and hyperparameter selection

For each dataset, a grid search was first performed (using the Lookahead solver Zhang et al. 2019; Wright 2019) over batch sizes (16, 32, 64, 128, 256), learning rates (0.1, 0.01) and weight decay (1E-3, 5E-5) for each dataset using a single expert to determine appropriate optimisation parameters. Next, an experiment on MSR-VTT compared several choices for the dimensionality of the projection operation applied to the features (described in Sec. 3.1) (choosing among 512, 768 and 1024 dimensions), which suggested that 768 was most effective. This was then fixed for all remaining experiments (this represents a difference from the original paper, in which 512 was used). Further ablations (provided below) indicate that performance is not sensitive to this hyperparameter. Next, Asynchronous Hyperband Li et al. 2018 was used to select all remaining hyperparameters on MSR-VTT by partially evaluating 1k configurations on the validation sets for each dataset. These hyperparameters consisted of: the number of VLAD clusters and ghost clusters Zhong et al. 2018 used for different experts, the zero-padding length applied to variable-length experts, the margin hyperparameter mm in Eq. 3, the Collaborative Gating architecture (whether to use batch normalization Ioffe and Szegedy 2015, the number of layers used to form the MLP, and the choice of activation function). The architecture choices were then fixed for all datasets. Note that to ensure a fair comparison on MSR-VTT with the MoEE method of Miech et al. 2018 in Tab. 6, MoEE was also provided with a budget of 1k sampled configurations. To determine zero-padding, margin and VLAD clusters for DiDeMo, MSVD and LSMDC further Asynchronous Hyperband searches were conducted, each with a budget of 500 sampled configurations. Since, differently from the other datasets with available validation and test sets, the validation set itself is used to assess performance on ActivityNet, hyperparameters were copied from the DiDeMo configuration. The configurations, experts, pretrained models and logs for each of the experiments reported in this paper are made available as part of the updated open-source implementation at www.robots.ox.ac.uk/~vgg/research/collaborative-experts/.

A.4 Ablation Studies - Full Tables

A.5 Implementation Details

Object frame-level embeddings of the visual data are generated with two models, Obj(IN) and Obj(IG). Obj(IN) is an SENet-154 model Hu et al. 2019 (pretrained on ImageNet for the task of image classification) from frames extracted at 25 fps, where each frame is resized to 224 ×\times 224 pixels. Obj(IG) is a ResNext-101 Xie et al. 2017 pretrained on Instagram hashtags Mahajan et al. 2018, using the same frame preparation as Obj(IN). Features are collected from the final global average pooling layer of both models, and have a dimensionality of 2048. Action embeddings are similarly generated from two models, Action(KN) and Action(IG). Action(KN) is an I3D inception model that computes features following the procedure described by Carreira and Zisserman 2017. Frames extracted at 25fps and processed with a window length of 64 frames and a stride of 25 frames. Each frame is first resized to a height of 256 pixels (preserving aspect ratio), before a 224 ×\times 224 centre crop is passed to the model. Each temporal window produces a (1024x7)-matrix of features. Action(IG) is a 34-layer R(2+1)D model Tran et al. 2018 trained on IG-65m Ghadiyaram et al. 2019 which processes clips of 8 consecutive 112 ×\times 112 pixel frames, extracted at 30 fps (we use the implementation provided by Daniel 2019). Face embeddings are extracted in two stages: (1) Each frame (also extracted at 25 fps) is resized to 300 ×\times 300 pixels and passed through an SSD face detector Liu et al. 2016; Bradski 2000 to extract bounding boxes; (2) The image region of each box is resized such that the minimum dimension is 224 pixels and a centre crop is passed through a ResNet50 He et al. 2016 that has been trained for task of face classification on the VGGFace2 dataset Cao et al. 2018, producing a 512-dimensional embedding for each detected face. Audio embeddings are obtained with a VGGish model, trained for audio classification on the YouTube-8m dataset Hershey et al. 2017. To produce the input for this model, the audio stream of each video is re-sampled to a 16kHz mono signal, converted to an STFT with a window size of 25ms and a hop of 10ms with a Hann window, then mapped to a 64 bin log mel-spectrogram. Finally, the features are parsed into non-overlapping 0.96s collections of frames (each collection comprises 96 frames, each of 10ms duration), which is mapped to a 128-dimensional feature vector. Scene embeddings of 2208 dimensions are extracted from 224×\times224 pixel centre crops of frames extracted at 1fps using a DenseNet-161 Huang et al. 2017 model pretrained on Places365 Zhou et al. 2017. Speech to Text The audio stream of each video is re-sampled to a 16kHz mono signal. We then obtained transcripts of the spoken speech for MSR-VTT, MSVD and ActivityNet using the Google Cloud Speech to Text API https://cloud.google.com/speech-to-text/ from the resampled signal. The language for the API is specified as English. For reference, of the 10,000 videos contained in MSR-VTT, 8,811 are accompanied by audio streams. Of these, we detected speech in 5,626 videos. Optical Character Recognition is extracted in two stages: (1) Each frame is resized to 800×400800\times 400 pixels) and passed through Pixel Link Deng et al. 2018 text detection model to extract bounding boxes for texts; (2) The image region of each box is resized to 32×25632\times 256 and then pass through a model Liu et al. 2018; Shi et al. 2017 that has been trained for text of scene text recognition on the Synth90K datasetJaderberg et al. 2014, producing a character sequence for each detect box. They are then encoded via a pretrained word2vec embedding model Mikolov et al. 2013. Text We encode each word using the Google News GoogleNews-vectors-negative300.bin.gz found at: https://code.google.com/archive/p/word2vec/ trained word2vec word embeddings Mikolov et al. 2013. All the word embeddings are then pass through a pretrained OpenAI-GPT model to extract the context-specific word embeddings (i.e., not only learned based on word concurrency but also the sequential context). Finally, all the word embeddings in each sentence are aggregated using NetVLAD.