Just Ask: Learning to Answer Questions from Millions of Narrated Videos

Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, Cordelia Schmid

Introduction

Answering questions about videos requires a detailed understanding of the visual content and its association with the natural language. Indeed, given the large diversity of questions, methods for Video Question Answering (VideoQA) should reason about scenes, objects and human actions as well as their complex temporal interactions.

Current approaches to VideoQA rely on deep fully-supervised models trained on manually annotated datasets with question and answer pairs . Collecting and annotating VideoQA datasets, however, is cumbersome, time consuming, expensive and therefore not scalable. As a result, current VideoQA datasets are relatively small (see Figure 2). This limitation hinders the progress in the field as state-of-the-art VideoQA models often require a large amount of training data.

In this work, we address the scale issue with a new approach for automatically generating a VideoQA dataset, see Figure 1 for examples. The idea is to leverage cross-modal supervision together with text-only tools for question generation and to automatically annotate VideoQA from a large amount of readily-available narrated videos. Inspired by the recent progress in language generation using transformer-based language models , we leverage transformers trained on a question-answering text corpus to generate a diverse set of non-scripted questions and corresponding open-vocabulary answers from text. By applying these transformers to speech transcripts of narrated videos from the large-scale HowTo100M dataset , we create HowToVQA69M, an open-ended VideoQA dataset with 69 million video-question-answer triplets and a diverse set of more than 16M unique answers (see Figure 3). As shown in Figure 2, our HowToVQA69M is two orders of magnitude larger compared to prior VideoQA datasets.

Given the limited diversity of existing datasets, current methods typically reduce video question answering to a classification problem, where frequent answers are assigned to unique classes. Typically, up to 5K unique possible answers are considered. Such an approach, however, does not scale to the open vocabulary of 16M different answers in our dataset. To address this problem and to enable video question answering with highly diverse questions and answers, we introduce a training procedure based on contrastive learning between a video-question multi-modal transformer and an answer transformer that can handle free-form answers. This bypasses the need to define a discrete set of answer classes.

The goal of our work is to advance truly open-ended and generic solutions to VideoQA. To evaluate generalization, we propose a new zero-shot VideoQA task where we prohibit any manual supervision of visual data during training. Our VideoQA model, trained on HowToVQA69M, demonstrates excellent zero-shot results on multiple existing datasets, especially for rare answers. Moreover, when finetuned on target datasets, our model significantly outperforms the state of the art on MSRVTT-QA , MSVD-QA ActivityNet-QA , and How2QA .

Initial experiments showed that existing benchmarks for open-ended VideoQA contain a language bias , i.e., their questions can often be answered without looking at the video. To better evaluate the impact of visual information in VideoQA, we introduce a new open-ended VideoQA dataset (iVQA) with manually collected questions and answers, where we exclude questions that could be answered without watching the video. Moreover, to account for multiple possible answers, iVQA contains five independently collected answers for each question.

In summary, our work proposes the following three contributions:

We introduce an approach to automatically generate a large-scale VideoQA dataset, HowToVQA69M. Relying on cross-modal supervision, we use transformers trained on an existing text-only question-answering corpus and generate video-question-answer triplets from videos and transcribed narrations.

We train a VideoQA model on HowToVQA69M with contrastive learning between a multi-modal video-question transformer and an answer transformer. We show the efficiency of our model in the new zero-shot VideoQA task and outperform the state of the art in four existing VideoQA benchmarks: MSRVTT-QA, MSVD-QA, ActivityNet-QA and How2QA.

Finally, we introduce a new manually annotated open-ended VideoQA benchmark iVQA that excludes non-visual questions and contains multiple possible answers for each question.

Code, datasets and trained models are available at .

Related Work

Visual Question Answering (VQA). VQA is typically tackled by classifying the image-question (or video-question) representation into a fixed vocabulary of answers. Various approaches to combine spatial image representations and sequential question representations have been proposed . More specifically to the video domain (VideoQA), spatio-temporal video representations in terms of motion and appearance have been used in .

Methods above are limited to pre-defined vocabularies of answers and are difficult to apply outside of specific datasets. To address this problem, Hu et al. propose a joint embedding where image-question representations can be matched with free-form answers. Our VideoQA model follows this idea, but instead of relying on manually annotated datasets of limited scale, we train it on a large-scale VideoQA dataset that we automatically generate. In contrast to some previous works using additional video features such as subtitles , our video representation is exclusively based on visual information, as we focus on the visual understanding of videos.

To evaluate the generalization of VQA models, Teney and Hengel define zero-shot VQA by answering previously unseen questions, which is a related but less challenging task compared to the zero-shot VQA task we propose in Section 6.2. Vatashsky and Ullman address VQA using COCO image annotations , while our zero-shot model is trained with no manual annotations. Our proposed zero-shot VQA task is analogous to zero-shot video retrieval or zero-shot action recognition .

Visual question generation (VQG) has been introduced in . The methods in and propose to jointly learn VQG and VQA to improve the image VQA task. However, these works do not generate questions to obtain additional training data, but use visual data annotation for question generation as an additional loss.

VideoQA datasets. Manually collecting and annotating video-question-answer triplets is cumbersome, costly and difficult to scale. As result, current VideoQA datasets are limited in size, as the largest, TGIF-QA , contains only 72K annotated clips (see Figure 2 for more details). To address this issue, several works have explored leveraging manually annotated video descriptions for automatic generation of VideoQA datasets, using rule-based approaches.

Instead, we propose to use video narrations that are available at large-scale with no manual supervision. Moreover, rule-based generation requires the manual creation of rules by experts which is expensive, and has also been recently outperformed by neural question generation as used in our approach.

Large-scale pretraining for vision and language. Several recent methods pretrain multi-modal vision-language representations, such as transformers, using datasets with image captions, e.g., COCO , Conceptual Captions and Visual Genome . These methods are often optimized using generic objectives such as masked language losses and losses for text-image matching and image caption generation. In our work, we pretrain models using large amounts of narrated videos. In contrast to task-agnostic pretraining in the previous work, we show the benefits of task-specific pretraining for our target VideoQA task.

Learning from narrated videos. In this work, we exploit noisy correlations between videos and narrations in unlabeled instructional videos from the recent HowTo100M dataset . Methods using such readily-available data have shown significant improvements on several tasks including video retrieval, action localization, action recognition and video captioning , sometimes outperforming fully-supervised baselines. Some recent works use narrated videos for VideoQA. Amrani et al. propose a text-video pretraining approach and finetune for VideoQA. Li et al. propose HERO, a pretraining approach restricted to multiple-choice VideoQA, for which question and answer are treated as a single text stream. Seo et al. propose a pretraining approach based on next utterance prediction and finetune for VideoQA. Differently to these methods with task-agnostic pretraining, we propose a pretraining approach specifically dedicated for VideoQA using automatically generated question and answer pairs from narrated videos, and show in Section 6 the superiority of our approach.

Large-scale generation of VideoQA data

This section presents our approach to generate a large-scale VideoQA dataset from videos and transcribed narrations describing the content of the videos. Section 3.1 presents our proposed generation procedures. Section 3.2, then, describes the resulting HowToVQA69M dataset.

We tackle the task of generating video-question-answer triplets from a large-scale instructional video dataset with transcribed spoken narration . This is a challenging task because of transcription errors and lack of punctuation. We also wish to obtain highly diverse data. To address these issues, we propose to leverage powerful language models trained on text data. Our approach is illustrated in Figure 3 and details are given next.

We first present details about the generation procedure. Let ss be the transcribed speech data obtained with automatic speech recognition (ASR). First, we use a recurrent neural network pp, to infer punctuation in the transcribed speech data. We denote the punctuated transcript as p(s)p(s). We extract video clips vv temporally aligned with the inferred sentences p(s)p(s) using the ASR timestamps. We found that the generation works significantly better when applied to sentences rather than the original sentence fragments from the HowTo100M dataset, see Table 1. Second, for each sentence, we apply a transformer TaT_{a}, to extract a set of potential answers: a=Ta(p(s))a=T_{a}(p(s)). Third, we use another transformer TqT_{q} to generate a question given each transcript sentence and each extracted answer such that: q=Tq(a,p(s))q=T_{q}(a,p(s)). The output is a set of video-question-answer triplets (v,q,a)(v,q,a).

We now explain details about the language models and their training procedure. For ASR, we follow and use the readily-available ASR data provided by YouTube. For punctuation pp, we use the BRNN model from and the weights available at trained on IWSLT2011 . For TaT_{a} and TqT_{q}, we use the transformer-based T5-small and T5-base models , respectively. We follow and use the weights available at trained for answer span extraction and answer-aware question generation, respectively, on SQuADv1 . SQuADv1 is a text-only question-answering dataset consisting of questions for which the answer is a segment of text extracted from a paragraph.

2 HowToVQA69M: large-scale VideoQA dataset

We have applied the previously described procedure to all 1.2M original videos from the HowTo100M dataset . The result is HowToVQA69M, a dataset of 69,270,581 video clip, question and answer triplets (v,q,a)(v,q,a). HowToVQA69M is two orders of magnitude larger than any of the currently available VideoQA datasets (see Figure 2). On average, each original video results in 43 video clips, where each clip lasts 12.1 seconds and is associated to 1.2 question-answer pairs. Questions and answers contain 8.7 and 2.4 words on average respectively. HowToVQA69M is highly diverse and contains over 16M unique answers, where over 2M unique answers appear more than once and over 300K unique answers appear more than ten times. Examples of (v,q,a)(v,q,a) triplets from the HowToVQA69M dataset are illustrated in Figure 4.

Manual evaluation of HowToVQA69M. As shown in Figure 4, HowToVQA69M annotations are noisy, which can be attributed to: (i) errors in speech transcription, (ii) speech not describing the video content, or (iii) errors in question-answer generation. We manually evaluated the quality of 100 randomly sampled (v,q,a)(v,q,a) triplets in HowToVQA69M, collected 5 different annotations for each triplet to reduce variance, and reported results in Table 1. Among 100 triplets generated by our method we find 30 to be correctly generated and matching well to the video content, 31 are incorrectly generated and 39 are correctly generated but unrelated to the video content. To demonstrate the influence of the different components of our automatic question-answer generation procedure, we compare it with (i) a variant of our approach that does not split transcribed narrations into sentences using a punctuator, and (ii) a rule-based approach for question-answer generation. Table 1 confirms the importance of punctuation and demonstrates the superior performance of our generation method compared to . Inter-rater agreement statistics, and more details for the generated dataset are provided in Appendix A. Further comparison with is given in Section 6.5. We describe next how we use HowToVQA69M to train our VideoQA model.

VideoQA model and training procedure

This section presents our VideoQA model in Section 4.1 and describes its training procedure in Section 4.2. Figure 5 gives an overview of the model.

Word representation. The question and answer are separately tokenized with the WordPieces embedding and fed to DistilBERT . DistilBERT is a light version of BERT pretrained in a self-supervised fashion on English Wikipedia and the Toronto Book Corpus .

Video representation. We use a frozen S3D pretrained on HowTo100M using MIL-NCE . This model is pretrained from scratch on HowTo100M only.

2 Training procedure

This section describes the training of our VideoQA model on the HowToVQA69M dataset and its finetuning on downstream VideoQA datasets.

We wish to make a pair of video and question (v,q)(v,q) close to its correct answer aa measured by the dot product of their embeddings, f(v,q)⊤g(a)f(v,q)^{\top}g(a). Conversely, the incorrect answers should be far, i.e., the dot product with their embeddings should be small. Formally, this can be done by maximizing the following contrastive objective:

where (vi,qi,ai)(v_{i},q_{i},a_{i}) represents a triplet of generated (video clip, question, answer) from HowToVQA69M. Given a specific positive triplet (vi,qi,ai)(v_{i},q_{i},a_{i}), we construct the set Ni\mathcal{N}_{i} of negative triplets by concatenating incorrect answers aja_{j} within the training batch to the video-question pair (vi,qi)(v_{i},q_{i}) as: (vi,qi,aj)(v_{i},q_{i},a_{j}) with aj≠aia_{j}\neq a_{i}. In particular, if the same negative answer aja_{j} is present multiple times in a batch, we only count it once. We found that sampling the same negative answer multiple times leads to worse results (see Section 6.6), which we believe is due to different distributions of answers in the pretraining and downstream datasets. Removing duplicate negatives helps to mitigate this difference.

Finetuning on downstream VideoQA datasets.

We leverage the model pretrained on HowToVQA69M and finetune it on a downstream VideoQA dataset that typically has a smaller vocabulary of answers VV (e.g. ∣V∣∼4000|V|\sim 4000). To this end, we adapt the training objective in (1) by constructing the negative set Ni\mathcal{N}_{i} from all incorrect answers in VV. Note that in such setting (1) becomes equivalent to optimizing the standard cross-entropy objective. In the specific case of multiple-choice VideoQA, the set of negatives Ni\mathcal{N}_{i} is the set of incorrect answers for each sample.

Masked Language Modeling (MLM).

In addition to the contrastive loss (1) we apply the masking loss to question tokens during both pretraining and finetuning. We found this to have a positive regularization effect when finetuning the DistilBERT weights (see Section 6.6).

iVQA: new dataset for VideoQA evaluation

In this section we present our Instructional VQA dataset (iVQA). We start from a subset of HowTo100M videos and manually annotate video clips with questions and answers. We aim to (i) provide a well-defined evaluation by including five correct answer annotations per question and (ii) avoid questions which can be answered without watching the video. The dataset is described below and more details are given in Appendix C and E.3.

iVQA videos are obtained by randomly sampling 7-30 sec. video clips from the HowTo100M dataset . We avoid overlap between datasets and make sure iVQA and HowToVQA69M have no videos in common. Each clip is manually annotated with one question and 5 answers on Amazon Mechanical Turk. We ask workers to annotate questions about objects and scenes in the video and remove videos that could not be annotated. The correctness of annotations is manually verified by the authors. Moreover, we manually reduce the language bias by excluding questions that could be answered without watching the video. To increase diversity, each question is answered by 5 different workers. The answers are restricted to 4 words and are complemented by a confidence level. Questions that receive multiple answers with low confidence are removed.

Statistical Analysis.

iVQA contains 10,000 video clips with one question and five corresponding answers per clip. We split the dataset into 60%/20%/20% train/validation/test subsets. On average, questions and answers contain 7.6 and 1.1 words respectively. The average duration of video clips is 18.6 seconds. The majority of questions have at least 2 annotators providing the same answer. Similarly to , this motivates us to define the following accuracy measure for a given answer aa: acc(a)=min⁡(#ground truth answers = a2,1).acc(a)=\min(\frac{\#\textrm{ground truth answers =}\,a}{2},1). This metric assigns 100% accuracy to answers confirmed by at least 2 annotators, 50% accuracy to answers confirmed by only 1 annotator and 0% otherwise. Note that this definition is specific to multiple ground truth answers per question.

Experiments

This section demonstrates the benefits of training using our generated HowToVQA69M dataset and compares our method to the state of the art. We first outline the used datasets, baseline methods and implementation details in Section 6.1. We then present results for the novel zero-shot VideoQA task in Section 6.2. The comparison to the state of the art in VideoQA and alternative training strategies is given in Section 6.3. Section 6.4 presents results for rare answers. Finally, we compare our VideoQA generation approach to previous methods in Section 6.5 and present ablation studies in Section 6.6.

We use two datasets for training and five datasets for evaluation as described below. We follow previous evaluation protocols for open-ended settings and use a fixed vocabulary of training answers. Unless stated otherwise, we report top-1 test accuracy and use original splits for training, validation and test.

For training we use our new HowToVQA69M dataset introduced in Section 3.2 with 90% and 10% videos in training and validation subsets. For comparison, we also train our model using a large-scale text-video dataset, HowTo100M , that contains videos with transcribed narrations but no video-question-answer triplets. Test and validation videos of downstream datasets are excluded from HowTo100M and HowToVQA69M.

We evaluate results on four open-ended VideoQA downstream datasets: MSRVTT-QA , MSVD-QA , ActivityNet-QA and our new iVQA dataset (see Section 5). We also evaluate on a multiple-choice VideoQA dataset How2QA where each question is associated with one correct and three incorrect answers.

Baselines.

To evaluate the contribution of the visual modality, we compare our VQA-T model with its language-only variant QA-T. QA-T does not use video input, i.e. we set the input vv of the video-question transformer to zero (see Figure 5). To evaluate our generated dataset, we also compare VQA-T trained on HowToVQA69M and on HowTo100M. Since HowTo100M has no (v,q,a)(v,q,a) triplets, we only train the ff branch of VQA-T on HowTo100M using the standard masking and cross-modal matching losses . In the zero-shot setting we evaluate VQA-T trained on HowTo100M by computing f(v,[q,a])f(v,[q,a]) for concatenated pairs of questions and answers [q,a][q,a]. During finetuning we also initialize the gg branch of VQA-T with parameters of the text encoding obtained from ff (see further details in Appendix B).

Implementation details.

For the training on HowToVQA69M we use the Adam optimizer and mini-batches with 4096 video clips sampled from 128 random videos. The optimization over 10 epochs lasts 2 days on 8 Tesla V100 GPUs. Further details are included in Appendix D.

2 Zero-shot VideoQA

In this section, we address the zero-shot VideoQA task where we prohibit any manual supervision of visual data during training. We explore this setup to evaluate the generalization of VQA-T trained on HowToVQA69M to unseen downstream datasets. For consistency, we use the vocabulary of answers from downstream datasets during testing (see Section 6.1).

Zero-shot results are presented in Table 2. We first observe that the use of visual cues by VQA-T outperforms QA-T when both models are trained on HowToVQA69M. This demonstrates the importance of the cross-modality in HowToVQA69M despite the VideoQA annotation being exclusively generated from text-only methods. Since HowToVQA69M has been generated using no manual annotation of visual data, our approach is scalable and can lead to further improvements by increasing the dataset size, as we discuss in Section 6.6.

Training on HowToVQA69M significantly outperforms the training on HowTo100M and the random baseline. This confirms the advantage of our HowToVQA69M dataset for the VideoQA task over other generic text-video datasets that do not contain video-question-answer triplets. We emphasize that our training does not use any information about target VideoQA datasets. Qualitative results for zero-shot VideoQA are presented for our approach and compared with baselines in Figure 6. We observe that QA-T (trained on HowToVQA69M) provides plausible but video-unrelated answers to the questions. Moreover, VQA-T (trained on HowTo100M) is able to associate visual content with related answers, but fails to have a complex multi-modal understanding. Our VQA-T model trained on HowToVQA69M, on the other hand, correctly understands questions and uses information in the video to provide correct answers, confirming results in Table 2.

3 Benefits of HowToVQA69M pretraining

This section evaluates the effect of VQA-T pretraining in combination with finetuning on target datasets. As shown in Table 3, pretraining on HowToVQA69M provides consistent and significant improvements for all datasets when compared to pretraining on HowTo100M and no pretraining. In particular, we observe the largest improvement for our new iVQA dataset which comes from the same domain as HowToVQA69M. Hence, the automatic generation of training data for other domains using our method can lead to further improvements on other datasets.

We compare our pretrained model to the state-of-the-art in VideoQA in Tables 4-5. Notably, VQA-T pretrained on HowToVQA69M outperforms previous methods on all tested datasets. In particular, our method improves over the recent CoMVT approach that has been pretrained on HowTo100M. These strong results show the importance of our proposed HowToVQA69M dataset.

4 Results for rare answers

Training on downstream VideoQA datasets typically leads to particularly large improvements for questions with most frequent answers. As shown in Table 6, our approach brings significant improvements both for common and rare answers compared to models trained from scratch or pretrained on HowTo100M. Interestingly, for the most rare answers in iVQA (Q3 and Q4) our model without finetuning (zero-shot mode) outperforms finetuned models that have not been pretrained on HowToVQA69M. We make similar observations for rare answers in other datasets and report corresponding results in Appendix E.2. We conclude that VideoQA specific pretraining on additional large-scale, diverse data helps improve generalization of VideoQA models.

5 Comparison of VideoQA generation methods

In this section, we compare our question-answer generation approach to Heilman et al. , that was notably used in to generate VideoQA data from video descriptions. We run the method of on sentences extracted from HowTo100M, apply our pretraining method on the generated data and show results in Table 7. Note that we do not choose MSRVTT-QA and MSVD-QA as downstream datasets for this comparison because their evaluation sets were automatically generated using Heilman et al. . We find that our generation method leads to significantly better performance both in zero-shot and finetuning settings. We also provide a qualitative comparison in Appendix A, further demonstrating the benefit of our transformer-based question-answer generation approach compared to previous methods. We also show the benefit of our generated HowToVQA69M dataset by comparing our results to cross-dataset transfer using existing VideoQA datasets in Appendix E.1.

6 Ablation studies

Pretraining losses. As shown in Table 8, removing duplicate negative answers in our contrastive loss, as discussed in Section 4.2, is beneficial notably in the zero-shot setting. Moreover, adding the MLM loss at pretraining improves the downstream results for both zero-shot and finetuning when used in combination with our contrastive learning strategy. These results motivate our proposed pretraining approach.

Importance of scale. Results of our method after pretraining on different fractions of HowToVQA69M are shown in Table 9. We construct these subsets such that larger subsets include the smaller ones. These results suggest that the scale is an important factor and that we can expect further improvements with additional pretraining data, both in the zero-shot and finetuning settings.

Conclusion

We propose a novel and scalable approach for training VideoQA models without manually annotated visual data. We automatically generate HowToVQA69M – a large-scale VideoQA training dataset generated from narrated videos with readily-available speech transcripts, significantly exceeding existing datasets by size and diversity. We demonstrate several benefits of pretraining on HowToVQA69M. We are the first to demonstrate zero-shot VideoQA results without the use of any manually annotated images or videos. Furthermore, finetuning our HowToVQA69M pretrained model on downstream tasks outperforms the state of the art on MSRVTT-QA, MSVD-QA, ActivityNet-QA and How2QA. We further validate our approach on a new iVQA benchmark we manually collect.

Acknowledgements. This work was granted access to the HPC resources of IDRIS under the allocation 2020-101267 made by GENCI. The work was funded by a Google gift, the French government under management of Agence Nationale de la Recherche as part of the ”Investissements d’avenir” program, reference ANR-19-P3IA-0001 (PRAIRIE 3IA Institute), the Louis Vuitton ENS Chair on Artificial Intelligence, the European Regional Development Fund under project IMPACT (reg. no. CZ.02.1.01/0.0/0.0/15 003/0000468) and A. Miech’s Google PhD fellowship. We thank P.-L. Guhur and M. Tapaswi for advice on using Amazon Mechanical Turk, E. Berthier, Q. Le Lidec and E. Chane-Sane for the manual evaluation of generated VideoQA data, and I. Rocco for proofreading.

References

Appendix

In this Appendix, we start by giving additional analysis and examples of our proposed HowToVQA69M dataset in Section A. We, then, provide additional architecture details for our VideoQA model in Section B. Next, we present additional statistics and details of the collection procedure for our manually collected iVQA evaluation benchmark in Section C. We describe additional implementation details in Section D and present experiments including cross-dataset transfer, results per answer quartile and per question type in Section E.

Appendix A Analysis of HowToVQA69M dataset

Figure 7 shows the statistics of the HowToVQA69M dataset in terms of the question length, answer length and video clip duration. Overall, HowToVQA69M contains longer answers than downstream open-ended VideoQA datasets like MSRVTT-QA, MSVD-QA or ActivityNet-QA. The distribution of clip duration has a peak at around seven seconds with a long tail of longer clips. These statistics demonstrate the diversity of our HowToVQA69M dataset, both in terms of videos and answers.

Word cloudsTo generate the word clouds, we used https://github.com/amueller/word_cloud. for questions and answers in HowToVQA69M are shown in Figure 8 and illustrate the diverse vocabulary in HowToVQA69M as well as the presence of speech-related words such as as okay, right, oh. In Figure 10 we illustrate the diversity and the noise in the automatically obtained annotations in the HowToVQA69M dataset.

We show quantitative comparisons of our question-answer generation models with in Section 6.5, and supplement it here with a qualitative comparison shown in Figure 9. We found that compared to our generation method provides higher quality as well as higher diversity of question-answer pairs when applied to the uncurated sentences extracted from speech in narrated videos.

In Section 3.2 we present a manual evaluation of the quality of the automatically generated video-question-answer triplets for our method and two other baselines. We complement this analysis here with inter-rater agreement statistics. For the 300 generated video-question-answer triplets (100 for each generation method), 94 were in an agreement of all 5 annotators, 198 in an agreement of at least 4 annotators, and 299 in an agreement of at least 3 annotators. This high agreement of annotators demonstrates the reliability of the results in Table 1.

We further manually classify the 100 video-question-answer triplets obtained with our method by the question type (“Attribute”, “Object”, “Action”, “Counting”, “Place”, “People”, or “Other”), evaluate the quality of generated triplets for different question types and report results in Table 10. Out of the 6 most common categories, we observe that questions related to “Action” lead to the best annotations, “Counting” questions lead to the highest number of QAs unrelated to the video content, and questions related to “Place” lead to the highest number of QA generation errors. Qualitatively, we found that actions are often depicted in the video, while counted quantities (e.g. time, weight, length) mentioned in the speech are hard to guess from the video only.

Appendix B VideoQA architecture

Our architecture, shown in Figure 11, has two main modules: (i) a video-question multi-modal transformer (top) and (ii) an answer transformer (bottom). Details are given next, and further implementation details are given in Section D.

Appendix C Details of the iVQA dataset

The Amazon Mechanical Turk interfaces used for collecting the question and answer annotations, are shown in Figure 14. An emphasis was placed on collecting visually grounded questions about objects and scenes that could not be easily guessed without watching the video, and collecting short answers in order to maximize the chance for consensus between annotators, i.e., having multiple annotators giving exactly the same answer.

C.2 Statistical Analysis

Word clouds for questions and answers in iVQA, shown in Figure 12, demonstrate the relation of iVQA to the domains of cooking, hand crafting and gardening. These word clouds also indicate that questions in iVQA often require spatial reasoning (behind, front, right, left) and temporal understanding (first, end, left, beginning) of the video. The most frequent answer (spoon) in iVQA corresponds to 2% of all answers in the dataset. In contrast, the most frequent answers in other VideoQA datasets account for more than 9% of all answers in these datasets (we have verified this for MSRVTT-QA, MSVD-QA and ActivityNet-QA). As a consequence, the most frequent answer baseline is significantly lower for our iVQA dataset compared to other VideoQA datasets. Figure 13 shows the distributions of question length, answer length, clip duration and clip relative start time in the original video. Clip duration and start time distributions are almost uniform because we randomly sampled them to obtain the clips, which results in a high video content diversity. Answers are in great majority one or two words as a result of our collection procedure.

We observe that 27.0% of questions lead to a perfect consensus among the five answer annotators, 48.4% of questions lead to a consensus among at least four annotators, and 77.3% lead to a consensus among at least three annotators, while only six questions do not lead to a consensus between at least two annotators, justifying the defined accuracy metric. Additionally, 27.5% of questions have two different answers that had a consensus between at least two annotators.

Appendix D Additional experimental details

VideoQA generation. The input sequence to the answer extractor and question generation transformers are truncated and padded up to a maximum of 32 tokens. The question decoding is done with the beam search keeping track of the 4 most probable states at each level of the search tree. We have used the original captions (including stop words) from the HowTo100M dataset and removed word repetitions from adjacent clips.

VideoQA model. We use the following hyperparameters: l=20l=20, t=20t=20, m=10m=10, d=512d=512, dh=2048d_{h}=2048, N=2N=2, H=8H=8, pd=0.1p_{d}=0.1, dq=da=768d_{q}=d_{a}=768, dv=1024d_{v}=1024. The video features are sampled at equally spaced timestamps, and padded to length tt. Sequences of question and answer tokens are truncated and padded to length ll and mm, respectively. Attention is computed only on non-padded sequential video and question features.

VideoQA datasets. For MSRVTT-QA and MSVD-QA, we follow and use a vocabulary made of the top 40004000 training answers for MSRVTT-QA, and all 18521852 training answers for MSVD-QA. For our iVQA dataset and ActivityNet-QA, we consider all answers that appear at least twice in the training set, resulting in 23482348 answers for iVQA and 16541654 answers for ActivityNet-QA.

Training. We use a cosine annealing learning rate schedule with initial values of 5×10−55\times 10^{-5} and 1×10−51\times 10^{-5} for pretraining and finetuning, respectively. For finetuning, we use the Adam optimizer with batch size of 256 and training runs for 20 epochs. The final model is selected by the best performance on the validation set.

Masked Language Modeling. For the masked language modeling objective, a token is corrupted with a probability 15%, and replaced 80% of the time with [MASK], 10% of the time with the same token and 10% of the time with a randomly sampled token. To guess which token is masked, each sequential question output QiQ_{i} of the multi-modal transformer is classified in a vocabulary of 30,522 tokens, and we use a cross-entropy loss.

Pretraining on HowTo100M. For video-text cross-modal matching, we sample one video negative and one text negative per (positive) video-text pair, and use a binary cross-entropy loss. The cross-modal matching module is used to perform zero-shot VideoQA for the variant VQA-T trained on HowTo100M, by computing scores for f(v,[q,a])f(v,[q,a]) for all possible answers aa, for each video-question pair (v,q)(v,q). We aggregate adjacent clips from HowTo100M to have at least 10 second clips and at least 10 narration words.

Appendix E Additional experiments

We define cross-dataset transfer as a procedure where we pretrain our VideoQA model on a VideoQA dataset and then finetune and test it on another VideoQA dataset. The training follows the procedure described for finetuning in Section 4.2. We report results for cross-dataset transfer in Table 11. Note that we do not use MSVD-QA as downstream dataset as its test set has been automatically generated with the same method as MSRVTT-QA. As can be observed, our approach with pretraining on HowToVQA69M significantly outperforms cross-dataset transfer models using the previously largest VideoQA dataset (MSRVTT-QA), or the largest manually annotated VideoQA dataset (ActivityNet-QA), both for the zero-shot and finetuning settings, on all four downstream datasets. We emphasize that our dataset is generated relying on text-only annotations, while MSRVTT-QA was generated using manually annotated video descriptions and ActivityNet-QA was manually collected. These results further demonstrate the benefit of our HowToVQA69M dataset.

E.2 Results for rare answers and per question type

Results for different answers frequencies are presented for the iVQA dataset in Section 6.4. Here, we show results for MSRVTT-QA, MSVD-QA and ActivityNet-QA datasets in Table 12. As for iVQA, we observe that our model pretrained on our HowToVQA69M dataset, after finetuning, shows the best results for quartiles corresponding to rare answers (Q3 and Q4), notably in comparison with the model trained from scratch or the model pretrained on HowTo100M. We also find that our pretrained model, in the zero-shot setting, performs similarly across the different quartiles, with the exception of ActivityNet-QA, which includes in its most common answers yes, no. Note that in order to have a consistent evaluation with other experiments, we keep the same train vocabulary at test time. This implies that a significant part of answers in the test set is considered wrong because the answer is not in the vocabulary. This represents 16% of answers in iVQA, 3% of answers in MSRVTT-QA, 6% for MSVD-QA and 19% for ActivityNet-QA. Note, however, that our joint embedding framework could allow for different vocabularies to be used at the training and test time.

We also present results per question type for MSRVTT-QA, MSVD-QA and ActivityNet-QA in Tables 13 and 14. Compared to the model trained from scratch or the model pretrained on HowTo100M, we observe consistent improvements for most categories.

E.3 Comparison between QA-T and VQA-T on different datasets.

We show in Table 15 that QA-T is a strong baseline compared to VQA-T on existing VideoQA datasets, when both are trained from scratch. However, on iVQA, VQA-T improves more over QA-T than in other datasets, as measured by absolute improvement in top-1 accuracy. This suggests that the visual modality is more important in iVQA than in other VideoQA datasets.