Look Before you Speak: Visually Contextualized Utterances

Paul Hongsuck Seo, Arsha Nagrani, Cordelia Schmid

Introduction

Imagine that you are cooking an elaborate meal, but forget the next step in the recipe – or fixing your car and uncertain about which tool to pick up next. Developing an intelligent dialogue systemoften used interchangeably with the term ‘conversational AI’ that not only emulates human conversation, but also predicts and suggests future actions – not to mention is able to answer questions on complex tasks and topics – has long been a moonshot goal for the AI community. Conversational AI allows humans to interact with systems in free-form natural language, in the same way that we would communicate with one another. This has led to an outpouring of research in NLP focused on conversational agents, ranging from goal-oriented systems for helping with reservations to chit-chat models , both of which are found in modern virtual assistants such as Alexa, Google Assistant and Siri.

Such works, however, are limited to linguistic interactions only. In contrast, human interaction in the physical world is facilitated through multiple modalities (e.g. verbal, visual, haptic), each modality often complementing the other seamlessly. While doing a task, it is often easier to show another person your progress, than to describe it verbally. Hence we argue that a truly intelligent dialogue system would have knowledge of both visual and textual contexts before making its next utterance. Unfortunately, a major challenge for incorporating visual context is a lack of suitable data. Most traditional conversational datasets are solely text based, and notoriously difficult to collect, relying on narrowly constructed ontologies and highly specific domains . More importantly, they do not contain visual information of the surrounding physical environment.

In an attempt to incorporate visual context to dialogue systems, the task of visual dialog was proposed, which requires an AI agent to hold a meaningful dialog with humans given an image or a video . In these works, a dialog history and question are artificially created for each image/video in a dataset, and the goal is then to infer context from history and answer the question accurately. Such datasets, while valuable, are created at great manual effort and contain artificially contrived scenarios, where the dialog history is not naturally present in video. Such datasets are also limited in size.

Unlike such works , we propose to use online videos to learn from naturally co-occurring vision and dialogue in a scalable manner. We note that certain video domains such as narrated instructional videos and lifestyle vlogs are available in huge numbers (e.g. online on video sharing platforms) and are likely to contain narration explicitly linked to the visual content. Given the availability of high quality ASR, this gives us a large amount of readily available paired visual and textual data. We begin by proposing a future prediction task, where the goal is to predict the next utterance in an instructional video, given both visual and textual contexts (Figure 1). The labels for such a task are freely available from the video itself.

As we show in this work, solving such a task requires knowledge of both visual and textual contexts. Leveraging recent advances in multimodal learning, we do so with a two-stream co-attentional transformer based model, each stream encoding a different modality. Our co-attentional model effectively attends to features within each modality, as well as across modalities through lateral self-attention blocks. We demonstrate that using both visual and textual information leads to a large performance gain over using text alone, and additionally, our two-stream co-attentional model outperforms single stream multimodal models. In addition, we show that our model, trained for this future prediction task, can be transferred to other conversational tasks, achieving state-of-the-art performance on various VideoQA benchmarks.

Concretely, we make the following four contributions: (i) We formulate a future utterance prediction (FUP) task which uses both dialogue and vision; (ii) We re-purpose freely available online instructional video datasets to create training and testing benchmarks for this task (HowToFUP); (iii) We propose a new two-stream multimodal video transformer based architecture (CoMVT) which effectively attends jointly over words in text and visual objects and scenes to learn visual-dialogue context; and finally (iv) We show that our model trained on unlabelled instructional videos is also, perhaps surprisingly, able to achieve state-of-the-art performance on a number of downstream vision-language QA datasets, including MSRVTT-QA , MSVD-QA , ActivityNet-QA , and How2QA .

Related Work

Future Utterance Prediction. Predicting future utterances from textual data alone has been widely explored in the NLP community for conversational AI systems. Approaches include hand-crafted rules , example-based agents and modern neural networks , and aim to generate realistic responses for goal-oriented dialog systems or chatbots. Future prediction has also been used as an unsupervised pretraining objective for text corpora, e.g. next sentence prediction in BERT . Unlike these works, we focus on jointly learning from visual context as well as text. Related to our work is the task of scene-aware dialog prediction , where the goal is to answer questions grounded to a video clip input, given a manually created dialog history. A number of works show promising results on this task , however they rely on manually created VQA datasets. Vision and Language Tasks. Popular vision and language tasks include visual question answering , visual dialog , visual captioning , visual grounding and video-text retrieval . There have also been attempts to use transcribed speech in videos as a source of weak supervision , where the goal is to learn a good visual encoder, and consequently such works are largely evaluated only on downstream tasks that involve unimodal video frame inputs. In contrast, we learn an encoder that can effectively learn to co-attend to both vision and text, and is useful for downstream tasks that involve both modalities.

Multimodal Vision-Text Architectures. A large number of multimodal architectures focus on late fusion of modalities, with popular choices being summation, concatenation and canonical correlation analysis . encodes multimodal inputs in a hierarchical structure, where the local context of a video frame is captured by a Cross-modal Transformer via multimodal fusion, and global video context is captured by a Temporal Transformer. Here cross-modal interactions are limited to a single segment only (where the input modalities are aligned), covering a short timespan. This does not allow multimodal interactions between non-aligned inputs - and we note that the content of human speech is not always precisely aligned with its corresponding visual contexts in time . In contrast, our method allows global cross-modal interactions, unconstrained by temporal alignment. More recent works explore deeper interactions between video frame features and features from other modalities, by fusing modalities earlier, at the input level itself. In these works, however, inputs from multiple modalities are fed into a single transformer, with ‘modality specific’ encodings to distinguish between the modalities. In contrast, our two stream transformer decouples within-modality interactions in individual modality streams and allows cross-modality interactions with lateral self-attention blocks. Another differentiater is the fact that all these works operate on scene-level features, while we focus on objects. Thanks to off-the-shelf high-quality object detection and simple bounding box representations, object-centric features have been widely used in a number of image and language problems . In particular, ViLBERT feeds object features and language inputs to a co-attentional transformer. Extending such methods to video, however, is non-trivial, given the number of frames in a video. While object detectors work well on single frames, obtaining reliable tracklet based video features is still an open problem.

Future Utterance Prediction

We begin by proposing a new future utterance prediction task. The goal is to predict the next utterance in a video, given the previous multimodal context (Figure 1). While next utterance prediction can be evaluated as a generative task , we simplify the problem to be selection among pre-collected candidates. Precisely speaking, given a video clip V=(F,W)\mathcal{V}=\left(F,W\right) where F={fi}i=1NfF=\left\{f_{i}\right\}_{i=1}^{N_{f}} is a sequence of video frames and W={wi}i=1NwW=\left\{w_{i}\right\}_{i=1}^{N_{w}} is a sequence of transcribed words, our goal is to select the true next utterance uTu_{\mathcal{T}} from a set of candidates U={ui}i=1MU=\left\{u_{i}\right\}_{i=1}^{M} where T\mathcal{T} is the index of the true element in UU and we set M=100M=100 (Figure 1). Performance is then assessed by ranking the candidates and using popular retrieval metrics.

Our reasons for opting for ranking rather than generation are two fold: (i) Metrics for evaluating language generation (e.g. BLEU , METEOR ) are focused on local matches (n-grams, longest matching sequences, etc). By definition, such metrics are limited to local contexts, and struggle to account for complex sentence structures and word semantics. It is therefore widely accepted that these metrics do not align well with human ratings ; and (ii) The output distribution in sentence generation is multimodal, i.e. the same information can be paraphrased in multiple ways, leading to many correct answers. There is also inherent ambiguity in the future – given the observation of the present, multiple predictions about the future are possible . In language generation tasks, this is handled by collecting multiple ground truth answers, a strategy which is expensive and difficult to scale. In contrast, popular ranking metrics such as recall@kk are better able to assess model performance in tasks with a multimodal output distribution .

We note here that unlabelled videos can be used to generate data for our task in a scalable manner. A list of future candidate utterances can be created automatically, with the positive sample being the next utterance in the video, and making the assumption that randomly sampled utterances from different video clips are likely to be negatives .

In this work, we use the videos from the HowTo100M dataset , as this is a large dataset of 1.2M instructional videos where the speech is usually explicitly linked to the visual content in the video. Examples of future prediction candidates for this dataset can be seen in Figure 3. We use 90% of the videos in HowTo100M for training, and reserve 5% each for validation and test respectively. We name this benchmark How2FUP (more details are provided in Section 5.1).

Model

Our goal is to effectively learn from both vision and text in a video. We propose a Co-attentional Multimodal Video Transformer (CoMVT), which given a video clip V\mathcal{V}, extracts contextualized word embeddings and visual features from transcribed words and video frames respectively, and fuses the extracted features to form a multimodal video feature using a co-attentional transformer. We first describe our network architecture, and then the losses used to train this model to solve the task described above (both network and losses are visually depicted in Figure 2).

Given a sequence of transcribed wordsWe use WordPiece tokenization from the BERT vocabulary. WW, we extract Nw=∣W∣N_{w}=|W| contextualized word embeddings eie_{i} using BERT .

1.2 Visual Input Features

We extract two types of visual features – multi-frame scene level features, and object level features extracted per frame.

Scene-level: We first extract Nf′N_{f}^{\prime} multi-frame features mim_{i} by feeding F={fi}i=1NfF=\left\{f_{i}\right\}_{i=1}^{N_{f}} frames into S3D , a 3D CNN model which has been used in previous approaches in multimodal representation learning . Note that Nf′(≤Nf)N_{f}^{\prime}(\leq N_{f}) is the number of extracted multi-frame features (determined by the stride and temporal downsampling rate of S3D). Similarly to , for every non-overlapping 1-second-long segment of the video, we sample 30 frames and obtain a single feature vector by applying global average-pooling spatiotemporally to the feature activations before the final classifier. This gives us one multi-frame feature mim_{i} per second.

Object-level: While multi-frame features are effective at capturing spatiotemporal dynamics, they limit access to individual concepts or objects by squeezing information into a single vector. To overcome this, we also extract object-level features corresponding to single visual objects in each frame. We first subsample Nf′N_{f}^{\prime} frames building F′={fi′}i=1Nf′F^{\prime}=\left\{f_{i}^{\prime}\right\}_{i=1}^{N_{f}^{\prime}} where each frame fi′f_{i}^{\prime} is temporally aligned with a multi-frame feature mim_{i}. Single-frame object features are then extracted from top-scoring bounding box proposals in each fi′f_{i}^{\prime}. Following , bounding boxes are proposed by a region proposal network (RPN) in and featurized using Graph-Regularized Image Semantic Embeddings (Graph-RISE) . That is, the single-frame object features {oij}j=1L\left\{o_{ij}\right\}_{j=1}^{L} are extracted from each frame fi′f_{i}^{\prime} by

where Bi={bij}j=1LB_{i}=\left\{b_{ij}\right\}_{j=1}^{L} is a set of top LL bounding box in fi′f_{i}^{\prime} proposed by RPN.

Note that oijo_{ij} is object-specific but lacks temporal information whereas mim_{i} encodes temporal dynamics without allowing object-specific access. Therefore, we construct combined spatiotemporal visual features that provide both temporal information and object-specific access by merging these two types of features:

1.3 Co-attentional Transformers

A CoTRM block consists of two streams, each built by stacking two TRM blocks. The first TRM block in each stream takes two multimodal inputs sets: one for queries and the other for keys and values alternating their roles in each stream. The second TRM block is independant within a modality stream.

Formally, given two sets of input features V(s)V^{(s)} and E(s)E^{(s)}, the visual features in V(s)V^{(s)} at sths_{\text{th}} CoTRM block are contextualized by

The two stream nature of CoTRMs inherently treats each modality separately allowing modality-specific operations and representations through different parameterizations of TRMs in the streams.

2 Training Objectives

We train our model with the following two losses: 1) Next Utterance Prediction Loss: Here we treat the textual modality as the main modality and treat e1(S)e^{(S)}_{1} as the embedding of all multimodal inputs. Note that e1(S)e^{(S)}_{1} corresponds to the contextualized embedding of the special ‘[CLS]’ token added to the input. Since our goal is to choose the true next utterance uTu_{\mathcal{T}} from a set of candidate utterances U={ui}i=1MU=\left\{u_{i}\right\}_{i=1}^{M}, we first embed each candidate utterance using an additional BERT encoder and predict the probability of uiu_{i} being the true next utterance P(ui∣e1(S),U)P(u_{i}|e^{(S)}_{1},U) by

2) Masked Language Modelling: In addition to our next utterance prediction loss, we implement the masking scheme and loss function introduced in and mask out some of input words. We apply this loss on downstream evaluations as well. This appears to have a regularisation effect, similar to dropout . We additionally explored visual input masking as in , but found little changes in performance.

Implementation details. We use the BERT base model for the contextualized word embedding extraction and test the proposed networks with S∈{1,2,4}S\in\left\{1,2,4\right\}. RPN , GRAPH-Rise and S3D are initialized and fixed with pretrained weights; all the other parameters are updated during training in all experiments. We set the maximum lengths for the transcribed words NwN_{w} and the downsampled video frames Nf′N_{f}^{\prime} to 128 and 30, respectively, and truncate longer sequences keeping the last elements.

Experiments

We first train our model for Future Utterance Prediction and show results on two datasets, HowToFUP and Coin-FUP. We then take the model pretrained on this task and demonstrate that it generalises well to VideoQA datasets, achieving state-of-the-art results. The input/output configurations for these tasks are described in Appendix B. The next section describes all the datasets used in this work, and then delves into experimental details.

HowToFUP: We repurpose HowTo100M , a large-scale dataset of 1.2M instructional videos for the task of Future Utterance Prediction. Transcripts are obtained using the YouTube ASR API , however these are noisy (Figure B in the Appendix shows an example). Videos that have been taken down from YouTube are not used. We then divide these videos into shorter segments, henceforth referred to as video clips. The duration of video clips is determined as follows: we start with a single ASR sentence and then iteratively expand the length of the video clip backwards by adding previous sentences until the segment is longer than 5 seconds. Each video clip therefore contains full sentences in the ASR (no sentences are cut-off mid way). This process results in 35M training examples and 2M examples each in the validation and test splits. In order to create diverse validation and test sets, we then further reduce the number of clips in each by randomly subsampling 6% of the clips (as many clips contain redundant input contexts). The final validation and test splits consist of 120K clips, and are used for testing all models.

For each video clip, we then create a list of M=100M=100 future utterance candidates through random sampling. M−1M-1 negative candidates are sampled from the entire answer pool to build UU for the test and validation splits.

We note that this dataset is an order of magnitude larger in number of datapoints than existing video captioning datasets, as well as Conceptual Captions , the largest publicly released image captioning dataset widely used for pretraining vision-text models in the image domain , however is noisier due to (i) errors caused by imperfect ASR and (ii) given the ASR is generated from continuous narration, it often consists of incomplete sentences that lack punctuation. A further analysis in provided in Appendix D. COIN-FUP: We also repurpose COIN , another dataset of instructional videos to evalute the task of future utterance prediction. This dataset is smaller, with 12K videos. We follow the same clip generation pipeline used for HowToFUPHowever we do not subsample in the validation and test sets, to maintain a reliable number of samples for evaluation. to create COIN-FUP, resulting in 78K, 5K and 4K examples in the train, validation and test splits respectively.

1.2 Next Step Prediction

COIN-NSP: Unlike HowTo100M, COIN also contains additional manually annotated categorical steps labelled for each video. Hence we also investigate the performance of a related, albeit slightly different task on this dataset – next step prediction (NSP). Unlike FUP, where the goal is to select from a list of utterances in free form natural language, NSP focuses on predicting the next step from a list of pre-defined action categories. Similarly to FUP example generation, we automatically construct COIN-NSP by iterating over each step annotation and extract its precedent multimodal video segment generating 18K training examples and 1K examples for both validation and test splits with 735 step classes (24.5 examples per class on average in the entire dataset). Note that COIN-NSP is smaller than COIN-FUP in the number of examples since there are fewer manual step annotations than total number of utterances. At inference time, we simply replace the candidate selection component in our model with a softmax classifier. CrossTask-NSP: CrossTask is another dataset that contains manual annotations of steps for instructional videos of 18 pre-selected tasks. We create CrossTask-NSP following the COIN-NSP construction process resulting in 14K training examples and 2K validation/test examples with 105 target classes.

1.3 Downstream VideoQA Benchmarks

MSRVTT-QA and MSVD-QA: MSRVTT-QA and MSVD-QA are popular video question answering benchmarks introduced in . We use publicly available features and follow the standard train, val and test splits used in : 158K, 12K and 73K QA pairs for MSRVTT-QA and 31K, 6K and 13K pairs for MSVD-QA.

ActivityNet-QA: ActivityNet-QA contains 58K open-ended QA annotations where the train, val and test splits have 32K, 18K and 8K QA pairs, respectively.

How2QA: How2QA consists of QA annotations for the HowTo100M dataset. It contains 35K train and 3K publicly available val samples. Each question has three negative answers and one correct answer.

2 Baselines

We compare our model to a number of single and multiple modality baselines. S3D - visual only: We use the S3D model pretrained on Kinetics applied to video frames. Text only Baseline: For the text only baseline, we use BERT , which is the winning model in the Eighth Dialog System Technology Challenge for response prediction in text only dialog benchmarks . Single-Stream Multimodal Baseline: We also implement a single stream transformer operating on a single multimodal input stream, which is the most widely used framework for video encoding with multimodal inputs . We adopt the architecture used in and train the network using the same next utterance prediction loss for FUP. Note that this architecture is slightly different from that of BERT, and hence we cannot use pre-trained BERT weights. In addition to our full model, we also show results without the object level features referred to as ‘CoMVT (Scene feats only)’ in Table 1, as this is more similar to previous multimodal models , as well as show the effects without BERT pretraining for the text stream and the MLM loss. For each model including the baselines, we perform grid search on learning rates and report the test performance of the best models in the validation set. On HowToFUP, every network is trained for 2M with a batch size of 512. The learning rate is warmed up for 10K iterations and is continuously decayed per every 30K iterations by the factor of 0.95. On the other datasets, due to the small sizes of the datasets, the models are trained for 20K iterations with a 50 iteration warm-up period and 1K decay length.

3 Results

Table 1 shows the recall at k∈{1,5}k\in\{1,5\} (R@kk) on HowToFUP. All multimodal models outperform text only baselines showing the value of visual inputs for this task. Our best model results in an 8% improvement in R@1. The gain due to visual input is also demonstrated by the examples in Figure 3 and Appendix C.

We next ablate various aspects of our model and training setup. Architecture Components: Table 1 shows the incremental value of different components in our architecture. Our two stream model surpasses the single stream multimodal baseline model while using the same scene-level features. It is also interesting that both BERT pretraining and the masked language modeling loss improve performance even though we train on a large-scale dataset with more than 35M training examples. Finally, the use of the combined features in our full model and additional CoTRM blocks show additional gains. Efficiency: We also analyze the efficiency gains of our compact feature set extraction module (see Table 2 for flopsFlops are measured per sample by profiling evaluation steps on TPUs. and R@kk results). Using our compact feature set module significantly reduces flops by 20.2% and 39.3% with S=2S=2 and 44, respectively, while maintaining performance with S=2S=2. For S=4S=4, we are unable to obtain R@kks without compact feature set extraction on our TPU configurations due to significant memory consumption during training. Effect of Training Data: We also perform an ablation study analysing the effect of training data size on performance. We train on different fractions of the HowToFUP training set (results in Figure 4), and show steep performance drops when the size of the training set is reduced. We also note that the trend points towards linear improvements in R@1 as the training set size is doubled. Given that the performance does not seem to be saturated yet, we hypothesise that further performance gains are possible by scaling up with instructional video data beyond HowTo100M.

Future Utterance Prediction results for COIN are shown in Table 3. Similar trends hold, however we also note that pretraining on HowToFUP provides a massive boost in performance (over 30% value increase in R@1). This significant gain can be explained by the relatively small size of COIN-FUP.

3.2 Next Step Prediction

The results for Next Step Prediction (a classification task) on COIN-NSP and CrossTask-NSP are provided in Table 4. Interestingly, using visual inputs shows more of an improvement over the text only baseline for this task compared to FUP. Note that NSP contains more examples of humans activities (while FUP has a more diverse set of utterances). In addition to the baselines described in Section 5.2, we also compare to two state-of-the-art vision-only models for next step prediction, RULSTM and TAR We reimplemented these models to use our extracted features to provide a fair comparison.

Pretraining on BERT and using the MLM loss particularly help prevent our multimodal model from overfitting on this small dataset. The largest gain comes from pretraining on HowToFUP, which also allows us to increase the value of SS in the model without suffering from overfitting. We also show results of our model using visual inputs only (with pretrained weights), and find that we obtain decent results (this is achieved by feeding in a dummy text input with an empty sentence, i.e., a sequence containing a [CLS] token and a [SEP] token). A qualitative result is shown in Figure 5.

3.3 Transfer Learning to Video QA

We additionally show results of our model pretrained on HowToFUP and then fine-tuned on 4 popular video QA benchmarks, in Table 5 for MSRVTT-QA and MSVD-QA, Table 6 for ActivityNet-QA, and Table 7 for How2QA. For all datasets, our network architecture trained from scratch already outperforms or performs comparably to the existing state of the art; and finetuning from pretrained weights on HowToFUP provides a further boost. We note that our model is pretrained without any QA supervision at all, and generalises well to video QA.

Conclusion

We propose a new visually conditioned Future Utterance Prediction (FUP) learning task, where the goal is to predict the next utterance in an instructional video using both visual frames and transcribed speech. We set benchmarks on both the HowTo100M and COIN datasets, and show state-of-the-art results on downstream video QA benchmarks. We hope that this work will increase interest in the exciting field of visually contextualized dialogue systems.

References

Appendix

We describe further experiments ablating our model (Section A), provide further insight into the experimental set up of the tasks used in the paper (Section B), display more qualitative results (Section C), and analyse the effect of ASR noise in HowToFUP with a brief manual study (Section D).

Appendix A Further Ablations on HowToFUP

We ablate our model described in Section 4 of the main paper, varying (i) the number of CoTRM blocks SS (described in Section 4.1.3 of the main paper), (ii) the MLM loss, and (iii) the visual input feature type (scene features only vs. combined features) in Table A. With scene features only, and without the MLM loss, performance degrades rapidly as SS is increased (almost 4% drop). Using combined features (object and scene) however, prevents this performance drop, as does adding in the MLM loss. Adding both together, gives the best performance with S=4S=4, suggesting that the gains are complementary.

Appendix B Configurations in Different Tasks

In the main paper, we show results for 3 different tasks, Future Utterance Prediction, Next Step Prediction, and Video Question Answering. Here we describe the different setups for each one, in particular the inputs and outputs of our model (depicted visually in Figure A).

In the default configuration for FUP, our model ingests video frames and transcribed speech, and expects a set of next utterance candidates. Our model then returns a score for each candidate computed by a dot product of the input multimodal feature with each candidate.

Since this task is formulated as a classification task, instead of encoding a set of FUP candidates, we use a two-layered classifier for prediction. Inputs are video frames and transcribed speech, and output is a softmax 735-way classifier prediction. Note that the new classifier module in this task cannot be initialized with the pretrained weights on HowToFUP and is trained from scratch.

The goal here is to answer a question given an input video. Compared to the other tasks, there is an input question in addition to video frames and corresponding transcribed speech. We simply concatenate this additional input question to the transcript and feed the concatenated string as a single textual input to our model. Note that some videos do not contain any speech and, in such cases, the model simply takes the question only.

It is common to formulate VideoQA as a classification task using the most frequent answers as target classes. We instead adopt the formulation of answer ranking as in FUB where all possible answers in the training set are encoded using the candidate encoder and scored by the softmax normalized dot-product. In other words, we extract all possible answers from the training set, treat them as candidate answers, and select the best candidate using the same type of a candidate encoder as in FUP.

Appendix C Further Qualitative Results

We present additional qualitative examples for HowToFUP in Figure B and C.

Appendix D ASR error analysis in HowToFUP

It is a well known fact that the HowTo100M dataset is noisy, and because ASR is obtained via an automatic method, there is the potential for error. We briefly investigate the quality of transcripts by manually correcting ASR mistakes in the next utterances of 100 random samples from HowToFUP. We observe a word error rate of 3.3%, and note that at a sentence level - 80.0% of the automatic transcripts are correct while the others contain only small (nominal) mistakes (e.g., want it →\rightarrow wanted). To quantify the impact of these ASR errors on our model performance, we evaluate our model on the 100 samples with and without these corrected transcripts for the task of future utterance prediction. The model selects the identical candidates in both settings except for 2 samples (98 correct). This indicates that the few errors introduced by ASR do not make a significant difference, particularly at scale.