Video Graph Transformer for Video Question Answering

Junbin Xiao, Pan Zhou, Tat-Seng Chua, Shuicheng Yan

Introduction

Since the 1960s, the very beginning of Artificial Intelligence (AI), long efforts and steady progresses have been made towards machine systems that can demonstrate their understanding of the dynamic visual world by responding to humans’ natural language queries in the context of videos which directly reflect our physical surroundings. In particular, since 2019 , we have been witnessing a drastic advancement in such multi-disciplinary AI where computer vision, natural language processing as well as knowledge reasoning are coordinated for accurate decision making. This advancement stems, in part from the success of multi-modal pretraining on web-scale vision-text data , and in part from the unified deep neural network that can well model both vision and natural language data, i.e., transformer . As a typical multi-disciplinary AI task, Video Question Answering (VideoQA) has benefited a lot from these developments which helps to propel the field steadily forward over the use of purely conventional techniques .

Despite the excitement, we find that the advances made by such transformer-style models mostly lie in answering questions that demand the holistic recognition or description of video contents . The problem of answering questions that challenge real-world visual relation reasoning, especially the causal and temporal relations that feature video dynamics , is largely under-explored. Cross-modal pretraining seems promising . Yet, it requires the handling of prohibitively large-scale video-text data , or otherwise the performances are still inferior to the state-of-the-art (SoTA) conventional techniques . In this work, we reveal two major reasons accounting for the failure: 1) Video encoders are overly simplistic. Current video encoders are either 2D neural networks (CNNs or Transformers ) operated over sparse frames or 3D neural networks operated over short video segments. Such networks encode the videos holistically, but fail to explicitly model the fine-grained details, i.e., spatio-temporal interactions between visual objects. Consequently, the resulting VideoQA models are weak in reasoning and require large-scale video data for learning to compensate for such weak forms of input. 2) Formulation of VideoQA problem is sub-optimal. Often, in multi-choice QA, the video, question, and each candidate answer are appended (or fused) into one holistic token sequence and fed to a cross-modal Transformer to gain a global representation for answer classification . Such a global representation is weak in disambiguating the candidate answers, because the video and question portions are the same and large, which may overwhelm the short answer and dominate the overall representation. In open-ended QA (popularly formulated as a multi-class classification problem ), answers are treated as class indexes and their word semantics (which are helpful for QA.) are ignored. The insufficient information modelling exacerbates the data-hungry issue and leads to sub-optimal performance as well.

To improve visual relation reasoning and also reduce the data demands for video question answering, we propose the Video Graph Transformer (VGT) model. VGT addresses the aforementioned problems and advances over previous transformer-style VideoQA models mainly in two aspects: 1) For video encoder, it designs a dynamic graph transformer module which explicitly captures the objects and relations as well as their dynamics to improve visual reasoning in dynamic scenario. 2) For problem formulation, it exploit separate vision and text transformers to encode video and text respectively for similarity (or relevance) comparison instead of using a single cross-modal transformer to fuse the vision and text information for answer classification. Vision-text communication is done by additional cross-modal interaction modules. Through more sufficient video information modelling and more reasonable QA problem solution, we show that VGT can achieve much better performances on benchmarks featuring dynamic relation reasoning than previous arts including those pretrained on million-scale vision-text data. Such strong performance comes even without using external data to pretrain. When pretraining VGT with a small amount of data, we can observe further and non-trivial performance improvements. The results clearly demonstrate VGT’s effectiveness and superiority in visual reasoning, as well as its potential for more data-efficientThe model demands on less training data to achieve good performance. video-language pretraining.

To summarize our contributions: 1) We propose Video Graph Transformer (VGT) that advances VideoQA from shallow description to in-depth reason. 2) We design a dynamic graph transformer module which shows strength for visual reasoning. The module is task-agnostic and can be easily applied to other video-language tasks. 3) We achieve SoTA results on NExT-QA and TGIF-QA that task visual reasoning of dynamic visual contents. Also, our structured video representation gives a promise for data-efficient video-language pretraining.

Related Work

Conventional Techniques for VideoQA. Prior to the success of Transformer for vision-language tasks, various techniques, e.g., cross-modal attention , motion-appearance memory , and graph neural networks , have been proposed to model informative videos contents for answering questions. Yet, most of them leverage frame- or clip-level video representations as information source. Recently, graphs constructed over object-level representations have demonstrated superior performance, especially on benchmarks that emphasize visual relation reasoning . However, these graph methods either construct monolithic graphs that do not disambiguate between relations in 1) space and time, 2) local and global scopes , or build static graphs at frame-level without explicitly capturing the temporal dynamics . The monolithic graph is cumbersome to long videos where multiple objects interact in space-time. Besides, the static graphs may lead to incorrect relations (e.g., hug vs. fight) or fail to capture dynamic relations (e.g., take away). In this work, we model video as a local-to-global dynamic visual graph, and design graph transformer module to explicitly model the objects, their relations, and dynamics, for exploiting object and relations in adjacent frames to calibrate the spurious relations obtained at static frame-level. Importantly, we also integrate strong language models and explore cross-modal pretraining techniques to learn the structured video representations in a self-supervised manner.

Transformer for VideoQA. Pioneer works learn generalizable representations from HowTo100M by either applying various proxy tasks , or curating more tailored-made supervisions (e.g., future utterance and QA pairs ) for VideoQA. However, they focus on answering questions that demand the holistic recognition or shallow description , and their performances on visual relation reasoning remains unknown. Furthermore, recent works reveal that these models may suffer from performance lose on open-domain questions due to the heavy noise and limited data scope of HowTo100M. Recent efforts tend to use open-domain vision-text data for end-to-end learning. ClipBERT takes advantage of image-caption data for pretraining, but it only has limited performance improvement on temporal reasoning tasks , as the temporal relations are hard to learn from static images. In addition, ClipBERT relies on human annotated descriptions which are expensive to annotate and hard to scale up. More recent works collect million-scale user-generated (vastly abundant on the Web) vision-text data for pretraining, but suffers from huge computational cost to train on such large-scale datasets. Two latest works reveal the potential of Transformers for learning on the target datasets (relatively small scale). While promising, they either target at revealing the single-frame bias of benchmark datasets by using image-text pretrained features (e.g. from CLIP ), or only demonstrate the model’s effectiveness on synthesized data . Overall, the poor-dynamic-reasoning and data-hungry problems in existing transformer-style video-language models largely motivate this work. To alleviate these problems, we explicitly model the objects and relations for dynamic visual reasoning and incorporate structure priors (or relational inductive bias ) into transformer architectures to reduce the demand on data.

Graph Transformer. The connection between graph neural networks and Transformer has earned increasing attention . Nonetheless, the major advancements are made in modelling natural graph data (e.g. social connections) by either incorporating graph expertise (e.g., node degrees) into self-attention block of Transformer , or designing transformer-style convolution blocks to fuse information from heterogeneous graphs . A recent work combines graphs and Transformers for video dialogues. Yet, it simply applies global transformer over pooled graph representations built from static frames and does not explicitly encode object and relation dynamics. Our work differs from it by designing and learning dynamic visual graph over video objects and using transformers to capture the temporal dynamics at both local and global scopes.

Method

Given a video v and a question q, VideoQA aims to combine the two stream information v and q to predict the answer a. Depending on the task settings, a can be given in multiple choices along with each question for multi-choice QA, or it is given in a global answer set for open-ended QA. In this work, we handle both types of VideoQA by optimizing the following objective:

in which A\mathcal{A} can be Amc\mathcal{A}_{mc} corresponding to the candidate answers of each question in multi-choice QA, or Aoe\mathcal{A}_{oe} corresponding to the global answer set in open-ended QA. FW\mathcal{F}_{W} denotes the mapping function with learnable parameters WW.

To solve the problem, we design a video graph transformer (VGT) model to perform the mapping FW\mathcal{F}_{W} in Eqn. (1). As illustrated in Fig. 1, at the visual part (Orange), VGT takes as input visual object graphs, and drives a global feature fqvf^{qv} with the integration of textual information, to represent the query-relevant video content. At the textual part (Blue), VGT extracts the feature

representations FAF^{\mathcal{A}} for all the candidate answers via a language model (e.g., BERT ). The final answer a∗a^{*} is determined by returning the candidate answers with maximal similarity (relevance score) between fqvf^{qv} and fa∈FAf^{a}\in F^{\mathcal{A}} via dot-product. At the heart of the model is the dynamic graph transformer module (DGT). The module clip-wisely reasons over the input graphs, and aggregates them into a sequence of feature representations FDGTF^{\text{DGT}} which are then fed to a global transformer to achieve fqvf^{qv}. During training, the whole framework is end-to-end optimized with Softmax cross-entropy loss. For pretraining with weakly-paired video-text data, we adopt cross-modal matching as the major proxy task and optimize the model in a contrastive manner along with masked language modelling .

2 Video Graph Representation

Given a video, we sparsely sample lvl_{v} frames in a way analogous to . The lvl_{v} frames are evenly distributed into kk clips of length lc=lvkl_{c}=\frac{l_{v}}{k}. For each sampled frame (see Fig. 2), we extract nn RoI-aligned features as object appearance representations Fr ⁣= ⁣{fri}i=1nF_{r}\!=\!\{f_{r_{i}}\}_{i=1}^{n} along with their spatial locations B ⁣= ⁣{bri}i=1nB\!=\!\{b_{r_{i}}\}_{i=1}^{n} with a pretrained object detector , where rir_{i} represents the ii-th object region in a frame. Additionally, we obtain an image-level feature FI ⁣= ⁣{fIt}t=1lvF_{I}\!=\!\{f_{I_{t}}\}_{t=1}^{l_{v}} for all the sampled frames with a pretrained image classification model . FIF_{I} serve as global contexts to augment the graph representations aggregated from the local objects.

To find the same object across different frames within a clip, we define a linking score ss by considering their appearance and spatial location:

where ψ\psi denotes the cosine similarity between two detected objects ii and jj in adjacent frames. Intersection-over-union (IoU) computes the location overlap of objects ii and jj. Our experiments always set λ\lambda as one. The nn detected objects in the first frame of each clip are designated as anchor objects. Detected objects in consecutive frames are then linked to the anchor objects by greedily maximizing ss frame by frameWe assume that the group of objects do not change in a short video clip.. By aligning objects within a clip, we ensure the consistency of the node and edge representations for the graphs constructed at different frames.

Next, we concatenate the object appearance frf_{r} and location flocf_{loc} representations and project the combined feature into the dd-dimensional space via

where [;][;] denotes feature concatenation and flocf_{loc} is obtained by applying a 1×11\times 1 convolution over the relative coordinates as in . The function ϕWo\phi_{W_{o}} denotes a linear transformation with parameters WoW_{o}. With Fo ⁣= ⁣{foi}i=1nF_{o}\!=\!\{f_{o_{i}}\}_{i=1}^{n}, the relations in the tt-th frame can be initialized as pairwise similarities:

3 Dynamic Graph Transformer

As illustrated in Fig. 3, the temporal graph transformer unit takes as input a set of graphs GinG_{in} and outputs a new set of graphs GoutG_{out} by mining the temporal dynamics among them via a node transformer (NTrans) and an edge transformer (ETrans). For completeness, we briefly recap the self-attention in Transformer . It uses a multi-head self-attention (MHSA) to fuse a sequence of input features Xin={xint}t=1lX_{in}=\{x_{in}^{t}\}_{t=1}^{l}:

where ϕWc\phi_{W_{c}} is a linear transformation with parameters WcW_{c}, and

where ϕWiq\phi_{W_{i_{q}}}, ϕWik\phi_{W_{i_{k}}} and ϕWiv\phi_{W_{i_{v}}} denote the linear transformations of the query, key, and value vectors of the ii-th self-attention (SA) head respectively. ee denotes the number of self-attention heads, and SA is defined as:

in which dkd_{k} is the dimension of the key vector. Finally, a skip-connection with layer normalization (LN) is applied to the output sequence X=LN(Xout+Xin)X=LN(X_{out}+X_{in}). XX can undergo more MHSAs depending on the number of transformer layers.

In temporal graph transformer, we apply HH self-attention blocks to enhance the node (or object) representations by aggregating information from other nodes of the same object from all adjacent frames within a clip:

transformer is that it models the change of single object behaviours and thus infer the dynamic actions (e.g. bend down). Also, it is helpful in improving the objects’ appearance feature in the cases where the object at certain frames suffer from motion blur or partial occlusion.

Based on the new nodes Fo′={Foi′}i=1nF^{\prime}_{o}=\{F^{\prime}_{o_{i}}\}_{i=1}^{n}, we update the relation matrix RR via Eqn. (4). Then, to explicitly model the temporal relation dynamics, we apply an edge transformer on the updated relation matrices:

3.2 Spatial Graph Convolution

The temporal graph transformer focuses on temporal relation reasoning. To reason over the object spatial interactions, we apply a UU-layer graph attention convolution on all the lvl_{v} graphs:

where W(u)W^{(u)} is the graph parameters at the uu-th layer. II is the identity matrix for skip connections. Fo′(u){F^{\prime}_{o}}^{(u)} are initialized by the output node representations Fo′F^{\prime}_{o} as aforementioned. The index tt is omitted for brevity. A last skip-connection: Foout=Fo′+Fo′(U)F_{o_{out}}=F^{\prime}_{o}+{F^{\prime}_{o}}^{(U)} is used to obtain the final node representations.

3.3 Hierarchical Aggregation

The node representations so far have explicitly token into account the objects’ spatial and temporal interactions. But such interactions are mostly atomic. To aggregate these atomic interactions into higher-level video elements, we adopt a hierarchical aggregation strategy in Fig. 4.

First, we aggregate the graph nodes at each frame by a simple attention:

The set of kk clips are finally represented by FDGT ⁣= ⁣{fcDGT}c=1kF^{\text{DGT}}\!=\!\{f_{c}^{\text{DGT}}\}_{c=1}^{k}.

4 Cross-modal Interaction

To find the informative visual contents with respect to a particular text query, a cross-model interaction between the visual and textual nodes is essential. Given a set of visual nodes denoted by XvX^{v}, we integrate textual information Xq={xmq}m=1MX^{q}=\{x_{m}^{q}\}_{m=1}^{M} into the visual nodes via a simple cross-modal attention:

where MM is the number of tokens in the text query. In principle, the XvX^{v} can be visual representations from different levels of the DGT module similar to . In our experiment, we explore performing the cross-modal interaction with visual representations at the object-level (FOF_{O} in Eqn. (3)), frame-level (FGF_{G} in Eqn. (12)), and clip-level (FDGTF^{DGT} in Eqn. (13)). We find that the results vary among different datasets. As a default, we perform cross-modal interaction at the clip-level outputs (i.e., the outputs of the DGT module Xv:=FDGTX^{v}:=F^{\text{DGT}}), since the number of nodes at this stage is much smaller, and the node representations have already absorbed the information from the preceding layers. For the text node XqX^{q}, we obtain them by a simple linear projection on the token outputs of a language model :

5 Global Transformer

The global transformer has two major advantages: 1) It retains the overall hierarchical structure which progressively drives the video elements at different granularity as in . 2) It improves the feature compatibility of vision and text, which may benefit cross-modal comparison.

6 Answer Prediction

To obtain a global representation for a particular answer candidate, we mean-pool its token representations from BERT by fA=MPool(XA),f^{A}=\text{MPool}(X^{A}), where XAX^{A} denotes a candidate answer’s token representations, and is obtained in a way analogous to Eqn. (15). Its similarity with the query-aware video representation fqvf^{qv} is then obtained via a dot-product. Consequently, the candidate answer of maximal similarity is returned as the final prediction:

in which ⊙\odot is element-wise product. During training, we maximize the ⟨\langleVQ, A⟩\rangle similarity corresponding to the correct answer of a given sample by optimizing the Softmax cross entropy loss function. L=−∑i=1∣A∣yilog⁡si,\mathcal{L}=-\sum\nolimits_{i=1}^{|\mathcal{A}|}y_{i}\log s_{i}, where sis_{i} is the matching score for the ii-th sample. yi=1y_{i}=1 if the answer index corresponds to the ii-th sample’s ground-truth answer and 0 otherwise.

7 Pretraining with Weakly-Paired Data

For cross-model matching, we encourage the representation of each video-text interacted representation fqvf^{qv} to be closer to that of its paired description fqf^{q} and be far away from that of negative descriptions which are randomly collected from other video-text pairs in each training iteration. This is formally achieved by maximizing the following contrastive objective:

where Ni\mathcal{N}_{i} denotes the representations of all the negative video-description pairs of the ii-th sample. The parameters to be optimized are hidden in the process of calculating fqvf^{qv} and fqf^{q} as introduced above. For negative sampling, we sample them from the whole training set at each iteration. For masked language modelling, we only corrupt the positive description of each video for efficiency.

Experiment

We conduct experiments on benchmarks whose QAs feature temporal dynamics: 1) NExT-QA is a manually annotated dataset that features causal and temporal object interaction in space-time. 2) TGIF-QA features short GIFs; it asks questions about repeated action recognition, temporal state transition and frame QA which invokes a certain frame for answer. For better comparison, we also experiment on MSRVTT-QA which challenges a holistic visual recognition or description. Other data statistics are presented in Appendix 0.A.

We decode the video into frames following , and then sparsely sample lv=32l_{v}=32 frames from each video. The frames are distributed into k=8k=8 clips whose length lc=4l_{c}=4. For each frame, we detect and keep N=20N=20 regions of high confidence for NExT-QA (Top-5 are used in the pretraining-free experiments, refer to our analysis in Appendix 0.C.2 ), and N=10N=10 for the other datasets, using the object detection model provided by . The dimension of the models’ hidden states is d=512d=512. The default number of layers and self-attention heads in transformer are H=1H=1 and e=8e=8 (e=5e=5 for edge transformer in DGT) respectively. Besides, the number of graph layers is U=2U=2. For training, we use Adam optimizer with initial learning rate 1×10−51\times 10^{-5} of a cosine annealing schedule. The batch size is set to 64, and the maximum epoch varies from 10 to 30 among different datasets. Our pretraining data (∼\sim 0.180.18M) are collected from WebVid . More details are presented in Appendix 0.B.

2 Sate-of-the-Art Comparison

In Table 1, we compare VGT with the prior arts on NExT-QA . The results show that VGT surpasses the previous SoTAs by clear margins on both the val and test sets, improving the overall accuracy by 1.6% and 1.9% respectively. VGT even outperforms a latest work ATP which is based on CLIP features (VGT vs. ATP: 55.02% vs. 54.3%), and thus sets the new SoTA results. In particular, we note that such strong results come without considering large-scale cross-modal pretraining. When pretraining VGT with (relatively) small amount of data, we can further increase the results to 56.9% and 55.7% on NExT-QA val and test sets respectively (refer to our analysis of Table 5 in Sec. 4.4).

Compared with VQA-T which also formulates VideoQA as problem of similarity comparison instead of classification, VGT outperforms it almost in all metrics. The strong results could be due to that VGT explicitly models the object interactions and dynamics for visual reasoning, instead of holistically encoding video clips with S3D . For a better analysis, we further replace the S3D encoder in VQA-T with our DGT module. As shown in Table 2 (S3D →\rightarrow DGT), our DGT encoder significantly improves VQA-T’s result by 4.7%, in which most of the improvements are from answering reasoning

type of questions. Aside from the DGT module, we encode the candidate answers in the context of the corresponding question with a single language model, whereas VQA-T encodes Q and A independently with two language models . Our method improves answer encoding with contexts and reduces the model size (or parameters), as shown in Table 2 (VGT (DistilBERT)). Finally, VQA-T adopts cross-modal transformer to fuse the video-question pair, whereas we design light-weight cross-modal interaction module. The module is more parameter efficient but has little impact on the performances (CMTrans→\rightarrowCM in Table 2).

Compared with other graph based methods , VGT enjoys several advantages: 1) It explicitly model the temporal dynamics of both objects and their interactions. 2) It solves VideoQA by explicit similarity comparison between the video and text instead of classification. 3) It represents both visual and textual data with Transformers which may improve the feature compatibility and benefit cross-modal interaction and comparison . 4) VGT uses much few frames for training and inference (e.g., VGT vs. HQGA : 32 vs. 256), which benefits efficiency for video encoding. The detailed analyses are given in Sec. 4.3.

In Table 3, we compare VGT with previous arts on the TGIF-QA and MSRVTT-QA datasets. The results show that VGT performs pretty well on the tasks of repeating action recognition and state transition that feature temporal dynamics, surpassing the previous pretraining-free SoTA results significantly by 10.6% (VGT vs. MASN : 95.0% vs. 84.4%) and 6.8% (VGT vs. MHN : 97.6% vs. 90.8%) respectively. It even beats the pretraining SoTA (i.e. MERLOT ) by about 1.0%, yet without using external data for cross-modal pretraining. On TGIF-QA-R which is curated by making the negative answers in TGIF-QA more challenging, we can also observe remarkable improvements. Besides, VGT also achieves competitive results on normal descriptive QA tasks as defined in FrameQA and MSRVTT-QA though they are not our focus.

3 Model Analysis

DGT. The middle block of Table 4 shows that removing the DGT module (w/o DGT) (i.e. directly summarizing the object representations in each clip) leads to clear performance drops (∼\sim2.0%) on all tasks that challenge spatio-temporal reasoning. We then study the temporal graph transformer module (w/o TTrans) by removing both NTrans and ETrans. It shows better results than removing the whole DGT module. Yet, its performances on tasks featuring temporal dynamics are still weak. We further ablate the temporal graph transformer module to investigate the independent contribution of the node transformer (NTrans) and edge transformer (ETrans). The results (w/o NTrans and w/o ETrans) demonstrate that both transformers benefit temporal dynamic modelling. Finally, the ablation study on the global frame feature FIF_{I} reveals its vital role to DGT.

Similarity Comparison vs. Classification. We study a model variant by concatenating the outputs of the DGT module with the token representations from BERT in a way analogous to ClipBERT . The formed text-video representation sequence is fed to a cross-modal transformer for information fusion. Then, the output of the ‘[CLS]’ token is fed to a ∣A∣|\mathcal{A}|-way classifier in open-ended QA or a 11-way classifier for binary relevance in multi-choice QA following . As can be seen from the bottom part of Table 4, this classification model variant (Comp →\rightarrow CLS) leads to drastic performance drops. To be complete, we also conduct additional experiments on the FrameQA task which is set as open-ended QA. Again, we find that the accuracy drops from 61.6% to 56.9%. A detailed analysis of the performances on the training and validation sets (see Appendix 0.C.1) reveals that the CLS-model suffers from serious over-fitting on the target datasets. The experiment demonstrates the superiority of solving QA by relevance comparison instead of answer classification.

Cross-modal Interaction. Fig. 5 investigates several implementation variants of the cross-modal interaction module as depicted in Sec. 3.4. The results

suggest that it is better to integrate textual information at both the frame- and clip-level outputs (CM-CF) for TGIF-QA, while our default interaction at the clip-level outputs (CM-C) brings the optimal results on NExT-QA. Compared with the baselines that do not use cross-modal interaction, all three kinds of interactions improve the performances. We notice that the cross-modal interaction improves the accuracy on TGIF-QA by more than 10%. A possible reason is that the GIFs are trimmed short videos that only contain the QA-related visual contents. This greatly eases the challenge in spatial-temporal grounding of the positive answers, especially when most of the negative answers are not presence in the short GIFs. Thus, the cross-modal interaction performs more effectively on this dataset. The videos in NExT-QA are not trimmed, thereby the improvements are relatively smaller. Base on these observations, we perform cross-modal interaction at both the frame- and clip-level outputs for the temporal reasoning tasks in TGIF-QA, and keep the default implementation for other datasets.

4 Pretraining and Finetuning

Table 5 presents a comparison between VGT with and without pretraining. We can see that pretraining can steadily boost the QA performance, especially on NExT-QA. The relatively smaller improvements on TGIF-QA could be due to that TGIF-QA dataset is large, and has enough annotated data for fine-tuning. As such, pretraining helps little . Besides, we find that finetuning with masked language modelling (MLM) can improve the generalization from val to test set, and thus achieves the best overall accuracy (i.e. 55.7%) on NExT-QA test set. Fig. 6 studies the QA performances on NExT-QA val set with respect to different amounts of pretraining data. Generally, there is a clear tendency of performance improvements for the overall accuracy (Acc@All) when more data is available. A more detailed analysis shows that these improvements mostly come from a stronger performance in answering causal (Acc@C) and descriptive (Acc@D) questions. For temporal questions, it seems that pretraining with more data does not help much. Therefore, to boost performance, it is promising to add more data or explore a better way to handle temporal languages.

5 Qualitative Analysis

In Fig. 7, we qualitatively analyze the benefits of both dynamic graph transformer and pretraining. The example in (a) shows that the model without the DGT module is prone to predicting atomic or contact actions (e.g. ‘grab’) that can be captured at static frame-level. (b) shows that the model without pretraining fails to predict the answer that is highly abstract (e.g. ‘adjust’). Finally, we show a failure case in (c). It indicates that our model tends to predict distractor answers that are semantically close to the questions when the object of interests in the video are small and the detector fails to detect it. Keeping more detected regions could be helpful, but one needs to carefully balance the graph complexity as well as the inference efficiency. Another alternative is to perform modulated detection as in , we leave it for future exploration.

Conclusions

We presented video graph transformer which explicitly exploits the objects, their relations, and dynamics, to improve visual reasoning and alleviate the data-hungry issue for VideoQA. Our extensive experiments show that VGT can achieve superior performances as compared with previous SoTA methods on tasks that challenge temporal dynamic reasoning. The performance even surpasses those methods that are pretrained on large-scale vision-text data. To study the learning capacity of VGT, we further explored pretraining on weakly-paired video-text data and obtained promising results. With careful and comprehensive analyses of the model, we hope this work can encourage more efforts in designing effectiveness models to alleviate the burden of handling large-scale data, and also promote VQA research that goes beyond a holistic recognition/description to reason about the fine-grained video details.

Acknowledgements

This research is supported by the Sea-NExT joint Lab. Major work was done when Junbin was a research intern at Sea AI Lab. We greatly thank Angela Yao as well as the anonymous reviewers for their thoughtful comments towards a better work.

References

Appendix 0.A Data Statistics

The statistical details of the experimented datasets are presented in Table 6. For better comparison with previous works, we focus on the multi-choice QA task in NExT-QA though it has also defined open-ended QA. For TGIF-QA , we also conduct experiments on a latest version which generates more challenging negative answers for each question in the multi-choice tasks. In particular, we further fix the ‘redundant answer’ issue as we find that there are about 10% of questions have redundant candidate answers and some of the candidate answers are even identical to the correct one. The rectified annotations will be released along with the code.

Appendix 0.B Implementation Details

For training with QA annotations, we firstly train the whole model (except for the object detection model) end-to-end, and then freeze BERT to fine-tune the other parts of the best model obtained at the 11st stage. The best results in the two stages are determined as final results. Note that our hyper-parameters are mostly searched on the NExT-QA validation set and kept unchanged for other datasets. The maximum epoch varies from 10 to 30 among different datasets. For pretraining with data crawled from the Web, we randomly select 0.18M video-text data (less than 10%) from WebVid2.5M https://m-bain.github.io/webvid-dataset/ . The videos are then extracted at 5 frames per second and are processed in the same way as for QA. We then optimize the model with an initial learning rate of 5×10−55\times 10^{-5} and batch size 64. The number of negative descriptions of a video for cross-modal matching is set to 63, and they are randomly selected from the descriptions of other videos in the whole training set. Besides, a text token is corrupted at a probability of 15% in masked language modelling. Following , a corrupted token will be replaced with 1) the ‘[MASK]’ token by a chance of 80%, 2) a random token by a chance of 10%, and 3) the same token by a chance of 10%. We train the model by maximal 2 epochs which gives to the best generalization results, and it takes about 2 hours.

Appendix 0.C Additional Model Analysis

To study the reason for the poor performance of the classification model variant described in Sec. 4.3 of the main text, we visualize the training and validation accuracy with regard to different training epochs in Fig. 8. The results indicate that the classification model variant suffers from serious over-fitting issues, especially on NExT-QA whose QA contents are relative complex but with less training data. To study whether the problem comes from the classification formulation or the cross-modal transformer, we further substitute the cross-modal transformer (CM-Trans) with our cross-modal interaction (CM) module introduced in Sec. 3.4 of the main text. We find that such a substitution can slightly alleviate the problem. For example, on NExT-QA val set, the accuracy increases from 45.82% to 46.98%. Nevertheless, the performance is still much worse than a comparison-based model implementation (i.e. 55.02%). This experiment reveals two facts: 1) Formulating QA problem as classification is the major cause for the weak performance. 2) The cross-modal transformer exacerbates the over-fitting problem, possibly because it involves additional parameters.

C.2 Study of Video Sampling

In Fig. 9, we study the effect of sampled video clips and region proposals on NExT-QA test set. Regarding the number of sampled video clips, we find that the setting of 8 clips steadily wins on 4 clips. This is understandable as the videos in NExT-QA are relatively long. As for the sampled regions, when learning the model from scratch, the setting of 5 regions gives relatively better result, e.g., 53.68%. Nonetheless, when pretraining are considered, the setting of 20 regions gives better result, e.g., 55.70%. Such difference could be due to that learning with more regions can yield over-fitting issues when the dataset is not large enough, since the constructed graph become much larger and more complex. Our speculation is also supported by the fact that the accuracy increases with the number of sampled regions when we only sample 4 video clips and thus less number of total graph nodes.

C.3 Model Efficiency

We compare VGT with VQA-T in Tab. 7 for better understanding of the memory and time cost. Experiments are done on 1 Tesla V100 GPU with batch size 64. We use 1 example to report inference FLOPs. Memory: VGT has less training parameters (133.7M vs. 156.5M) and thus smaller model size than VQA-T (511M vs. 600M). The BERT encoder in VGT takes 82% of the parameters, the vision part is lightweight with only 24M parameters. VGT needs more GPU memory for training. Yet, the memory for inference are fairly small and close to that of VQA-T. We also implement a smaller version of VGT by replacing BERT with DistilBERT as in VQA-T. With nearly 0.6×\times number of VQA-T’s parameters (90.5/156.5M), we can still achieve strong performances (i.e. 53.46%). Time: Our FLOPs on 1 example is ∼\sim2.9×\times that of VQA-T and ∼\sim1.6×\times if we use DistilBERT. However, VGT converges much faster and needs much fewer epochs (total FLOPs) to get results superior to VQA-T when training with the same data. For example, on NExT-QA, VGT’s result at epoch 2 (50.16%) already significantly surpasses VQA-T’s best result (45.30%) achieved at epoch 8. Also, VGT’s result without pretraining can surpasses that of VQA-T pretrained with million-scale data. In this sense, VGT needs much fewer total FLOPs than VQA-T and other similar pretrained models for visual reasoning.