DualVGR: A Dual-Visual Graph Reasoning Unit for Video Question Answering
Jianyu Wang, Bing-Kun Bao, Changsheng Xu
I Introduction
VideoQA is a challenging and high-level multimedia task , which requires the agents to understand videos and perform relational reasoning according to questions based on visual, textual, as well as spatial-temporal contents. Since the input of VideoQA is a sequence of frames, there are two differences between ImageQA and VideoQA: (1) In addition to appearance information, VideoQA also needs to understand the motion information to answer the questions. (2) VideoQA requires to perform spatio-temporal reasoning over the objects, while ImageQA only requires spatial reasoning over the objects. Therefore, scene graphs and neural-symbolic reasoning frameworks , which are used in ImageQA, are hard to be implemented in VideoQA as the agents have to address the issues of comprehensive representation (e.g. appearance and motion information) and multi-step reasoning.
In order to solve the above challenges, this work focuses on performing multi-step reasoning via Graph Networks for VideoQA. Previous reasoning-based methods can be divided into four categories based on their frameworks. The first group implements spatial and temporal attention mechanism to iteratively select useful information to answer the questions. The second group focuses on memory-based network, which is quite popular in TextQA. However, these methods neglect the visual relation information when performing multi-step reasoning. The third group aims to perform relational reasoning via a simple module, like relation network. However, this module can only model a limited number of objects. The fourth group aims to use graph neural networks to integrate relation information into their frameworks by considering GNN’s powerful representation ability on relation modeling. GNN is a novel relation encoder that captures the inter-object relations beyond static object/region detection, thus enabling reasoning with rich relational information. Compared with relation network, graph neural network is more flexible and powerful in relational reasoning, thus we follow this group in our work.
However, existing graph-based methods neglect two key attributes of VideoQA task when performing reasoning: (1) Not all video shots or objects are correlated to the question. As illustrated in Fig. 1., to answer Question1: Who is a dog that looks like a panda walking with?, only question-related video clips are needed to infer the answer. (2) Appearance and motion features are associated and complementary to each other in each reasoning step. Using Fig. 1 as an example again, when answering Question2: What is a small dog which looks exactly like a panda running rapidly on?, it is necessary for agents to utilize both appearance and motion information. Specifically, in order to understand a small dog which looks exactly like a panda, agents need to rely on appearance information to find the dog. Then, they have to analyze the motion information to infer the action of the dog: running rapidly on the road. In short, The appearance information could provide clues to pay attention to certain motion information to answer the questions, vice versa.
Motivated by these observations, we devise a novel graph-based reasoning unit named Dual-Visual Graph Reasoning Unit (DualVGR), which is stacked in an iterative manner to perform multi-step reasoning. At first, we design Query Punishment Module to generate query-guided masks to allow a limited number of question-related video features into relational reasoning. Then, in order to fully capture the multi-view visual relation information, we propose Video-based Multi-view Graph Attention Network to reveal the relation between appearance and motion channels. The proposed graph network includes two graphs for each visual channel. The first graph aims to learn the underlying complementary relation within each specific visual space, and the other is to learn the concomitant and correlated relation between appearance and motion features. The losses of video-based multi-view graph network are consistency constraint loss and disparity constraint loss. Consistency constraint loss is to enhance the commonality between appearance and motion features, while disparity constraint loss is to enhance the heterogeneity between them.
The main contributions of this work can be summarized as follows. (1) We propose a DualVGR unit, a multimodal reasoning unit enabling the represention of rich interactions between question and video clips. This unit includes two components. One is an explainable Query Punishment Module to filter out irrelevant visual-based information. It has been demonstrated good results in both short and long compositional questions during multiple cycles of reasoning. The other is a Video-based Multi-view Graph Attention Network to perform spatial-temporal relational reasoning such that the relations between appearance and motion can be adaptively captured; (2) We incorporate our DualVGR unit into a full DualVGR network by stacking it in an iterative manner. Through multi-step reasoning, our method achieves state-of-the-art or competitive results on several mainstream datasets: MSVD-QA , MSRVTT-QA and SVQA .
II Related Work
In this section, we review the recent studies related to visual question answering, which can be divided into two sub-categories: image question answering and video question answering.
There are three research directions about image question answering. The first line, namely monolithic method , locates the most relevant visual region of images based on attention mechanism, then projects visual features and textual features together into a common latent space through a single step. Antol et al. combine the visual features and the textual features via multimodal pooling, such as addition and concatenation, then map them into a unified space. However, multimodal pooling methods do not well capture the complex associations between two modalities due to their different distributions. Therefore, some approaches related to Bilinear Pooling, which has been used to integrate different CNN features for fine-grained image recognition , have been proposed. However, Bilinear Pooling needs a huge number of parameters and the dimensionality of output feature is usually too high. Some extended versions are proposed to handle these issues, such as Multimodal Compact Bilinear Pooling (MCB) , Hadamard Product for Low-rank Bilinear Pooling (MLB) , Multimodal Factorized Bilinear Pooling (MFB) and Multimodal Factorized High-order pooling (MFH) . Monolithic methods have been demonstrated useful in some cases, but failed to perform well in long and compositional questions.
The second line focuses on modeling the multi-step interaction process through a recurrent cell . Nguyen and Okatani , Yang et al. stack the attention layers to model the multi-step interaction process. Xiong et al. introduce a novel dynamic memory network to iteratively retrieve meaningful visual contents. Meanwhile, some approaches aim to integrate relational reasoning in each step, especially graph-based methods. Li et al. model multi-type inter-object relations via a graph attention mechanism to learn question-adaptive relation representations.
The third line is neural-symbolic reasoning. This kind of methods decompose the whole task into some subtasks, and design several Neural Module Networks (NMN) to solve those subtasks. Andreas et al. firstly utilize an off-the-shelf parser to decompose the compositional question into logical expressions, then construct a monolithic network for each subtask. Hu et al. claim that a brittle off-the-shelf semantic parser would lead to bad performance, hence they design a seq-to-seq RNN to end-to-end predict instance-specific layout. These kinds of methods receive good performances on synthetic datasets , as the questions are easy to parse into subtasks. Besides, Yi et al. propose a CLEVRER dataset for exploring the problem of temporal and causal reasoning in videos. This dataset further promotes the development of the neuro-symbolic reasoning research in video-related tasks.
II-B Video Question Answering
Different from ImageQA, VideoQA is more challenging as videos contain much more complex patterns than a single image. To tackle this problem, agents have to fully comprehend the temporal structure of videos. Jang et al. encode appearance feature and motion feature with ResNet and C3D respectively, then design a spatial and temporal attention mechanism to select different regions during multi-step reasoning. Later, with the success of Dynamic Memory Network in ImageQA, some work applies memory-based network to VideoQA tasks. Xu et al. propose an Attention Memory Unit (AMU), which gradually refines its attention over the appearance and motion features by using question as guidance. Considering that motion and appearance features are correlated in the reasoning process, Gao et al. propose a motion-appearance co-memory network and use a temporal convolutional and deconvolutional neural network to generate multi-level contextual facts. However, their methods do not consider the relations between different objects.
Relation information is obviously an important clue in the reasoning process. Le et al. propose a Conditional Relation Network (CRN) to model the relations between visual objects, and stack CRNs to perform multi-step relational reasoning. However, the limitation of relation network-based methods is that they can only process a limited number of objects at a time. Therefore, graph neural network comes into researchers’ sight. Huang et al. point out that previous work neglects the interactions among objects in each frame and design a location-aware graph convolutional network to learn a relation-enhanced visual feature encoder. Jiang et al. propose an undirected heterogeneous graph with each video snippet and question word as node to integrate correlations of both inter- and intra-modality in an uniform module. In these graph-based methods, all the objects, even those irrelevant to the question, are engaged into graph construction without discrimination, which could bring lots of uninformative noises into relational reasoning.
III Our Approach
The task of VideoQA can be described as follow. Given a video and the corresponding question , agents aim to infer the answer correctly. Formally, the prediction ã is given by classification scores:
where is our trainable model parameters.
The proposed DualVGR framework is depicted in Fig. 2. Firstly, each video is divided into several clips, and we extract appearance and motion features for each clip. Meanwhile, the corresponding question is embed with BiLSTM. Secondly, appearance features, motion features and the corresponding question features are input into the our stacked DualVGR unit iteratively. We use each DualVGR unit to determine the attention within video guided by the question, and model the relational reasoning by stacking the DualVGR units. Thirdly, we fuse appearance and motion features of each clip by Multimodal Factorized Bilinear pooling (MFB) , and utilize attention mechanism to fuse features of all clips to obtain the final visual vector. Finally, we concatenate the visual vector with our question vector to predict the answers.
III-B The proposed DualVGR
To answer the question based on the given video, the designed agent needs to have the ability of locating on the question-related video clips, understanding the video with both appearance and motion channels, and performing multi-step reasoning. The proposed DualVGR Unit determines the question-related video clips, mines the relation within/between appearance and motion, and combines complementing relation features with corresponding visual features, then stacks the DualVGR units for multi-step relational reasoning. It consists of two modules, Query Punishment Module and Video-based Multi-view Graph Attention Network, as shown in Fig. 3. Query Punishment Module is designed to filter out the irrelevant video clips by employing query-guided mask on video features. Video-based Multi-view Graph Attention Network is to reveal the underlying complementary relations within appearance and motion by constructing appearance-specific and motion-specific graphs, and reveal the correlated relations between appearance and motion by constructing appearance-motion correlation and motion-appearance correlation graphs.
Only query-related information is needed to infer the correct answer. Moreover, when perform multi-step reasoning, people have a propensity to pay attention to different textual parts of the question, especially the long and compositional questions. Therefore, we propose a Query Punishment Module to mimic human’s step-to-step reasoning behavior, and filter out irrelevant visual features in each reasoning step. Specifically, based on previous word-level contextual word embedding and initialized word embedding , we employ a self-attention mechanism to obtain the current step’s question feature .
III-B2 Video-based Multi-view Graph Attention Network
As mentioned before, both appearance and motion channels are crucial for video understanding. In order to fully reveal the complemented information from these two channels, we need not only to extract the appearance and motion features from video clip itself, but also to consider the relations among video clips within each channel as well as the relations between two channels for each video clip. To this end, our work tries to aggregate the within-channel and between-channel relations into appearance and motion features. Inspired by GAT , which updates node representation over its neighbors with self-attention mechanisms, and has the ability to deal with both transductive learning and inductive learning problems, especially arbitrarily structured graph learning problems, we construct appearance graph with node as appearance feature of each clip, and motion graph with node as motion feature of each clip, then follow the proposed self-attention strategy in GAT to encode the neighborhood-relation into node representation for these two graphs. In this way, the within-channel relations are aggregated into appearance and motion features respectively. For between-channel relations, we follow the idea of AM-GCN , which is a new type of GCNs for graph classification task that can optimally integrate node features and topological structures by adaptively fusing specific and common embeddings from node features, topological structures and their combinations rather than simple GNN whose capability in extracting deep correlation information between topological structures and node features is distant from optimal, to seek one specific embedding and one common embedding for each channel by performing graph convolution operation. The specific embedding of one channel remains the specific characters of this channel, which should be disparity from that of the other channel. Two common embeddings model the correlated information from both channels, thus they should be consistency with each other.
For appearance channel, we construct two undirected complete graph networks, including appearance independent graph (AIG) and appearance-motion correlation graph (AMC), with each query punished clip-based appearance feature as a node. AIG aims to learn specific spatial-temporal contextual embeddings within appearance channel. AMC is designed to extract the correlated relation features shared by appearance and motion channels. For motion channel, we also utilize two graph networks to learn the contextual embeddings, including motion independent graph (MIG) and motion-appearance correlated graph (MAC). Their goals are the same with graphs in appearance channel. According to these two motion graphs, we can get the specific embedding and the correlated embedding . Each graph is implemented with a multi-head Graph Attention Network (GAT) to model the relations between clip-features. Specifically, the attention score of each head between two nodes is given by:
Once the multi-head attention scores for each node are obtained, we update the representation of each node by:
Finally, we propose two losses to enhance the multi-view representation ability of our graphs, that is, consistency constraint and disparity constraint. Consistency constraint is to enhance the commonality between the correlated contextual information of appearance and motion spaces. Disparity constraint to enhance the independence of specific embeddings and correlated embeddings. The formulations of these two constraints are listed as follows.
Consistency Constraint: For two output embeddings and , we design a consistency constraint for our multi-view task, which could enhance their commonality. We first normalize the final embedding matrix and into and respectively. For simplicity, here we use to represent , and to represent . Then, the similarity of embeddings can be calculated by:
The consistency indicates that the two similarity matrices should be similar as much as possible. The consistency loss of a single unit can be represented as:
Then the consistency loss of the whole network is:
where is the number of iteration steps.
Disparity Constraint: Since specific embeddings and common embeddings, such as and , are learned from graphs with the same topology, to make sure that these two embeddings could capture different information, we employ the Hilbert-Schmidt Independence Criterion (HSIC) to enhance the independence of these two embeddings. HSIC, a simple yet effective measure of independence, has been implemented to several machine learning tasks . For appearance, we use to represent and to represent for simplicity. The HSIC constraint of and in a single unit is defined as:
Similarly, the disparity constraint of embeddings and in motion space can be given by:
Finally, the disparity constraint for the whole network is given by:
where is the number of iteration steps.
III-B3 Attention
Now we have two embeddings and in appearance space, and two embeddings and in motion space. In order to adaptively capture the contextual information for each space, we utilize attention mechanism to fuse two embeddings in appearance space, as well as those in motion space. Their corresponding attention importance scores and are given by:
Similarly, we can get the attention score for and the attention score for . Then, we normalize the attention values with softmax function to get the final score:
Finally, we obtain the final embedding in appearance space by combining these two embeddings and , and the final embedding in motion space by combining and :
Then, a residual connection is used to avoid the vanishing gradient problem. This operation can also be considered as an amalgam of complementing factors including appearance, motion and query-related visual relation information. Each clip-based feature is updated by:
III-B4 Multi-step Reasoning
Finally, the DualVGR unit is stacked as a chain to perform the final DualVGR network:
Through multiple steps of iteration, our DualVGR’s final visual representation contains complementary information about the question-related clips, including appearance, motion and the corresponding relations between them.
III-C Video Representation Fusion
III-D Answer Decoder
Following , we adopt the same answer decoders of these open-ended questions for the fair comparison:
III-E Total Loss
We cast our open-ended VideoQA task as a classification task. Hence, we use cross-entropy loss for this task. Then, the total loss is:
where and are parameters of the consistency and disparity constraint terms. The consistency constraint is implemented as Eqn. (10) and disparity constraint is implemented as Eqn. (14).
IV Experiments
1) MSVD-QA : There are trimmed videos collected from the Microsoft Research Video Description (MSVD) Corpus . It contains 50,500 QA pairs automatically generated by the NLP algorithm in total, which contains five general types of questions, including what, how, when, where and who. The average video length is approximately 10 seconds, and the average question length is approximately 6 words. Therefore, it is a small-scale dataset with short questions. We conduct experiments on this dataset to test our model’s generalization ability in short videos in real worlds. The details of MSVD-QA dataset are illustrated in Table I.
2) MSRVTT-QA : Compared with MSVD-QA, MSR-VTT contains longer videos and more complex scenes. We conduct experiments on this dataset to test our model’s performance for longer videos of real datasets. It contains 10,000 trimmed videos from MSR-VTT dataset and 243,000 QA pairs generated by the NLP algorithm. The average video length is approximately 15 seconds, and the average question length is approximately 7 words. The details of MSRVTT-QA are illustrated in Table II.
3) SVQA : It is a large-scale synthetic dataset which contains 12,000 synthetic videos and around 120K QA pairs. Specifically, videos are generated from Unity3D, and each video length is the same as 10 seconds. Meanwhile, QA pairs are generated from question templates automatically, and the questions are generated exclusively long with an average length of 20 words. Furthermore, each question can be decomposed into human readable logical chain or tree layout easily. The goal of this dataset is to test the reasoning ability of VideoQA systems. Table III illustrates the statistics of SVQA dataset.
IV-B Implementation Details
For each video in MSVD-QA, we segment the video into clips. For MSRVTT-QA, each video is divided into clips. Besides, in SVQA dataset, each video is splitted into clips. The numbers of divided clips are determined by grid search method from . For all the datasets, each clip contains frames by default. The video appearance and question encoders are one-layer BiLSTMs. The dimension , , and the number of the multi-heads is . The iteration step for MSVD-QA is set to , the iteration step for MSRVTT-QA is , and the iteration step for SVQA is . For loss function, we use the grid search method to select the best coefficients. Specifically, the value set of is , while the value set of the remaining coefficient is . After searching, we obtain the best coefficients for all the datasets: = , = in MSVD-QA, = , = in MSRVTT-QA and = , = in SVQA. Our framework is implemented in PyTorch, and the network is trained by Adam optimizer with a fixed learning rate . The batch size is set to . All experiments are terminated after epochs and the results are reported at the epoch which has the best validation accuracy.
IV-C Comparison with the State-of-the-art
We compare our proposed model with state-of-the-art methods (SOTA) on aforementioned datasets. For MSVD-QA and MSRVTT-QA, we compare with most resent SOTA, including HME , HGA , HCRN and TSN .
HME is a model equipped with memory network. It uses a redesigned question memory to improve the question representation in each step. The visual representations, including appearance and motion features, are mapped into the heterogeneous memory. Then, multi-step reasoning is performed with self-updated attention in this memory network.
HGA is a model with graph network. It represents all video shots and question words as the graph to perform cross-modal reasoning.
HCRN is a model with stacked clip-based relation networks. The CRN takes input as an array of tensorial objects and a conditioning feature, and output as an array of relation information within them. By hierarchically stacking these blocks, HCRN performs multi-step relational reasoning.
TSN is a model containing several modules to perform multi-step reasoning. For example, since appearance and motion play different roles in multi-step reasoning, a switch module has been proposed to adaptively choose appearance or motion channel as the primary channel, guiding the reasoning process.
The results are summarized in Table IV for MSVD-QA and MSRVTT-QA. It is clear that our framework consistently outperforms or is competitive with SOTA models on all tasks for short questions (average question length words). For MSVD-QA dataset, our DualVGR achieves overall accuracy, which is improvement over previous SOTA methods. It also achieves quite high performance in each question types. For instance, as shown in Table IV, DualVGR achieves approximately 3% improvement for the question of “what” and “who”. The improvement of DualVGR compared with HME and TSN confirms the effectiveness of relational reasoning in VideoQA tasks. Furthermore, the improvement of DualVGR compared with HGA shows that the critical role of query punishment module in our framework.
For MSRVTT-QA, our model achieves accuracy, which is only lower than the current state-of-the-art performance. Compared with SOTA method HCRN which conducts the question-aware frame-level feature and the question-aware clip-level feature in a hierarchical structure, our method only extracts question-aware clip-level feature, which reduces the computational cost, with only a minimum drop of performance even in dataset with complicated scenes. To sum up, MSVD-QA and MSRVTT-QA results imply that our model can handle real-world videos well.
We further compare our methods with SOTA methods on synthetic dataset, SVQA. This dataset contains many compositional questions, which require agents to be able to perform multi-step relational reasoning to infer the answer. For SVQA, three video question answering models as well as ours are used for comparisons:
Unified-Attn is a model with two attention mechanisms, including sequential video attention and temporal question attention mechanisms.
SA+TA-GRU is a model with a refined GRU whose hidden state transfer process is associated with temporal attention to strengthen long-term temporal dependency.
STRN is a model with spatio-temporal relational network, which aims to model temporal changes in both the interactions among different objects and the motion-dynamics of individual objects.
The results are summarized in Table V. From the results, we can observe that our proposed model DualVGR outperforms the state-of-the-art methods. Specifically, our framework achieves the best overall accuracy , leading to approximately improvement over the best compared method STRN. Our method achieves promising results on all question types in SVQA dataset, especially improvement on “Count” questions. On the contrary, previous methods perform poorly on “Count” questions, which indicates the powerful generalization ability of our method for multi-step reasoning. The aforementioned results indicate the effectiveness of the DualVGR for long and compositional questions. Furthermore, the method STRN is implemented with relation network, which means that DualVGR unit with graph network may be more suitable for relational reasoning task in VideoQA.
IV-D Ablation Study
To prove the effectiveness of the essential components of DualVGR unit and to provide more detailed parameter analysis, we conduct extensive ablation studies on MSVD-QA test set and SVQA test set with a wide range of configurations. The results of the component analysis are reported in Table VI.
The whole architecture of our DualVGR unit mainly incorporates two essential components: Query Punishment Module and Video-based Multi-view Graph Attention Network. We consider the following ablation models to verify the importance of each component in VideoQA:
AG: this model extracts appearance feature from video clips, and constructs them as the appearance graph. Then, multi-head GAT is implemented to get the contextual representation. At last, graph attention readout mechanism is used to get the final visual vector, and fuse it with the question vector to infer the answer.
MG: compared with AG, this model extracts motion feature from video clips, and constructs them as the motion graph.
FG: compared with AG, this model extracts both appearance and motion features from video clips, and fuses them into a new clip-based visual vector with MLP. Then, the visual graph is constructed with them.
Bigraph: compared with AG, this model constructs both appearance graph and motion graph. Then, we apply two multi-head GAT to capture contextual representations of both visual spaces in each unit. Next, appearance and motion features of each clip are fused with MFB.
MVgraph: compared with Bigraph, this model uses Video-based Multi-view Graph Attention Network to learn the contextual representations of appearance and motion features.
PFG: compared with FG, this model implements Query Punishment Module to keep the relevant visual features. Then, multi-head GAT is utilized to represent the contextual information.
DualVGR: this is our proposed model. Compared with MVgraph, this model adds Query Punishment Module in our unit.
sharedVGR: compared with DualVGR, this model use one shared multi-head GAT to learn the contextual representations of both visual spaces, including appearance and motion spaces.
The results are shown in Table VI, and the key observations and conclusions are as follows: (1) Two-stream Visual Features: AG and MG just consider one type of visual information, such as appearance and motion features. FG combines the appearance and motion features into a new visual space, and learn the contextual representations of this space. After considering the appearance and motion features, it increases by and on SVQA dataset, and and on MSVD-QA dataset. This illustrates the effectiveness of utilizing two-stream visual features. Furthermore, as we can see, AG outperforms MG in MSVD-QA and MG outperforms AG in SVQA. This can be attributed to the differences between MSVD-QA and SVQA. For instance, questions in SVQA always require agents to understand the spatial-temporal relationships between all clips, hence motion features might be more important than appearance features. Questions in MSVD-QA are usually very simple, which only require agents to analyze certain clips to answer the questions. Therefore, motion features would be less important in MSVD-QA. (2) Multi-view Graph Attention Network: Bigraph utilizes two multi-head GAT to learn the contextual representations of visual features. MVgraph implements our Multi-view Graph Attention Network to learn the contextual representations. The results imply that our Multi-view Graph Attention Network is more suitable for this task, especially multi-view relation learning. However, we observe that MVgraph underperforms the FG in both datasets. The detailed analysis will be illustrated later. (3) Query Punishment Module: PFG further considers Query Punishment Module in our FG. DualVGR is our whole unit with Query Punishment Module and Video-based Multi-view Graph Attention Network. After the consideration of involving Query Punishment Module to our unit, it increases and on SVQA dataset, and and on MSVD-QA dataset, which proves the advantage of our Query Punishment Module to filter our irrelevant information. Moreover, DualVGR outperforms PFG, which is distinct from what was discussed above. This could be putted down to that MVgraph and FG just represent all video clips as useful node to propogate information with GAT which could lead to much noise. Consequently, with more graphs in MVgraph, it contains more noise than FG, which leads to worse performance. With Query Punishment Module, our DualVGR successfully achieves better performance than PFG. shareDVGR only uses one shared Multi-head GAT to extract the common contextual information in both visual spaces. We can observe from the results that with two common graphs in the unit, we obtain better results than one shared graph. (4) Self-Attention for Questions: simpleDualVGR is the DualVGR framework without self-attention mechanism in each reasoning step. We can observe that with self-attention mechanism, our DualVGR obtains performance gain in SVQA dataset. It verify the effectiveness of self-attention mechanism when performing multi-step reasoning for long and compositional questions.
IV-D2 The detailed analysis of iteration steps T𝑇T
We perform the detailed analysis of the iterative process in Fig. 4. We train our DualVGR network on all datasets with the configuration as our best model. First, questions in MSVD-QA dataset are usually very short, which implies that agents do not need to perform very complex relational reasoning to fully understand the questions. Consequently, the overall accuracy on MSVD-QA decreases when performing more steps of relational reasoning. Next, the test accuracy on MSRVTT-QA dataset increases from to when the number of reasoning iteration increases from to . Since MSRVTT-QA dataset has more complex questions and longer videos than MSVD-QA, it could benefit from a higher number of reasoning iterations over our DualVGR unit. Then, when the iteration step is set to , the model’s performance becomes very stable. At last, SVQA is a dataset designed for testing the relational reasoning ability of models. Therefore, the questions are compositional and complex which require agents to perform multi-step relational reasoning. The results reveal that the accuracy of SVQA dataset also benefits from a higher number of iteration steps. Network with four steps provides a gain of in overall test accuracy on SVQA over the network with a single step. Then, when the step increases by , the model’s performance becomes very stable.
IV-D3 The detailed analysis of the number of clips N𝑁N
Since questions in SVQA are compositional, requiring agents to understand the whole video contents and spatial-temporal relations between relevant objects in videos, the key information of visual facts may distribute evenly in the videos. Therefore, we also perform a detailed analysis of the number of divided clips of SVQA dataset. We train several networks in three iteration steps on SVQA, the results are illustrated in Fig. 6.
Networks with and clips respectively provide a gain of and in overall accuracy on SVQA test set over the network with clips. This implies that the performance of spatial-temporal relational modeling via Graph Neural Network could benefit from a larger number of divided clips. With more divided clips, the key information in videos would distribute more evenly, that is, we explore relational reasoning of video content in a more fine-grained manner. Finally, since each video in SVQA has the same length of frames, network with clips would contain much noise in each clip representation. Therefore, their performance degrades as the number increases.
IV-E Qualitative Analysis
To better understand the contributions of our Query Punishment Module in DualVGR, we provide the visualization examples of the attention results of query features and the query-guided mask values on MSVD-QA and SVQA datasets in Fig. 5. In the first example of one step in real dataset MSVD-QA, this question intends to ask who is showing a gun in a box? Our model mainly focuses on the right part of the question “showing a gun”. This question requires agents to focus on the motion information to infer the answer. We can observe from the example that the motion-based mask values are higher than the appearance-based mask values in all clips, which proves the effectiveness of our Query Punishment Module in real datasets. Besides, the motion-based mask values of the first two clips are higher than others. It further demonstrates the correctness of our model to pay more attention to the relevant video clips to predict the answer. For the second example of SVQA dataset, this video is extremely simple as it only contains four objects. The gray cube is rotating all the time. Then, the cyan object starts to rotate for a while. At the same time, the green cylinder begins to move upwards. In iterative step , the question pays more attention to “rotating present”, then the motion-based mask values are higher than the appearance-based mask values. In iterative step , the model pays much attention to “cyan objects”, then appearance-based mask values are higher than the motion-based mask values to learn the contextual representations of cyan objects. Finally, in iterative step , the model considers both of them to infer the final answer. The mask values are becoming closer to each other than previous steps. Through multi-step message passing, our DualVGR progressively finds out much more question-related visual semantics to infer the answer. The visualization results further verify the effectiveness of our design.
V CONCLUSION
In this paper, we propose a Dual-Visual Graph Reasoning Unit (DualVGR) for Video Question Answering, which could be stacked iteratively to model the question-related rich spatial-temporal interactions between video clips. Specifically, in our DualVGR unit, a Query Punishment Module is proposed to filter out irrelevant visual information through multiple cycles of reasoning. Then, a Video-based Multi-view Graph Attention Network is designed to learn the contextual representations of visual features to perform relational reasoning. With multi-step iterations, our DualVGR network achieves state-of-the-art or competitive performance in three mainstream VideoQA datasets, MSVD-QA, MSRVTT-QA and SVQA.