Location-aware Graph Convolutional Networks for Video Question Answering
Deng Huang, Peihao Chen, Runhao Zeng, Qing Du, Mingkui Tan, Chuang Gan
Introduction
Recently, deep learning has witnessed a great process (?; ?; ?; ?; ?). Video question answering (video QA) has become an emerging task in computer vision and has drawn increasing interests over the past few years due to its vast potential applications in artificial question answering system and robot dialogue, video retrieval, etc. In this task, a robot is required to answer a question after watching a video. Unlike the well-studied Image Question Answering (image QA) task which focuses on understanding static images (?; ?; ?), video QA is more practical since the input visual information often change dynamically, as shown in Figure 1.
Compared with image QA, video QA is much more challenging due to several reasons. (1) Visual content is more complex in a video since it may contain thousands of frames, as shown in Figure 1. More importantly, some frames may be dominated with strong background content which however is irrelevant to questions. (2) Videos often contain multiple actions, but only a part of them are of interest to questions. (3) Questions in video QA task often contain queries related to temporal cues, which implies we should consider both temporal location of objects and complex interaction between them for answer reasoning. For example in Figure 1, to answer the question “What does the man do before spinning bucket?”, the robot should not only recognize the actions “spin laptop” and “spin bucket” by understanding the interaction between the man and objects (i.e., laptop and bucket) in different frames, but also find out the temporal order of actions (e.g., before/after) for answer reasoning along time axis.
Taking video frames as inputs, most existing methods (?; ?) employ some spatio-temporal attention mechanism on frame features to ask the network “where and when to look”. However, these methods are often not robust due to complex background content in videos. Lei et al. (?) tackle this problem by detecting the objects in each frame and then processing the sequence of object features via an LSTM. However, the order of the input object sequence, which may affect the performance, is difficult to arrange. More importantly, processing the objects in a recurrent manner will inevitably neglect the direct interaction between nonadjacent objects. This is critical for video QA (see experiments in Section 4.4).
In this paper, we introduce a simple yet powerful network named Location-aware Graph Convolutional Networks (L-GCN) to model the interaction between objects related to questions. We propose to represent the content in a video as a graph and identify actions through graph convolution. Specifically, the objects of interest are first detected by an off-the-shelf object detector. Then, we construct a fully-connected graph where each node is an object and the edges between nodes represent their relationship. We further incorporate both spatial and temporal object location information into each node, letting the graph be aware of the object locations. When performing graph convolution on the object graph, the objects directly interact with each other by passing message through edges. Last, the output of GCNs and question features are fed into a visual-question interaction module to predict a answer. Extensive experiments demonstrate the effectiveness of the proposed location-aware graph. We achieve state-of-the-art results on TGIF-QA, Youtube2Text-QA and MSVD-QA datasets.
The main contributions of the proposed method are as follows: (1) we propose to explore actions for video QA task through learning interaction between detected objects such that irrelevant background content can be explicitly excluded; (2) we propose to model the relationships between objects through GCNs such that all objects are able to interact with each other directly; (3) we propose to incorporate object location information into graph such that the network is aware of the location of a specific action; (4) our method achieves state-of-the-art performance on TGIF-QA, Youtube2Text-QA and MSVD-QA datasets.
Related Work
Visual Question Answering (VQA) is a task to answer the given question based on the input visual information.
Based on the visual sources, we can classify the VQA tasks into two categories: image QA (?; ?) and video QA (?; ?). Image QA focuses on spatial information. Most image QA models adopt attention mechanism to capture spatial area that related to question words. Yang et al. (?) proposed a multi-layer Stacked Attention Network (SAN) which uses questions as query to extract the image region related to the answer. Anderson et al. (?) combined bottom-up and top-down attention which connect questions to specific objects detected by Faster-RCNN. After that, associating feature vector with visual regions becomes a popular framework in the VQA research (i.e. Pythia (?)). Xiong et al. (?) introduced the dynamic memory network (DMN) architecture to image QA, which strengthens the reasoning ability of network.
In video QA task, understanding untrimmed videos (?; ?) is important. To this end, Jang et al. (?) utilized both motion (i.e. C3D) and appearance (i.e., ResNet (?)) features to better represent the video. Li et al. (?) replaced RNN with self-attention together with location encoding to model long-range dependencies. However, all the existing methods neglect the interaction between objects, which is vital for video QA task.
Graph-based reasoning has been popular in recent years (?; ?) and shown to be powerful for relation reasoning. To dynamically learn graph structures, CGM (?) applied a cutting plane algorithm to iteratively activate a group of cliques. Recently, Graph Convolution Networks (GCNs) (?) have been used for semi-supervised classification. In text-based tasks, such as machine translation and sequence tagging, GCNs breaks the sequence restriction between each word and learns the graph weight by attention mechanism, which makes it work better in modeling longer sequence than LSTM. Some methods (?; ?; ?) took into consideration the object position for image QA tasks. In video recognition, Wang et al. (?) proposed to use GCNs to capture relations between objects in videos, where objects are detected by an object detector pre-trained on extra training data. Despite their success, there is no efficient graph model for video QA task.
Attention mechanism has been leveraged in various tasks. Several works (?; ?) used attention model to improve the performance on video recognition. Vaswani et al. (?) utilized self-attention mechanism for language translation and (?) proposed Co-Attention which can be stacked to form a hierarchy for multi-step interactions between visual and language features. Jang et al. (?) proposed a simple baseline which uses both spatial and temporal attention to reason the video and answer the question. In our proposed method, we use attention mechanism to fuse video and question modalities.
Proposed Method
In this paper, we focus on video QA task, which requires the model to answer questions related to a video. This task is challenging as video contents are complex with strong irrelevant backgrounds. Besides, most QA pairs in video QA task are related to more than one action with temporal cues. To answer the question correctly, the model is required not only to recognize the actions correctly from complex contents but also to be aware of their temporal order.
2 General Scheme
The general scheme of our method is shown in Figure 2, which consists of two streams. The first stream is regarding a question encoder, which processes queries with a Bi-LSTM. The second stream is related to a video encoder, which focuses on understanding video contents by exploiting a location-aware graph built on objects. The outputs of two streams are then combined by a visual-question (VQ) interaction module, which employs an attention mechanism to explore which question words are more relevant to the visual representation. Last, the answer is predicted by applying an FC layer on top of the VQ interaction module.
In this paper, the location-aware graph plays a critical role. Specifically, we use an object graph to model the relationships between objects in a video. Note the temporal ordering of actions in the video is important for answer reasoning w.r.t. a question in a video QA task. We thus propose to integrate the spatial and temporal location information into the object features of each node in a graph (See details in Section 3.4). In this way, we can exploit both spatial and temporal order information of actions for temporally related answer reasoning.
For convenience, we present the overall training process in Algorithm 1. In the following, we first describe the question encoder. Then we depict the construction of the location-aware graph and the graph convolution for message passing, followed by description of visual encoder. After that, we detail the visual-question interaction module. Last, we present the answer reasoning and loss functions.
3 Question Encoder Stream
In the optimization, the word embedding function is initialized with a pre-trained 300-dimension GloVe (?), and the character embedding function is randomly initialized. Given the character and word embeddings, the question embedding can be represented by a two-layer highway network (?), which is proven to be effective to solve the training difficulties, that is:
where the character embedding is further processed by a which consists of a 2D convolutional layer.
To better encode the question, we feed the question embedding into a bi-directional LSTM (Bi-LSTM). Then we obtain the question feature by stacking the hidden states of the Bi-LSTM from both directions at each time step.
4 Location-aware Graph Construction
Given a video with detected objects for each frame, we seek to represent the video into a graph. Noting that actions can be inferred from the interaction between objects, we thus construct a fully-connected graph on the detected objects. We may use object features to represent each node. However, this node type ignores the location information of objects, which is vital for temporally related answer reasoning. To address this, we will describe how to encode the location information with so-called location features. With location features, we are able to construct a location-aware graph, namely, we concatenate both object appearance and location features as node features.
where is represented by the top-left coordinate and the width and the height of detected objects.
Moreover, we also encoder temporal location feature of objects using sine and cosine functions of different frequencies (?) as follows:
where is the -th entry of the temporal location feature , and is its dimension. Then, the feature of each graph node can be defined as:
where concatenates three vectors into a longer vector. In this way, each node in the graph contains not only the object appearance features but also the location information.
5 Reasoning with Graph Convolution
Given the constructed location-aware graph, we perform graph convolution to obtain the regional features. In our implementation, we build -layer graph convolutions. Specifically, for the -th layer (1 ), the graph convolution can be formally represented as:
where is the hidden features of the -th layer; is the input node features in Eq. (5); is the adjacency matrix calculated from the node features in the -th layer; and is the trainable weight matrix. Let be the output of the last layer of the -layer GCNs. Then, we define the regional features as:
This can be considered as a skip connection of input and output , and it helps to improve the training performance, similar to ResNet (?). In our method, the adjacency matrix is a learnable matrix, which is able to simultaneously infer a graph by learning the weight of all edges. We calculate the adjacency matrix by:
where and are projection matrices. The softmax operation is performed in the row axis.
6 Visual Encoder Stream
The visual encoder is to model video contents via object interaction for video QA. Given a -frame video, we extract frame features using a fixed feature extractor (e.g., ResNet-152). At the same time, bounding boxes are detected for each frame by an off-the-shelf object detector. The object features are obtained using RoIAlign (?) on top of the image features, followed by an FC layer and ELU activation function (?) to reduce dimension.
Given the detected object set , we construct a location-aware graph on the objects. Then, we perform graph convolution to enable the message passing between objects through edges, which can be formally represented as:
where indicates the concatenation of vectors and denotes for any mapping function, e.g., multi-layer perceptron (MLP). The output of GCNs is termed as regional features . Besides, in order to introduce the context information, we apply global average pooling on the frame features to generate global features .
7 Visual-question Interaction Module
Specifically, we first calculate similarity matrix between and via dot product together with a softmax function applying along each row, that is:
where means the element-wise product operation. To yield the final representation for answer prediction, we leverage a Bi-LSTM followed by a max pooling layer across the dimension .
8 Answer Reasoning and Loss Function
The questions for video QA can be summarized as three types: multiple-choice, open-ended and counting. In this subsection, we will describe how to predict answers for each question type given cross modality features .
Multiple-choice question: for this kind of questions, there exist choices and the model is required to choose the correct one. We first embed the content of each choice in the same way as question encoding described in Section 3.3, leading to independent answer features . Then, each answer feature is interacted with visual features in the way described in Section 3.7, where we replace the question feature by answer question, yielding the weighted answer features . Then, the cross modality representation in Eq. (12) is constructed as . We leverage an identical FC layer on cross modality representations to predict scores . The scores are processed by a softmax function. We use cross entropy loss as the loss function:
where if answer is the right choice, otherwise . We take the choice with the highest score as the prediction.
Open-ended question: for these questions, the model is required to choose a correct word as answer from the pre-defined answer set of candidate words in total. We predict the scores of each candidate word using an FC layer together with a softmax layer. Also, we use the cross entropy loss as the loss function:
where if answer is the right answer, otherwise . We take the word with the highest score as our prediction.
Counting question: for these questions, the model is required to predict a number ranging from 0 to 10. We leverage an FC layer upon to predict the number. We use mean square error loss to train the model:
where is the predicted number, is the ground truth. During the testing, the prediction is rounded to the nearest integer and clipped within 0 to 10.
Experiments
In this section, we first introduce three benchmark datasets and implementation details. Then, we compare the performance of our model with the state-of-the-art methods. Last, we perform ablation studies to understand the effect of each component.
We evaluate our method on three video QA datasets. The statistics of the datasets are listed in Table 1. More details are given below.
TGIF-QA (?) consists of 165K QA pairs from 72K animated GIFs. The QA-pairs are splited into four tasks: 1) Action: a multiple-choice question recognizing action repeated certain times; 2) Transition (Trans.): a multiple-choice question asking about the state transition; 3) FrameQA: an open-ended question that can be inferred from one frame in videos; 4) Count: an open-ended question counting the number of repetition of an action. The multiple-choice questions in this dataset have five options and the open-ended questions are with a pre-defined answer set of size 1,746.
Youtube2Text-QA (?) includes the videos from MSVD video set (?) and the question-answer pairs collected from Youtube2Text (?) video description corpus. It consists of open-ended and multiple-choice questions, which are divided into three types (i.e., what, who and others).
MSVD-QA (?) is based on MSVD video set. It consists of five types of questions, including what, who, how, when and where. All questions are open-ended with a pre-defined answer set of size 1,000.
2 Implementation Details
(1) For the “Count” task in TGIF-QA dataset, we adopt the Mean Square Error (MSE) between the predicted answer and the ground truth answer as the evaluation metric. (2) For all other tasks in our experiments, we use accuracy to evaluate the performance.
Training details.
We convert all the words in the question and answer to lower cases, and then transform each word to a 300-dimension vector with a pre-trained GloVe model (?). For fair comparisons, we adopt the same feature extractors as those are used in the compared methods. More details can be found in Table 1. We use Mask R-CNN (?) as object detector and select detected objects with the highest score for each frame. By default, is set to 5. The number of GCNs layers is set to 2. We employ a Adam optimizer (?) to train the network with an initial learning rate of 1e-4. We set the batch size to 64 and 128 for multiple-choice and open-ended tasks, respectively.
3 Comparison with State-of-the-art Results
We compare our L-GCN with the state-of-the-art methods, including ST-VQA (?), Co-Men (?), PSAC (?) and HME (?). From Table 2, our L-GCN achieves the best performance on four tasks. It is worth noting that our method outperforms HME, ST-VQA and Co-Mem by a large margin even if they use additional features (i.e., C3D features (?) and optical flow feature) to model actions. These results demonstrate the effectiveness of leveraging an object graph to capture the object-object interaction and perform reasoning.
Results on Youtube2Text-QA.
For further comparison, we test our model on a more challenging dataset Youtube2Text-QA. This dataset consists of open-ended and multiple-choice questions, which are divided into three categories (i.e., what, who and others). We consider two state-of-the-art baseline methods (HME and r-ANL (?)), and report the results in Table 3.
From Table 3, compared with the baselines, our method achieves better performance in overall accuracy in both multi-choice and open-ended questions. More specifically, for multiple-choice questions, we achieve the best performance on what and who tasks. The relatively poor performance on others task cannot represent the ability of different models because this kind of questions only occupies 2% of all QA pairs. For open-ended questions, our L-GCN significantly improves the accuracy from 29.4% to 53.2% on who task, where most questions are related to the subject of actions. This demonstrates the superiority of leveraging object features, which explicitly localizes the object for video QA task.
Results on MSVD-QA.
In Table 4, we compare our L-GCN with ST-VQA, Co-Mem, AMU (?) and HME on MSVD-QA dataset. From Table 4, our L-GCN achieves the most promising performance in overall accuracy, which demonstrates the superiority of the proposed method on the non-trivial scenarios.
4 Ablation Study
We first construct a simple variant of the proposed method as baseline, which uses only the global frame features to generate visual features via Eq. (10). Then, the object features, GCNs, and location features will be incorporated into the baseline progressively to generate visual features in higher quality, and we denote them as “OF”, “GCNs” and “Loc”, respectively. “FC”and “LSTM” represent the models where GCNs are replaced by two Fully-Connected (FC) layers or a 2-layer LSTM, respectively. “Loc_T”and “Loc_S” represent the location features which only consist of temporal or spatial location information, respectively.
We show the results on TGIF-QA dataset in Table 5. (1) Compared with the baseline, incorporating object features boosts the performance in all tasks consistently, demonstrating the effectiveness of using detected objects for video QA task. We speculate that the detected objects explicitly help the model exclude irrelevant background. (2) Applying GCNs on object features further boosts the performance, demonstrating the importance of modeling relationships between objects through GCNs. On the other hand, using FC layer or LSTM only brings minor increases or even drops the performance. This is not surprising because the model cannot learn object-object relationship when applying FC layer on each object separately. Besides, objects in different spatial locations cannot be regarded as a sequence and thus LSTM is not suitable for modeling their relationship. (3) Adding location features further increases the performance. Especially, the improvements on the task of transition and count are more significant. One possible reason is that these two tasks are more sensitive to the knowledge of event’s order, where the transition task asks about the action transition and the count task asks the number of repetition of an action. We also try to only incorporate temporal or spatial location information into L-GCN. The performance decreases compared to the variant using both location types, demonstrating that these two location information are complementary and both vital for video QA task.
Impact of #GCNs layers and detected objects.
In this paper, we propose to leverage GCNs on detected objects to learn actions. Here, we conduct ablation studies on the depth of GCNs and the number of the objects in each frame. From Table 6, GCNs with two layers performs best on three tasks. Considering the efficiency and performance, we leverage 2-layer GCNs by default. Besides, as shown in Table 7, GCNs with 5 detected objects achieves the best performance on three tasks. It is not surprising that the network with 2 detected objects performs worst because the network may neglect some important objects. Additionally, as most of the question answering pairs in TGIF-QA dataset are only relative to a few salient objects, feeding too many objects into network may cripple the performance. By default, we leverage 5 detected objects in experiments.
5 Qualitative Analysis
We demonstrate the similarity matrix in the GCNs using two examples in Figure 3. We draw two conclusions from these examples: 1) Almost all salient objects which are related to question answering pair have been detected beforehand, such as the airplane and the boy in example 1, the man and the motorcycle in example 2, etc. These detected objects explicitly help the network avoid the influence from complex irrelevant background content. 2) Our graph not only captures relationships between similar objects in different frames but also focuses on semantic similarity. For the first example, the airplane is correlative to not only itself in different frames but also the little boy. This is helpful to recognize the action of “airplane running over boy”.
Conclusion
In this paper, we have proposed a location-aware graph to model the relationships between detected objects for video QA task. Compared with existing spatio-temporal attention mechanism, L-GCN is able to explicitly get rid of the influences from irrelevant background content. Moreover, our network is aware of the spatial and temporal location of events, which is important for predicting correct answer. Our method outperforms state-of-the-art techniques on three benchmark datasets.
Acknowledgment
This work was partially supported by Guangdong Provincial Scientific and Technological Funds under Grants 2018B010107001, National Natural Science Foundation of China (NSFC) 61602185, key project of NSFC (No. 61836003), Program for Guangdong Introducing Innovative and Entrepreneurial Teams 2017ZT07X183, Tencent AI Lab Rhino-Bird Focused Research Program (No. JR201902), Natural Science Foundation of Guangdong Province under Grant 2016A030310423, Fundamental Research Funds for the Central Universities D2191240.