Bridge to Answer: Structure-aware Graph Interaction Network for Video Question Answering
Jungin Park, Jiyoung Lee, Kwanghoon Sohn
Introduction
Video question answering (VideoQA) is a task to answer the question regarding a given video in a natural language form. Over the past few years, several methods have been focused on manipulating spatio-temporal visual representations conditioned by linguistic cues for VideoQA . However, because of its specificities such as dynamic spatiotemporal dependencies of the video and sophisticated compositional semantics of the question, the VideoQA still remains a challenging problem.
Recent works have adopted the encoder-decoder structure. Typically, LSTM-based encoders are used to encode the representations of video frames and a question into the visual and word sequence. The encoded representations are then incorporated to provide the answer with an attention mechanism. The several types of attention have shown promising results by learning the temporal relations between video frames , the spatial relations between regions in every single frame , or spatiotemporal relations using appearance and motion representations . Although these methods have suggested how to use the visual relationship for VideoQA, they still rely on learning positional relationships, not on in-depth semantic meaning, which makes capturing sophisticated appearance-motion or visual-question relations difficult.
Meanwhile, methods to understand cross-modal relationships have been proposed for vision-language interaction tasks, such as image-text matching or video-text matching , exploiting global or local representations for visual and textual information. Similar approaches have also been adopted in VideoQA. The global question representation has been used as a condition to learn a question-specific visual representation . For example, the global question feature vector was used to update the memory network to learn attention that attributed to the question in . Le et al. proposed a hierarchical architecture that transforms a sequence of objects into a new array conditioned on the global question feature. The word-level features of the question have been treated as sequential data in the local approach . These approaches leveraged each word representation to learn visual attention or co-attention by fusing visual and word representations. However, the global approaches have learned coarse relations that frequently fail to capture video-word relations. The local approaches associate visual and word information based on co-occurrence statistics, not compositional semantics of the question. For instance, without semantic relations, the word “woman” of the question in \figreffig:1-(a) can incorrectly be correlated with all women in the video. On the other hand, compositional semantics clearly indicate from the phrase “in the red” that the question point to the left woman.
To address these limitations, the consideration of grammatical dependencies between sentence words has been raised. For visual question answering (VQA), Teney et al. proposed structured representations that the input image and question are encoded as graphs to leverage compositional semantics of the question. For image-text matching, Liu et al. proposed a graph-structured network that construct graphs for the image and corresponding captions to find the fine-grained image-text correspondences using node-level and structure-level matching. Although the effectiveness of structured representations for image-text relations has been extensively demonstrated, it is still underexplored in VideoQA.
In this paper, we propose a novel method, called Bridge to Answer, that formulates structure-aware interaction for semantic relation modeling between crossmodal information, including appearance, motion, and question. Contrary to existing approaches , we construct not only appearance and motion graphs for video but also the question graph that represents compositional semantic relations between words. We perform question-to-video (Q2V) interactions that propagate the question node to its relevant visual nodes to learn question conditioned visual representations with visual-question relations, as shown in \figreffig:1-(b). Also, we apply visual-to-visual (V2V) interactions to visual graphs delivering each visual node to nodes in the relative visual graph to model appearance-motion relations. To utilize compositional semantic structure of the question, we use the question graph as an intermediate bridge, as shown in \figreffig:1-(c). We demonstrate the capability of the proposed method through extensive ablation studies and comparison with state-of-the-art methods on three datasets, including TGIF-QA , MSVD-QA , and MSRVTT-QA .
Related Works
Video question answering (VideoQA) has attracted intense attention over the past few years due to its applicability to human-robot interaction and video retrieval, etc. The existing methods have mostly been proposed to learn visual representations by leveraging video-question interactions. We summarize recent works according to the types of utilized interactions. Typically, the temporal attention has been learned by exploiting relationships between the appearance and question . Li et al. proposed to learn co-attention between the appearance and question, and Li et al. enhanced co-attention by using self-attention mechanism. Some researchers have presented to capture fine-grained appearance-question interactions. Jin et al. introduced object-aware temporal attention that learns object-question interactions. Huang etal also utilize frame and object features to enhance co-attention between the appearance and question.
Since Jang et al. proposed a two-stream architecture using appearance and motion features, researchers have focused on learning spatiotemporal attention that leverages motion, appearance, and question interactions. Developments of spatiotemporal attention have successfully been applied to various approaches including a multimodal fusion memory , co-memory attention , hierarchical attention , multi-head attention , and multi-step progressive attention . The hierarchical structure that capture appearance-question and motion-question relations from the frame-level to segment-level have also been proposed by Zhao et al. and Le et al. .
Although they have suggested methods that effectively learn relations between appearance and question or even motion, they still rely on positional relationships . Moreover, there have not been presented for capturing the relationship between appearance and motion with compositional semantics of the question. To address these limitations, we explicitly model appearance, motion, and the question as graphs. Our model learns question conditioned visual representations and mutually enhances appearance and motion representations by leveraging compositional semantics of the question.
Graph-structured vision-language interaction has recently been studied to learn semantic relations between visual and textual information. Teney et al. firstly proposed to learn graph-structured representations of the input image and question for visual question answering (VQA). More recently, Li et al. proposed to learn relationships between regions in the input image using graph convolutional network (GCN) and capture image-phrase correspondence for image-text matching. To enable fine-grained image-text matching, Liu et al. constructed a visual graph for the input image and textual graph with the compositional semantics of the caption, respectively. They successfully achieved state-of-the-art performance by learning correspondences between two structured graphs. Similarly, Chen et al. recently proposed a hierarchical graph reasoning method that learns fine-grained video-text correspondence. They composed hierarchical graphs for video and caption according to semantic roles, and performed global and local matching between two graphs.
While these works have suggested methods that effectively learn visual-text relations with structured representations, they cannot be directly applied to VideoQA. To our knowledge, our work is the first attempt to perform relation reasoning between appearance and motion information of the video with compositional semantics of the question.
Method
Given a video and a question , the VideoQA problem is generally formulated as follows:
With appearance and motion representations, we construct an appearance graph and a motion graph as undirected fully-connected graphs. The frames and clips are set to nodes, and each node is connected with all the other nodes in each graph with edges. The weight matrices and , which represent node connections and their edge weights are computed by the affinity between node representations of and as
where and indicate the edge weights between -th and -th node in each graph. is a scaling factor.
Linguistic representations and question graph.
To take compositional semantic structure of the question into the question graph , we identify the semantic dependency within the question (and answer candidates) using Stanford CoreNLP . This parser is used to analyze the components in a sentence (\eg, nouns, verbs, or quantifiers) and parse their semantic dependencies (\eg, nominal subject or adjectival modifier). For example, given a question “What is the woman in the red holding in her hand?”, “What”, “red”, and “holding” are semantically dependent with “woman”. Based on these dependencies, we set each word as a node and connect two nodes if they are semantically dependent. To obtain the weight matrix of the question graph, we compute the affinity matrix of the question representation as
where indicates the affinity between the -th and -th question node and is a scaling factor. Then the weight matrix is represented by a Hadamard product between and the adjacency matrix , followed by normalization, such that,
where the adjacency matrix represents the connectivity of the question graph.
2 Question-to-Visual Interactions
The goal of question-to-visual (Q2V) interactions is to learn question conditioned visual representations by associating the question nodes with visual nodes and propagating question representations along visual edges through graph convolution layers . The Q2V interactions are performed as question-to-appearance (Q2A) and question-to-motion (Q2M) interaction, respectively. Since Q2V interactions are symmetric operations on each graph except for the number of nodes, we describe Q2A interaction in detail and then roughly depict that on Q2M interaction. Specifically, we first obtain an interaction matrix by applying softmax function to the affinity matrix between and along the question axis, such that . The interaction value represents how much the -th question node is associated with the -th appearance node. All the question nodes are aggregated to the corresponding visual node with , followed by a fully connected (FC) layer, so that the aggregated appearance node representation is formulated as
where is the -th node representation of the aggregated appearance graph, and are the learnable parameters of FC layer, and is an activate function such as ReLU.
Subsequently, we apply consecutive graph convolution layers that take and the weight matrix as inputs to propagate the aggregated node to its neighborhoods along the appearance edges. Formally, the output of Q2A interaction is represented as
where is the set of parameters of graph convolution layers.We denote consecutive graph convolution layers as a feed-forward process .
where and is the set of parameters of graph convolution layers.
3 Visual-to-Visual Interactions
One of the most important capabilities for VideoQA is to capture and incorporate the relations between appearance and motion information. To realize this, we present visual-to-visual (V2V) interaction that learns semantic relationships between appearance and motion. Different from previous works that appearance and motion information directly interact, we use the question graph as a bridge to leverage compositional semantics of the question. Since the structure of the question graph reflects semantic dependencies between words, the question conditioned visual node can effectively be delivered to the relative visual nodes along the question edges.
The bridged representation is propagated to its neighbors along the question edges through graph convolution layers and aggregated to the question representation as
where is the set of trainable parameters of graph convolution layers. This form of the aggregated question graph enables the representation to have motion and question information simultaneously.
Finally, the aggregated question node is delivered to the appearance graph to obtain a question conditioned appearance representation attributed to motion. The output of M2A interaction can be formulated by following equation:
where is the -th node representation of the final appearance graph, and are the parameters of FC layer.
As a symmetric process, the node representation of the final motion graph can be obtained with A2M interaction by following equations:
where and are the weight parameters of graph convolution layers and FC layer, respectively.
We apply an average pooling to the final visual node representations along the temporal axis to vectorize the representations, and concatenate them to make the incorporated visual representation :
where denotes concatenation operation.
4 Answer Decoder and Loss Functions
where , , , and are of learnable parameters of the decoder. We employ the cross-entropy loss for open-ended questions.
We treat repetition count task as a linear regression problem that the decoder takes in \equrefeq:y as an input and applies a rounding function for integer count results. The Mean Squared Error (MSE) is employed as the loss function.
For multiple choice question types (\ie, repeating action and state transition), answer candidates are used to make a set of visual representations corresponding to each candidate in the same way with the question. As a result, we have a set of final visual representations, conditioned by the question, and conditioned by answer candidates. The classifier for multiple choice question takes , , , and answer candidates to output probabilities for candidates as follows:
where , , , and are of learnable parameters the decoder. Then, the candidate with the largest value is selected as the answer such that,
We employ the hinge loss between incorrect answer score and correct answer score , , as the loss function.
Experiment
dataset contains animated GIF files and question answer pairs. The dataset provides four kinds of tasks that address the unique properties of videos. Repetition Count is to retrieve number of occurrences of an action. Repeating action is a task to identify an action repeated for a given number of times among multiple choices. State Transition is a multiple choice task to identify an action regarding the temporal order of action state. Frame QA is to find a particular frame in a video that can answer the questions.
MSVD-QA [38]
dataset contains short clips and question answer pairs. The questions are composed of five types, including what, who, how, when, and where.
MSRVTT-QA [39]
dataset contains videos and question answer pairs. While types of questions are the same with MSVD-QA dataset, the contents of the videos in MSRVTT-QA are more complex and the lengths of the videos are much longer from 10 to 30 seconds.
For the evaluation metrics, we use Mean Squared Error (MSE) for repetition count on TGIF-QA dataset and use accuracy for all the other experiments.
2 Implementation Details
We divide the video into 8 clips containing 16 frames in each clip by default. Following the previous work , the long videos in MSRVTT-QA are additionally divided into 24 clips. We train our model for 25 epochs with a batch size of 16 for TGIF-QA and MSVD-QA datasets, and of 4 for MSRVTT-QA dataset. The learning rate is set to and decayed by half for every 5 epochs. The reported results are at the epoch showing the best validation accuracy.
3 Experimental Results
We compare our model with several state-of-the-art methods on TGIF-QA, MSVD-QA, and MSRVTT-QA datasets.Reported values of the other methods are taken from the original papers and For TGIF-QA dataset, we compare our model with in \tabreftab:tgif-qa. We display evaluation results over four tasks, including repeating action, state transition, frameQA, and repetition count. The results show that our model achieves state-of-the-art performance and outperforms the existing methods on all tasks except for FrameQA task.
For MSVD-QA and MSRVTT-QA datasets, our model is compared with in \tabreftab:msvd-qa. Since these datasets provide open-ended questions, they are referred to as highly challenging benchmarks compared to the TGIF-QA dataset. Our model achieves and accuracy, outperforming the existing approaches by and for accuracy, respectively.
We provide qualitative results for two challenging examples in \figreffig:qual1. The first example shows that a sudden transition of the scene causes the problem of capturing semantic relations. Our model handles this problem by learning in-depth semantic relations, not positional relations. The second example reflects the case in which a long and complex question has givenGroundtruth is probably incorrectly annotated.. Our model successfully analyzes this long and complex question by explicitly modeling the compositional semantics of the question.
4 Ablation Studies
To validate the effectiveness of components within our model, we conduct extensive ablation studies on TGIF-QA , as shown in \tabreftab:ablation_tgif. Ablation studies widely cover the results according to input conditioning, interactions, question bridge, and the value of the parameter . The overall result verify that all components affect performance, and even any direction of interaction contribute to the performance improvement. The detailed analyzes are described below.
We study the effect of the input condition with following settings: : Remove appearance feature from full model. Q2A and V2V interaction also been removed. Since the frames are not used in this setting, we do not measure the performance for FrameQA. : Remove motion feature from full model. Q2M and V2V interaction also been removed. We find that while the absence of either appearance or motion feature is critical to the performance, motion contributes more to the performance in action-related tasks. The results of FrameQA show that motion information does not play an important role in tasks where appearance information of the frame is important.
Effect of interactions.
Overall we find that the absence of any direction of interaction significantly degrades the performance on all tasks. Specifically, Q2V works as a primary component for VideoQA. This is not surprising given that learning question conditioned visual representations is one of the most important ingredient for VideoQA. The notable performance degradation due to the ablation of Q2M is shown in all the tasks except for FrameQA.
We also find that A2M and M2A are complemented each other from the results that promising performance can be achieved with only one-way interaction. To analyze V2V interaction, we display the visualization for the connectivity of motion-question and question-appearance as shown in \figreffig:qual2. We depict M2A interaction only and represent the clips as frames sampled at each clip for visibility. The clips and words are placed according to temporal and word order, respectively, and corresponding frames of each word are placed regardless of temporal order. The connections indicate that two nodes are associated with the maximum interaction value. For example, the first clip is associated with the word “man” by the interaction value of 0.15, and the word “man” is related to the fourth frame by the interaction value of 0.15. Note that, when all the nodes are connected with uniformly distributed weights, the motion-question interaction value and the question-appearance interaction value is 0.09 and 0.008, respectively. The results show that the relavant nodes are connected with relatively high interaction values, and the question node is also connected with the appearance node through semantic relation.
Effect of the question bridge.
The question bridge is a key component to leverage the compositional semantics of the question. To verify the effectiveness, we conduct ablation study for the question bridge. The V2V interactions without the question bridge are performed by directly aggregating the relative node representations based on the affinity value between appearance and motion nodes. For instance, the output of the M2A interaction without the bridge is obtained by
The results at each task demonstrate the advantage of the bridged architecture with the performance improvement of , , for accuracy, and for MSE value, respectively.
Effect of λ𝜆\mathbf{\lambda}.
The scaling parameter adjusts the relative weight of different nodes in Q2V and V2V interactions and the edge weight of graphs. The large value of distills nodes highly correlated to the specific node and filters out irrelevant nodes. Contrary to this, the small value of is difficult to distinguish relevant nodes. Therefore, it is important to properly set the value of . To investigate the performance with various values, we measure VideoQA performance by setting the as 1, 5, 10, 20. Not surprisingly, Lager shows better performance compared to .
We additionally evaluate our model according to values on MSVD-QA and MSRVTT-QA datasets. As shown in \tabreftab:lambda_msvd, our model yield highest performance when on MSVD-QA. The results on MSRVTT-QA show that brings out the best performance. The different optimal value of on two datasets might be caused by different lengths of videos.
Conclusion
We proposed a novel method for VideoQA, called Bridge to Answer, that constructs heterogeneous multimodal graphs and learns relations between visual and question graphs to learn question conditioned visual representations attributed to appearance and motion. In the process, in-depth semantic relations between visual and question graphs are encapsulated to visual representations using question-visual interactions. The relations between appearance and motion graphs are modulated by compositional semantics of the question as a bridge to effectively enhance each relative visual representation. This bridged structure allows a model robust to the scene composition and sophisticated structure of the question. Our model was evaluated on several VideoQA benchmarks, including TGIF-QA, MSVD-QA, and MSRVTT-QA, achieving state-of-the-art performance.
Acknowledgement
This work was supported by Institute of Information communications Technology Planning & Evaluation (IITP) grant funded by the Korea government(MSIT) (No.2020-0-00056, To create AI systems that act appropriately and effectively in novel situations that occur in open worlds)