Bridge to Answer: Structure-aware Graph Interaction Network for Video Question Answering

Jungin Park, Jiyoung Lee, Kwanghoon Sohn

Introduction

Video question answering (VideoQA) is a task to answer the question regarding a given video in a natural language form. Over the past few years, several methods have been focused on manipulating spatio-temporal visual representations conditioned by linguistic cues for VideoQA . However, because of its specificities such as dynamic spatiotemporal dependencies of the video and sophisticated compositional semantics of the question, the VideoQA still remains a challenging problem.

Recent works have adopted the encoder-decoder structure. Typically, LSTM-based encoders are used to encode the representations of video frames and a question into the visual and word sequence. The encoded representations are then incorporated to provide the answer with an attention mechanism. The several types of attention have shown promising results by learning the temporal relations between video frames , the spatial relations between regions in every single frame , or spatiotemporal relations using appearance and motion representations . Although these methods have suggested how to use the visual relationship for VideoQA, they still rely on learning positional relationships, not on in-depth semantic meaning, which makes capturing sophisticated appearance-motion or visual-question relations difficult.

Meanwhile, methods to understand cross-modal relationships have been proposed for vision-language interaction tasks, such as image-text matching or video-text matching , exploiting global or local representations for visual and textual information. Similar approaches have also been adopted in VideoQA. The global question representation has been used as a condition to learn a question-specific visual representation . For example, the global question feature vector was used to update the memory network to learn attention that attributed to the question in . Le et al. proposed a hierarchical architecture that transforms a sequence of objects into a new array conditioned on the global question feature. The word-level features of the question have been treated as sequential data in the local approach . These approaches leveraged each word representation to learn visual attention or co-attention by fusing visual and word representations. However, the global approaches have learned coarse relations that frequently fail to capture video-word relations. The local approaches associate visual and word information based on co-occurrence statistics, not compositional semantics of the question. For instance, without semantic relations, the word “woman” of the question in \figreffig:1-(a) can incorrectly be correlated with all women in the video. On the other hand, compositional semantics clearly indicate from the phrase “in the red” that the question point to the left woman.

To address these limitations, the consideration of grammatical dependencies between sentence words has been raised. For visual question answering (VQA), Teney et al. proposed structured representations that the input image and question are encoded as graphs to leverage compositional semantics of the question. For image-text matching, Liu et al. proposed a graph-structured network that construct graphs for the image and corresponding captions to find the fine-grained image-text correspondences using node-level and structure-level matching. Although the effectiveness of structured representations for image-text relations has been extensively demonstrated, it is still underexplored in VideoQA.

In this paper, we propose a novel method, called Bridge to Answer, that formulates structure-aware interaction for semantic relation modeling between crossmodal information, including appearance, motion, and question. Contrary to existing approaches , we construct not only appearance and motion graphs for video but also the question graph that represents compositional semantic relations between words. We perform question-to-video (Q2V) interactions that propagate the question node to its relevant visual nodes to learn question conditioned visual representations with visual-question relations, as shown in \figreffig:1-(b). Also, we apply visual-to-visual (V2V) interactions to visual graphs delivering each visual node to nodes in the relative visual graph to model appearance-motion relations. To utilize compositional semantic structure of the question, we use the question graph as an intermediate bridge, as shown in \figreffig:1-(c). We demonstrate the capability of the proposed method through extensive ablation studies and comparison with state-of-the-art methods on three datasets, including TGIF-QA , MSVD-QA , and MSRVTT-QA .

Related Works

Video question answering (VideoQA) has attracted intense attention over the past few years due to its applicability to human-robot interaction and video retrieval, etc. The existing methods have mostly been proposed to learn visual representations by leveraging video-question interactions. We summarize recent works according to the types of utilized interactions. Typically, the temporal attention has been learned by exploiting relationships between the appearance and question . Li et al. proposed to learn co-attention between the appearance and question, and Li et al. enhanced co-attention by using self-attention mechanism. Some researchers have presented to capture fine-grained appearance-question interactions. Jin et al. introduced object-aware temporal attention that learns object-question interactions. Huang etal also utilize frame and object features to enhance co-attention between the appearance and question.

Since Jang et al. proposed a two-stream architecture using appearance and motion features, researchers have focused on learning spatiotemporal attention that leverages motion, appearance, and question interactions. Developments of spatiotemporal attention have successfully been applied to various approaches including a multimodal fusion memory , co-memory attention , hierarchical attention , multi-head attention , and multi-step progressive attention . The hierarchical structure that capture appearance-question and motion-question relations from the frame-level to segment-level have also been proposed by Zhao et al. and Le et al. .

Although they have suggested methods that effectively learn relations between appearance and question or even motion, they still rely on positional relationships . Moreover, there have not been presented for capturing the relationship between appearance and motion with compositional semantics of the question. To address these limitations, we explicitly model appearance, motion, and the question as graphs. Our model learns question conditioned visual representations and mutually enhances appearance and motion representations by leveraging compositional semantics of the question.

Graph-structured vision-language interaction has recently been studied to learn semantic relations between visual and textual information. Teney et al. firstly proposed to learn graph-structured representations of the input image and question for visual question answering (VQA). More recently, Li et al. proposed to learn relationships between regions in the input image using graph convolutional network (GCN) and capture image-phrase correspondence for image-text matching. To enable fine-grained image-text matching, Liu et al. constructed a visual graph for the input image and textual graph with the compositional semantics of the caption, respectively. They successfully achieved state-of-the-art performance by learning correspondences between two structured graphs. Similarly, Chen et al. recently proposed a hierarchical graph reasoning method that learns fine-grained video-text correspondence. They composed hierarchical graphs for video and caption according to semantic roles, and performed global and local matching between two graphs.

While these works have suggested methods that effectively learn visual-text relations with structured representations, they cannot be directly applied to VideoQA. To our knowledge, our work is the first attempt to perform relation reasoning between appearance and motion information of the video with compositional semantics of the question.

Method

Given a video V\mathcal{V} and a question qq, the VideoQA problem is generally formulated as follows:

With appearance and motion representations, we construct an appearance graph Gv\mathcal{G}_{v} and a motion graph Gm\mathcal{G}_{m} as undirected fully-connected graphs. The frames and clips are set to nodes, and each node is connected with all the other nodes in each graph with edges. The weight matrices Wv\mathbf{W}^{v} and Wm\mathbf{W}^{m}, which represent node connections and their edge weights are computed by the affinity between node representations of V^\hat{\mathbf{V}} and M^\hat{\mathbf{M}} as

where wijvw_{ij}^{v} and wijmw_{ij}^{m} indicate the edge weights between ii-th and jj-th node in each graph. λ\lambda is a scaling factor.

Linguistic representations and question graph.

To take compositional semantic structure of the question into the question graph Gq\mathcal{G}_{q}, we identify the semantic dependency within the question (and answer candidates) using Stanford CoreNLP . This parser is used to analyze the components in a sentence (\eg, nouns, verbs, or quantifiers) and parse their semantic dependencies (\eg, nominal subject or adjectival modifier). For example, given a question “What is the woman in the red holding in her hand?”, “What”, “red”, and “holding” are semantically dependent with “woman”. Based on these dependencies, we set each word as a node and connect two nodes if they are semantically dependent. To obtain the weight matrix of the question graph, we compute the affinity matrix E\mathbf{E} of the question representation U\mathbf{U} as

where eije_{ij} indicates the affinity between the ii-th and jj-th question node and λ\lambda is a scaling factor. Then the weight matrix Wq\mathbf{W}^{q} is represented by a Hadamard product between E\mathbf{E} and the adjacency matrix A\mathbf{A}, followed by L2L_{2} normalization, such that,

where the adjacency matrix A\mathbf{A} represents the connectivity of the question graph.

2 Question-to-Visual Interactions

The goal of question-to-visual (Q2V) interactions is to learn question conditioned visual representations by associating the question nodes with visual nodes and propagating question representations along visual edges through graph convolution layers . The Q2V interactions are performed as question-to-appearance (Q2A) and question-to-motion (Q2M) interaction, respectively. Since Q2V interactions are symmetric operations on each graph except for the number of nodes, we describe Q2A interaction in detail and then roughly depict that on Q2M interaction. Specifically, we first obtain an interaction matrix Sv\mathbf{S}^{v} by applying softmax function to the affinity matrix between V^\hat{\mathbf{V}} and U\mathbf{U} along the question axis, such that Sv=softmaxU(λV^UT)\mathbf{S}^{v}=\text{softmax}_{\mathbf{U}}(\lambda\hat{\mathbf{V}}\mathbf{U}^{T}). The interaction value sijvs_{ij}^{v} represents how much the jj-th question node is associated with the ii-th appearance node. All the question nodes are aggregated to the corresponding visual node with Sv\mathbf{S}^{v}, followed by a fully connected (FC) layer, so that the aggregated appearance node representation is formulated as

where v^i′\hat{\mathbf{v}}_{i}^{\prime} is the ii-th node representation of the aggregated appearance graph, Wfv\mathbf{W}_{f}^{v} and bb are the learnable parameters of FC layer, and σ(⋅)\sigma(\cdot) is an activate function such as ReLU.

Subsequently, we apply consecutive graph convolution layers that take V^′\hat{\mathbf{V}}^{\prime} and the weight matrix Wv\mathbf{W}^{v} as inputs to propagate the aggregated node to its neighborhoods along the appearance edges. Formally, the output of Q2A interaction is represented as

where Wgv\mathbf{W}_{g}^{v} is the set of parameters of graph convolution layers.We denote consecutive graph convolution layers as a feed-forward process F\mathcal{F}.

where Sm=softmaxU(λM^UT)\mathbf{S}^{m}=\text{softmax}_{\mathbf{U}}(\lambda\hat{\mathbf{M}}\mathbf{U}^{T}) and Wgm\mathbf{W}_{g}^{m} is the set of parameters of graph convolution layers.

3 Visual-to-Visual Interactions

One of the most important capabilities for VideoQA is to capture and incorporate the relations between appearance and motion information. To realize this, we present visual-to-visual (V2V) interaction that learns semantic relationships between appearance and motion. Different from previous works that appearance and motion information directly interact, we use the question graph as a bridge to leverage compositional semantics of the question. Since the structure of the question graph reflects semantic dependencies between words, the question conditioned visual node can effectively be delivered to the relative visual nodes along the question edges.

The bridged representation is propagated to its neighbors along the question edges through graph convolution layers and aggregated to the question representation as

where Wgbm\mathbf{W}_{gb}^{m} is the set of trainable parameters of graph convolution layers. This form of the aggregated question graph enables the representation to have motion and question information simultaneously.

Finally, the aggregated question node is delivered to the appearance graph to obtain a question conditioned appearance representation attributed to motion. The output of M2A interaction can be formulated by following equation:

where vif\mathbf{v}_{i}^{f} is the ii-th node representation of the final appearance graph, Wbv\mathbf{W}_{b}^{v} and bb are the parameters of FC layer.

As a symmetric process, the node representation of the final motion graph Mf\mathbf{M}^{f} can be obtained with A2M interaction by following equations:

where Wgbv\mathbf{W}_{gb}^{v} and Wbm\mathbf{W}_{b}^{m} are the weight parameters of graph convolution layers and FC layer, respectively.

We apply an average pooling to the final visual node representations along the temporal axis to vectorize the representations, and concatenate them to make the incorporated visual representation o\mathbf{o}:

where [⋅;⋅][\cdot;\cdot] denotes concatenation operation.

4 Answer Decoder and Loss Functions

where W1\mathbf{W}_{1}, W2\mathbf{W}_{2}, Wy\mathbf{W}_{y}, and Wy′\mathbf{W}_{y^{\prime}} are of learnable parameters of the decoder. We employ the cross-entropy loss for open-ended questions.

We treat repetition count task as a linear regression problem that the decoder takes y′y^{\prime} in \equrefeq:y as an input and applies a rounding function for integer count results. The Mean Squared Error (MSE) is employed as the loss function.

For multiple choice question types (\ie, repeating action and state transition), ∣A∣|\mathcal{A}| answer candidates are used to make a set of visual representations corresponding to each candidate in the same way with the question. As a result, we have a set of final visual representations, o\mathbf{o} conditioned by the question, and {oia}i=1∣A∣\{\mathbf{o}_{i}^{a}\}_{i=1}^{|\mathcal{A}|} conditioned by answer candidates. The classifier for multiple choice question takes o\mathbf{o}, oia\mathbf{o}_{i}^{a}, uˉ\bar{\mathbf{u}}, and answer candidates aˉi\bar{\mathbf{a}}_{i} to output probabilities for candidates as follows:

where W1\mathbf{W}_{1}, Wa\mathbf{W}_{a}, Wy\mathbf{W}_{y}, and Wy′\mathbf{W}_{y^{\prime}} are of learnable parameters the decoder. Then, the candidate with the largest ss value is selected as the answer such that,

We employ the hinge loss between incorrect answer score sns^{n} and correct answer score sps^{p}, max⁡(0,1+sn−sp)\max(0,1+s^{n}-s^{p}), as the loss function.

Experiment

dataset contains 72K72K animated GIF files and 165K165K question answer pairs. The dataset provides four kinds of tasks that address the unique properties of videos. Repetition Count is to retrieve number of occurrences of an action. Repeating action is a task to identify an action repeated for a given number of times among multiple choices. State Transition is a multiple choice task to identify an action regarding the temporal order of action state. Frame QA is to find a particular frame in a video that can answer the questions.

MSVD-QA [38]

dataset contains 1,9701,970 short clips and 50,50550,505 question answer pairs. The questions are composed of five types, including what, who, how, when, and where.

MSRVTT-QA [39]

dataset contains 10K10K videos and 243K243K question answer pairs. While types of questions are the same with MSVD-QA dataset, the contents of the videos in MSRVTT-QA are more complex and the lengths of the videos are much longer from 10 to 30 seconds.

For the evaluation metrics, we use Mean Squared Error (MSE) for repetition count on TGIF-QA dataset and use accuracy for all the other experiments.

2 Implementation Details

We divide the video into 8 clips containing 16 frames in each clip by default. Following the previous work , the long videos in MSRVTT-QA are additionally divided into 24 clips. We train our model for 25 epochs with a batch size of 16 for TGIF-QA and MSVD-QA datasets, and of 4 for MSRVTT-QA dataset. The learning rate is set to 10−410^{-4} and decayed by half for every 5 epochs. The reported results are at the epoch showing the best validation accuracy.

3 Experimental Results

We compare our model with several state-of-the-art methods on TGIF-QA, MSVD-QA, and MSRVTT-QA datasets.Reported values of the other methods are taken from the original papers and For TGIF-QA dataset, we compare our model with in \tabreftab:tgif-qa. We display evaluation results over four tasks, including repeating action, state transition, frameQA, and repetition count. The results show that our model achieves state-of-the-art performance and outperforms the existing methods on all tasks except for FrameQA task.

For MSVD-QA and MSRVTT-QA datasets, our model is compared with in \tabreftab:msvd-qa. Since these datasets provide open-ended questions, they are referred to as highly challenging benchmarks compared to the TGIF-QA dataset. Our model achieves 37.2%37.2\% and 36.9%36.9\% accuracy, outperforming the existing approaches by 1.1%1.1\% and 1.3%1.3\% for accuracy, respectively.

We provide qualitative results for two challenging examples in \figreffig:qual1. The first example shows that a sudden transition of the scene causes the problem of capturing semantic relations. Our model handles this problem by learning in-depth semantic relations, not positional relations. The second example reflects the case in which a long and complex question has givenGroundtruth is probably incorrectly annotated.. Our model successfully analyzes this long and complex question by explicitly modeling the compositional semantics of the question.

4 Ablation Studies

To validate the effectiveness of components within our model, we conduct extensive ablation studies on TGIF-QA , as shown in \tabreftab:ablation_tgif. Ablation studies widely cover the results according to input conditioning, interactions, question bridge, and the value of the parameter λ\lambda. The overall result verify that all components affect performance, and even any direction of interaction contribute to the performance improvement. The detailed analyzes are described below.

We study the effect of the input condition with following settings: ▶w/o appearance\blacktriangleright\textit{w/o appearance}: Remove appearance feature from full model. Q2A and V2V interaction also been removed. Since the frames are not used in this setting, we do not measure the performance for FrameQA. ▶w/o motion\blacktriangleright\textit{w/o motion}: Remove motion feature from full model. Q2M and V2V interaction also been removed. We find that while the absence of either appearance or motion feature is critical to the performance, motion contributes more to the performance in action-related tasks. The results of FrameQA show that motion information does not play an important role in tasks where appearance information of the frame is important.

Effect of interactions.

Overall we find that the absence of any direction of interaction significantly degrades the performance on all tasks. Specifically, Q2V works as a primary component for VideoQA. This is not surprising given that learning question conditioned visual representations is one of the most important ingredient for VideoQA. The notable performance degradation due to the ablation of Q2M is shown in all the tasks except for FrameQA.

We also find that A2M and M2A are complemented each other from the results that promising performance can be achieved with only one-way interaction. To analyze V2V interaction, we display the visualization for the connectivity of motion-question and question-appearance as shown in \figreffig:qual2. We depict M2A interaction only and represent the clips as frames sampled at each clip for visibility. The clips and words are placed according to temporal and word order, respectively, and corresponding frames of each word are placed regardless of temporal order. The connections indicate that two nodes are associated with the maximum interaction value. For example, the first clip is associated with the word “man” by the interaction value of 0.15, and the word “man” is related to the fourth frame by the interaction value of 0.15. Note that, when all the nodes are connected with uniformly distributed weights, the motion-question interaction value and the question-appearance interaction value is 0.09 and 0.008, respectively. The results show that the relavant nodes are connected with relatively high interaction values, and the question node is also connected with the appearance node through semantic relation.

Effect of the question bridge.

The question bridge is a key component to leverage the compositional semantics of the question. To verify the effectiveness, we conduct ablation study for the question bridge. The V2V interactions without the question bridge are performed by directly aggregating the relative node representations based on the affinity value between appearance and motion nodes. For instance, the output of the M2A interaction without the bridge is obtained by

The results at each task demonstrate the advantage of the bridged architecture with the performance improvement of 0.5%0.5\%, 0.9%0.9\%, 0.6%0.6\% for accuracy, and 0.120.12 for MSE value, respectively.

Effect of λ𝜆\mathbf{\lambda}.

The scaling parameter λ\lambda adjusts the relative weight of different nodes in Q2V and V2V interactions and the edge weight of graphs. The large value of λ\lambda distills nodes highly correlated to the specific node and filters out irrelevant nodes. Contrary to this, the small value of λ\lambda is difficult to distinguish relevant nodes. Therefore, it is important to properly set the value of λ\lambda. To investigate the performance with various λ\lambda values, we measure VideoQA performance by setting the λ\lambda as 1, 5, 10, 20. Not surprisingly, Lager λ\lambda shows better performance compared to λ=1\lambda=1.

We additionally evaluate our model according to λ\lambda values on MSVD-QA and MSRVTT-QA datasets. As shown in \tabreftab:lambda_msvd, our model yield highest performance when λ=10\lambda=10 on MSVD-QA. The results on MSRVTT-QA show that λ=20\lambda=20 brings out the best performance. The different optimal value of λ\lambda on two datasets might be caused by different lengths of videos.

Conclusion

We proposed a novel method for VideoQA, called Bridge to Answer, that constructs heterogeneous multimodal graphs and learns relations between visual and question graphs to learn question conditioned visual representations attributed to appearance and motion. In the process, in-depth semantic relations between visual and question graphs are encapsulated to visual representations using question-visual interactions. The relations between appearance and motion graphs are modulated by compositional semantics of the question as a bridge to effectively enhance each relative visual representation. This bridged structure allows a model robust to the scene composition and sophisticated structure of the question. Our model was evaluated on several VideoQA benchmarks, including TGIF-QA, MSVD-QA, and MSRVTT-QA, achieving state-of-the-art performance.

Acknowledgement

This work was supported by Institute of Information communications Technology Planning & Evaluation (IITP) grant funded by the Korea government(MSIT) (No.2020-0-00056, To create AI systems that act appropriately and effectively in novel situations that occur in open worlds)

References