Joint learning of object graph and relation graph for visual question answering

Hao Li, Xu Li, Belhal Karimi, Jie Chen, Mingming Sun

Introduction

VQA tasks require a model to answer a free-form natural language question using visual information from an image. Scene graph (SG) reasoning is an essential instance of VQA tasks . The model extracts objects’ names, attributes, and relations from the input images and organizes them into a graph representation to generate the scene graph.

SG representation modeling displays several virtues over classical VQA techniques since the features in SG are presented in plain and free text form and the graph structures of SG have better interpretability . In this contribution, two reasoning methods on scene graphs are proposed: (i) consider scene graphs as probabilistic graphs and iteratively update nodes’ probabilities using soft instructions extracted from questions ; (ii) apply Graph Neural Network (GNN) into scene graphs to learn joint representations of nodes and their relations, and then feed these representations into a predictor to get the answer. Scene graph reasoning frameworks are useful in VQA . However, there still remain imperfections dealing with complex reasoning questions.

First, existing models tend to predict wrong answers for complex reasoning questions with attributes. Consider the “why” question in Fig. 1(a) as an example, false attribute selection occurs because the model cannot associate “off” relation with “dark monitor” object. The attribute selections require comprehensive supervision from objects, relations, and attributes, but existing methods focus too much on objects while ignoring attributes. Generally, information from objects and relations connected to them are reconstructed into object features in GNN-based methods . However, these encoding methods lack information from objects’ attributes. NSM methods use soft instructions to update answer possibilities, but they treat attributes as secondary information.

Second, existing approaches answer poorly for complex questions that require information about relations. For instance, for localization questions, as in the “where” type of questions, in Fig. 1(a), we observe that missing relation occurs because the model cannot capture the relation information “Behind”. Existing models have a strong bias towards object features, while considering relation as references. The unbalanced focus on objects and relations makes the models fail to learn discriminative representations for relations.

To improve the balance of all kinds of information in scene grpahs, we propose the Dual Message-passing enhanced Graph Neural Network (DM-GNN) for VQA, introducing a novel scene graph reasoning model that extracts balanced feature maps from objects, attributes, and relations information in scene graphs. Concretely, as shown in Fig. 1(b), our DM-GNN model is composed of a scene graph generator, a question encoder, dual graph encoders, and a fusion module. Besides, to balance the importance of objects and relations, we transform scene graphs into a relation-significant modality, where nodes represent relations and edges represent objects, and an object-significant modality, in which nodes represent objects and edges represent relations. After receiving scene graphs in two modalities, dual graph encoders can produce feature maps focusing on relations and objects.

Furthermore, to enhance the information transfer between objects, relations and attributes, we modify the gated graph neural network (GGNN) structure in our DM-GNN by adding the message-passing module. It is a bidirectional GRU that guides the internal information flow. The encoder captures information from nodes, edges, and adjacent nodes that connect to them. In Fig. 1(b), the output feature map of the encoder passes through the fusion module, where the attribute features are explicitly modeled into the feature map to increase the information weight from attributes. Then the feature map passes through multi-head attention layers using question features extracted from the question encoder. Hence, the model dynamically focuses on the critical parts of the questions and uses the most similar part of the scene graph as the most adequate answer. Our main contributions are as follows:

We analyse that existing models answer imperfectly for complex reasoning questions with attributes or relations due to the unbalance focus on three information types in scene graphs, which contain objects, relations and attributes.

We propose a novel DM-GNN model containing a dual encoder structure and a message-passing module. Our model can obtain a balanced representation by properly encoding multi-scale scene graph information.

Experimental results on various datasets show that DM-GNN effectively improves the reasoning accuracy on semantically complicated questions.

Related Work

Visual Question Answering. Most VQA approaches use sequential models to encode questions and CNN-based pretrained models to encode images. Then they use attention methods to fuse features from images and questions. Transformer models achieve outstanding performances on VQA tasks, yet they are heavy to train and hard to explain . Instead, the scene graph model stands for an alternative that is more lightly and explainable.

Scene Graph Generation and Reasoning. Scene graph generation (SGG) methods use object detection methods to extract region proposals from images. Scene graph can promote explainable reasoning for downstream multimodal tasks such as VQA . In typical scene graph reasoning models, NSM performs sequential reasoning over the scene graph by iteratively traversing its nodes. Other models use GGNN based model to encode scene graphs. However, previous works are hard to fully utilize the attribute information and learn the comprehensive representation of SG.

Graph Neural Network. GNN is designed to infer on data described by graphs. apply GNN-based models on knowledge graphs, which are similar to scene graphs. However, existing GNN-based models cannot effectively process graphs with node attributes and complicated labels. Our DM-GNN model can learn a comprehensive and balanced representation using full-scale scene graph information from objects, attributes, and relations to overcome these problems.

DM-GNN Methodology

Our proposed architecture is illustrated in Fig. 2. We use the scene graph generator from . In the question encoder, semantic questions are first projected into an embedding space using GLOVE pretrained word embedding model . Then we use long short-term memory (LSTM) networks to generate questions representation q∈Rdimq\in R^{dim}, where dim is the dimension of the question representation. We introduce our dual graph encoders and fusion module in following subsections.

We organize scene graphs into object-significant graphs and relation-significant graphs.

Object-Significant Graph. We define the object significant graph as Gobj\emph{G}_{obj}, where each node represents an object in the image and each edge represents a relation between two objects. Define N as the node set and E as the edge set. For ni,nj∈Nn_{i},n_{j}\in\emph{N}, ek∈Ee_{k}\in\emph{E}, <ni<n_{i} - eke_{k} - nj>n_{j}> denotes the relation tuple that represents relation eke_{k} from object nin_{i} to object njn_{j}.

Relation-Significant Graph. We define relation significant modality as Grel\emph{G}_{rel}, where each node represents a relation between objects in the image and each edge represents an object, which is completely opposed to the object-significant modality. For ei,ej∈Ee_{i},e_{j}\in\emph{E}, nk∈Nn_{k}\in\emph{N}, <ei−nk−ej><e_{i}-n_{k}-e_{j}> represents the relations eie_{i} and eje_{j} have a shared object nkn_{k}.

Attribute types. Define L as attribute types. For each node ni∈Nn_{i}\in\emph{N}, we define a set of L+1\emph{L}+1 property variables {nil}l=0L{\left\{n_{i}^{l}\right\}}_{l=0}^{L}, where ni0n_{i}^{0} represents nin_{i}’s name embedding and niln_{i}^{l} represents the embedding of node nin_{i}’s lthl^{th} attribute.

2 Dual Encoders

We apply two GGNN-based encoders for object significant graph and relation significant graph. The encoder for object graph focuses on object features and the encoder for relation graph focuses on relation features. The dual structure can balance the importance of relations and objects.

Prior to encoding, every input scene graph is transformed into an information tuple (N,E,Ain,Aout\emph{N},\emph{E},\emph{A}_{in},\emph{A}_{out}): N and E are collections of node embeddings and edge embeddings. Ain\emph{A}_{in} and Aout\emph{A}_{out} are the adjacency matrix of incident and output edges.

Let hith_{i}^{t} is the hidden state of node nin_{i} in the encoder at timestep t, then at t=0\emph{t}=0, we initialize hi0h_{i}^{0} as the GLOVE embedding of nin_{i} with zero padding:

Message-passing Module. To enhance the information transfer from edges and adjacent nodes to the updating nodes, we use the message-passing module (MP) in Fig. 3(a). MP module comes as a replacement of the fully-connected layers from the original GNN model. Consider a tuple <ni,ek,nj><{n}_{i},{e}_{k},{n}_{j}> as the processing sample of the message-passing module. The embedding state, noted ek\textit{e}_{k}, of the edge ek{e}_{k} and neighbor node nj{n}_{j}’s hidden state hj\textit{h}_{j} are injected into a bidirectional GRU network as input sequence while the node ni{n}_{i}’s hidden state hi\textit{h}_{i} is injected as the GRU’s initial hidden state. The output of the GRU represents the updating information for hidden state hi\textit{h}_{i}, which corresponds to the key information from edge ek{e}_{k} and node nj{n}_{j} that is related to node ni{n}_{i}. The sum of every GRU output is ni{n}_{i}’s total information gain from ni{n}_{i}’s adjacent nodes and edges. We detail the message-passing module formula as follows:

where MPi(Ain)MP_{i}(A_{in}) is ni{n}_{i}’s incident information gain, and MPi(Aout)MP_{i}(A_{out}) is ni{n}_{i}’s output information gain.

Propagation Module. In Fig. 3(b), at timestep tt, the hidden states of all nodes are updated by the following gated propagator module:

where kitk_{i}^{t} is the node ni{n}_{i}’s representation from all its incident edges, output edges and adjacent nodes.

Then, we incorporate information from adjacent nodes and from the previous timestep leading to an update of each node’s hidden state:

where W,UzW,U^{z} and UrU^{r} are referred to as the trainable weight matrices. At timestep tt, zitz_{i}^{t} and ritr_{i}^{t} are the update and reset gates, respectively. Then we have:

Here, U1U_{1} and U2U_{2} denote the trainable parameters of the linear layers, the operator ⊙\odot is the element-wise multiplication. σ\sigma is the ReLU function. After TT steps, the encoder generates the final hidden state map GG of the graph. Finally, we compute the graph embedding gi∈Gg_{i}\in G for node ni{n}_{i} as follows:

where f(hiT,ni)\emph{f}(h_{i}^{T},n_{i}) is multi-layer perceptron (MLP) which receives the concatenation of hiTh_{i}^{T} and nin_{i}.

3 Fusion Module and Answer Predictor

Once the dual encoders, embedded in our model, output the node and relation features, we first fuse the attributes into feature maps. For node feature map GNG^{N} and relation feature map GEG^{E}, the fusion feature map FNF^{N} and FEF^{E} are defined as

where FiNF_{i}^{N} indicates the fusion features of node ii and giNg_{i}^{N} is node ii’s representation from the encoder. niln_{i}^{l} is the attribute embedding of node ii. FjEF_{j}^{E} corresponds to the fusion feature of edge jj. gjEg_{j}^{E} is edge jj’s representation from the encoder. eje_{j} is jj-th edge original embedding. The full-scale feature map, noted FF, is obtained by concatenating FNF^{N} and FEF^{E}.

Then, the question embedding q and the full-scale feature map FF are fed into a multi-head attention layer. The reasoning vector, noted r, and which stems from the graph and the question, is computed using a weighted sum of the feature map using the scores output from the attention layer,

In answer predictor module, we adopt a two-layer MLP noted by f(⋅)f(\cdot). This MLP can be viewed as a classifier over the set of candidate answers. The input of the answer predictor is the concatenation vector (q,r)(\emph{q},\emph{r}). Such a classifier has been applied in many VQA models . The answer a^\hat{a} reads:

Experiments

Results on VG dataset. Table 1 reports the results on the test sets of the VG ground truth dataset and the motif-VG dataset. Compared to the baseline models, we can observe that our DM-GNN model outperforms the others at 3%-4%. In addition, we provide detailed results on the VG dataset and motif-VG dataset with different question types. Compared to the other scene graph based VQA models, our model performs well in “what”, “where”, “who” and “why” types. On the VG dataset, our model has 6.5% accuracy improvement in “why” type questions, which highly requires VQA models’ ability to jointly exploit objects, relations and attributes.

Results on GQA dataset. We report in Table 3 the detailed results on the test sets of the GQA dataset. Our DM-GNN model outperforms baselines. We also evaluate our model and other baselines across GQA dataset’s various metrics, where “Binary” and “Open” stand for binary-answer and open domain questions. “Distribution” corresponds to the distance between prediction distribution and standard answer distribution. For open domain questions which are difficult for reasoning, our model outperforms the others by 9.63%. For distribution metric, our model also achieves 2nd score. However, the “Binary” question is challenging for DM-GNN although it achieves SOTA performances. Specifically, the advantage of DM-GNN is to correctly locate key objects, relations and attributes. That is why it works well for the open domain (“what, why, how”) questions. But for “yes/no” questions, whose answers are not explicit in scene graphs, DM-GNN is easy to locate but hard to correctly answer.

2 Ablation Study

We compare several ablated forms of DM-GNN with our complete model on the VG dataset. The accuracy for each variant of DM-GNN are reported in Table 3. We use the raw GGNN network as the Base model. The Base-Obj and Base-Rel model represent the original GGNN network processing object-significant graph and relation-significant graph. The models with +MP contain the message-passing module. The models with +Dual apply the dual encoder structure.

Effect of dual encoder structure. We first validate the efficacy of applying dual structure to balance the importance of relations and objects by splitting our DM-GNN into two single models (Base-Obj and Base-Rel). Both single models perform poorly at 35.3%±0.1%35.3\%\pm 0.1\%. This also shows that both relations and objects are vital to VQA performance. Absence of any of those modules leads to severe accuracy recession. Adding dual encoder structure (+Dual) leads to an empirical gain of 32.6% accuracy upward, which shows that the dual structure is significant in balancing relations and objects.

Effect of the message-passing module. We validate the effectiveness of applying message-passing structure to learn a more comprehensive representation for scene graphs than the raw GGNN structure. Comparing with the Base-Obj and +MP model, we note that after adding the message-passing structure, there is an improvement of 3.9%. Comparing with the Base-Rel and +MP model, we observe an improvement of 3.6%. The +Dual model is DM-GNN without message-passing module. Comparing with DM-GNN, there is a 7.3% improvement after adding the message-passing module. These empirical results show that message-passing structure can successfully improve the representation quality of scene graphs.

Effect of explicit modeling. The w/o attr model and the w/o rela model remove the explicit attribute modeling and relation modeling part. Comparing w/o attr, w/o rela with DM-GNN, removing attribute modeling has a 3.6% decrease in accuracy and removing relation modeling has 0.7% decrease.

3 Visualization

Fig. 4 (a) shows visualization on “what” question type. Three “what” question examples aimed at retrieving either object, relation or attribute information. Comparing row 1, row 2 with row 3, Obj and Rel models have strong attention bias toward objects and relations, while their combination Obj+Rel, balances the attention on both sides and captures correct answers. Comparing row 3 with row 4, the message-passing module increases the score of correct answers.

Fig. 4 (b) shows the visualization on “why”, which need models to jointly exploit objects, relations and attributes to infer answers. With dual encoders and the message-passing module, our DM-GNN achieves 96.1% on “why” questions.

Conclusion

We propose DM-GNN, which encodes each scene graph into feature representations via an object encoder and a relation encoder generating a balanced and full-scale feature map using objects, attributes, and relations information, and demonstrate our model can effectively boost performances on GQA, VG and Motif-VG datasets .

References