Graph-Structured Referring Expression Reasoning in The Wild

Sibei Yang, Guanbin Li, Yizhou Yu

Introduction

Grounding referring expressions aims to locate in an image an object referred to by a natural language expression, and the object is called the referent. It is a challenging problem because it requires understanding as well as performing reasoning over semantics-rich referring expressions and diverse visual contents including objects, attributes and relations.

Analyzing the linguistic structure of referring expressions is the key to grounding referring expressions because they naturally provide the layout of reasoning over the visual contents. For the example shown in Figure 1, the composition of the referring expression “the girl in blue smock across the table” (i.e., triplets (“the girl”, “in”, “blue smock”) and (“the girl”, “across” , “the table”)) reveals a tree-structured layout of finding the blue smock, locating the table and identifying the girl who is “in” the blue smock and meanwhile is “across” the table. However, nearly all the existing works either neglect linguistic structures and learn holistic matching scores between monolithic representations of referring expressions and visual contents or neglect syntactic information and explore limited linguistic structures via self-attention mechanisms .

Consequently, in this paper, we propose a Scene Graph guided modular network (SGMN) to fully analyze the linguistic structure of referring expressions and enable reasoning over visual contents using neural modules under the guidance of the parsed linguistic structure. Specifically, SGMN first models the input image with a structured representation, which is a directed graph over the visual objects in the image. The edges of the graph encode the semantic relations among the objects. Second, SGMN analyzes the linguistic structure of the expression by parsing it into a language scene graph using an external parser, including the nodes and edges of which correspond to noun phrases and prepositional/verb phrases respectively. The language scene graph not only encodes the linguistic structure but is also consistent with the semantic graph representation of the image. Third, SGMN performs reasoning on the image semantic graph under the guidance of the language scene graph by using well-deigned neural modules including AttendNode, AttendRelation, Transfer, Merge and Norm. The reasoning process can be explicitly explained via a graph attention mechanism.

In addition to methods, datasets are also important for making progress on grounding referring expressions, and various real-world datasets have been released . However, recent work indicates dataset biases exist and they may be exploited by the methods. And methods accessing the images only achieve marginally higher performance than a random guess. Existing datasets also have other limitations. First, the samples in the datasets have unbalanced levels of difficulty. Many expressions in the datasets directly describe the referents with attributes due to the annotation process. Such an imbalance makes models learn shallow correlations instead of achieving joint image and text understanding, which defeats the original intention of grounding referring expressions. Second, evaluation is only conducted on final predictions but not on the intermediate reasoning process , which does not encourage the development of interpretable models . Thus, a synthetic dataset over simple 3D shapes with attributes is proposed in to address these limitations. However, the visual contents in this synthetic dataset are too simple, which is not conducive to generalizing trained models on the synthetic dataset to real-world scenes.

To address the aforementioned limitations, we build a large-scale real-world dataset, named Ref-Reasoning. We generate semantically rich expressions over the scene graphs of images using diverse expression templates and functional programs, and automatically obtain the ground-truth annotations at all intermediate steps during the modularized generation process. Furthermore, we carefully balance the dataset by adopting uniform sampling and controlling the distribution of expression-referent pairs over the number of reasoning steps.

In summary, this paper has the following contributions:

A scene graph guided modular neural network is proposed to perform reasoning over a semantic graph and a scene graph using neural modules under the guidance of the linguistic structure of referring expressions, which meets the fundamental requirement of grounding referring expressions.

A large-scale real-word dataset, Ref-Reasoning, is constructed for grounding referring expressions. Ref-Reasoning includes semantically rich expressions describing objects, attributes, direct and indirect relations with a variety of reasoning layouts.

Experimental results demonstrate that the proposed method not only significantly surpasses existing state-of-the-art algorithms on the new Ref-Reasoning dataset, but also outperforms state-of-the-art structured methods on common benchmark datasets. In addition, it can provide interpretable visual evidences of reasoning.

Related Work

A referring expression normally not only directly describes the appearance of the referent, but also its relations to other objects in the image, and its reference information depends on the meanings of its constituent expressions and the rules used to compose them . However, most of the existing works neglect linguistic structures and learn holistic representations for the objects in the image and the expression. Recently, there are some works which involve the expression analysis into their models, and learn the components of expression and visual inference from end to end. The methods in softly decompose the expression into different semantic components relevant to different visual evidences, and compute a matching score for every component. They use fixed semantic components, e.g. subject-relation-object triplets and subject-location-relation components , which are not feasible for complex expressions. DGA analyzes linguistic structures for complex expressions by iteratively attending their constituent expressions. However, they all resort to self-attention on the expression to explore its linguistic structure but neglect its syntactic information. Another work grounds the referent using a parse tree, where each node of the tree is a word (or phrase) which can be a noun, preposition or verb.

2 Dataset Bias and Solutions

Recently, the dataset bias began to be discussed for grounding referring expressions . The work in reveals even the linguistically-motivated models tend to learn shallow correlations instead of making use of linguistic structures because of the dataset bias. In addition, expression-independent models can achieve high performance. The dataset bias can have a significantly negative impact on the evaluation of a model’s readiness for joint understanding and reasoning for language and vision.

In order to address the above problem, the work in proposes a new diagnostic dataset, called CLEVR-Ref+. Same as CLEVR in visual question answering, it contains rendered images and automatically generated expressions. In particular, the objects in the images are simple 3D shapes with attributes (i.e., color, size and material), and the expressions are generated using designed templates which include spatial and same-attribute relations. However, the models trained on this synthetic dataset cannot be easily generalized to real-world scenes because the visual contents (i.e., simple 3D shapes with attributes and spatial relations) are too simple to jointly reason about language and vision.

Thanks for the scene graph annotations of real-world images provided in the Visual Genome datasets and further cleaned in the GQA dataset , we generate semantically rich expressions over the scene graphs with objects, attributes and relations using carefully designed templates along with functional programs.

Approach

We now present the proposed scene graph guided modular network (SGMN). As illustrated in Figure 2, given an input expression and an input image with visual objects, our SGMN first builds a pair of semantic graph and scene graph representations for the image and expression respectively, and then performs structured reasoning over the graphs using neural modules.

Scene graph based representations form the basis of our structured reasoning. In particular, the image semantic graph flexibly captures and represents all the visual contents needed for grounding referring expressions in the input image while the language scene graph explores the linguistic structure of the input expression, which defines the layout of the reasoning process. In addition, these two types of graphs have consistent structures, where the nodes and edges of the language scene graph respectively correspond to a subset of the nodes and edges of the image semantic graph.

Given an image with objects O={oi}i=1N\mathcal{O}=\{o_{i}\}_{i=1}^{N} , we define the image semantic graph over the objects O\mathcal{O} as a directed graph, Go=(Vo,Eo)\mathcal{G}^{o}=(\mathcal{V}^{o},\mathcal{E}^{o}), where Vo={vio}i=1N\mathcal{V}^{o}=\{v_{i}^{o}\}_{i=1}^{N} is the set of nodes and node viov_{i}^{o} corresponds to object oio_{i}; Eo={eijo}i,j=1N\mathcal{E}^{o}=\{e_{ij}^{o}\}_{i,j=1}^{N} is the set of directed edges, and eijoe_{ij}^{o} is the edge from vjov_{j}^{o} to viov_{i}^{o}, which denotes the relation between objects ojo_{j} and oio_{i}.

For each node viov_{i}^{o}, we obtain two types of features, visual feature vio\mathbf{v}_{i}^{o} extracted from a pretrained CNN model and spatial feature pio=[xi,yi,wi,hi,wihi]\mathbf{p}_{i}^{o}=[x_{i},y_{i},w_{i},h_{i},w_{i}h_{i}], where (xi,yi)(x_{i},y_{i}), wiw_{i} and hih_{i} are the normalized top-left coordinates, width and height of the bounding box of node viv_{i} respectively. For each edge eijoe_{ij}^{o}, we compute the edge feature eijo\mathbf{e}_{ij}^{o} by encoding the relative spatial feature lijo\mathbf{l}_{ij}^{o} between viov_{i}^{o} and vjov_{j}^{o} and the visual feature vjo\mathbf{v}_{j}^{o} of node vjov_{j}^{o} together because relative spatial information between objects along with their appearance information is the key indicator of their semantic relation . Specifically, the relative spatial feature is represented as lijo=[xj−xciwi,yj−ycihi,xj+wj−xciwi,yj+hj−ycihi,wjhjwihi]\mathbf{l}_{ij}^{o}=[\frac{x_{j}-{x_{c}}_{i}}{w_{i}},\frac{y_{j}-{y_{c}}_{i}}{h_{i}},\frac{x_{j}+w_{j}-{x_{c}}_{i}}{w_{i}},\frac{y_{j}+h_{j}-{y_{c}}_{i}}{h_{i}},\frac{w_{j}h_{j}}{w_{i}h_{i}}], where (xci,yci)({x_{c}}_{i},{y_{c}}_{i}) are the normalized center coordinates of the bounding box of node viov_{i}^{o}. And eijo\mathbf{e}_{ij}^{o} is the concatenation of an encoded version of lijo\mathbf{l}_{ij}^{o} and vjo\mathbf{v}_{j}^{o}, i.e., eijo=[WoTlijo,vjo]\mathbf{e}_{ij}^{o}=[\mathbf{W}_{o}^{T}\mathbf{l}_{ij}^{o},\mathbf{v}_{j}^{o}], where Wo\mathbf{W}_{o} is a learnable matrix.

1.2 Language Scene Graph

Given an expression SS, we first use an off-the-shelf scene graph parser to parse the expression into an initial language scene graph, where a node and an edge of the graph correspond to an object and the relation between two objects mentioned in SS respectively, and the object is represented as an entity with a set of attributes.

We define the language scene graph over SS as a directed graph G=(V,E)\mathcal{G}=(\mathcal{V},\mathcal{E}), where V={vm}m=1M\mathcal{V}=\{v_{m}\}_{m=1}^{M} is a set of nodes and node vmv_{m} is associated with a noun or noun phrase, which is a sequence of words from SS; E={ek}k=1K\mathcal{E}=\{e_{k}\}_{k=1}^{K} is a set of edges and edge ek=(vks,rk,vko)e_{k}=({v_{k}}_{s},r_{k},{v_{k}}_{o}) is a triplet of subject node vks∈V{v_{k}}_{s}\in\mathcal{V}, object node vko∈V{v_{k}}_{o}\in\mathcal{V} and relation rkr_{k}, the direction of which is from vko{v_{k}}_{o} to vks{v_{k}}_{s}. Relation rkr_{k} is associated with a preposition/verb word or phrase from SS, and eke_{k} indicates that subject node vks{v_{k}}_{s} is modified by object node vko{v_{k}}_{o}.

2 Structured Reasoning

We perform structured reasoning on the nodes and edges of graphs using neural modules under the guidance of the structure of language scene graph G\mathcal{G}. In particular, we first design the inference order and reasoning rules for its nodes V\mathcal{V} and edges E\mathcal{E}. Then, we follow the inference order to perform reasoning. For each node, we adopt the AttendNode module to find its corresponding node in graph Go\mathcal{G}^{o} or use the Merge module to combine information from its incident edges. For each edge, we execute specific reasoning steps using carefully designed neural modules, including AttendNode, AttendRelation and Transfer.

In this section, we first introduce the inference order, and then present specific reasoning steps on the nodes and edges respectively. In general, for every node in language scene graph G\mathcal{G}, we learn its attention map over the nodes of image semantic graph Go\mathcal{G}^{o} on the basis of its connections.

Given a language scene graph G\mathcal{G}, we locate the node with zero out-degree as its referent node vrefv_{ref} because the referent is usually modified by other entities rather than modifying other entities in a referring expression. Then, we perform breadth-first traversal of the nodes in graph G\mathcal{G} from the referent node vrefv_{ref} by reversing the direction of all edges, meanwhile, push the visited nodes into a stack which is initially empty. Next, we iteratively pop one node from the stack and perform reasoning on the popped node. The stack determines the inference order for the nodes, and one node can reach the top of the stack only after all of its modifying nodes have been processed. This inference order essentially converts graph G\mathcal{G} into a directed acyclic graph. Without loss of generality, suppose node vmv_{m} is popped from the stack in the present iteration, and we carry out reasoning on node vmv_{m} on the basis of its connections to other nodes. There are two different situations: 1) If the in-degree of vmv_{m} is zero, vmv_{m} is a leaf node, which means node vmv_{m} is not modified by any other nodes. Thus, node vmv_{m} should be associated with the nodes of image semantic graph Go\mathcal{G}^{o} independently; 2) otherwise, if node vmv_{m} has incident edges Em∈E\mathcal{E}_{m}\in\mathcal{E} starting from other nodes, vmv_{m} is an intermediate node, and its attention map over Vo\mathcal{V}^{o} should depend on the attention maps of its connected nodes and the edges between them.

. We learn an embedding for the words associated with the nodes of the language scene graph G\mathcal{G} in advance. Then, for node vmv_{m}, suppose its associated phrase consists of words {wt}t=1T\{w_{t}\}_{t=1}^{T}, and the embedded feature vectors for these words are {ft}t=1T\{\mathbf{f}_{t}\}_{t=1}^{T}. We use a bi-directional LSTM to compute the context of every word in this phrase, and define the concatenation of the forward and backward hidden vectors of a word wtw_{t} as its context, denoted as ht\mathbf{h}_{t}. Meanwhile, we represent the whole phrase using the concatenation of the last hidden vectors of both directions, denoted as h\mathbf{h}. In a referring expression, an individual entity is often described by its appearance and spatial location. Therefore, we learn feature representations for node vmv_{m} from both appearance and spatial location. In particular, inspired by self-attention in , we first learn the attention over each word on the basis of its context, and obtain feature representations vmlook\mathbf{v}^{look}_{m} and vmloc\mathbf{v}^{loc}_{m} at node vmv_{m} by aggregating attention weighted word embedding as follows,

where Wlook\mathbf{W}_{look} and Wloc\mathbf{W}_{loc} are learnable parameters, and vmlook\mathbf{v}^{look}_{m} and vmloc\mathbf{v}^{loc}_{m} correspond to the appearance and spatial location of node vmv_{m}. Then, we feed these two features into the AttendNode neural module to compute attention maps {λn,mlook}n=1N\{\lambda^{look}_{n,m}\}_{n=1}^{N} and {λn,mloc}n=1N\{\lambda^{loc}_{n,m}\}_{n=1}^{N} over the nodes of image semantic graph Go\mathcal{G}^{o}. Finally, we combine these two attention maps to obtain the final attention map for node vmv_{m}. A noun phrase may place emphasis on appearance, spatial location or both of them. We flexibly adapt to the variations of noun phrases by learning a pair of weights at node vmv_{m} for the attention maps related to appearance and spatial location. The weights (i.e. βlook\beta^{look} and βloc\beta^{loc}) and the final attention map {λn,m}n=1N\{\lambda_{n,m}\}_{n=1}^{N} for node vmv_{m} are computed as follows,

where W0T\mathbf{W}^{T}_{0}, b0b_{0}, W1T\mathbf{W}^{T}_{1} and b1b_{1} are learnable parameters, and the Norm module is used to constrain the scale of the attention map.

. As an intermediate node, vmv_{m} is connected to other nodes that modify it, and such connections are actually a subset of edges, Em∈E\mathcal{E}_{m}\in\mathcal{E}, incident to vmv_{m}. We compute an attention map over the edges of image semantic graph Go\mathcal{G}^{o} for each edge in this subset, then transfer and combine all these attention maps to obtain a final attention map for node vmv_{m}.

For each edge ek=(vks,rk,vko)e_{k}=({v_{k}}_{s},r_{k},{v_{k}}_{o}) in Em\mathcal{E}_{m} (where vks{v_{k}}_{s} is exactly vmv_{m}), we first form a sentence associated with eke_{k} by concatenating the words or phrases associated with vks{v_{k}}_{s}, rkr_{k} and vko{v_{k}}_{o}. Then, we obtain the embedded feature vectors {ft}t=1T\{\mathbf{f}_{t}\}_{t=1}^{T} and word contexts {ht}t=1T\{\mathbf{h}_{t}\}_{t=1}^{T} for the words {wt}t=1T\{w_{t}\}_{t=1}^{T} in this sentence and the feature representation of the whole sentence by following the same computation for leaf nodes. Next, we compute the attention map for node vks{v_{k}}_{s} from two different aspects, i.e. subject description and relation-based transfer, because eke_{k} not only directly describes subject vks{v_{k}}_{s} itself but also its relation to object vko{v_{k}}_{o}. From the aspect of subject description, same as the computation for leaf nodes, we obtain attention maps corresponding to the appearance and spatial location of vks{v_{k}}_{s} ( i.e. {λn,kslook}n=1N\{\lambda^{look}_{n,k_{s}}\}_{n=1}^{N} and {λn,ksloc}n=1N\{\lambda^{loc}_{n,k_{s}}\}_{n=1}^{N}) and weights (i.e. βkslook\beta^{look}_{k_{s}} and βksloc\beta^{loc}_{k_{s}}) to combine them. From the aspect of relation-based transfer, we first compute a relational feature representation for edge eke_{k} as follows,

where Wrel\mathbf{W}_{rel} is a learnable parameter. Then we feed the relational representation rk\mathbf{r}_{k} to the AttendRelation neural module to attend the relation rk\mathbf{r}_{k} over the edges Eijo\mathcal{E}_{ij}^{o} of graph Go\mathcal{G}^{o}, and the computed attention weights are denoted as {γij,k}i,j=1N\{\gamma_{ij,k}\}_{i,j=1}^{N}. Moreover, we use the Transfer module and the Norm module to transfer the attention map {λn,ko}n=1N\{\lambda_{n,k_{o}}\}_{n=1}^{N} for object node vko{v_{k}}_{o} to node vmv_{m} by modulating {λn,ko}n=1N\{\lambda_{n,k_{o}}\}_{n=1}^{N} with the attention weights on edges {γij,k}i,j=1N\{\gamma_{ij,k}\}_{i,j=1}^{N}, and the transferred attention map for node vmv_{m} is denoted as {λn,ksrel}n=1N\{\lambda^{rel}_{n,k_{s}}\}_{n=1}^{N}. It is worth mentioning that object node vko{v_{k}}_{o} has been accessed before and the attention map {λn,ko}n=1N\{\lambda_{n,k_{o}}\}_{n=1}^{N} for node vko{v_{k}}_{o} has been computed. Next, we estimate the weight of relation at edge eke_{k} and integrate the attention maps for node vks{v_{k}}_{s} related to subject description and relation-based transfer to obtain attention map {λn,ks}n=1N\{\lambda_{n,k_{s}}\}_{n=1}^{N} for node vks{v_{k}}_{s} contributed by edge eke_{k}, and {λn,ks}n=1N\{\lambda_{n,k_{s}}\}_{n=1}^{N} is defined as follows.

where W2\mathbf{W}_{2} and b2b_{2} are learnable parameters.

Finally, we combine the attention maps {{λn,ks}n=1N}\{\{\lambda_{n,k_{s}}\}_{n=1}^{N}\} for node vmv_{m} contributed by all edges in Em\mathcal{E}_{m} using the Merge module followed by the Norm module to obtain the final attention map {λn,m}n=1N\{\lambda_{n,m}\}_{n=1}^{N} for node vmv_{m}.

2.2 Neural Modules

We present a series of neural modules to perform specific reasoning steps, inspired by the neural modules in . In particular, the AttendNode and AttendRelation modules are used to connect the language mode with the vision mode. They receive feature representations of linguistic contents from the language scene graph and output attention maps of the features defined over visual contents in the image semantic graph. The Merge, Norm and Transfer modules are adopted to further integrate and transfer attention maps over the nodes and edges of the image semantic graph.

AttendNode [appearance query, location query]

module aims to find relevant nodes among the nodes of the image semantic graph Go\mathcal{G}^{o} given an appearance query and location query. It takes the query vectors of the appearance query and location query as inputs and generate attention maps {λnlook}n=1N\{\lambda^{look}_{n}\}_{n=1}^{N} and {λnloc}n=1N\{\lambda^{loc}_{n}\}_{n=1}^{N} over the nodes Vo\mathcal{V}^{o}, where every node vno∈Vov^{o}_{n}\in\mathcal{V}^{o} has two attention weights, i.e., λnlook∈\lambda^{look}_{n}\in and λnloc∈\lambda^{loc}_{n}\in. The query vectors are linguistic features at nodes of the language scene graph, denoted as vlook\mathbf{v}^{look} and vloc\mathbf{v}^{loc}. For node vnov^{o}_{n} in graph Go\mathcal{G}^{o}, its attention weights λnlook\lambda^{look}_{n} and λnloc\lambda^{loc}_{n} are defined as follows,

where MLP0()\text{MLP}_{0}(), MLP1()\text{MLP}_{1}(), MLP2()\text{MLP}_{2}() and MLP3()\text{MLP}_{3}() are multilayer perceptrons consisting of several linear and ReLU layers, L2Norm() is the L2 normalization, and vno\mathbf{v}^{o}_{n} and pno\mathbf{p}_{n}^{o} are the visual feature and spatial feature at node vnov^{o}_{n} respectively, which are mentioned in Section 3.1.1.

module aims to find relevant edges in the image semantic graph Go\mathcal{G}^{o} given a relation query. The purpose of a relation query is to establish connections between nodes in graph Go\mathcal{G}^{o}. Given query vector e\mathbf{e}, the attention weights {γij}i,j=1N\{\gamma_{ij}\}_{i,j=1}^{N} on edges {eijo}i,j=1N\{\mathbf{e}^{o}_{ij}\}_{i,j=1}^{N} are defined as follows,

where MLP5()\text{MLP}_{5}(), MLP5()\text{MLP}_{5}() are multilayer perceptrons, and the ReLU activation function σ\sigma ensures the attention weights are larger than zero.

module aims to find new nodes by passing attention weights {λn}n=1N\{\lambda_{n}\}_{n=1}^{N} on nodes that modify those new nodes along attended edges {γij}i,j=1N\{\gamma_{ij}\}_{i,j=1}^{N}. The updated attention weights {λnnew}n=1N\{\lambda^{new}_{n}\}_{n=1}^{N} are calculated as follows,

module aims to combine multiple attention maps generated from different edges of the same node, where the attention weights over edges are computed individually. Given the set of attention maps Λ\Lambda for a node, the merged attention map {λn}n=1N\{\lambda_{n}\}_{n=1}^{N} is defined as follows,

module aims to set the range of weights in attention maps to [−1, 1][-1,\,1]. If the maximum absolute value of an attention map is larger than 11, the attention map is divided by the maximum absolute value.

3 Loss Function

Once all the nodes in the stack have been processed, the final attention map for the referent node of the language scene graph is obtained. This attention map is denoted as {λn,ref}n=1N\{\lambda_{n,ref}\}_{n=1}^{N}. As in previous methods for grounding referring expressions , during the training phase, we adopt the cross-entropy loss, which is defined as

where pgtp_{gt} is the probability of the ground-truth object. During the inference phase, we predict the referent by choosing the object with the highest probability.

Ref-Reasoning Dataset

The proposed dataset is built on the scenes from the GQA dataset . We automatically generate referring expressions for every image on the basis of the image scene graph using a diverse set of expression templates.

. We generate referring expressions according to the ground-truth image scene graphs. Specifically, we adopt the scene graph annotations provided by the Visual Genome dataset and further normalized by the GQA dataset. In a scene graph annotation of an image, each node represents an object with about 1-3 attributes, and each edge represents a relation (i.e., semantic relation, spatial relation and comparatives) between two objects. In order to use the scene graphs for referring expression generation, we remove some unnatural edges and classes, e.g., “nose left of eyes”. In addition, we add edges between objects to represent same-attribute relations between objects, i.e., “same material”, “same color” and “same shape”. In total, there are 1,664 object classes, 308 relation classes and 610 attribute classes in the adopted scene graphs.

. In order to generate referring expressions with diverse reasoning layouts, for each specified number of nodes, we design a family of referring expression templates for each reasoning layout. We generate expressions according to layouts and templates using functional programs, and the functional program for each template can be easily obtained according to the layout. In particular, layouts are sub-graphs of directed acyclic graphs, where only one node (i.e., the root node) has zero out-degree and other nodes can reach the root node. The functional program for a layout provides a step-wise plan for reaching the root node from leaf nodes (i.e., the nodes with zero in-degree) by traversing all the nodes and edges in this layout, and templates are parameterized natural language expressions, where the parameters can be filled in. Moreover, we set the constraint that the number of nodes in a template ranges from one to five.

2 Generation Process

Given an image, we generate dozens of expressions from the scene graph of the image, and the generation process for one expression is summarized as follows,

Randomly sample the referent node and randomly decide the number of nodes, denoted as CC.

Randomly sample a sub-graph with CC nodes including the referent node in the scene graph.

Judge the layout of the sub-graph and randomly sample a referring expression template from the family of templates corresponding to the layout.

Fill in the parameters in the template using contents of the sub-graph, including relations and objects with randomly sampled attributes.

Execute the functional program with filled parameters and accept the expression if the referred object is unique in the scene graph.

Note that we perform extra operations during the generation process: 1) If there are objects that have same-attribute relations in the sub-graph, we avoid choosing the attributes that appear in such relations for these objects. This restriction intends to make the modified node identified by the relation edge instead of the attribute directly. 2) To help balance the dataset, during the process of random sampling, we decrease the chances of nodes and relations whose classes most commonly exist in the scene graphs. In addition, we increase the chances of multi-order relationships with C=3C=3 or C=4C=4 to reasonably increase the level of difficulty for reasoning. 3) We define a difficulty level for a referring expression. We find its shortest sub-expression which can identify the referent in the scene graph, and the number of objects in the sub-expression is defined as the difficulty level. For example, if there is only one bottle in an image, the difficulty level of “the bottle on a table beside a plate” is still one even though it describes three objects and their relations. Then, we obtain the balanced dataset and its final splits by randomly sampling expressions of images according to their difficulty level and the number of nodes described by them.

Experiments

We have conducted extensive experiments on the proposed Ref-Reasoning dataset as well as on three commonly used benchmark datasets (i.e., RefCOCO, RefCOCO+ and RefCOCOg). Ref-Reasoning contains 791,956 referring expressions in 83,989 images. It has 721,164, 36,183 and 34,609 expression-referent pairs for training, validation and testing, respectively. Ref-Reasoning includes semantically rich expressions describing objects, attributes, direct relations and indirect relations with different layouts. RefCOCO and RefCOCO+ datasets includes short expressions collected from an interactive game interface. RefCOCOg collects from a non-interactive settings and it has longer complex expressions.

2 Implementation and Evaluation

The performance of grounding referring expressions is evaluated by accuracy, i.e., the fraction of correct predictions of referents.

For the Ref-Reasoning dataset, we use a ResNet-101 based Faster R-CNN as the backbone, and adopt a feature extractor which is trained on the training set of GQA with an extra attribute loss following . Visual features of annotated objects are extracted from the pool5 layer of the feature extractor. For the three common benchmark datasets (i.e., RefCOCO, RefCOCO+ and RefCOCOg), we follow CMRIN to extract the visual features of objects in images. To keep the image semantic graph sparse and reduce computational cost, we connect each node in the image semantic graph to its five nearest nodes based on the distances between their normalized center coordinates. We set the mini-batch size to 64. All the models are trained by the Adam optimizer with the learning rate set to 0.0001 and 0.0005 for the Ref-Reasoning dataset and other benchmark datasets respectively.

3 Comparison with the State of the Art

We conduct experimental comparisons between the proposed SGMN and existing state-of-the-art methods on both the collected Ref-Reasoning dataset and three commonly used benchmark datasets.

We evaluate two baselines (i.e., a CNN model and a CNN+LSTM model), two state-of-the-art methods (i.e., CMRIN and DGA ) and the proposed SGMN on the Ref-Reasoning dataset. The CNN model is allowed to access objects and images only. The CNN+LSTM model embeds objects and expressions into a common feature space and learns matching scores between them. For CMRIN and DGA, we adopt their default settings in our evaluation. For a fair comparison, all the models use the same visual object features and the same setting in LSTMs.

Table 1 shows the evaluation results on the Ref-Reasoning dataset. The proposed SGMN significantly outperforms the baselines and existing state-of-the-art models, and it consistently achieves the best performance on all the splits of the testing set, where different splits need different numbers of reasoning steps. The CNN model has a low accuracy of 12.15%, which is much lower than the accuracy (i.e., 41.1% ) of the image-only model for the RefCOCOg dataset, which demonstrates that joint understanding of images and text is required on Ref-Reasoning. The CNN+LSTM model achieves a high accuracy of 75.29% on the split where expressions directly describe the referents. This is because relation reasoning is not required in this split and LSTM may be qualified to capture the semantics of expressions. Compared with the CNN+LSTM model, DGA and CMRIN achieve higher performance on the two-, three- and four-node splits because they learn a language-guided contextual representation for objects.

Quantitative evaluation results on RefCOCO, RefCOCO+ and RefCOCOg datasets are shown in Table 2. The proposed SGMN consistently outperforms existing structured methods across all the datasets, and it improves the average accuracy over the testing sets achieved by the best performing existing structured method by 0.92%, 2.54% and 2.96% respectively on the RefCOCO, RefCOCO+ and RefCOCOg datasets. Moreover, it also surpasses all the existing models on the RefCOCOg dataset which has relatively longer complex expressions with an average length 8.43, and achieves a performance comparable to the best performing holistic method on the other two common benchmark datasets. Note that holistic models usually have higher performance than structured models on the common benchmark datasets because those datasets include many simple expressions describing the referents without relations, and holistic models are prone to learn shallow correlations without reasoning and may exploit this dataset bias . In addition, the inference mechanism of holistic methods has poor interpretability.

4 Qualitative Evaluation

Visualizations of two examples along with their language scene graphs and attention maps over the objects in images at every node of the language scene graphs are shown in Figure 3. This qualitative evaluation results demonstrate that the proposed SGMN can generate interpretable visual evidences of intermediate steps in the reasoning process. In Figure 3(a), SGMN parses the expression into a tree structure and finds the referred “woman” who is walking “a dog” and meanwhile is with “pink and black hair”. Figure 3(b) shows a more complex expression which describes four objects and their relations. SGMN first successfully changes from the initial attention map (bottom-right) to the final attention map (top-right) at the node “a chair” by performing relational reasoning along the edges (i.e., triplets (“a chair”, “to the left of” , “lamp”) and (“lamp”, “on”, “a floor”)), and then identifies the target “blanket” on that chair.

5 Ablation Study

To demonstrate the effectiveness of reasoning under the guidance of scene graphs inferred from referring expressions as well as the design of neural modules, we train four additional models for comparison. The results are shown in Table 3. All the models have similar performance on the split of expressions directly describing the referents. For the other splits, SGMN without the Transfer module and SGMN without the Norm module have much lower performance than the original SGMN because the former treats the referent as an isolated node without performing relation reasoning while the latter unfairly treats different relational edges and the nodes connected by them. Next, we explore different options of the function (i.e., max, min and sum) used in the Merge module. Compared to SGMN with sum-merge, its performance with min-merge and max-merge drops because max-merge only captures the most significant relation for each intermediate node and min-merge is sensitive to parsing errors and recognition errors.

Conclusion

In this paper, we present a scene graph guided modular network (SGMN) for grounding referring expressions. It performs graph-structured reasoning over the constructed graph representations of the input image and expression using neural modules. In addition, we propose a large-scale real-world dataset for structured referring expression reasoning, named Ref-Reasoning. Experimental results demonstrate that SGMN not only significantly outperforms existing state-of-the-art algorithms on the new Ref-Reasoning dataset, but also surpasses state-of-the-art structured methods on commonly used benchmark datasets. Moreover, it can generate interpretable visual evidences of reasoning via a graph attention mechanism.

References