LinkNet: Relational Embedding for Scene Graph
Sanghyun Woo, Dahun Kim, Donghyeon Cho, In So Kweon
Introduction
Current state-of-the-art recognition models have made significant progress in detecting individual objects in isolation . However, we are still far from reaching the goal of capturing the interactions and relationships between these objects. While objects are the core elements of an image, it is often the relationships that determine the global interpretation of the scene. The deeper understating of visual scene can be realized by building a structured representation which captures objects and their relationships jointly. Being able to extract such graph representations have been shown to benefit various high-level vision tasks such as image search , question answering , and 3D scene synthesis .
In this paper, we address scene graph generation, where the objective is to build a visually-grounded scene graph of a given image. In a scene graph, objects are represented as nodes and relationships between them as directed edges. In practice, a node is characterized by an object bounding box with a category label, and an edge is characterized by a predicate label that connects two nodes as a subject-predicate-object triplet. As such, a scene graph is able to model not only what objects are in the scene, but how they relate to each other.
The key challenge in this task is to reason about inter-object relationships. We hypothesize that explicitly modeling inter-dependency among the entire object instances can improve a model’s ability to infer their pairwise relationships. Therefore, we propose a simple and effective relational embedding module that enables our model to jointly represent connections among all related objects, rather than focus on an object in isolation. This significantly benefits main part of the scene graph generation task: relationship classification.
We further improve our network by introducing global context encoding module and geometrical layout encoding module. It is well known that fusing global and local information plays an important role in numerous visual tasks . Motivated by these works, we build a module that can provide contextual information. In particular, the module consists of global average pooling and binary sigmoid classifiers, and is trained for multi-label object classification. This encourages its intermediate features to represent all object categories present in an image, and supports our full model. Also, for the geometrical layout encoding module, we derive inspiration from the fact the most relationships in general are spatially regularized, implying that subject-object relative geometric layout can thus be a powerful cue for inferring the relationship in between. Our novel architecture results in our final model LinkNet, of which the overall architecture is illustrated in Fig. 1.
On the Visual Genome dataset, LinkNet obtains state-of-the-art results in scene graph generation tasks, revealing the efficacy of our approach. We visualize the weight matrices in relational embedding module and observe that inter-dependency between objects are indeed represented(see Fig. 2).
Contribution. Our main contribution is three-fold.
We propose a simple and effective relational embedding module in order to explicitly model inter-dependency among entire objects in an image. The relational embedding module improves the overall performance significantly.
In addition, we introduce global context encoding module and geometrical layout encoding module for more accurate scene graph generation.
The final network, LinkNet, has achieved new state-of-the art performance in scene graph generation tasks on the large-scale benchmark . Extensive ablation studies demonstrate the effectiveness of the proposed network.
Related Work
Relational reasoning has been explicitly modeled and adopted in neural networks. In the early days, most works attempted to apply neural networks to graphs, which are a natural structure for defining relations . Recently, the more efficient relational reasoning modules have been proposed . Those can model dependency between the elements even with the non-graphical inputs, aggregating information from the feature embeddings at all pairs of positions in its input (e.g., pixels or words). The aggregation weights are automatically learned driven by the target task. While our work is connected to the previous works, an apparent distinction is that we consider object instances instead of pixels or words as our primitive elements. Since the objects have variations in scale/aspect ratio, we use ROI-align operation to generate fixed 1D representations, easing the subsequent relation computations.
Moreover, relational reasoning of our model has a link to an attentional graph neural network. Similar to ours, Chen et.al. uses a graph to encode spatial and semantic relations between regions and classes and passes information among them. To do so, they build a commonsense knowledge graph( i.e.adjacency matrix) from relationship annotations in the set. However, our approach does not require any external knowledge sources for the training. Instead, the proposed model generates soft-version of adjacency matrix(see Fig. 2) on-the-fly by capturing the inter-dependency among the entire object instances.
The task of recognizing objects and the relationships has been investigated by numerous studies in a various form. This includes detection of human-object interactions , localization of proposals from natural language expressions , or the more general tasks of visual relationship detection and scene graph generation .
Among them, scene graph generation problem has recently drawn much attention. The challenging and open-ended nature of the task lends itself to a variety of diverse methods. For example: fixing the structure of the graph, then refining node and edge labels using iterative message passing ; utilizing associative embedding to simultaneously identify nodes and edges of graph and piece them together ; extending the idea of the message passing from with additional RPN in order to propose regions for captioning and solve tasks jointly ; staging the inference process in three-step based on the finding that object labels are highly predictive of relation labels ;
In this work, we utilize relational embedding for scene graph generation. It utilizes a basic self- attention mechanism within the aggregation weights. Compared to previous models that have been proposed to focus on message passing between nodes and edges, our model explicitly reasons about the relations within nodes and edges and predicts graph elements in multiple steps , such that features of a previous stage provides rich context to the next stage.
Proposed Approach
A scene graph is a topological representation of a scene, which encodes object instances, corresponding object categories, and relationships between the objects. The task of scene graph generation is to construct a scene graph that best associates its nodes and edges with the objects and the relationships in an image, respectively.
2 LinkNet
An overview of LinkNet is shown in Fig. 1. To generate a visually grounded scene graph, we need to start with an initial set of object bounding boxes, which can be obtained from ground-truth human annotation or algorithmically generated. Either cases are somewhat straightforward; In practice, we use a standard object detector, Faster R-CNN , as our bounding box model (). Given an image , the detector predicts a set of region proposals . For each proposal , it also outputs a ROI-align feature vector and an object label distribution .
We build upon these initial object features , and design a novel scene graph generation network that consists of three modules. The first module is a relational embedding module that explicitly models inter-dependency among all the object instances. This significantly improves relationship classification(). Second, global context encoding module provides our model with contextual information. Finally, the performance of predicate classification is further boosted by our geometric layout encoding.
In the following subsections, we will explain how each proposed modules are used in two main steps of scene graph generation: object classification, and relationship classification.
3 Object Classification
Then, for a given image, we can obtain N object proposal features =1,…,N. Here, we consider object-relational embedding that computes the response for one object region by attending to the features from all N object regions. This is inspired by the recent works for relational reasoning . Despite the connection, what makes our work distinctive is that we consider object-level instances as our primitive elements, whereas the previous methods operate on pixels or words .
3.2 Global Context Encoding
Here we describe the global context encoding module in detail. This module is designed with the intuition that knowing contextual information in prior may help inferring individual objects in the scene.
4 Relationship Classification
Finally, the -way relationship classification is optimized on the resulting feature as:
4.2 Geometric Layout Encoding
We hypothesize that relative geometry between the subject and object is a powerful cue for inferring the relationship between them. Indeed, many predicates have straightforward correlation with the subject-object relative geometry, whether they are geometric (e.g., ’behind’), possessive (e.g.,’has’), or semantic (e.g.,’riding’).
To exploit this cue, we encode the relative location and scale information as :
5 Loss
The whole network can be trained in an end-to-end manner, allowing the network to predict object bounding boxes, object categories, and relationship categories sequentially (see Fig. 1). Our loss function for an image is defined as:
By default, we set and as 1, and thus all the terms are equally weighted.
Experiments
We conduct experiments on Visual Genome benchmark .
Since the current work in scene graph generation is largely inconsistent in terms of data splitting and evaluation, we compared against papers that followed the original work . The experimental results are summarized in Table. 1.
The LinkNet achieves new state-of-the-art results in Visual Genome benchmark , demonstrating its efficacy in identifying and associating objects. For the scene graph classification and predicate classification tasks, our model outperforms the strong baseline by a large margin. Note that predicate classification and scene graph classification tasks assume the same perfect detector across the methods, whereas scene graph detection task depends on a customized pre-trained detector.
2 Ablation Study
In order to evaluate the effectiveness of our model, we conduct four ablation studies based on the scene graph classification task as follows. Results of the ablation studies are summarized in Table. 2 and Table. 3.
The first row of Table. 2a shows the results of more relational embedding modules. We argue that multiple modules can perform multi-hop communication. Messages between all the objects can be effectively propagated, which is hard to do via standard models. However, too many modules can arise optimization difficulty. Our model with two REMs achieved the best results. In second row of Table. 2a, we compare performance with four different reduction ratios. The reduction ratio determines the number of channels in the module, which enables us to control the capacity and overhead of the module. The reduction ratio 2 achieves the best accuracy, even though the reduction ratio 1 allows higher model capacity. We see this as an over-fitting since the training losses converged in both cases. Overall, the performance drops off smoothly across the reduction ratio, demonstrating that our approach is robust to it.
Here we construct an input() of edge-relational embedding module by combining an object class representation() and a global contextual representation(). The operations are inspired by the recent work that contextual information is critical for the relationship classification of an objects. To do so, we turn of object label probabilities into one-hot vectors via an argmax operation(committed to a specific object class label) and we concatenate it with an output() which passed through the relational embedding module(contextualized representation). As shown in the Table. 2b, we empirically confirm that both operations contribute to the performance boost.
We perform an ablation study to validate our modules in the network, which are relation embedding module, geometric layout encoding module, and global context encoding module. We remove each module to verify the effectiveness of utilizing all the proposed modules. As shown in Exp 1, 2, 3, and Ours, we can clearly see the performance improvement when we use all the modules jointly. This shows that each module plays a critical role together in inferring object labels and their relationships. Note that Exp 1 already achieves state-of-the-art result, showing that utilizing relational embedding module is crucial while the other modules further boost performance.
We conduct additional analysis to see how the network performs with the use of geometric layout encoding module. We select top-10 predicates with highest recall increase in scene graph classification task. As shown in right side of Table. 3, we empirically confirm that the recall value of geometrically related predicates are significantly increased, such as using, carrying, riding. In other words, predicting predicates which have clear subject-object relative geometry, was helped by the module.
In this experiment, we conduct an ablation study to compare row-wise operation methods in relational embedding matrix: softmax and sigmoid; As we can see in Exp 4 and Ours, softmax operation which imposes competition along the row dimension performs better, implying that explicit attention mechanism which emphasizes or suppresses relations between objects helps to build more informative embedding matrix.
In this experiment, we investigate two commonly used relation computation methods: dot product and euclidean distance. As shown in Exp 1 and 5, we observe that dot-product produces slightly better result, indicating that relational embedding behavior is crucial for the improvement while it is less sensitive to computation methods. Meanwhile, Exp 5 and 6 shows that even we use euclidean distance method, geometric layout encoding module and global context encoding module further improves the overall performance, again showing the efficacy of the introduced modules.
3 Qualitative Evaluation
We visualize our relational embedding of our network in Fig. 2. For each example, the bottom-left is the ground-truth binary triangular matrix where its entry is filled as: only if there is a non-background relationship(in any direction) between the -th and -th instances, and otherwise. The bottom-right is the trained weights of an intermediate relational embedding matrix (Eq. (4)), folded into a triangular form. The results show that our relational embedding represents inter-dependency among all object instances, being consistent with the ground-truth relationships. To illustrate, in the first example, the ground-truth matrix refers to the relationships between the ’man’(1) and his body parts(2,3); and the ’mountain’(0) and the ’rocks’(4,5,6,7), which are also reasonably captured in our relational embedding matrix. Note that our model infers relationship correctly even there exists missing ground-truths such as cell(7,0) due to sparsity of annotations in Visual Genome dataset. Indeed, our relational embedding module plays a key role in scene graph generation, leading our model to outperform previous state-of-the-art methods.
Qualitative examples of scene graph detection of our model are shown in Fig. 3. We observe that our model properly induces scene graph from a raw image.
Conclusion
We addressed the problem of generating a scene graph from an image. Our model captures global interactions of objects effectively by proposed relational embedding module. Using it on top of basic Faster R-CNN system significantly improves the quality of node and edge predictions, achieving state-of-the-art result. We further push the performance by introducing global context encoding module and geometric layout encoding module, constructing a LinkNet. Through extensive ablation experiments, we demonstrate the efficacy of our approach. Moreover, we visualize relational embedding matrix and show that relations are properly captured and utilized. We hope LinkNet become a generic framework for scene graph generation problem.
This research is supported by the Study on Deep Visual Understanding funded by the Samsung Electronics Co., Ltd (Samsung Research)