Graphical Contrastive Losses for Scene Graph Parsing

Ji Zhang, Kevin J. Shih, Ahmed Elgammal, Andrew Tao, Bryan Catanzaro

Introduction

Given an image, the aim of scene graph parsing is to infer a visually grounded graph comprising localized entity categories, along with predicate edges denoting their pairwise relationships. This is often formulated as the detection of ⟨subject,predicate,object⟩\langle subject,predicate,object\rangle triplets within an image, e.g. ⟨man,holds,guitar⟩\langle man,holds,guitar\rangle in Figure 2(b). Current state-of-the-art methods achieve this goal by a two-stage mechanism: first detecting entities, then predicting a predicate for each pair of entities.

We find that scene graph parsing models using such pipelines tend to struggle with two types of errors. The first is Entity Instance Confusion, in which the subject or object is related to one of many instances of the same class, and the model fails to distinguish between the target instance and the others. We show an example in Figure 4(a), in which the model identifies the man is holding a wine glass, but struggles to determine exactly which of the 3 visually similar wine glasses is being held. The incorrectly predicted wine glass is transparent and intersecting with the left arm, which makes it look like being held. The second type of error, Proximal Relationship Ambiguity, occurs when the image contains multiple subject-object pairs interacting in the same way, and the model fails to identify the correct pairing. An example can be seen in the multiple musicians ”playing” their respective instruments in Figure 4(b). Due to their close proximity, visual features for each musician-instrument pair overlap significantly, making it difficult for the scene graph models to identify the correct pairings.

The primary cause of these two failures lies in the inherent difficulty of inferring relationships like “hold” and “play” from visual cues. Concretely, which glass is being held is determined by the small part of the hand that covers the glass. Whether a player is playing the drum can only be inferred by very subtle visual cues such as his standing pose or where his fingers are placed. It is challenging for any model to learn to attend to these details precisely, and it would be impractical to specify which details to focus on for all kinds of relationships, let alone to learn all these details. These challenges motivate the need for a mechanism that can automatically learn fine details that determine visual relationships, and explicitly discriminate related entities from unrelated ones, for all types of relationships. This is the goal of our work.

In this paper we propose a set of Graphical Contrastive Losses to tackle these issues. The losses use the form of the margin-based triplet loss, but are specifically designed to address the two aforementioned errors. It adds additional supervision in the form of hard negatives specific to Entity Instance Confusion and Proximal Relationship Ambiguity. To demonstrate the effectiveness of our proposed losses, we design a relationship detection network named RelDN using the aforementioned pipeline with our losses. Figure 2 shows a result of RelDN with N-way cross-entropy loss only vs. with our additional contrastive losses. Our best model achieves 0.328 on the Private set of the OpenImages Relationship Detection Challenge, outperforming the winning model by a significant 4.7% (16.5% relative) margin. It also attains state-of-the-art performance on the Visual Genome and VRD datasets.

In this paper, we denote subject, predicate, object and attribute with s,pred,o,as,pred,o,a. We use “entity” to describe individual detected objects to distinguish from “object” in the semantic sense, and use “relationships” to describe the entire ⟨s,pred,o⟩\langle s,pred,o\rangle tuple, not to be confused with “predicate,” which is an element of said tuple.

Related Work

Scene Graph Parsing: A large number of scene graph parsing approaches have emerged during the last couple of years. They use the same pipeline that first either uses off-the-shelf detectors or detectors fine-tuned with relationship datasets to detect entities, then predicts the predicate using proposed methods. Most of them model the second step as a classification task that takes features of each entity pair as input and output a label independently from other pairs. instead learn embeddings for subjects, predicates and objects and use nearest neighbor searching during testing to predict predicates. Nevertheless, the prediction is still done on each entity pair individually. We show that this pipeline struggles with two major scenarios. We find that ignoring the intrinsic graph structure of relationships and predicting each predicate separately is the main cause. Our proposed losses compensate for such drawback by contrasting positive against negative edges for each node, providing global supervision to the classifier and significantly alleviating those two issues.

The scene graph parsing work most related to ours is Associative Embedding . They use use a push and pull contrastive loss to train embeddings for entities within a visual genome scene graph. Our work differs in that we propose to have different sets of hard negatives to target specific error types within scene graph parsing.

Phrase Grounding and Referring Expressions: Phrase grounding and referring expression models aim to localize the region described by a given expression, with the latter focusing more on cases of possible reference confusion . It can be abstracted as a bipartite graph matching problem, where nodes on the visual side are the regions and nodes on the language side are the expressions, and the goal is to find all matched pairs. In contrast, scene graphs are arbitrarily connected graphs whose nodes are visual entities and edges are predicates with rich semantic information. Our losses are designed to leverage that information to better discriminate between related and non-related entities.

Contrastive Training: Contrastive training using a triplet loss has wide application in both computer vision and natural language processing. Representative works include Negative Sampling and Noise Contrastive Sampling . More recent work also utilizes it to solve multi-modal tasks such as phrase grounding, image captioning, VQA, and vector embeddings . Our setting differs in that we define hard negative contrastive margins along the known structure of the annotated scene graph, allowing us to specifically target entity instance and proximal relationship confusion. By adding our losses as additional supervision on top of the N-way cross-entropy loss, we are able to improve the model by significant margins.

Graphical Contrastive Losses

Our Graphical Contrastive Losses encompass three types of loss, each addressing the two aforementioned issues in their own way: 1) Class Agnostic: contrasts positive/negative entity pairs regardless of their relation and adds contrastive supervision for generic cases; 2) Entity Class Aware: addresses the issue in Figure 4(a) by focusing on entities with the same class; 3) Predicate Class Aware: addresses the issue in Figure 4(b) by focusing on entity pairs with the same potential predicate. We define our contrastive losses over an affinity term Φ(s,o)\Phi(s,o), which can be interpreted as the probability that subject ss and object oo have some relationship or interaction. Given a model that outputs the distribution over predicate classes conditioned on a subject and object pair p(pred∣s,o)p(pred|s,o), we define Φ(s,o)\Phi(s,o) as:

where ∅\varnothing is the class symbol representing no_relationship. This is equivalent to summing over all predicate classes except ∅\varnothing.

Our first contrastive loss term aims to maximize the affinity of the lowest scoring positive pairing and minimize the affinity of the highest scoring negative pairing. For a subject indexed by ii and an object indexed by jj, the margins we wish to maximize can be written as:

where Vi+\mathcal{V}_{i}^{+} and Vi−\mathcal{V}_{i}^{-} represent sets of objects related to and not related to subject sis_{i}; Vj+\mathcal{V}_{j}^{+} and Vj−\mathcal{V}_{j}^{-} are defined similarly for object jj as the sets of subjects related to and not related to ojo_{j}.

The class agnostic loss for all sampled positive subjects and objects is written as:

where NN is the number of annotated entities and α1\alpha_{1} is the margin threshold.

This loss tries to contrast positive and negative (s,o)(s,o) pairs, ignoring any class information, and is similar to the triplet losses used referring expression and phrase-grounding literature. We found it works as well in our scenario and even better with the following class-aware losses, as shown in Table 1.

2 Entity Class Aware Loss

The Entity Class Aware loss deals with entity instance confusion, in which the model struggles to determine interactions between a subject (object) and multiple instances of a same-class object (subject). It can be viewed as an extension of the Class Agnostic loss where we further specify a class cc when populating the positive and negative sets V+\mathcal{V}^{+} and V−\mathcal{V}^{-}. We extend the formulation in equation (3) as:

where Vic+\mathcal{V}_{i}^{c+}, Vic−\mathcal{V}_{i}^{c-}, Vjc+\mathcal{V}_{j}^{c+} and Vjc−\mathcal{V}_{j}^{c-} are now constrained to instances of class cc.

The entity class aware loss for all sampled positive subjects and objects is defined as

where C()\mathcal{C}() returns the set of unique classes of the sets Vi+\mathcal{V}_{i}^{+} and Vj+\mathcal{V}_{j}^{+} as defined in the class agnostic loss. Compared to the class agnostic loss which maximizes the margins across all instances, this loss maximizes the margins between instances of the same class. It forces a model to disentangle confusing entities illustrated in Figure 4(a), where the subject has several potentially related objects with the same class.

3 Predicate Class Aware Loss

Similar to the entity class aware loss, this loss maximizes the margins within groups of instances determined by their associated predicates. It is designed to deal with the proximal relationship ambiguity as exemplified in Figure 4(b), where instances joined by the same predicate class are within close proximity of each other. In the context of Figure 4(b), this loss would encourage the correct pairing of who is playing which instrument by penalizing wrong pairing, i.e., “man plays drum” in the red box. Replacing the class groupings in equation (4) with predicate groupings restricted to predicate class ee, we define our margins to maximize as:

Here, we define the sets Vie+\mathcal{V}_{i}^{e+} and Vje+\mathcal{V}_{j}^{e+} as the sets of subject-object pairs where the ground truth predicate between sis_{i} and ojo_{j} is ee, anchored with respect to subject ii and object jj respectively. We define the sets Vie−\mathcal{V}_{i}^{e-} and Vje−\mathcal{V}_{j}^{e-} as is the set of instances where the model incorrectly predicts (via argmax) the predicate to be ee, anchored with respect to subject ii and object jj respectively.

The predicate class aware loss for all sampled positive subjects and objects is defined as

where E()\mathcal{E}() returns the set of unique predicates associated with the input (excluding ∅\varnothing). The final loss is expressed as:

where L0L_{0} is the cross-entropy loss over predicate classes.

4 Complexity Analysis

We look at the case where the subject sis_{i} is fixed and we vary object for positive/negative pairings. The reverse case (object fixed, subject varies) has the same complexity. All sampling is conducted on the entities of a single image per batch. The set of entities include ground truth bounding boxes, as well as any detector output with >=0.5>=0.5 IOU to ground truth entities.

For the Class Agnostic Loss L1L_{1}, the computational complexity of the sampling procedure is O(N2)O(N^{2}), where NN is the upper bounded on number of sampled entities per image. In practice, for each subject, we randomly sample at most KK non-related objects (negative pairings), which makes the actual complexity O(NK)O(NK).

For the Entity Class Aware Loss L2L_{2}, the sampling procedure is the same as with L1L_{1}, except that we need to keep only those non-related objects that are of class cc, i.e., the object class of the current oo in the sampled (s,o)(s,o) pair. This involves a filtering operation on the KK objects which takes O(K)O(K) time, therefore the overall complexity is still O(NK)O(NK).

The analysis for the Predicate Class Aware Loss L3L_{3} is similar to that of L2L_{2}, except that the filtering operation looks at the predicate class ee instead of the object class cc. The overall complexity is also O(NK)O(NK).

We set N=512N=512 and K=64K=64 per batch in practice.

RelDN

We demonstrate the efficacy of our proposed losses with our Relationship Detection Network (RelDN). The RelDN follows a two stage pipeline: it first identifies a proposal set of likely subject-object relationship pairs, then extracts features from these candidate regions to perform a fine-grained classification into a predicate class. We build a separate CNN branch for predicates (conv_body_rel) with the same structure as that of entity detector CNN (conv_body_det) to extract predicate features. The intuition for having a separate branch is that we want visual features for predicates to focus on the interactive areas of subjects and objects as opposed to individual entities. As Figure 7 illustrates, the predicate CNN clearly learns better features which concentrate on regions that strongly imply relationships.

The first stage of the RelDN exhaustively returns bounding box regions containing every pair. In the second stage, it computes three types of features for each relationship proposal: semantic, visual, and spatial. Each feature is used to output a set of class logits, which we combine via element-wise addition, and apply softmax normalization to attain a probability distribution over predicate classes. See Figure 5 for our model pipeline.

Semantic Module: The semantic module conditions the predicate class prediction on subject-object class co-occurrence frequencies. It is inspired by Zeller, et al. which introduced a frequency baseline that performs reasonably well on Visual Genome by counting frequencies of predicates given subject and object. Its motivation is that in general, the combination of relationships between two entities is usually very limited, e.g., the relationship between a person-horse subject-object pairing is most likely to be “ride”, “walk”, or “feed”, and unlikely to be “stand on” or “wear”. For each training image, we count the occurrences of predicate class predpred given subject and object classes ss and oo in the ground truth annotations. This gives us an empirical distribution p(pred∣s,o)p(pred|s,o). We assume that the test set is also drawn from the same distribution.

Spatial Module: The spatial module conditions the predicate class predictions on the relative positions of the subject and object. One of the major predicate types are about positions, for example, “on”, “under”, or “inside_of.” These predicate types can often be inferred using only relative spatial information. We capture spatial information by encoding the box coordinates of subjects and objects using the box delta and normalized coordinates.

We define the delta feature between two sets of bounding box coordinates as follows:

where b1b_{1} and b2b_{2} are two coordinate tuples in the form of (x,y,w,h)(x,y,w,h).

We then compute the normalized coordinate features for a bounding box bb as follows:

where wimgw_{img} and himgh_{img} are the width and height dimensions of the image. Our spatial feature vector for the subject, object, and predicate bounding boxes bsb_{s}, bob_{o}, bpredb_{pred} is represented as:

Note that bpredb_{pred} is the tightest bounding box around bsb_{s} and bob_{o}. This feature vector is fed through an MLP to attain predicate class logit scores.

Visual Module: The visual module produces a set of class logits conditioned ROI feature maps, as in the fast-RCNN pipeline. We extract subject and object ROI features from the entity detector’s convolution layers (conv_body_det in Figure 5) and extract predicate ROI features from the relationship convolution layers (conv_body_rel in Figure 5). The subject, object, and predicate feature vectors are concatenated and passed through an MLP to attain the predicate class logits.

We also include two skip-connections projecting subject-only and object-only ROI features to the predicate class logits. These skip connections are inspired by the observation that many relationships, such as human interactions , can be accurately inferred by the appearance of only the subjects or objects. We show an improvement from adding these skip connections in 6.4.

Module Fusion: As illustrated in Figure 5, we obtain the final probability distribution over predicate classes by adding the three scores followed by softmax normalization:

where fvis,fspt,fsem\textbf{f}_{vis},\textbf{f}_{spt},\textbf{f}_{sem} are unnormalized class logits from the visual, spatial, semantic modules.

Implementation Details

We train the entity detector CNN (conv_body_det) independently using entity annotations, then fix it when training our model. While previous works claim it is beneficial to fine-tune the entity detector end-to-end with the second stage of the pipeline, we opt to freeze our entity detector weights for simplicity. We initialize the predicate CNN (conv_body_rel) with the entity detector’s weights and fine-tune it end-to-end with the second stage.

During training, we independently sample positive and negative pairs for each loss, subject to their respective constraints. For L0L_{0}, we sample 512 pairs in total where 128 of them are positive. For our class-agnostic loss, we sample 128 positive subjects, then for each of them sample the two closet contrastive pairs according to Eq.2; we do the sampling symmetrically for objects. For our entity and predicate aware losses, we sample in the same way with class-agnostic except that negative pairs are grouped by entity and predicate classes, as described in Eq.4,6. We set λ1=1.0,λ2=0.5,λ3=0.1\lambda_{1}=1.0,\lambda_{2}=0.5,\lambda_{3}=0.1, determined by cross-validations, for all experiments.

During testing, we take up to 100 outputs from the entity detector and exhaustively group all pairs as relationship proposals/entity pairs. We rank relationship proposals by multiplying the predicted subject, object, predicate probabilities as pdet(s)⋅ppred(pred)⋅pdet(o)\textbf{p}^{det}(s)\cdot\textbf{p}^{pred}(pred)\cdot\textbf{p}^{det}(o) where pdet(s),pdet(o)\textbf{p}^{det}(s),\textbf{p}^{det}(o) are the probabilities of the predicted subject and object classes from the entity detector, and ppred(pred)\textbf{p}^{pred}(pred) is the probability of the predicted predicate class from the result of Eq.12.

To match the architectures of previous state-of-the-art methods, We use ResNeXt-101-FPN as our OpenImages backbone and VGG-16 on Visual Genome (VG) and Visual Relationship Detection (VRD).

Experiments

We present experimental results on three datasets: OpenImages (OI) , Visual Genome (VG) and Visual Relationship Detection (VRD) . We first report evaluation settings, followed by ablation studies and finally external comparisons.

OpenImages: The full train and val sets contains 53,953 and 3,234 images, which takes our model 2 days to train. For quick comparisons, we sample a “mini” subset of 4,500 train and 1,000 validation images where predicate classes are sampled proportionally with a minimum of one instance per class in train and val. We first conduct parameter searches on the mini set, then train and compare with the top model of the OpenImages VRD Challenge on the full set. We show two types of results, one using the same entity detector from the top model, and the other using a detector trained by our own initialized by COCO pre-trained weights.

In the OpenImages Challenge, results are evaluated by calculating Recall@50 (R@50), mean AP of relationships (mAPrel), and mean AP of phrases (mAPphr). The final score is obtained by score=0.2×R@50+0.4×mAPrel+0.4×mAPphr\text{score}=0.2\times R@50+0.4\times mAP_{rel}+0.4\times mAP_{phr}. The mAPrel evaluates AP of s,pred,os,pred,o triplets where both the subject and object boxes have an IOU of at least 0.5 with ground truth. The mAPphr is similar, but applied to the enclosing relationship boxMore details of evaluation can be found on the official page: https://storage.googleapis.com/openimages/web/vrd_detection_metric.html. In practice, we find mAPrel and mAPphr to suffer from extreme predicate class imbalance. For example, 64.48% of the relationships in val have the predicate “at”, while only 0.03% of them are “under”. This means a single “under” relationship is worth much more than the more common “at” relationships. We address this by scaling each predicate category by their relative ratios in the val set, which we refer to as the weighted mAP (wmAP). We use wmAP in all of our ablation studies (Table 1-4), in addition to reporting scorewtd which replaces mAP with wmAP in the score formula.

We compare with other top models on the official evaluation server. The official test set is split into a Public and Private set with a 30%/70% split. The Public set is used as a dev set. We present individual results for both, as well as their weighted average under Overall in Table 9.

Visual Genome: We follow the same train/val splits and evaluation metrics as . We train our entity detector initialized by COCO pre-trained weights. Following , we conduct three evaluations: scene graph detection(SGDET), scene graph classification (SGCLS), and predicate classification (PRDCLS). We report results for these tasks with and without the Graphical Contrastive Losses.

VRD: We evaluate our model with entity detectors initialized by ImageNet and COCO pre-trained weights. We use the same evaluation metrics as in , which reports R@50 and R@100 for relationship predictions at 1, 10, and 70 predicates per entity pair.

2 Loss Analysis

Loss Combinations: We now look at whether our proposed losses reduce two aforementioned errors without affecting the overall performance, and whether all three losses are necessary. Results in Table 1 show that combination of all the three losses with the N-way cross-entropy loss (L0+L1+L2+L3L_{0}+L_{1}+L_{2}+L_{3}) has consistently superior performance over just L0L_{0}. Notably, APrel on “holds” improves by from 41.84 to 43.09 (+1.3). It improves even more significantly from 36.04 to 41.04 (+5.0) on “plays” and from 40.43 to 44.16 (+3.7) on “interacts_with” respectively. These three classes suffer the most from the two aforementioned problems. Our results also show that any subset of the losses is worse than the entire ensemble. We see that L0+L1L_{0}+L_{1}, L0+L2L_{0}+L_{2} and L0+L3L_{0}+L_{3} are inferior to L0+L1+L2+L3L_{0}+L_{1}+L_{2}+L_{3}, especially on “holds”, “plays”, and “interacts_with”, where the largest margin is 3.87 (L0+L2L_{0}+L_{2} vs. L0+L1+L2+L3L_{0}+L_{1}+L_{2}+L_{3} on “play”).

To better verify the isolated impact of our losses, we carefully sample a subset of 100 images containing five predicates that significantly suffer from the two aforementioned problems, selected via visual inspection on a random set of images. The five predicates are “at”, “holds”, “plays”, “interacts_with”, and “wears”. We sample them by looking at the raw images and select those with either entity instance confusion or proximal relationship ambiguity. Example images can be found in Figure 13. Table 2 shows comparison of our losses with L0L_{0} only on this subset. The overall gap is 1.4 and the largest gap is 4.1 at APrel on “holds”.

Figure 9 shows two examples from this subset, one containing entity instance confusion and the other containing proximal relationship ambiguity. In Figure 9(a) the model with only L0L_{0} fails to identify the wine glass being held, while by adding our losses, the area surrounding the correct wine glass lights up. In Figure 9(b) ⟨woman,plays,drum⟩\langle woman,plays,drum\rangle is incorrectly predicted since the L0L_{0}-only model mistakenly pairs the unplayed drum with the singer – a reasonable error considering the amount of person-play-drum examples as well as the relative proximities between the singer and the drum. Our losses successfully suppress that region and attend to the correct microphone being held, demonstrating the effectiveness of our hard-negative sampling strategies.

Margin Thresholds: We study the effects of various values of the margin thresholds α1,α2,α3\alpha_{1},\alpha_{2},\alpha_{3} used in Eq.3,5,7. For each experiment, we set α1=α2=α3=m\alpha_{1}=\alpha_{2}=\alpha_{3}=m while varying mm. As shown in Table 4, we observe similar results with previous work that m=0.1m=0.1 or m=0.2m=0.2 achieves the best performance. Note that m=1.0m=1.0 is the largest possible margin, as our affinity scores range from 0 to 1.

3 Loss Analysis with the Official mAP metrics

Here, we show our ablation studies using the official uniform-class-weighting evaluation metrics, mAPrel, mAPphr and score. We also include mAPrel*{}_{rel}^{\textbf{*}}, mAPphr*{}_{phr}^{\textbf{*}} and score*{}^{\textbf{*}}, which is the standard mAP and score excluding “under” and “hits” in the evaluation. Table 5 presents ablation study results on loss components. Table 6 shows comparison between the L0L_{0}-only model against the model with our losses on the 100 selected images. In Table 5 the variation of numbers using mAP and score demonstrates the necessity of de-emphasizing the extremely infrequent classes. Note that the mAP*-based columns show a similar trend to our wmAP-based results from the paper. In Table 6, the model with our losses is still better than the L0L_{0}-only model by a non-trivial margin, mainly because the former outperform the latter on almost every per-class AP metric for those 5 selected classes. Note that since “under” and “hits” are not in the 100 image subset, there is no need to evaluate with mAPrel*{}_{rel}^{\textbf{*}}, mAPphr*{}_{phr}^{\textbf{*}} and score*{}^{\textbf{*}}.

4 Model Analysis

We conduct an effectiveness evaluation on the three modules of the RelDN. For the visual module, we also investigate the two skip-connections. As Table 3 shows, the semantic module alone cannot solve relationship detection by using language bias only. By adding the basic visual feature, i.e., the ⟨\langleS,P,O⟩\rangle concatenation, we see a significant 4.7 gain, which is further improved by adding additional separate S,O skip-connections, especially at “plays” (+3.1), “interacts_with” (+1.0), “wears” (+2.0) where subjects’ or objects’ appearance and poses are highly representative of the interactions. Finally, adding the spatial module gives the best results, and the most obvious gaps are at spatial relationships, i.e., “at” (+0.2), “on” (+0.2), “inside_of” (+2.4).

5 Comparison to State of the Art

OpenImages: We present results compared with top 5 models from the Challenge in Table 9. We surpass the 1st place Seiji by 4.7% on Private set and 2.9% on the full set, which is in fact a significant margin considering the low absolute scores and the large amount of test images (99,999 in total). Even using the same entity detector as Seiji, we noticeable gaps (1.4% and 0.8%) on the two sets.

Visual Genome: Table 7 shows that our model is better than state-of-the-arts on all metrics. It outperforms the previous best, MotifNet-LeftRight, by a 2.4% gap on Scene Graph Detection (SGDET) with Recall@100 and by a 12.7% gap on Predicate Classification (PRDCLS) with Recall@50. Note that although our entity detector is better than MotifNet-LeftRight on mAP at 50% IoU (25.5 vs. 20.0), our implementation of Frequency+Overlap baseline (Recall@20: 16.2, Recall@50: 19.8, Recall@100: 21.5) is not better than their version (Recall@20: 21.0, Recall@50: 26.2, Recall@100: 30.1), indicating that our better relationship performance mostly comes from our model design.

We also observe that our losses achieve smaller gains over the standard cross-entropy loss setup than it does on OpenImages_mini. The reasons are two-fold: 1) One of the few dominant relationship types in the Visual Genome dataset is possessive, e.g., “ear of man”, which has much fewer entity confusion issues; 2) The Recall@kRecall@k metric is less strict than mAP. If there is an image with only one ground truth, then Recall@100 will always be 100% as long as this ground truth target is within the top 100 model predictions, regardless of the ranking of the 100 outputs. As such, the small improvements in ranking the top 100 will not affect the score. Nevertheless, the improvements from our loss is still non-trivial and consistent on all metrics under different values of kk.

In addition, we also show results using a better backbone, ResNeXt-101-FPN , for the entity detector in Table 7.

VRD: Table 8 presents results on VRD compared with state-of-the-art methods. Note that only specifically states that they use ImageNet pre-trained weights while others remain unknown. Therefore, we show results for pre-training on either ImageNet or COCO. Our model is competitive with those methods when pre-trained on ImageNet, but significantly outperforms when pre-trained on COCO. The gap between L0L_{0} only and the full model is smaller when pre-trained on ImageNet than on COCO. We believe the stronger localization features from pre-training on COCO is much easier for our model and losses to leverage.

6 Qualitative Results

In Figure 11 we provide four example images where our losses correct the false predictions made by the L0L_{0} only model. Both the Entity Instance Confusion and the Proximal Relationship Ambiguity issues are included here. In the fourth row, the L0L_{0} only model is confused between two entity instances, i.e., which person is holding the microphone, while our losses manage to refer to the correct one. In the third row the relationship between the guitar player and the drum is ambiguous. Here, the L0L_{0} only model fails by predicting a false-positive, but our model trained with all losses correctly detects no relationship there.

Conclusion

In this work we present methods to overcome two major issues in scene graph parsing: Entity Instance Confusion and Proximal Relationship Ambiguity. We show that traditional multi-class cross-entropy loss does not take advantage of intrinsic knowledge of structured scene graphs and is therefore insufficient to handle these two issues. To address that, we propose Graphical Contrastive Losses which effectively utilize semantic properties of scene graphs to contrast positive relationships against hard negatives. We carefully design three types of losses to solve the issues in three aspects. We demonstrate efficacy of our losses by adding it to a model built with the same pipeline, and we achieve state-of-the-art results on three datasets.

References