Large-Scale Visual Relationship Understanding

Ji Zhang, Yannis Kalantidis, Marcus Rohrbach, Manohar Paluri, Ahmed Elgammal, Mohamed Elhoseiny

Introduction

Scale matters. In the real world, people tend to describe visual entities with open vocabulary, e.g., the raw ImageNet (Deng et al., 2009) dataset has 21,841 synsets that cover a vast range of objects. The number of entities is significantly larger for relationships since the combinations of ⟨\langlesubject, relation, object⟩\rangle are orders of magnitude more than objects (Lu et al., 2016; Plummer et al., 2017; Zhang et al., 2017c). Moreover, the long-tailed distribution of objects can be an obstacle for a model to learn all classes sufficiently well, and such challenge is exacerbated in relationship detection because either the subject, the object, or the relation could be infrequent, or their triple might be jointly infrequent. Figure 1 shows an example from the Visual Genome dataset, which contains commonly seen relationship (e.g., ⟨\langleman,wearing,glasses⟩\rangle) along with uncommon ones (e.g., ⟨\langledog,next to,woman⟩\rangle).

Another challenge is that object categories are often semantically associated (Deng et al., 2009; Krishna et al., 2017; Deng et al., 2014), and such connections could be more subtle for relationships since they are conditioned on the contexts. For example, an image of ⟨\langleperson,ride,horse⟩\rangle could look like one of ⟨\langleperson,ride,elephant⟩\rangle since they both belong to the kind of relationships where a person is riding an animal, but ⟨\langleperson,ride,horse⟩\rangle would look very different from ⟨\langleperson,walk with,horse⟩\rangle even though they have the same subject and object. It is critical for a model to be able to leverage such conditional connections.

In this work, we study relationship recognition at an unprecedented scale where the total number of visual entities is more than 80,000. To achieve that we use a continuous output space for objects and relations instead of discrete labels. We demonstrate superiority of our model over competitive baselines on a large and imbalanced benchmark based of Visual Genome that comprises 53,000+53,000+ objects and 29,000+29,000+ relations. We also achieve state-of-the-art performance on the Visual Relationship Detection (VRD) dataset (Lu et al., 2016), and the scene graph dataset (Xu et al., 2017).

Related Work

Visual Relationship Detection A large number of visual relationship detection approaches have emerged during the last couple of years. Almost all of them are based on a small vocabulary, e.g., 100 object and 70 relation categories from the VRD dataset (Lu et al., 2016), or a subset of VG with the most frequent object and relation categories (Zhang et al., 2017a; Xu et al., 2017; Zhang et al., 2018b, a, 2019).

In one of the earliest works, Lu et al. (2016) utilize the object detection output of an an R-CNN detector and leverage language priors from semantic word embeddings to fine-tune the likelihood of a predicted relationship. Very recently, Zhuang et al. (2017) use language representations of the subject and object as “context” to derive a better classification result for the relation. However, similar to Lu et al. (2016) their language representations are pre-trained. Unlike these approach, we fine-tune subject and object representations jointly and employ the interaction between branches also at an earlier stage before classification.

In Yu et al. (2017), the authors employ knowledge distillation from a large Wikipedia-based corpus and get state-of-the-art results for the VRD (Lu et al., 2016) dataset. In ViP-CNN (Li et al., 2017), the authors pose the problem as a classification task on limited classes and therefore cannot scale to the open-vocabulary scenarios. In our model we exploit co-occurrences at the relationship level to model such knowledge. Our approach directly targets the large category scale and is able to utilize semantic associations to compensate for infrequent classes, while at the same time achieves competitive performance in the smaller and constrained VRD (Lu et al., 2016) dataset.

Very recent approaches like Zhao et al. (2017); Plummer et al. (2017) target open-vocabulary for scene parsing and visual relationship detection, respectively. In Plummer et al. (2017), the related work closest to ours, the authors learn a CCA model on top of different combinations of the subject, object and union regions and train a Rank SVM. They however consider each relationship triplet as a class and learn it as a whole entity, thus cannot scale to our setting. Our approach embeds the three components of a relationship separately to the independent semantic spaces for object and relation, but implicitly learns connections between them via visual feature fusion and semantic meaning preservation in the embedding space.

Semantically Guided Visual Recognition. Another parallel category of vision and language tasks is known as zero-shot/few-shot, where class imbalance is a primary assumption. In Frome et al. (2013), Norouzi et al. (2014) and Socher et al. (2013), word embedding language models (e.g., Mikolov et al. (2013)) were adopted to represent class names as vectors and hence allow zero-shot recognition. For fine-grained objects like birds and flowers, several works adopted Wikipedia Articles to guide zero-shot/few-shot recognition (Lei Ba et al., 2015; Elhoseiny et al., 2017). However, for relations and actions, these methods are not designed with the capability of locating the objects or interacting objects for visual relations. Several approaches have been proposed to model the visual-semantic embedding in the context of the image-sentence similarity task (e.g., Kiros, Salakhutdinov, and Zemel (2014); Faghri et al. (2018); Wang, Li, and Lazebnik (2016); Gong et al. (2014)). Most of them focused on leaning semantic connections between the two modalities, which we not only aim to achieve, but with a manner that does not sacrifice discriminative capability since our task is detection instead of similarity-based retrieval. In contrast, visual relationship also has a structure of ⟨\langlesubject, relation, object⟩\rangle and we show in our results that proper design of a visual-semantic embedding architecture and loss is critical for good performance.

Note: in this paper we use “relation” to refer to what is also known as ‘predicate” in previous works, and “relationship” or “relationship triplet” to refer to a ⟨\langlesubject, relation, object⟩\rangle tuple.

Method

Figure 2 shows the work flow of our model. We take an image as input to the visual module and output three visual embeddings xs,xp,x^{s},x^{p}, and xox^{o} for subject, relation, and object. During training we take word vectors of subject, relation, object as input to the semantic module and output three semantic embeddings ys,yp,yoy^{s},y^{p},y^{o}. We minimize the loss by matching the visual and semantic embeddings using our designed losses. During testing we feed word vectors of all objects and relations and use nearest neighbor searching to predict relationship labels. The following sections describe our model in details.

The design logic of our visual module is that a relation exists when its subject and object exist, but not vice versa. Namely, relation recognition is conditioned on subject and object, but object recognition is independent from relations. The main reason is that we want to learn embeddings for subject and object in a separate semantic space from the relation space. That is, we want to learn a mapping from visual feature space (which is shared among subject/object and relation) to the two separate semantic embedding spaces (for objects and relations). Therefore, involving relation features for subject/object embeddings would have the risk of entangling the two spaces. Following this logic, as shown in Figure 2 an image is fed into a CNN (conv1_1conv1\_1 to conv5_3conv5\_3 of VGG16) to get a global feature map of the image, then the subject, relation and object features zsz^{s}, zpz^{p}, zoz^{o} are ROI-pooled with the corresponding regions RS\mathcal{R_{S}}, RP\mathcal{R_{P}}, RO\mathcal{R_{O}}, each branch followed by two fully connected layers which output three intermediate hidden features h2sh^{s}_{2}, h2ph^{p}_{2}, h2oh^{o}_{2}. For the subject/object branch, we add another fully connected layer w3sw^{s}_{3} to get the visual embedding xsx^{s}, and similarly for the object branch to get xox^{o}. For the relation branch, we apply a two-level feature fusion: we first concatenate the three hidden features h2sh^{s}_{2}, h2ph^{p}_{2}, h2oh^{o}_{2} and feed it to a fully connected layer w3pw^{p}_{3} to get a higher-level hidden feature h3ph^{p}_{3}, then we concatenate the subject and object embeddings xsx^{s} and xox^{o} with h3ph^{p}_{3} and feed it to two fully connected layers w4pw^{p}_{4} w5pw^{p}_{5} to get the relation embedding xpx^{p}.

Semantic Module

On the semantic side, we feed word vectors of subject, relation and object labels into a small MLP of one or two fcfc layers which outputs the embeddings. As in the visual module, the subject and object branches share weights while the relation branch is independent. The purpose of this module is to map word vectors into an embedding space that is more discriminative than the raw word vector space while preserving semantic similarity. During training, we feed the ground-truth labels of each relationship triplet as well as labels of negative classes into the semantic module, as the following subsection describes; during testing, we feed the whole sets of object and relation labels into it for nearest neighbors searching among all the labels to get the top kk as our prediction.

A good word vector representation for object/relation labels is critical as it provides proper initialization that is easy to fine-tune on. We consider the following word vectors:

Pre-trained word2vec embeddings (wiki). We rely on the pre-trained word embeddings provided by Mikolov et al. (2013) which are widely used in prior work. We use this embedding as a baseline, and show later that by combining with other embeddings we achieve better discriminative ability.

Relationship-level co-occurrence embeddings (relco). We train a skip-gram word2vec model that tries to maximize classification of a word based on another word in the same context. As is in our case we define context via our training set’s relationships, we effectively learn to maximize the likelihoods of P(P∣S,O)P(P|S,O) as well as P(S∣P,O)P(S|P,O) and P(O∣S,P)P(O|S,P). Although maximizing P(P∣S,O)P(P|S,O) is directly optimized in Yu et al. (2017), we achieve similar results by reducing it to a skip-gram model and enjoy the scalability of a word2vec approach.

Node2vec embeddings (node2vec). As the Visual Genome dataset further provides image-level relation graphs, we also experimented with training node2vec embeddings as in Grover and Leskovec (2016). These are effectively also word2vec embeddings, but the context is determined by random walks on a graph. In this setting, nodes correspond to subjects, objects and relations from the training set and edges are directed from S→PS\rightarrow P and from P→OP\rightarrow O for every image-level graph. This embedding can be seen as an intermediate between image-level and relationship level co-occurrences, with proximity to the one or the other controlled via the length of the random walks.

Training Loss

To learn the joint visual and semantic embedding we employ a modified triplet loss. Traditional triplet loss (Kiros, Salakhutdinov, and Zemel, 2014) encourages matched embeddings from the two modalities to be closer than the mismatched ones by a fixed margin, while our version tries to maximize this margin in a softmax form. In this subsection we review the traditional triplet loss and then introduce our triplet-softmax loss in a comparable fashion. To this end, we denote the two sets of triplets for each positive visual-semantic pair by (xl,yl)(\textbf{x}^{l},\textbf{y}^{l}):

where l∈{s,p,o}l\in\{s,p,o\}, and the two sets trix,triytri_{\textbf{x}},tri_{\textbf{y}} correspond to triplets with negatives from the visual and semantic space, respectively.

where NN is the number of positive ROIs, KK is the number of negative samples per positive ROI, mm is the margin between the distances of positive and negative pairs, and s(⋅,⋅)s(\cdot,\cdot) is a similarity function.

We can observe from Equation (3) that as long as the similarity between positive pairs is larger than that between negative ones by margin mm, [m+s(xi,xij−)−s(xi,yi)]≤0[m+s(\textbf{x}_{i},\textbf{x}_{ij}^{-})-s(\textbf{x}_{i},\textbf{y}_{i})]\leq 0, and thus max(0,⋅)max(0,\cdot) will return zero for that part. That means, during training once the margin is pushed to be larger than mm, the model will stop learning anything from that triplet. Therefore, it is highly likely to end up with an embedding space where points are not discriminative enough for a classification-oriented task.

It is worth noting that although theoretically traditional triplet loss can pushes the margin as much as possible when m=1m=1, most previous works (e.g., Kiros, Salakhutdinov, and Zemel (2014); Faghri et al. (2018); Gordo and Larlus (2017)) adopted a small mm to allow slackness during training. It is also unclear how to determine the exact value of mm given a specific task. We follow previous works and set m=0.2m=0.2 in all of our experiments.

Triplet-Softmax loss. The issue of triplet loss mentioned above can be alleviated by applying softmax on top of each triplet, i.e. missing:

where s(⋅,⋅)s(\cdot,\cdot) is the same similarity function (we use cosine similarity in this paper). All the other notations are the same as above. For each positive pair (xi,yi)(\textbf{x}_{i},\textbf{y}_{i}) and its corresponding set of negative pairs (xi,yij−)(\textbf{x}_{i},\textbf{y}_{ij}^{-}), we calculate similarities between each of them and put them into a softmax layer followed by multi-class logistic loss so that the similarity of positive pairs would be pushed to be 11, and otherwise. Compared to triplet loss, this loss always tries to enlarge the margin to its largest possible value (i.e. missing, 1), thus has more discriminative power than the traditional triplet loss.

Visual Consistency loss. To further force the embeddings to be more discriminative, we add a loss that pulls closer the samples from the same category while pushes away those from different categories, i.e.:

where NN is the number of positive ROIs, C(l)\mathcal{C}(l) is the set of positive ROIs in the same class of xi\textbf{x}_{i}, KK is the number of negative samples per positive ROI and mm is the margin between the distances of positive and negative pairs. The interpretation of this loss is: the minimum similarity between samples from the same class should be larger than any similarity between samples from different classes by a margin. Here we utilize the traditional triplet loss format since we want to introduce slackness between visual embeddings to prevent embeddings from collapsing to the class centers.

Empirically we found it the best to use triplet-softmax loss for Ly\mathcal{L}_{\textbf{y}} while using triplet loss for Lx\mathcal{L}_{\textbf{x}}. The reason is similar with that of the visual consistency loss: mode collapse should be prevented by introducing slackness. On the other hand, there is no such issue for yy since each label yy is a mode by itself, and we encourage all modes of yy to be separated from each other. In conclusion, our final loss is:

where we found that α=β=1\alpha=\beta=1 works reasonably well for all scenarios.

Implementation details. For all the three datasets, we train our model for 77 epochs using 8 GPUs. We set learning rate as 0.0010.001 for the first 55 epochs and 0.00010.0001 for the rest 22 epochs. We initialize each branch with weights pre-trained on COCO Lin et al. (2014). For the word vectors, we used the gensim library Řehůřek and Sojka (2010) for both word2vec and node2vechttps://github.com/aditya-grover/node2vec Grover and Leskovec (2016). For the triplet loss, we set m=0.2m=0.2 as the default value.

For the VRD and VG200 datasets, we need to predict whether a box pair has relationship, since unlike VG80k where we use ground-truth boxes, here we want to use general proposals that might contain non-relationships. In order for that, we add an additional “unknown” category to the relation categories. The word “unknown” is semantically dissimilar with any of the relations in these datasets, hence its word vector is far away from those relations’ vectors.

There is a critical factor that significantly affects our triplet-softmax loss. Since we use cosine similarity, s(⋅,⋅)s(\cdot,\cdot) is equivalent to dot product of two normalized vectors. We empirically found that simply feeding normalized vector could cause gradient vanishing problem, since gradients are divided by the norm of input vector when back-propagated. This is also observed in Bell et al. (2016) where it is necessary to scale up normalized vectors for successful learning. Similar with Bell et al. (2016), we set the scalar to a value that is close to the mean norm of the input vectors and multiply s(⋅,⋅)s(\cdot,\cdot) before feeding to the softmax layer. We set the scalar to 3.23.2 for VG80k and 3.03.0 for VRD in all experiments.

ROI Sampling. One of the critical things that powers Fast-RCNN is the well-designed ROI sampling during training. It ensures that for most ground-truth boxes, each has 3232 positive ROIs and 128−32=96128-32=96 negative ROIs, where positivity is defined as overlap IoU >=0.5>=0.5. In our setting, ROI sampling is similar for the subject/object branch, while for the relation branch, positivity is defined as both subject and object IoUs >=0.5>=0.5. Accordingly, we sample 6464 subject ROIs with 3232 unique positives and 3232 unique negatives, and do the same thing for object ROIs. Then we pair all the 6464 subject ROIs with 6464 object ROIs to get 40964096 ROI pairs as relationship candidates. For each candidate, if both ROIs’ IoU >=0.5>=0.5 we mark it as positive, otherwise negative. We finally sample 3232 positive and 9696 negative relation candidates and use the union of each ROI pair as a relation ROI. In this way we end up with a consistent number of positive and negative ROIs for the relation branch.

Experiments

Datasets. We present experiments on three datasets, the original Visual Genome (VG80k) (Krishna et al., 2017), the version of Visual Genome with 200 categories (VG200) (Xu et al., 2017), and Visual Relationship Detection (VRD) dataset (Lu et al., 2016).

VRD. The VRD dataset (Lu et al., 2016) contains 5,000 images with 100 object categories and 70 relations. In total, VRD contains 37,993 relation annotations with 6,672 unique relations and 24.25 relationships per object category. We follow the same train/test split as in Lu et al. (2016) to get 4,000 training images and 1,000 test images. We use this dataset to demonstrate that our model can work reasonably well on small dataset with small category space, even though it is designed for large-scale settings.

VG200. We also train and evaluate our model on a subset of VG80k which is widely used in previous methods (Xu et al., 2017; Newell and Deng, 2017; Zellers et al., 2018; Yang et al., 2018). There are totally 150150 object categories and 5050 predicate categories in this dataset. We use the same train/test splits as in Xu et al. (2017). Similarly with VRD, the purpose here is to show our model is also state-of-the-art in large-scale sample but small-scale category settings.

VG80k. We use the latest version of Visual Genome (VG v1.4) (Krishna et al., 2017) that contains 108,077108,077 images with 2121 relationships on average per image. We follow Johnson, Karpathy, and Fei-Fei (2016) and split the data into 103,077103,077 training images and 5,0005,000 testing images. Since text annotations of VG are noisy, we first clean it by removing non-alphabet characters and stop words, and use the autocorrect library to correct spelling. Following that, we check if all words in an annotation exist in the word2vec dictionary (Mikolov et al., 2013) and remove those that do not. We run this cleaning process on both training and testing set and get 99,96199,961 training images and 4,8714,871 testing images, with 53,30453,304 object categories and 29,08629,086 relation categories. We further split the training set into 97,96197,961 training and 2,0002,000 validation images.We will release the cleaned annotations along with our code.

Evaluation protocol. For VRD, we use the same evaluation metrics used in Yu et al. (2017), which runs relationship detection using non-ground-truth proposals and reports recall rates using the top 50 and 100 relationship predictions, with k=1,10,70k=1,10,70 relations per relationship proposal before taking the top 50 and 100 predictions.

For VG200, we use the same evaluation metrics used in Zellers et al. (2018), which uses three modes: 1) predicate classification: predict predicate labels given ground truth subject and object boxes and labels; 2) scene graph classification: predict subject, object and predicate labels given ground truth subject and object boxes; 3) scene graph detection: predict all the three labels and two boxes. Recalls under the top 20, 50, 100 predictions are used as metrics. The mean is computed over the 3 evaluation modes over R@50 and R@100 as in Zellers et al. (2018).

For VG80k, we evaluate all methods on the whole 53,30453,304 object and 29,08629,086 relation categories. We use ground-truth boxes as relationship proposals, meaning there is no localization errors and the results directly reflect recognition ability of a model. We use the following metrics to measure performance: (1) top1, top5, and top10 accuracy, (2) mean reciprocal ranking (rr), defined as 1M∑i=1M1ranki\frac{1}{M}\sum_{i=1}^{M}\frac{1}{rank_{i}}, (3) mean ranking (mr), defined as 1M∑i=1Mranki\frac{1}{M}\sum_{i=1}^{M}{rank_{i}}, smaller is better.

We first validate our model on VRD dataset with comparison to state-of-the-art methods using the metrics presented in Yu et al. (2017) in Table 1. Note that there is a variable kk in this metric which is the number of relation candidates when selecting top50/100. Since not all previous methods specified kk in their evaluation, we first report performance in the “free kk” column when considering kk as a hyper-parameter that can be cross-validated. For methods where the kk is reported for 1 or more values, the column reports the performance using the best kk. We then list all available results with specific kk in the right two columns.

For fairness, we split the table in two parts. The top part lists methods that use the same proposals from Lu et al. (2016), while the bottom part lists methods that are based on a different set of proposals, and ours uses better proposals obtained from Faster-RCNN as previous works. We can see that we outperform all other methods with proposals from Lu et al. (2016) even without using message-passing-like post processing as in Li et al. (2017); Dai, Zhang, and Lin (2017), and also very competitive to the overall best performing method from Yu et al. (2017). Note that although spatial features could be advantageous for VRD according to previous methods, we do not use them in our model in concern of large-scale settings. We expect better performance if integrating spatial features for VRD, but for model consistency we do experiments without it everywhere.

Scene Graph Classification & Detection on VG200

We present our results in Table 2. Note that scene graph classification isolates the factor of subject/object localization accuracy by using ground truth subject/object boxes, meaning that it focuses more on the relationship recognition ability of a model, and predicate classification focuses even more on it by using ground truth subject/object boxes and labels. It is clear that the gaps between our model and others are higher on scene graph/predicate classification, meaning our model displays superior relation recognition ability.

Relationship Recognition on VG80k

Baselines. Since there is no previous method that has been evaluated in our large-scale setting, we carefully design 3 baselines to compare with. 1) 3-branch Fast-RCNN: an intuitively straightforward model is a Fast-RCNN with a shared conv1conv1 to conv5conv5 backbone and 3 fcfc branches for subject, relation and object respectively, where the subject and object branches share weights since they are essentially an object detector; 2) our model with softmax loss: we replace our loss with softmax loss; 3) our model with triplet loss: we replace our loss with triplet loss.

Results. As shown in Table 3, we can see that our loss is the best for the general case where all instances from all classes are considered. The baseline has reasonable performance but is clearly worse than ours with softmax, demonstrating that our visual module is critical for efficient learning. Ours with triplet is worse than ours with softmax in the general case since triplet loss is not discriminative enough among the massive data. However it is the opposite for tail classes (i.e., #occurrence≤1024\#occurrence\leq 1024), since recognition of infrequent classes can benefit from the transferred knowledge learned from frequent classes, which the softmax-based model is not capable of. Another observation is that although the 3-branch Fast-RCNN baseline works poorly in the general case, it is better than our model with softmax. Since the main difference of them is with and without visual feature concatenation, it means that integrating subject and object features does not necessarily helps infrequent relation classes. This is because subject and object features could lead to strong prior on the relation, resulting in lower chance of predicting infrequent relation when using softmax. For example, when seeing a rare image where the relationship is “dog ride horse”, subject being “dog” and object being “horse” would give very little probability to the relation “ride”, even though it is the correct answer. Our model alleviates this problem by not mapping visual features directly to the discrete categorical space, but to a continuous embedding space where visual similarity is preserved. Therefore, when seeing the visual features of “dog”, “horse” and the whole “dog ride horse” context, our model is able to associate them with a visually similar relationship “person ride horse” and correctly output the relation “ride”.

Ablation Study

Variants of our model. We explore variants of our model in 4 dimensions: 1) the semantic embeddings fed to the semantic module; 2) structure of the semantic module; 3) structure of the visual module; 4) the losses. The default settings of them are 1) using wiki + relco; 2) 2 semantic layer; 3) with both visual concatenation; 4) with all the 3 loss terms. We fix the other 3 dimensions as the default settings when exploring one of them.

The scaling factor before softmax. As mentioned in the implementation details, this value scales up the output by a value that is close to the average norm of the input and prevents gradient vanishing caused by the normalization. Specifically, for Eq(7) in the paper we use s(x,y)=λxTy∣∣x∣∣∣∣y∣∣s(\textbf{x},\textbf{y})=\lambda\frac{\textbf{x}^{T}\textbf{y}}{||\textbf{x}||||\textbf{y}||} where λ\lambda is the scaling factor. In Table 5 we show results of our model when changing the value of the scaling factor applied before the softmax layer. We observe that when the value is close to the average norm of all input vectors (i.e., 5.0), we achieve optimal performance, although slight difference of this value does not change results too much (i.e., when it is 4.0 or 6.0). It is clear that when the scaling factor is 1.0, which is equivalent to training without scaling, the model is not sufficiently trained. We therefore pick 5.0 for this scaling factor for all the other experiments on VG80k.

Which semantic embedding to use? We explore 4 settings: 1) wiki and 2) relco use wikipedia and relationship-level co-occurrence embedding alone, while 3) wiki + relco and 4) wiki + node2vec use concatenation of two embeddings. The intuition of concatenating wiki with relco and node2vec is that wiki contains common knowledge acquired outside of the dataset, while relco and node2vec are trained specifically on VG80k, and their combination provides abundant information for the semantic module. As shown in Table 4, fusion of wiki and relco outperforms each one alone with clear margins. We found that using node2vec alone does not perform reasonably, but wiki + node2vec is competitive to others, demonstrating the efficacy of concatenation.

Number of semantic layers. We also study how many, if any, layers are necessary to embed the word vectors. As it is shown in Table 4, directly using the word vectors (0 semantic layers) is not a good substitute of our learned embedding; raw word vectors are learned to represent as much associations between words as possible, but not to distinguish them. We find that either 1 or 2 layers give similarly good results and 2 layers are slightly better, though performance starts to degrade when adding more layers.

Are both visual feature concatenations necessary? In Table 4, “early concat” means using only the first concatenation of the three branches, and “late concat” means the second. Both early and late concatenation boost performance significantly compared to no concatenation, and it is the best with both. Another observation is that late concatenation is better than early alone. We believe the reason is, as mentioned above, relations are naturally conditioned on and constrained by subjects and objects, e.g., given “man” as subject and “chair” as object, it is highly likely that the relation is “sit on”. Since late concatenation is at a higher level, it integrates features that are more semantically close to the subject and object labels, which gives stronger prior to the relation branch and affects relation prediction more than the early concatenation.

Do all the losses help? In order to understand how each loss helps training, we trained 3 models of which each excludes one or two loss terms. We can see that using Ly+Lx\mathcal{L}_{y}+\mathcal{L}_{x} is similar with Ly\mathcal{L}_{y}, and it is the best with all the three losses. This is because Lx\mathcal{L}_{x} pulls positive xx pairs close while pushes negative xx away. However, since (x,y)(x,y) is a many-to-one mapping (i.e., multiple visual features could have the same label), there is no guarantee that all xx with the same yy would be embedded closely, if not using Lc\mathcal{L}_{c}. By introducing Lc\mathcal{L}_{c}, xx with the same yy are forced to be close to each other, and thus the structural consistency of visual features is preserved.

The margin m in triplet loss We show results of triplet loss with various values for the margin mm in Table 6. As described earlier, this value allows slackness in pushing negative pairs away from positive ones. We observe similar results with previous works (Kiros, Salakhutdinov, and Zemel, 2014; Faghri et al., 2018) that it is the best to set m=0.1m=0.1 or m=0.2m=0.2 in order to achieve optimal performance. It is clear that triplet loss is not able to learn discriminative embeddings that are suitable for classification tasks, even with larger mm that can theoretically enforce more contrast against negative labels. We believe that the main reason is that in a hinge loss form, triplet loss treats all negative pairs equally “hard” as long as they are within the margin mm. However, as shown by the successful softmax models, “easy” negatives (e.g., those that are close to positives) should be penalized less than those “hard” ones, which is a property our model has since we utilize softmax for contrastive training.

The VG80k has densely annotated relationships for most images with a wide range of types. In Figure 6 there are interactive relationships such as “boy flying kite”, “batter holding bat”, positional relationships such as “glass on table”, “man next to man”, attributive relationships such as “man in suit” and “boy has face”. Our model is able to cover all these kinds, no matter frequent or infrequent, and even for those incorrect predictions, our answers are still semantic meaningful and similar to the ground-truth, e.g., the ground-truth “lamp on pole” v.s. the predicted “light on pole”, and the ground-truth “motorcycle on sidewalk” v.s. the predicted “scooter on sidewalk”.

Conclusions

In this work we study visual relationship detection at an unprecedented scale and propose a novel model that can generalize better on long tail class distributions. We find it is crucial to integrate subject and object features at multiple levels for good relation embeddings and further design a loss that learns to embed visual and semantic features into a shared space, where semantic correlations between categories are kept without hurting discriminative ability. We validate the effectiveness of our model on multiple datasets, both on the classification and detection task, and demonstrate the superiority of our approach over strong baselines and the state-of-the-art. Future work includes integrating a relationship proposal into our model that would enable end-to-end training.

References