Zoom-Net: Mining Deep Feature Interactions for Visual Relationship Recognition
Guojun Yin, Lu Sheng, Bin Liu, Nenghai Yu, Xiaogang Wang, Jing Shao, Chen Change Loy
Introduction
Visual relationship recognition aims at interpreting rich interactions between a pair of localized objects, i.e., performing tuple recognition in the form of subject-predicate-object as shown in Fig. 1(a). The fundamental challenge of this task is to recognize various vaguely defined relationships given diverse spatial layouts of objects and complex inter-object interactions. To complement visual-based recognition, a promising approach is to adopt a linguistic model and learn relationships between object and predicate labels from language. This strategy has been shown effective by many existing methods . These language-based methods either apply statistical inference to the tuple label set, establish a linguistic graph as the prior, or mine linguistic knowledge from external billion-scale textual data (e.g., Wikipedia).
In this paper, we explore a novel perspective beyond the linguistic-based paradigm. In particular, contemporary approaches typically recognize the tuple subject-predicate-object via separate convolutional neural network (CNN) branches. We believe that by enhancing message sharing and feature interactions among these branches, the participating objects and their visual relationship can be better recognized. To this end, we formulate a new spatiality-aware contextual feature learning model, named as Zoom-Net. Differing from previous studies that learn appearance and spatial features separatelyTheir spatiality-streams simply apply the union region , binary masks or centroid coordinates as the abstraction of the spatial features., Zoom-Net propagates spatiality-aware object features to interact with the predicate features and broadcasts predicate features to reinforce the features of subject and object.
The core of Zoom-Net is a Spatiality-Context-Appearance Module, abbreviated as SCA-M. It consists of two novel pooling cells that permit deep feature interactions between objects and predicates, as shown in Fig. 1(b). The first cell, Contrastive ROI Pooling Cell, facilitates predicate feature learning by inversely pooling object/subject features to a matching spatial context of predicate features via a unique deROI pooling. This allows all subject and object to fall on the same spatial ‘palette’ for spatiality-aware feature learning. The second cell is called Pyramid ROI Pooling Cell. It helps object/subject feature learning through broadcasting the predicate features to the corresponding object’s/subject’s spatial area. Zoom-Net stacks multiple SCA-Ms consecutively in an end-to-end network that allows multi-scale bidirectional message passing among subject, predicate and object. As shown in Fig. 1(c), the message sharing and feature interaction not only help recognize individual objects more accurately but also facilitate the learning of inter-object relation.
Another contribution of our work is an effective strategy of mitigating ambiguity and imbalanced data distribution in subject-predicate-object annotations. Specifically, we conduct our main experiments on the challenging Visual Genome (VG) dataset , which consists of over object categories, predicates, and relationship types. The large-scale ambiguous categories and extremely imbalanced data distribution in VG dataset (Tab. 1,2) prevent previous methods from predicting reliable relationships despite they succeed in the Visual Relationship Detection (VRD) dataset with only object categories, predicates and relationships. To alleviate the ambiguity and imbalanced data distribution in VG, we reformulate the conventional one-hot classification as a -hot multi-class hierarchical recognition via a novel Intra-Hierarchical trees (IH-trees) for each label set in the tuple subject-predicate-object.
Contributions. Our contributions are summarized as follows:
1) A general feature learning module that permits feature interactions - We introduce a novel SCA-M to mining intrinsic interactions between low-level spatial information and high-level semantical appearance features simultaneously. By stacking multiple SCA-Ms into a Zoom-Net, we achieve compelling results on VG dataset thanks to the multi-scale bidirectional message passing among subject, predicate and object.
2) Multi-class Intra-Hierarchical tree - To mitigate label ambiguity in large-scale datasets, we reformulate the visual relationship recognition problem to a multi-label recognition problem. The recognizability is enhanced by introducing an Intra-Hierarchical tree (IH-tree) for the object and predicate categories, respectively. We show that IH-tree can benefit other existing methods as well.
3) Large-scale relationship recognition - Extensive experiments demonstrate the respective effectiveness of the proposed SCA-M and IH-tree, as well as their combination on the challenging large-scale VG dataset.
It is noteworthy that the proposed method differs significantly from previous works as Zoom-Net neither models explicit nor implicit label-level interactions between subject-predicate-object. We show that feature-level interactions alone, which is enabled by SCA-M, can achieve state-of-the-art performance. We further demonstrate that previous state-of-the-arts that are based on label-level interaction can benefit from the proposed SCA-M and IH-trees.
Related work
Contextual Learning. Contextual information has been employed in various tasks , e.g., object detection, segmentation, and retrieval. For example, the visual features captured from a bank of object detectors are combined with global features in . For both detection and segmentation, learning feature representations from a global view rather than the located object itself has been proven effective in . Contextual feature learning for visual relationship recognition is little explored in previous works.
Class Hierarchy. In previous studies , class hierarchy that encodes diverse label relations or structures is used to improve performances on classification and retrieval. For instance, Deng et al. improve large-scale visual recognition of object categories through forming a semantic hierarchy that consists of many levels of abstraction. While object categories can be clustered easily by their semantic similarity given the clean and explicit labels of objects, building a semantic hierarchy for visual relationship recognition can be more challenging due to noisy and ambiguous labels. Moreover, the semantic similarity between some phrases and prepositions such as walking on a versus walks near the is not directly measurable. In our paper, we employ the part-of-speech tagger toolkit to extract and normalize the keywords of these labels, e.g. walk, on and near.
Visual Relationship. Recognizing visual relationship has been shown beneficial to various tasks, including action recogntion , pose estimation , recognition and object detection , and scene graph generation . Recent works show remarkable progress in visual relationship recognition, most of which focus on measuring linguistic relations with textual priors or language models. The linguistic relations have been explored for object recognition , object detection , retrieval , and caption generation . Yu et al. employ billions of external textual data to distill useful knowledge for triplet subject-predicate-object learning. These methods do not fully explore the potential of feature learning and feature-level message sharing for the problem of visual relationship recognition. Li et al. propose a message passing strategy to encourage feature sharing between features extracted from subject-predicate-object. However, the network does not capture the relative location of different objects thus it cannot capture valid contextual information between subject, predicate and object.
Zoom-Net: Mining Deep Feature Interactions
We propose an end-to-end visual relationship recognition model that is capable of mining feature-level interactions. This is beyond just measuring the interactions among the triplet labels with additional linguistic priors, as what previous studies considered.
As shown in Fig. 2(a), given the ROI-pooled features of the subject, predicate and object, we consider a question: how to learn good features for both object (subject) and predicate? We investigate three plausible modules as follows.
Appearance Module. This module focuses on the intra-dependencies within each ROI, i.e., the features of the subject, predicate and object branches are learned independently without any message passing. We term this network structure as Appearance Module (A-M), as shown in Fig. 2(a). No contextual and spatial information can be derived from such a module.
Context-Appearance Module. The Context-Appearance Module (CA-M) directly fuses pairwise features among three branches, in which subject/object features absorb the contextual information from the predicate features, and predicate features also receive messages from both subject/object features, as shown in Fig. 2(b). Nonetheless, these features are concatenated regardless of their relative spatial layout in the original image. The incompatibility of scale and spatiality makes the fused features less optimal in capturing the required spatial and contextual information.
Spatiality-Context-Appearance Module. The spatial configuration, e.g., the relative positions and sizes of subject and object, is not sufficiently represented in CA-M. To address this issue, we propose a Spatiality-Context-Appearance module (SCA-M) as shown in Fig. 2(c). It consists of two novel spatiality-aware feature alignment cells (i.e., Contrast ROI Pooling and Pyramid ROI Pooling) for message passing between different branches. In comparison to CA-M, the proposed SCA-M reformulates the local and global information integration in a spatiality-aware manner, leading to superior capability in capturing spatial and contextual relationships between the features of subject-predicate-object.
2 Spatiality-Context-Appearance Module (SCA-M)
We denote the respective regions of interest (ROIs) of the subject, predicate and object as , , and , where is the union bounding box that tightly covers both the subject and object. The ROI-pooled features for these three ROIs are , respectively. In this section, we present the details of SCA-M. In particular, we discuss how Contrastive ROI Pooling and Pyramid ROI Pooling cells, the two elements in SCA-M, permit deep feature interactions between objects and predicates.
Contrastive ROI Pooling denotes a pair of ROI, deROI operations that the objectSubject and object refer to the same concept, thus we only take object as the example for illustration. features are at first ROI pooled for extracting normalized local features, and then these features are deROI pooled back to the spatial palette of the predicate feature , so as to generate a spatiality-aware object feature with the same size as the predicate feature, as shown in Fig. 3(b) marked by the purple triangle. Note that the remaining region outside the relative object ROI in is set to . The spatiality-resumed local feature can thus influence the respective regions in the global feature map . In practice, the proposed deROI pooling can be considered as an inverse operation of the traditional ROI pooling (green triangle in Fig. 3), which is analogous to the top-down deconvolution versus the bottom-up convolution.
There are three Contrastive ROI pooling cells presented in the SCA-M module to integrate the feature pairs subject-predicate, subject-object and predicate-object, as shown in Fig. 3(b-d). Followed by several convolutional layers, the features from subject and object are spatially fused into the predicate feature for enhanced representation capability. The proposed ROI, deROI operations differ from conventional feature fusion operations (channel-wise concatenation or summation). The latter would introduce scale incompatibility between local subject/object features and global predicate features, which could hamper feature learning in subsequent convolutional layers.
3 Zoom-Net: Stacked SCA-M
By stacking multiple SCA-Ms, the proposed Zoom-Net is capable of capturing multi-scale feature interactions with dynamic contextual and spatial information aggregation. It enables a reliable recognition of the visual relationship triplet s-p-o, where the predicate indicates the relationships (e.g., spatiality, preposition, action and etc.) between a pair of localized subject and object .
As visualized in Fig. 4, we use a shared feature extractor with convolutional layers until conv3_3 to encode appearance features of different object categories. By indicating the regions of interests (ROIs) for subject, predicate and object, the associated features are ROI-pooled to the same spatial size and respectively fed into three branches. The features in three branches are at first independently fed into two convolutional layers (the conv4_1 and conv4_2 layers in VGG-16) for a further abstraction of their appearance features. Then these features are put into the first SCA-M to fuse spatiality-aware contextual information across different branches. After receiving the interaction-augmented subject, predicate and object features from the first SCA-M, , we continue to convolve these features with another two appearance abstraction layers (mimicking the structures of conv5_1 and conv5_2 layers in VGG-16) and then forward them to the second SCA-M, . After this module, the multi-scale interaction-augmented features in each branch are fed into three fully connected layers fc_s, fc_p and fc_o to classify subject, predicate and object, respectively.
Hierarchical Relational Classification
To thoroughly evaluate the proposed Zoom-Net, we adopt the Visual Genome (VG) datasetExtremely rare labels (fewer than samples) were pruned for a valid evaluation. for its large scale and diverse relationships. Our goal is to understand the a much broader scope of relationships with a total number of relationship types (as shown in Tab. 1), in comparison to the VRD dataset that focuses on only relationships. Recognizing relationships in VG is a non-trivial task due to several reasons:
(1) Variety - There are a total of object categories and predicates, tens times than those available in the VRD dataset.
(2) Ambiguity - Some object categories share a similar appearance, and multiple predicates refer to the same relationship.
(3) Imbalance - We observe long tail distributions both for objects and predicates.
To circumvent the aforementioned challenges, existing studies typically simplify the problem by manually removing a considerable portion of the data by frequency filtering or cleaning . Nevertheless, infrequent labels like “old man” and “white shirt” contain common attributes like “man” and “shirt” and are unreasonable to be pruned. Moreover, the flat label structure assumed by these methods is limited to describe the label space of the VG dataset with ambiguous and noisy labels.
To overcome the aforementioned issues, we propose a solution by establishing two Intra-Hierarchical trees (IH-tree) for measuring intra-class correlation within objectSubject and object refer to the same term in this paper, thus we only take the object as the example for illustration. and predicate, respectively. IH-tree builds a hierarchy of concepts that systematically groups rare, noisy and ambiguous labels together with those clearly defined labels. Unlike existing works that regularize relationships across the triplet -- by external linguistic priors, we only consider the intra-class correlation to independently regularize the occurrences of the object and predicate labels. During end-to-end training, the network employs the weighted Intra-Hierarchical losses for visual relationship recognition as , where hyper-parameters balance the losses with respect to subject , predicate and object . in our experiments. We introduce IH-tree and the losses next.
We build an IH-tree, , for object with a depth of three, where the base layer consists of the raw object categories.
(1) : is extracted from by pruning noisy labels with the same concept but different descriptive attributes or in different singular and plural forms. We employ the part-of-speech tagger toolkit from NLTK and NLTK Lemmatizer to filter and normalize the noun keyword, e.g., “man” from “old man”, “bald man” and “men”.
(2) : We observe that some labels have a close semantic correlation. As shown in the top panel of Fig. 5, labels with similar semantic concepts such as “shirt” and “jacket” are hyponyms of “clothing” and need to be distinguished from other semantic concepts like “animal” and “vehicle”. Therefore, we cluster labels in to the third level by semantical similarities computed by Leacock-Chodorow distance from NLTK. We find that a threshold of is well-suited for splitting semantic concepts.
The output of the subject/object branch is a concatenation of three independent softmax activated vectors corresponded to three hierarchical levels in the IH-tree. The loss () is thus a summation of three independent softmax losses with respect to these levels, encouraging the intra-level mutual label exclusion and inter-level label dependency.
The predicate IH-tree also has three hierarchy levels. Different from the object IR-tree that only handles nouns, the predicate categories include various part-of-speech types, e.g., verb (action) and preposition (spatial position). Even a single predicate label may contain multiple types, e.g., “are standing on” and “walking next to a”.
(1) : Similar to , is constructed aiming at extracting and normalizing keywords from predicates. We retain the keywords and normalize tenses with respective to three main part-of-speech types, i.e., verb, preposition and adjective, and abandon other pointless and ambiguous words. As shown in the bottom panel of Fig. 5, “wears a”, “wearing a yellow” and “wearing a pink” are mapped to the same keyword “wear”.
(2) : Different part-of-speech types own particular characteristics with various context representations, and hence a separate hierarchical structure for the verb (action) and preposition (spatial) is indispensable for better depiction. To this end, we construct for verb and preposition label independently, i.e., for action information and for spatial configuration. There are two cases in : (a) the label is in the form of phrase that consists of both verb and preposition (e.g. “stand on” and “walk next to”) and (b) the label is a single word (e.g., “on” and “wear”). For the first case, extracts the verb words from the two phrases while extracts the preposition words. It thus causes that a label might be simultaneously clustered into different partitions of . If the label is a single word , it would be normally clustered into the corresponding part-of-speech but remained the same in the opposite part-of-speech, as shown with the dotted line in the bottom panel of Fig. 5. The loss is constructed similarly to that for the object.
Experiments on Visual Genome (VG) Dataset
Dataset. We evaluate our method on the Visual Genome (VG) dataset (version 1.2). Each image is annotated with a triplet subject-predicate-object, where the subjects and objects are annotated with labels and bounding boxes while the predicates only have labels. The detailed statistics are stated in Sec. 4 and Tab. 1, 2. We randomly split the VG dataset into training and testing set with a ratio of . Note that both sets are guaranteed to have positive and negative samples from each object or predicate category. The details of data preprocessing and the source code will be released.
Evaluation Metrics. (1) Acc@. We adopt the Accuracy score as the major evaluation metric in our experiments. The metric is commonly used in traditional classification tasks. Specifically, we report the values of both Acc@ and Acc@ for subject, predicate, object and relationship, where the accuracy of relationship is calculated as the averaged accuracies of subject, predicate and object.
(2) Rec@. Following , we use Recall as another metric so as to handle incomplete annotations. Rec@ computes the ratio of the correct relationship instance that is covered in the top predictions per image. We report Rec@ and Rec@ in our experiments. For a fair comparison, we follow to evaluate Rec@ on three tasks, i.e., predicate recognition where both the labels and bounding boxes of the subject and object are given; phrase recognition that takes a triplet as a union bounding box and predicts the triple labels; relationship recognition, which also outputs triple labels but evaluates separate bounding boxes of subject and object. The recall performance is relative to the number of predicate per subject-object pair to be evaluated, i.e., top predictions. In the experiments on VG dataset, we adopt top for evaluation.
Training Details. We use VGG pre-trained on ImageNet as the network backbone. The newly introduced layers are randomly initialized. We set the base learning rate as and fix the parameters from conv1_1 to conv3_3. The implementations are based on Caffe , and the networks are optimized via SGD. The conventional feature fusion operations are implemented by channel-wise concatenation in SCA-M cells here.
SCA-Module. The advantage of Zoom-Net lies in its unique capability of learning spatiality-aware contextual information through the SCA-M. To demonstrate the benefits of learning visual features with spatial-oriented and context-aided cues, we compare the recognition performance of Zoom-Net with a set of variants achieved by removing each individual cue step by step, i.e.. the SCA-M without stacked structure, the CA-M that disregard the spatial layouts, and the vanilla A-M that does not perform message passing (see Sec. 3.1). Their accuracy and recall scores are reported in Tab. 3.
In comparison to vanilla A-M, both the CA-M and SCA-M obtain a significant improvement suggesting the importance of contextual information to individual subject, predicate, and object classification and their relationship recognition. Note that contemporary CNNs have already shown a remarkable performance on subject and object classification, i.e. it is not hard to recognize object via individual appearance information, and thus the gap () of subject is smaller than that of predicate () between A-M and SCA-M on Top- accuracy. Not surprisingly, since the key inherent problem of relationship recognition is to learning the interactions between subject and object, the proposed SCA-M module exhibit a strong performance, thanks to its capability in capturing correlation between spatiality and semantic appearance cues among different object. Its effectiveness can also be observed from qualitative comparisons in Fig. 6(a).
Intra-Hierarchical Tree. We use the two auxiliary levels of hierarchical labels and to facilitate the prediction of the raw ground truth labels for the subject, predicate and object, respectively. Here we show that by involving hierarchical structures to semantically cluster ambiguous and noisy labels, the recognition performance w.r.t. the raw labels of the subject, predicate, object as well as their relationships are all boosted, as shown in Tab. 3. Discarding one of two levels in IH-tree clearly hamper the performance, i.e., Zoom-Net without IH-tree experiences a drop of around on different metrics. It reveals that intra-hierarchy structures do provide beneficial information to improve the recognition robustness. Besides, Fig. 6(b) shows the Top- triple relationship prediction results of Zoom-Net with and without IH-trees. The novel design of the hierarchical label structure help resolves data ambiguity for both on object and predicate. For example, thanks to the hierarchy level introduced in Sec. 4, the predicates related to “wear” (e.g., “wearing” and “wears”) can be ranked in top predictions. Another example shows the contribution of designed for semantic label clustering, e.g. “sitting in”, which is grouped in the same cluster of the ground truth “in”, also appears in top ranking results.
2 Comparison with State-of-the-Art Methods
We summarize the comparative results on VG in Tab. 4 with two recent state of the arts . For a fair comparison, we implement both methods with the VGG-16 as the network backbone. The proposed Zoom-Net significantly outperforms these methods, quantitatively and qualitatively. Qualitative results are shown in the first row of Fig. 6(c). DR-Net exploits binary dual masks as the spatial configuration in feature learning and therefore loses the critical interaction between visual context and spatial information. ViP focuses on learning label interaction by proposing a phrase-guided message passing structure. Additionally, the method tries to capture contextual information by passing messages across triple branches before ROI pooling and thus fail to explore in-depth spatiality-aware feature representations.
Transferable SCA-M Module and IH-Tree. We further demonstrate the effectiveness of the proposed SCA-M module in capturing spatiality, context and appearance visual cues, and IH-trees for resolving ambiguous annotations, by plugging them into architectures of existing works. Here, we take the network of ViP as the backbone for its end-to-end training scheme and state-of-the-art results (Tab. 4). We compare three configurations, i.e., ViP+SCA-M, ViP+IH-tree and ViP+SCA-M+IH-tree. For a fair comparison, the ViP is modified by replacing the targeted components with SCA-M or IH-tree but with other components fixed. As shown in Tab. 4, the performance of ViP is improved by a considerable margin on all evaluation metrics after applying our SCA-M (i.e. ViP+SCA-M). The results again suggest the superiority of the proposed spatiality-aware feature representations to that of ViP. Note that the overall performance by adding both stacked SCA module and IH-tree (i.e., ViP+SCA-M+IH-tree) surpasses that of ViP itself. The ViP designs a phrase-guided message passing structure to learn textual connections among subject-predicate-object at label-level. On the contrary, we concentrate more on capturing contextual connections among subject-predicate-object at feature-level. Therefore, it’s not surprising that a combination of these two aspects can provide a better result.
Comparisons on Visual Relationship Dataset (VRD)
Settings. We further quantitatively compare the performance of the proposed method with previous state of the arts on the Visual Relationship Dataset (VRD) . The following comparisons keep the same settings as the prior arts. VRD dataset is widely used for its clean and accurate annotations, although it is much smaller and simpler than VG dataset as shown in Tab. 1. Since VRD has a clean annotation, we fine-tune the construction of IH-tree by removing the and , which aim at reducing data ambiguity and noise in VG (details in Sec. 4). For a fair comparison, object proposals are generated by RPN here and we use triplet NMS to remove redundant triplet candidates following the setting in due to its excellent performance.
Evaluation metrics. We follow to report Recall@ and Recall@ when . The IoU between the predicted bounding boxes and the ground truth is required above here. In addition, some previous works used for evaluation and thus we report our results with as well to compare these previous methods under the same conditions.
Results. The results listed in Tab. 5 show that the proposed Zoom-Net outperforms the state-of-the-art methods by significant gains on almost all the evaluation metrics Note that Yu et al. take external Wikipedia data with around billion and million sentences to distill linguistic knowledge for modeling the tuple correlation from label-aspect. It’s not surprising to achieve a superior performance. In this experiment, we only compare with the results without knowledge distillation.. In comparison to previous state-of-the-art approaches, Zoom-Net improves the recall of predicate prediction by Rec@ and Rec@ when . Besides, the Rec@50 on relationship and phrase prediction tasks are increased by and , respectively. Note that the result of predicate () only achieves comparable performance with some prior arts since these methods use the groundtruth of subject and object and only predict predicate while our method predicts subject, predicate, object together.
Among all prior arts designed without external data, CAI has achieved the best performances on predicate prediction ( Rec@) by designing a context-aware interaction recognition framework to encode the labels into semantic space. To demonstrate the effectiveness and robustness of the proposed SCA-M in feature representation, we replace the visual feature representation in CAI with our SCA-M (i.e. CAI + SCA-M). The performance improvements are significant as shown in Tab. 5 due to the better visual feature learned, e.g., predicate Rec@50 is increased by compared to . In addition, with neither language priors, linguistic models nor external textual data, the proposed method can still achieve the state-of-the-art performance on most of the evaluation metrics, thanks to its superior feature representations.
Conclusion
We have presented an innovative framework Zoom-Net for visual relationship recognition, concentrating on feature learning with a novel Spatiality-Context-Appearance module (SCA-M). The unique design of SCA-M, which contains the proposed Contrastive ROI Pooling and Pyramid ROI Pooling Cells benefits the learning of spatiality-aware contextual feature representation. We further designed the Intra-Hierarchical tree (IH-tree) to model intra-class correlations for handling ambiguous and noisy labels. Zoom-Net achieves the state-of-the-art performance on both VG and VRD datasets. We demonstrated the superiority and transferability of each component of Zoom-Net. It is interesting to explore the notion of feature interactions in other applications such as image retrieval and image caption generation.
This work is supported in part by the National Natural Science Foundation of China (Grant No. 61371192), the Key Laboratory Foundation of the Chinese Academy of Sciences (CXJJ-17S044) and the Fundamental Research Funds for the Central Universities (WK2100330002, WK3480000005), in part by SenseTime Group Limited, the General Research Fund sponsored by the Research Grants Council of Hong Kong (Nos. CUHK14213616, CUHK14206114, CUHK14205615, CUHK14203015, CUHK14239816, CUHK419412, CUHK14207-814, CUHK14208417, CUHK14202217), the Hong Kong Innovation and Technology Support Program (No.ITS/121/15FX).
References
Appendix
In our work, we use the proposed Intra-Hierarchical trees (IH-tree) to handle the ambiguous and noisy labels in Visual Genome (VG) dataset . Fig. 7 provides the wordle imageshttp://www.wordle.net/ to highlight the frequencies of object and predicate categories appeared in VG. Bigger font sizes suggest higher frequencies. The most frequent object is man and the most common predicate is on.
As shown in Fig. 5 in the main article, there are three levels in our Intra-Hierarchical trees, and . The class numbers in each level in our experiments are shown in Tab. 6. The labels in are the source labels in the datasets. In VG, there are only classes of objects after clustering the labels in by semantic similarity, which is much fewer than the original classes. Actually, there are nearly the same number of classes about verb-based and preposition-based predicates in both in the VG and VRD datasets. Since the annotation in VRD dataset is clean enough, we remove the intermediate layers and that aiming at reducing the label ambiguity and annotation noise.
2 Ablation Study
The parameters in Sec. 8.1 in the main body are to balance the scales of losses from three branches, so as to ensure balanced influences from subject, object and relationship during training. Since these branches are evenly interacted with each other through the feed-forward pass, the back-propagated gradients from any loss can update the network parameters in other branches, thus a slight variance of weights for different losses will not have dominant effect on the training. Tab. 7 shows the performance drop of using different scales of loss weights compared to equal weights. It is reasonable that the model may be sensitive to a large scale difference between predicate () and subject/object (/), while the small scale changes will not influence the results much.
Computational time per image. The computational cost of each component of Zoom-Net is listed in Tab. 8. The experiments are conducted on a single TITAN X GPU. A single SCA-M module (e.g. after conv4_3) only costs an additional 0.02s which make the whole framework efficient. If involving multiple SCA-M modules (e.g. after conv3_3), it may cause fewer shared layers and more time costs.
3 More Experiment Results of Zoom-Net
In this section, we show additional qualitative results on the VG dataset. The experiment settings and details can be found in Sec. 5 in the main body.
Scene graph generation can serve as the basis for a number of tasks, e.g. visual question answering and image retrieval The proposed Zoom-Net can also perform well on scene graph generation. The task here is to generate a directed graph for an image that captures objects and their relationships. Fig. 8 illustrates two scene graphs generated by the proposed Zoom-Net. The reported excellent performances come from the proposed effective and efficient visual relationship recognition.
3.2 Zero-shot Relationship Recognition.
Owing to the long tail distribution of relationship labels in VG dataset, even though each single object or predicate category can be guaranteed to appear both in the training and testing sets, it is hard to assure the distribution of their combination (i.e., tuple relationship). This results in a zero-shot relationship recognition problem. A couple of examples are shown in the first row of Fig. 9(c). They (e.g., water-in-window and vase-on-head) are not in the training set. Compared to the reference methods, these unseen relationships can be well inferred by our model using similar relationships (e.g., person-in-window and hat-on-head) learned from the training set.
3.3 Additional Results on Visual Genome
The additional qualitative comparisons are in visualized Fig. 9. Firstly, we compare the results among the different module configurations of the proposed Zoom-Net. Then we show the Top-10 triple relationship prediction results of Zoom-Net with and without IH-tree in Fig.9(b). Finally, Fig.9(c) shows the excellent performance of Zoom-Net, compared with the state-of-the-art methods, DR-Net and ViP . The details are depicted in Sec. 5.1 and Sec. 5.2 in the main body.