Zoom-Net: Mining Deep Feature Interactions for Visual Relationship Recognition

Guojun Yin, Lu Sheng, Bin Liu, Nenghai Yu, Xiaogang Wang, Jing Shao, Chen Change Loy

Introduction

Visual relationship recognition aims at interpreting rich interactions between a pair of localized objects, i.e., performing tuple recognition in the form of ⟨\langlesubject-predicate-object⟩\rangle as shown in Fig. 1(a). The fundamental challenge of this task is to recognize various vaguely defined relationships given diverse spatial layouts of objects and complex inter-object interactions. To complement visual-based recognition, a promising approach is to adopt a linguistic model and learn relationships between object and predicate labels from language. This strategy has been shown effective by many existing methods . These language-based methods either apply statistical inference to the tuple label set, establish a linguistic graph as the prior, or mine linguistic knowledge from external billion-scale textual data (e.g., Wikipedia).

In this paper, we explore a novel perspective beyond the linguistic-based paradigm. In particular, contemporary approaches typically recognize the tuple ⟨\langlesubject-predicate-object⟩\rangle via separate convolutional neural network (CNN) branches. We believe that by enhancing message sharing and feature interactions among these branches, the participating objects and their visual relationship can be better recognized. To this end, we formulate a new spatiality-aware contextual feature learning model, named as Zoom-Net. Differing from previous studies that learn appearance and spatial features separatelyTheir spatiality-streams simply apply the union region , binary masks or centroid coordinates as the abstraction of the spatial features., Zoom-Net propagates spatiality-aware object features to interact with the predicate features and broadcasts predicate features to reinforce the features of subject and object.

The core of Zoom-Net is a Spatiality-Context-Appearance Module, abbreviated as SCA-M. It consists of two novel pooling cells that permit deep feature interactions between objects and predicates, as shown in Fig. 1(b). The first cell, Contrastive ROI Pooling Cell, facilitates predicate feature learning by inversely pooling object/subject features to a matching spatial context of predicate features via a unique deROI pooling. This allows all subject and object to fall on the same spatial ‘palette’ for spatiality-aware feature learning. The second cell is called Pyramid ROI Pooling Cell. It helps object/subject feature learning through broadcasting the predicate features to the corresponding object’s/subject’s spatial area. Zoom-Net stacks multiple SCA-Ms consecutively in an end-to-end network that allows multi-scale bidirectional message passing among subject, predicate and object. As shown in Fig. 1(c), the message sharing and feature interaction not only help recognize individual objects more accurately but also facilitate the learning of inter-object relation.

Another contribution of our work is an effective strategy of mitigating ambiguity and imbalanced data distribution in ⟨\langlesubject-predicate-object⟩\rangle annotations. Specifically, we conduct our main experiments on the challenging Visual Genome (VG) dataset , which consists of over 5,3195,319 object categories, 1,9571,957 predicates, and 421,697421,697 relationship types. The large-scale ambiguous categories and extremely imbalanced data distribution in VG dataset (Tab. 1,2) prevent previous methods from predicting reliable relationships despite they succeed in the Visual Relationship Detection (VRD) dataset with only 100100 object categories, 7070 predicates and 6,6726,672 relationships. To alleviate the ambiguity and imbalanced data distribution in VG, we reformulate the conventional one-hot classification as a nn-hot multi-class hierarchical recognition via a novel Intra-Hierarchical trees (IH-trees) for each label set in the tuple ⟨\langlesubject-predicate-object⟩\rangle.

Contributions. Our contributions are summarized as follows:

1) A general feature learning module that permits feature interactions - We introduce a novel SCA-M to mining intrinsic interactions between low-level spatial information and high-level semantical appearance features simultaneously. By stacking multiple SCA-Ms into a Zoom-Net, we achieve compelling results on VG dataset thanks to the multi-scale bidirectional message passing among subject, predicate and object.

2) Multi-class Intra-Hierarchical tree - To mitigate label ambiguity in large-scale datasets, we reformulate the visual relationship recognition problem to a multi-label recognition problem. The recognizability is enhanced by introducing an Intra-Hierarchical tree (IH-tree) for the object and predicate categories, respectively. We show that IH-tree can benefit other existing methods as well.

3) Large-scale relationship recognition - Extensive experiments demonstrate the respective effectiveness of the proposed SCA-M and IH-tree, as well as their combination on the challenging large-scale VG dataset.

It is noteworthy that the proposed method differs significantly from previous works as Zoom-Net neither models explicit nor implicit label-level interactions between ⟨\langlesubject-predicate-object⟩\rangle. We show that feature-level interactions alone, which is enabled by SCA-M, can achieve state-of-the-art performance. We further demonstrate that previous state-of-the-arts that are based on label-level interaction can benefit from the proposed SCA-M and IH-trees.

Related work

Contextual Learning. Contextual information has been employed in various tasks , e.g., object detection, segmentation, and retrieval. For example, the visual features captured from a bank of object detectors are combined with global features in . For both detection and segmentation, learning feature representations from a global view rather than the located object itself has been proven effective in . Contextual feature learning for visual relationship recognition is little explored in previous works.

Class Hierarchy. In previous studies , class hierarchy that encodes diverse label relations or structures is used to improve performances on classification and retrieval. For instance, Deng et al. improve large-scale visual recognition of object categories through forming a semantic hierarchy that consists of many levels of abstraction. While object categories can be clustered easily by their semantic similarity given the clean and explicit labels of objects, building a semantic hierarchy for visual relationship recognition can be more challenging due to noisy and ambiguous labels. Moreover, the semantic similarity between some phrases and prepositions such as walking on a versus walks near the is not directly measurable. In our paper, we employ the part-of-speech tagger toolkit to extract and normalize the keywords of these labels, e.g. walk, on and near.

Visual Relationship. Recognizing visual relationship has been shown beneficial to various tasks, including action recogntion , pose estimation , recognition and object detection , and scene graph generation . Recent works show remarkable progress in visual relationship recognition, most of which focus on measuring linguistic relations with textual priors or language models. The linguistic relations have been explored for object recognition , object detection , retrieval , and caption generation . Yu et al. employ billions of external textual data to distill useful knowledge for triplet ⟨\langlesubject-predicate-object⟩\rangle learning. These methods do not fully explore the potential of feature learning and feature-level message sharing for the problem of visual relationship recognition. Li et al. propose a message passing strategy to encourage feature sharing between features extracted from ⟨\langlesubject-predicate-object⟩\rangle. However, the network does not capture the relative location of different objects thus it cannot capture valid contextual information between subject, predicate and object.

Zoom-Net: Mining Deep Feature Interactions

We propose an end-to-end visual relationship recognition model that is capable of mining feature-level interactions. This is beyond just measuring the interactions among the triplet labels with additional linguistic priors, as what previous studies considered.

As shown in Fig. 2(a), given the ROI-pooled features of the subject, predicate and object, we consider a question: how to learn good features for both object (subject) and predicate? We investigate three plausible modules as follows.

Appearance Module. This module focuses on the intra-dependencies within each ROI, i.e., the features of the subject, predicate and object branches are learned independently without any message passing. We term this network structure as Appearance Module (A-M), as shown in Fig. 2(a). No contextual and spatial information can be derived from such a module.

Context-Appearance Module. The Context-Appearance Module (CA-M) directly fuses pairwise features among three branches, in which subject/object features absorb the contextual information from the predicate features, and predicate features also receive messages from both subject/object features, as shown in Fig. 2(b). Nonetheless, these features are concatenated regardless of their relative spatial layout in the original image. The incompatibility of scale and spatiality makes the fused features less optimal in capturing the required spatial and contextual information.

Spatiality-Context-Appearance Module. The spatial configuration, e.g., the relative positions and sizes of subject and object, is not sufficiently represented in CA-M. To address this issue, we propose a Spatiality-Context-Appearance module (SCA-M) as shown in Fig. 2(c). It consists of two novel spatiality-aware feature alignment cells (i.e., Contrast ROI Pooling and Pyramid ROI Pooling) for message passing between different branches. In comparison to CA-M, the proposed SCA-M reformulates the local and global information integration in a spatiality-aware manner, leading to superior capability in capturing spatial and contextual relationships between the features of ⟨\langlesubject-predicate-object⟩\rangle.

2 Spatiality-Context-Appearance Module (SCA-M)

We denote the respective regions of interest (ROIs) of the subject, predicate and object as Rs\mathcal{R}_{s}, Rp\mathcal{R}_{p}, and Ro\mathcal{R}_{o}, where Rp\mathcal{R}_{p} is the union bounding box that tightly covers both the subject and object. The ROI-pooled features for these three ROIs are ft,t∈{s,p,o}\mathbf{f}_{t},t\in\{s,p,o\}, respectively. In this section, we present the details of SCA-M. In particular, we discuss how Contrastive ROI Pooling and Pyramid ROI Pooling cells, the two elements in SCA-M, permit deep feature interactions between objects and predicates.

Contrastive ROI Pooling denotes a pair of ⟨\langleROI, deROI⟩\rangle operations that the objectSubject and object refer to the same concept, thus we only take object as the example for illustration. features fo\mathbf{f}_{o} are at first ROI pooled for extracting normalized local features, and then these features are deROI pooled back to the spatial palette of the predicate feature fp\mathbf{f}_{p}, so as to generate a spatiality-aware object feature f^o\hat{\mathbf{f}}_{o} with the same size as the predicate feature, as shown in Fig. 3(b) marked by the purple triangle. Note that the remaining region outside the relative object ROI in f^o\hat{\mathbf{f}}_{o} is set to . The spatiality-resumed local feature f^o\hat{\mathbf{f}}_{o} can thus influence the respective regions in the global feature map fp\mathbf{f}_{p}. In practice, the proposed deROI pooling can be considered as an inverse operation of the traditional ROI pooling (green triangle in Fig. 3), which is analogous to the top-down deconvolution versus the bottom-up convolution.

There are three Contrastive ROI pooling cells presented in the SCA-M module to integrate the feature pairs subject-predicate, subject-object and predicate-object, as shown in Fig. 3(b-d). Followed by several convolutional layers, the features from subject and object are spatially fused into the predicate feature for enhanced representation capability. The proposed ⟨\langleROI, deROI⟩\rangle operations differ from conventional feature fusion operations (channel-wise concatenation or summation). The latter would introduce scale incompatibility between local subject/object features and global predicate features, which could hamper feature learning in subsequent convolutional layers.

3 Zoom-Net: Stacked SCA-M

By stacking multiple SCA-Ms, the proposed Zoom-Net is capable of capturing multi-scale feature interactions with dynamic contextual and spatial information aggregation. It enables a reliable recognition of the visual relationship triplet ⟨\langles-p-o⟩\rangle, where the predicate pp indicates the relationships (e.g., spatiality, preposition, action and etc.) between a pair of localized subject ss and object oo.

As visualized in Fig. 4, we use a shared feature extractor with convolutional layers until conv3_3 to encode appearance features of different object categories. By indicating the regions of interests (ROIs) for subject, predicate and object, the associated features are ROI-pooled to the same spatial size and respectively fed into three branches. The features in three branches are at first independently fed into two convolutional layers (the conv4_1 and conv4_2 layers in VGG-16) for a further abstraction of their appearance features. Then these features are put into the first SCA-M to fuse spatiality-aware contextual information across different branches. After receiving the interaction-augmented subject, predicate and object features from the first SCA-M, MSCA1\mathcal{M}^{1}_{\text{SCA}}, we continue to convolve these features with another two appearance abstraction layers (mimicking the structures of conv5_1 and conv5_2 layers in VGG-16) and then forward them to the second SCA-M, MSCA2\mathcal{M}^{2}_{\text{SCA}}. After this module, the multi-scale interaction-augmented features in each branch are fed into three fully connected layers fc_s, fc_p and fc_o to classify subject, predicate and object, respectively.

Hierarchical Relational Classification

To thoroughly evaluate the proposed Zoom-Net, we adopt the Visual Genome (VG) datasetExtremely rare labels (fewer than 1010 samples) were pruned for a valid evaluation. for its large scale and diverse relationships. Our goal is to understand the a much broader scope of relationships with a total number of 421,697421,697 relationship types (as shown in Tab. 1), in comparison to the VRD dataset that focuses on only 6,6726,672 relationships. Recognizing relationships in VG is a non-trivial task due to several reasons:

(1) Variety - There are a total of 5,3195,319 object categories and 1,9571,957 predicates, tens times than those available in the VRD dataset.

(2) Ambiguity - Some object categories share a similar appearance, and multiple predicates refer to the same relationship.

(3) Imbalance - We observe long tail distributions both for objects and predicates.

To circumvent the aforementioned challenges, existing studies typically simplify the problem by manually removing a considerable portion of the data by frequency filtering or cleaning . Nevertheless, infrequent labels like “old man” and “white shirt” contain common attributes like “man” and “shirt” and are unreasonable to be pruned. Moreover, the flat label structure assumed by these methods is limited to describe the label space of the VG dataset with ambiguous and noisy labels.

To overcome the aforementioned issues, we propose a solution by establishing two Intra-Hierarchical trees (IH-tree) for measuring intra-class correlation within objectSubject and object refer to the same term in this paper, thus we only take the object as the example for illustration. and predicate, respectively. IH-tree builds a hierarchy of concepts that systematically groups rare, noisy and ambiguous labels together with those clearly defined labels. Unlike existing works that regularize relationships across the triplet ⟨\langless-pp-oo⟩\rangle by external linguistic priors, we only consider the intra-class correlation to independently regularize the occurrences of the object and predicate labels. During end-to-end training, the network employs the weighted Intra-Hierarchical losses for visual relationship recognition as L=αLs+βLp+γLo\mathcal{L}=\alpha\mathcal{L}_{s}+\beta\mathcal{L}_{p}+\gamma\mathcal{L}_{o}, where hyper-parameters α,β,γ\alpha,\beta,\gamma balance the losses with respect to subject Ls\mathcal{L}_{s}, predicate Lp\mathcal{L}_{p} and object Lo\mathcal{L}_{o}. α=β=γ=1\alpha=\beta=\gamma=1 in our experiments. We introduce IH-tree and the losses next.

We build an IH-tree, Ho\mathcal{H}_{o}, for object with a depth of three, where the base layer Ho(0)\mathcal{H}_{o}^{(0)} consists of the raw object categories.

(1) Ho(0)→Ho(1)\mathcal{H}_{o}^{(0)}\rightarrow\mathcal{H}_{o}^{(1)}: Ho(1)\mathcal{H}_{o}^{(1)} is extracted from Ho(0)\mathcal{H}_{o}^{(0)} by pruning noisy labels with the same concept but different descriptive attributes or in different singular and plural forms. We employ the part-of-speech tagger toolkit from NLTK and NLTK Lemmatizer to filter and normalize the noun keyword, e.g., “man” from “old man”, “bald man” and “men”.

(2) Ho(1)→Ho(2)\mathcal{H}_{o}^{(1)}\rightarrow\mathcal{H}_{o}^{(2)}: We observe that some labels have a close semantic correlation. As shown in the top panel of Fig. 5, labels with similar semantic concepts such as “shirt” and “jacket” are hyponyms of “clothing” and need to be distinguished from other semantic concepts like “animal” and “vehicle”. Therefore, we cluster labels in Ho(1)\mathcal{H}_{o}^{(1)} to the third level Ho(2)\mathcal{H}_{o}^{(2)} by semantical similarities computed by Leacock-Chodorow distance from NLTK. We find that a threshold of 0.650.65 is well-suited for splitting semantic concepts.

The output of the subject/object branch is a concatenation of three independent softmax activated vectors corresponded to three hierarchical levels in the IH-tree. The loss Ls\mathcal{L}_{s} (Lo\mathcal{L}_{o}) is thus a summation of three independent softmax losses with respect to these levels, encouraging the intra-level mutual label exclusion and inter-level label dependency.

The predicate IH-tree also has three hierarchy levels. Different from the object IR-tree that only handles nouns, the predicate categories include various part-of-speech types, e.g., verb (action) and preposition (spatial position). Even a single predicate label may contain multiple types, e.g., “are standing on” and “walking next to a”.

(1) Hp(0)→Hp(1)\mathcal{H}_{p}^{(0)}\rightarrow\mathcal{H}_{p}^{(1)}: Similar to Ho(1)\mathcal{H}_{o}^{(1)}, Hp(1)\mathcal{H}_{p}^{(1)} is constructed aiming at extracting and normalizing keywords from predicates. We retain the keywords and normalize tenses with respective to three main part-of-speech types, i.e., verb, preposition and adjective, and abandon other pointless and ambiguous words. As shown in the bottom panel of Fig. 5, “wears a”, “wearing a yellow” and “wearing a pink” are mapped to the same keyword “wear”.

(2) Hp(1)→Hp(2)\mathcal{H}_{p}^{(1)}\rightarrow\mathcal{H}_{p}^{(2)}: Different part-of-speech types own particular characteristics with various context representations, and hence a separate hierarchical structure for the verb (action) and preposition (spatial) is indispensable for better depiction. To this end, we construct Hp(2)\mathcal{H}_{p}^{(2)} for verb and preposition label independently, i.e., Hp(2−1)\mathcal{H}_{p}^{(2-1)} for action information and Hp(2−2)\mathcal{H}_{p}^{(2-2)} for spatial configuration. There are two cases in Hp(1)\mathcal{H}_{p}^{(1)}: (a) the label is in the form of phrase that consists of both verb and preposition (e.g. “stand on” and “walk next to”) and (b) the label is a single word (e.g., “on” and “wear”). For the first case, Hp(2−1)\mathcal{H}_{p}^{(2-1)} extracts the verb words from the two phrases while Hp(2−2)\mathcal{H}_{p}^{(2-2)} extracts the preposition words. It thus causes that a label might be simultaneously clustered into different partitions of Hp(2)\mathcal{H}_{p}^{(2)}. If the label is a single word , it would be normally clustered into the corresponding part-of-speech but remained the same in the opposite part-of-speech, as shown with the dotted line in the bottom panel of Fig. 5. The loss Lp\mathcal{L}_{p} is constructed similarly to that for the object.

Experiments on Visual Genome (VG) Dataset

Dataset. We evaluate our method on the Visual Genome (VG) dataset (version 1.2). Each image is annotated with a triplet ⟨\langlesubject-predicate-object⟩\rangle, where the subjects and objects are annotated with labels and bounding boxes while the predicates only have labels. The detailed statistics are stated in Sec. 4 and Tab. 1, 2. We randomly split the VG dataset into training and testing set with a ratio of 8:28:2. Note that both sets are guaranteed to have positive and negative samples from each object or predicate category. The details of data preprocessing and the source code will be released.

Evaluation Metrics. (1) Acc@NN. We adopt the Accuracy score as the major evaluation metric in our experiments. The metric is commonly used in traditional classification tasks. Specifically, we report the values of both Acc@11 and Acc@55 for subject, predicate, object and relationship, where the accuracy of relationship is calculated as the averaged accuracies of subject, predicate and object.

(2) Rec@NN. Following , we use Recall as another metric so as to handle incomplete annotations. Rec@NN computes the ratio of the correct relationship instance that is covered in the top NN predictions per image. We report Rec@5050 and Rec@100100 in our experiments. For a fair comparison, we follow to evaluate Rec@NN on three tasks, i.e., predicate recognition where both the labels and bounding boxes of the subject and object are given; phrase recognition that takes a triplet as a union bounding box and predicts the triple labels; relationship recognition, which also outputs triple labels but evaluates separate bounding boxes of subject and object. The recall performance is relative to the number of predicate per subject-object pair to be evaluated, i.e., top kk predictions. In the experiments on VG dataset, we adopt top k=100k=100 for evaluation.

Training Details. We use VGG1616 pre-trained on ImageNet as the network backbone. The newly introduced layers are randomly initialized. We set the base learning rate as 0.0010.001 and fix the parameters from conv1_1 to conv3_3. The implementations are based on Caffe , and the networks are optimized via SGD. The conventional feature fusion operations are implemented by channel-wise concatenation in SCA-M cells here.

SCA-Module. The advantage of Zoom-Net lies in its unique capability of learning spatiality-aware contextual information through the SCA-M. To demonstrate the benefits of learning visual features with spatial-oriented and context-aided cues, we compare the recognition performance of Zoom-Net with a set of variants achieved by removing each individual cue step by step, i.e.. the SCA-M without stacked structure, the CA-M that disregard the spatial layouts, and the vanilla A-M that does not perform message passing (see Sec. 3.1). Their accuracy and recall scores are reported in Tab. 3.

In comparison to vanilla A-M, both the CA-M and SCA-M obtain a significant improvement suggesting the importance of contextual information to individual subject, predicate, and object classification and their relationship recognition. Note that contemporary CNNs have already shown a remarkable performance on subject and object classification, i.e. it is not hard to recognize object via individual appearance information, and thus the gap (4.96%4.96\%) of subject is smaller than that of predicate (12.25%12.25\%) between A-M and SCA-M on Top-11 accuracy. Not surprisingly, since the key inherent problem of relationship recognition is to learning the interactions between subject and object, the proposed SCA-M module exhibit a strong performance, thanks to its capability in capturing correlation between spatiality and semantic appearance cues among different object. Its effectiveness can also be observed from qualitative comparisons in Fig. 6(a).

Intra-Hierarchical Tree. We use the two auxiliary levels of hierarchical labels H(1)\mathcal{H}^{(1)} and H(2)\mathcal{H}^{(2)} to facilitate the prediction of the raw ground truth labels H(0)\mathcal{H}^{(0)} for the subject, predicate and object, respectively. Here we show that by involving hierarchical structures to semantically cluster ambiguous and noisy labels, the recognition performance w.r.t. the raw labels of the subject, predicate, object as well as their relationships are all boosted, as shown in Tab. 3. Discarding one of two levels in IH-tree clearly hamper the performance, i.e., Zoom-Net without IH-tree experiences a drop of around 1%∼4%1\%\sim 4\% on different metrics. It reveals that intra-hierarchy structures do provide beneficial information to improve the recognition robustness. Besides, Fig. 6(b) shows the Top-55 triple relationship prediction results of Zoom-Net with and without IH-trees. The novel design of the hierarchical label structure help resolves data ambiguity for both on object and predicate. For example, thanks to the hierarchy level H(1)\mathcal{H}^{(1)} introduced in Sec. 4, the predicates related to “wear” (e.g., “wearing” and “wears”) can be ranked in top predictions. Another example shows the contribution of H(2)\mathcal{H}^{(2)} designed for semantic label clustering, e.g. “sitting in”, which is grouped in the same cluster of the ground truth “in”, also appears in top ranking results.

2 Comparison with State-of-the-Art Methods

We summarize the comparative results on VG in Tab. 4 with two recent state of the arts . For a fair comparison, we implement both methods with the VGG-16 as the network backbone. The proposed Zoom-Net significantly outperforms these methods, quantitatively and qualitatively. Qualitative results are shown in the first row of Fig. 6(c). DR-Net exploits binary dual masks as the spatial configuration in feature learning and therefore loses the critical interaction between visual context and spatial information. ViP focuses on learning label interaction by proposing a phrase-guided message passing structure. Additionally, the method tries to capture contextual information by passing messages across triple branches before ROI pooling and thus fail to explore in-depth spatiality-aware feature representations.

Transferable SCA-M Module and IH-Tree. We further demonstrate the effectiveness of the proposed SCA-M module in capturing spatiality, context and appearance visual cues, and IH-trees for resolving ambiguous annotations, by plugging them into architectures of existing works. Here, we take the network of ViP as the backbone for its end-to-end training scheme and state-of-the-art results (Tab. 4). We compare three configurations, i.e., ViP+SCA-M, ViP+IH-tree and ViP+SCA-M+IH-tree. For a fair comparison, the ViP is modified by replacing the targeted components with SCA-M or IH-tree but with other components fixed. As shown in Tab. 4, the performance of ViP is improved by a considerable margin on all evaluation metrics after applying our SCA-M (i.e. ViP+SCA-M). The results again suggest the superiority of the proposed spatiality-aware feature representations to that of ViP. Note that the overall performance by adding both stacked SCA module and IH-tree (i.e., ViP+SCA-M+IH-tree) surpasses that of ViP itself. The ViP designs a phrase-guided message passing structure to learn textual connections among ⟨\langlesubject-predicate-object⟩\rangle at label-level. On the contrary, we concentrate more on capturing contextual connections among ⟨\langlesubject-predicate-object⟩\rangle at feature-level. Therefore, it’s not surprising that a combination of these two aspects can provide a better result.

Comparisons on Visual Relationship Dataset (VRD)

Settings. We further quantitatively compare the performance of the proposed method with previous state of the arts on the Visual Relationship Dataset (VRD) . The following comparisons keep the same settings as the prior arts. VRD dataset is widely used for its clean and accurate annotations, although it is much smaller and simpler than VG dataset as shown in Tab. 1. Since VRD has a clean annotation, we fine-tune the construction of IH-tree by removing the Ho(1)\mathcal{H}_{o}^{(1)} and Hp(1)\mathcal{H}_{p}^{(1)}, which aim at reducing data ambiguity and noise in VG (details in Sec. 4). For a fair comparison, object proposals are generated by RPN here and we use triplet NMS to remove redundant triplet candidates following the setting in due to its excellent performance.

Evaluation metrics. We follow to report Recall@5050 and Recall@100100 when k=70k=70. The IoU between the predicted bounding boxes and the ground truth is required above 0.50.5 here. In addition, some previous works used k=1k=1 for evaluation and thus we report our results with k=1k=1 as well to compare these previous methods under the same conditions.

Results. The results listed in Tab. 5 show that the proposed Zoom-Net outperforms the state-of-the-art methods by significant gains on almost all the evaluation metrics Note that Yu et al. take external Wikipedia data with around 44 billion and 450450 million sentences to distill linguistic knowledge for modeling the tuple correlation from label-aspect. It’s not surprising to achieve a superior performance. In this experiment, we only compare with the results without knowledge distillation.. In comparison to previous state-of-the-art approaches, Zoom-Net improves the recall of predicate prediction by 3.47%3.47\% Rec@5050 and 3.62%3.62\% Rec@100100 when k=70k=70. Besides, the Rec@50 on relationship and phrase prediction tasks are increased by 1.25%1.25\% and 6.46%6.46\%, respectively. Note that the result of predicate (k=1k=1) only achieves comparable performance with some prior arts since these methods use the groundtruth of subject and object and only predict predicate while our method predicts subject, predicate, object together.

Among all prior arts designed without external data, CAI has achieved the best performances on predicate prediction (53.59%53.59\% Rec@5050) by designing a context-aware interaction recognition framework to encode the labels into semantic space. To demonstrate the effectiveness and robustness of the proposed SCA-M in feature representation, we replace the visual feature representation in CAI with our SCA-M (i.e. CAI + SCA-M). The performance improvements are significant as shown in Tab. 5 due to the better visual feature learned, e.g., predicate Rec@50 is increased by 2.39%2.39\% compared to . In addition, with neither language priors, linguistic models nor external textual data, the proposed method can still achieve the state-of-the-art performance on most of the evaluation metrics, thanks to its superior feature representations.

Conclusion

We have presented an innovative framework Zoom-Net for visual relationship recognition, concentrating on feature learning with a novel Spatiality-Context-Appearance module (SCA-M). The unique design of SCA-M, which contains the proposed Contrastive ROI Pooling and Pyramid ROI Pooling Cells benefits the learning of spatiality-aware contextual feature representation. We further designed the Intra-Hierarchical tree (IH-tree) to model intra-class correlations for handling ambiguous and noisy labels. Zoom-Net achieves the state-of-the-art performance on both VG and VRD datasets. We demonstrated the superiority and transferability of each component of Zoom-Net. It is interesting to explore the notion of feature interactions in other applications such as image retrieval and image caption generation.

This work is supported in part by the National Natural Science Foundation of China (Grant No. 61371192), the Key Laboratory Foundation of the Chinese Academy of Sciences (CXJJ-17S044) and the Fundamental Research Funds for the Central Universities (WK2100330002, WK3480000005), in part by SenseTime Group Limited, the General Research Fund sponsored by the Research Grants Council of Hong Kong (Nos. CUHK14213616, CUHK14206114, CUHK14205615, CUHK14203015, CUHK14239816, CUHK419412, CUHK14207-814, CUHK14208417, CUHK14202217), the Hong Kong Innovation and Technology Support Program (No.ITS/121/15FX).

References

Appendix

In our work, we use the proposed Intra-Hierarchical trees (IH-tree) to handle the ambiguous and noisy labels in Visual Genome (VG) dataset . Fig. 7 provides the wordle imageshttp://www.wordle.net/ to highlight the frequencies of object and predicate categories appeared in VG. Bigger font sizes suggest higher frequencies. The most frequent object is man and the most common predicate is on.

As shown in Fig. 5 in the main article, there are three levels in our Intra-Hierarchical trees, Ho\mathcal{H}_{o} and Hp\mathcal{H}_{p}. The class numbers in each level in our experiments are shown in Tab. 6. The labels in H0\mathcal{H}^{0} are the source labels in the datasets. In VG, there are only 578578 classes of objects after clustering the labels in H2\mathcal{H}^{2} by semantic similarity, which is much fewer than the original 53195319 classes. Actually, there are nearly the same number of classes about verb-based and preposition-based predicates in H2\mathcal{H}^{2} both in the VG and VRD datasets. Since the annotation in VRD dataset is clean enough, we remove the intermediate layers Ho1\mathcal{H}^{1}_{o} and Hp1\mathcal{H}^{1}_{p} that aiming at reducing the label ambiguity and annotation noise.

2 Ablation Study

The parameters α,β,γ\alpha,\beta,\gamma in Sec. 8.1 in the main body are to balance the scales of losses from three branches, so as to ensure balanced influences from subject, object and relationship during training. Since these branches are evenly interacted with each other through the feed-forward pass, the back-propagated gradients from any loss can update the network parameters in other branches, thus a slight variance of weights for different losses will not have dominant effect on the training. Tab. 7 shows the performance drop of using different scales of loss weights compared to equal weights. It is reasonable that the model may be sensitive to a large scale difference between predicate (β\beta) and subject/object (α\alpha/γ\gamma), while the small scale changes will not influence the results much.

Computational time per image. The computational cost of each component of Zoom-Net is listed in Tab. 8. The experiments are conducted on a single TITAN X GPU. A single SCA-M module (e.g. after conv4_3) only costs an additional 0.02s which make the whole framework efficient. If involving multiple SCA-M modules (e.g. after conv3_3), it may cause fewer shared layers and more time costs.

3 More Experiment Results of Zoom-Net

In this section, we show additional qualitative results on the VG dataset. The experiment settings and details can be found in Sec. 5 in the main body.

Scene graph generation can serve as the basis for a number of tasks, e.g. visual question answering and image retrieval The proposed Zoom-Net can also perform well on scene graph generation. The task here is to generate a directed graph for an image that captures objects and their relationships. Fig. 8 illustrates two scene graphs generated by the proposed Zoom-Net. The reported excellent performances come from the proposed effective and efficient visual relationship recognition.

3.2 Zero-shot Relationship Recognition.

Owing to the long tail distribution of relationship labels in VG dataset, even though each single object or predicate category can be guaranteed to appear both in the training and testing sets, it is hard to assure the distribution of their combination (i.e., tuple relationship). This results in a zero-shot relationship recognition problem. A couple of examples are shown in the first row of Fig. 9(c). They (e.g., ⟨\langlewater-in-window⟩\rangle and ⟨\langlevase-on-head⟩\rangle) are not in the training set. Compared to the reference methods, these unseen relationships can be well inferred by our model using similar relationships (e.g., ⟨\langleperson-in-window⟩\rangle and ⟨\langlehat-on-head⟩\rangle) learned from the training set.

3.3 Additional Results on Visual Genome

The additional qualitative comparisons are in visualized Fig. 9. Firstly, we compare the results among the different module configurations of the proposed Zoom-Net. Then we show the Top-10 triple relationship prediction results of Zoom-Net with and without IH-tree in Fig.9(b). Finally, Fig.9(c) shows the excellent performance of Zoom-Net, compared with the state-of-the-art methods, DR-Net and ViP . The details are depicted in Sec. 5.1 and Sec. 5.2 in the main body.