UnionDet: Union-Level Detector Towards Real-Time Human-Object Interaction Detection

Bumsoo Kim, Taeho Choi, Jaewoo Kang, Hyunwoo J. Kim

Introduction

Recent advances in deep neural networks have achieved significant progress in detecting and recognizing individual objects from an image. However, to understand a scene, we need a deeper visual understanding that transcends the level of individual object detection. To understand what is happening in the image, not only do we have to accurately detect individual objects, but we also have to properly predict the interactions between the detected objects. Among the interactions, in this paper, we focus on human-object interaction (HOI) detection that involves the localization and classification of interactions between humans and surrounding objects. HOI detection has been formally defined in as the task to detect ⟨human,verb,object⟩\langle human,verb,object\rangle triplets within an image.

The main challenge of HOI detection boils down to a simple question: “How can we localize interactions?”. When asked to localize the area of “A person rides a horse.”, a human can naturally spot the tight area that covers both the person and the horse he/she is riding. This is the union region of the interacting objects that have been considered as a representation of visual relationships from previous works , and have been widely utilized in HOI detection . Ironically, no detector in the literature has been studied to directly capture the union region.

All the previous HOI detectors, therefore, incorporated a multi-stage and sequential pipeline that detects the individual objects first and ‘associate’ them to obtain the union region. This approach is far from intuitive and it makes HOI detectors inefficient. The sequential pipeline of object detection and interaction prediction makes end-to-end training impossible and creates a huge bottleneck in inference time for HOI detection. In standard object detection, one-stage detectors were able to speed up two-stage detectors by eliminating the second stage while yielding a competitive performance. Yet in HOI detection, previous multi-stage models mainly focused on performance (e.g., average precision) leaving the large gap between high-performance and real-time detection unexplored. In this work, our goal is to fill the gap between the performance and inference time of HOI detection with a fast, single-stage model.

To this end, we propose UnionDet: a one-stage meta-architecture powered by a novel union-level detector that captures the union region of human-object interaction. Instead of associating the object detection results by feeding each object pair into a separate neural network afterward, we directly detect interacting ⟨human,object⟩\langle human,object\rangle pairs with our novel union-level detection framework. This eliminates the need for heavy neural network inference after object detection and enables our model to detect interactions with minimal additional time on top of existing object detectors. Though the union-level detection sounds intuitive, detecting the union region is much more challenging than instance-level detection. In this paper, we study new challenges in union-level detection and address them by new techniques: (i) union anchor labeling, (ii) target object classification loss and (iii) union foreground focal loss. Based on these new methods, our proposed one-stage HOI detector achieves a 4×∼14×4\times\sim 14\times speed-up in additional inference time for interaction prediction while surpassing state-of-the-art performance on two HOI detection benchmark datasets: V-COCO (Verbs in COCO) and HICO-DET. The main contributions of our paper are threefold:

We study new technical challenges in union-level detection, including bias toward human regions, inaccuracy of standard IoU-based matching and union regions containing multiple interactions and more than two objects.

We propose a novel union-level detector that directly detects the interaction region. We study a new set of training techniques to address the new challenges of union-level detection.

We propose a meta-architecture UnionDet equipped with our union-level detector. It is a single-stage HOI detector achieving 4×∼14×4\times\sim 14\times speed-up in interaction prediction and the state-of-the-art performance in two public datasets.

Related Work

One-stage object detectors based on deep neural networks have formerly been proposed for faster detection . These detectors have achieved a significant speed-up but they often come with a considerable loss of accuracy. One known problem is the class imbalance problem. Since one-stage object detectors densely sample anchor boxes, foreground anchor boxes are relatively much rarer than background anchor boxes (or negative samples), unlike two-stage methods that classify only a few anchor boxes after RPN. One common technique to resolve the class imbalance is hard negative mining which samples a few hard anchor boxes for training . Later, RetinaNet introduced focal loss to address the issue in a fundamental way by modifying the loss function to reduce the effect of easy negatives. Including these efforts, various techniques have been proposed to enhance one-stage object detection frameworks . Recently, YOLACT has expanded the capacity of one-stage networks to perform instance segmentation.

2 Human-Object Interactions

Human-Object Interaction (HOI) detection has been initially proposed in . Later, human-object detectors have been improved using human body parts , human appearance , instance appearance and spatial relationship of human-object pairs . Especially, InteractNet extended an existing object detector by introducing an action-specific density map to localize target objects based on the appearance of a detected human. Note that interaction detection based on visual cues from individual boxes often suffers from the lack of contextual information. So iCAN proposed an instance-centric attention module that extracts contextual features complementary to the features from the localized objects/humans. GPNN proposes a Graph Parsing Neural Network for HOI recognition—a general framework that explicitly represents HOI structures with graphs and automatically infers the optimal graph structures. Deep Contextual Attention leverages contextual information by a contextual attention framework in HOI. Recent works in HOI have also explored external knowledge to improve the performance of HOI detection. Since the performance of HOI detection is dependent on how well we recognize the appearance of human actions, human pose information extracted from external models shows meaningful improvement in performance . Interactiveness Knowledge has also been implemented in previous works by adding an additional inference stage where the model learns the probability of interactiveness by combining multiple HOI training datasets. Linguistic priors and knowledge graphs are also utilized to improve HOI detection performance. These sources are either used directly as an additional feature or features to cluster the objects by their functions . However, all the previous methods are multi-stage detectors focusing on accuracy and they are not suitable for real-time applications.

Method

We now introduce our method to detect human-object-interaction. To be specific, the goal is to capture ⟨human,verb,object⟩\langle human,verb,object\rangle triplets from an image without any external knowledge. The standard HOI detection benchmarks (e.g., V-COCO and HICO-DET) require the localization and classification of interactions. In this paper, we propose a one-stage HOI detector powered by our union-level detector, which directly detects the union region of an interacting pair. Since standard benchmarks require the instance-level localization of humans and objects, we parallelly combine the union-level detector and an instance-level detector, which allows more accurate instance-level localization. We name this meta-architecture UnionDet shown in Figure 2. Our UnionDet is compatible with any one-stage object detectors such as SSD , RetinaNet , and STDN . For a fair comparison with baseline HOI detectors, in this paper, we implement our model based on RetinaNet with ResNet50-FPN since it’s performance is comparably similar to Faster-RCNN—the dominant backbone network in previous works on HOI in literature.

We discuss new challenges in union-level detection and how to address them by the components in our union-level detector, which is the union branch in UnionDet in Figure 2. We explain how to modify a standard instance-level detector in UnionDet for HOI detection and lastly, the details of training and inference are provided.

The union region of a pair of objects is an intuitive representation of visual relationships . Union-level detection looks similar to instance-level detection. But standard object detectors are not directly applicable due to the following technical challenges. First, a naive union-level detection often suffers from the large bias towards human regions since every union region of HOI has a human. The left figure in Figure 1 shows that union predictions (green bboxes) by a vanilla detector are densely distributed around a human. Second, the standard IoU is not an accurate metric for union bounding box matching. For instance, when one union region has two remote objects, a high IoU with the union region does not ensure that both human and target object is enclosed by the predicted region (see the left figure in Figure 1). Lastly, one union region (or anchor box) may contain multiple interactions and more than two objects. These are often observed especially when a human bounding box contains multiple interacting objects. In the following explanation of Union Branch that performs union-level detection, we show in detail how we address these issues.

2 Union-level Detector: Union Branch

Union Branch performs union-level detection which is the essence of our proposed meta-architecture, UnionDet. As in Figure 2, Union Branch consists of three sub-branches that share the backbone Feature Pyramid Network. Out of the three sub-branches, the Action Classification sub-branch and the Union Box Regression sub-branch are the main sub-branches that contribute to the inference stage. Action Classification sub-branch performs multi-class classification for the interactions that are related to the union region, and the Union Box Regression sub-branch performs action-agnostic bounding box regression to predict the final union region with multiple actions. Vanilla detection results for union regions can be obtained through these two sub-branches. However, union regions inherently accompany several technical challenges as mentioned above. To address these challenges, Union Branch is trained by new techniques: i) union anchor labeling ii) target object classification loss iii) foreground focal loss. This provides accurate union-level detections even in various distances, see Figure 4.

where tu,th,tot_{u},t_{h},t_{o} indicate thresholds for union IoU, human inclusion ratio, and object inclusion ratio. They are set to 0.50.5 in our experiments. If multiple union-level ground truths are matched, the union with the largest IoU is associated with the anchor box so that an anchor box has at most one ground truth.

After labeling each anchor according to Eq.1, we can build a basic loss function to train the Union Branch. Based on the positive anchor set A+⊆AA_{+}\subseteq A where {aj∣∑iUij=1}\{a_{j}|\sum_{i}{U_{ij}}=1\} and the negative anchor samples A−⊆AA_{-}\subseteq A where {aj∣∑iUij=0}\{a_{j}|\sum_{i}{U_{ij}}=0\}, the loss function Lu(θ˘)\mathcal{L}_{u}(\breve{\theta}) is written as

where G˘\breve{\mathcal{G}} denotes the ground truth union box set and θ˘\breve{\theta} denotes the Union Branch model parameters. Lijact(θ˘)=FL(a˘jact,g˘iact,θ˘)\mathcal{L}^{act}_{ij}(\breve{\theta})=FL(\breve{a}_{j}^{act},\breve{g}_{i}^{act},\breve{\theta}), Lijloc(θ˘)=smoothL1(a˘jloc,g˘iloc,θ˘)\mathcal{L}^{loc}_{ij}(\breve{\theta})=smooth_{L1}(\breve{a}_{j}^{loc},\breve{g}_{i}^{loc},\breve{\theta}), Ljbg=FL(a˘jact,0⃗,θ˘)\mathcal{L}_{j}^{bg}=FL(\breve{a}_{j}^{act},\vec{0},\breve{\theta}), where FLFL and smoothL1smooth_{L1} each denotes focal loss and Smooth L1 loss, respectively. After training the Union Branch with Eq.2, a vanilla prediction of union regions can be obtained. However, it suffers from 1) the prediction being biased toward the human region, and 2) the noisy learning caused when multiple union regions overlap over each other.

2.2 Target Object Classification Loss.

To address the first issue where the union prediction is biased toward the human region with a vanilla union-level detector, we design a pretext task, ‘target object classification’ from the detected union region. This encourages the union-level detector to focus more on target objects and helps the union-level detector to capture the region that encloses the target object. We add the target object classification loss to Eq.2 and the loss function Lu\mathcal{L}_{u} of Union Branch is given as

where Lijcls(θ˘)=BCE(a˘jcls,g˘icls,θ˘)\mathcal{L}^{cls}_{ij}(\breve{\theta})=BCE(\breve{a}_{j}^{cls},\breve{g}_{i}^{cls},\breve{\theta}) is the Binary Cross Entropy loss. Though we do not use the target classification score at inference in the final HOI score function, we observed that learning to classify the target objects during training improves the union region detection as well as overall performance (see, Table 3).

2.3 Union Foreground Focal Loss.

Union regions often overlap over each other when a single person interacts with multiple surrounding objects. The right subfigure in Fig.1 shows an extreme example of overlapping union regions where different interaction pairs have the exactly same union region. In such cases where large portion of union regions overlap with each other, applying vanilla focal loss Lijact(θ˘)\mathcal{L}_{ij}^{act}(\breve{\theta}) as in Eq.2 and Eq.3 might mistakenly give negative loss to the overlapped union actions (more detailed explanation of such cases will be dealt in our supplement). To address this issue, we deployed a variation of focal loss where we selectively calculate losses for only positive labels for foreground regions. This is implemented by simply multiplying g˘iact\breve{g}_{i}^{act} to Lijact\mathcal{L}^{act}_{ij}, thus our final loss function is written as:

3 Instance-level Detector: Instance Branch

HOI detection benchmarks require the localization of instances in interactions. For more accurate instance localization, we added Instance Branch to our architecture, see Fig. 2. The Instance Branch parallelly performs instance-level HOI detection: object classification, bbox regression, and action (or verb) classification.

The instance-level detector was built based on a standard anchor-based single-stage object detector that performs object classification and bounding box regression. For training, we adopt the focal loss to handle the class imbalance problem between the foreground and background anchors. The object detector is frozen for the V-COCO dataset and fine-tuned for the HICO-DET dataset. More discussion is available in the supplement.

3.2 Action Classification.

The instance-level detector was extended by another sub-branch for action classification. We treat the action of subjects TsT_{s} and objects ToT_{o} as different types of actions. So, the action classification sub-branch predicts (Ts+To)(T_{s}+T_{o}) action types at every anchor. This helps to recognize the direction of interactions and can be combined with the interaction prediction from the Union Branch. For action classification, we only calculate the loss at the positive anchor boxes where an object is located at. This leads to more efficient loss calculation and improvement accuracy.

4 Training UnionDet

The two branches of UnionDet shown in Figure 2 (i.g., the Union Branch and Instance Branch) are trained jointly. Our overall loss is the sum of the losses of both branches, Lu\mathcal{L}_{u} and Lg\mathcal{L}_{g}, where θ˘\breve{\theta} is the parameters for Union Branch and θ^\hat{\theta} is the parameters for Instance Branch (θ=θ˘∪θ^\theta=\breve{\theta}\cup\hat{\theta}). The final loss becomes L(θ)=Lu(θ˘)+Lg(θ^)\mathcal{L}(\theta)=\mathcal{L}_{u}(\breve{\theta})+\mathcal{L}_{g}(\hat{\theta}). For focal loss, we use α=0.25\alpha=0.25, γ=2.0\gamma=2.0 as in . Our model is trained with an Adam optimizer with a learning rate of 1e-5.

5 HOI Detection Inference

UnionDet at inference time parallelly performs the inference of Union Branch and Instance Branch and then seeks the highly-likely triplets using a summary score combining predictions from the subnetworks. Instance Branch performs object detection and action classification per anchor box. Non-maximum suppression with its object classification scores was performed. Union Branch directly detects the union region that covers the ⟨human,verb,object⟩\langle human,verb,object\rangle triplet. For Union Branch, non-maximum suppression was applied with union-level action classification scores. Instead of applying class-wise NMS as in ordinary object detection, we treated different action classes altogether to handle multi-label predictions of union regions.

As mentioned in section 3.2, IoU is not an accurate measure for union regions, especially in the case where the target object of the interaction is remote and small. To search for a solid union region that covers the given human box bhb_{h} and object box bob_{o}, we search for the union box bub_{u} with our proposed union-instance matching score defined as

where b1b2\frac{b_{1}}{b_{2}} is the ratio of the areas of two bounding boxes b1b_{1} and b2b_{2} and ⌈⋅⌋\lceil\cdot\rfloor stands for the tightest bounding box that covers the area. We use this union-instance matching score to calculate the HOI score instead of the standard IoU.

5.2 HOI Score.

The detections from Union Branch and Instance Branch are integrated. This further improves the accuracy of the final HOI detection. Our HOI score function combines union-level action score suas^{a}_{u} from Union Branch with the human category score shs_{h}, human action score shas^{a}_{h}, object class score sos_{o}, instance-level action score soas^{a}_{o} from Instance Branch. For each ⟨human,object⟩\langle human,object\rangle pair, we first identify the best union area with the highest union-instance matching score μu\mu_{u} in Eq. (5) and then calculate the HOI score Sh,oaS^{a}_{h,o} as

When the action classes do not involve target objects, or no union region is predicted, the score will be Sh,oa=sh⋅shaS^{a}_{h,o}=s_{h}\cdot s^{a}_{h} and Sh,oa=sh⋅sha+so⋅soaS^{a}_{h,o}=s_{h}\cdot s^{a}_{h}+s_{o}\cdot s^{a}_{o}, respectively.

The calculation of Eq.(6) has in principle O(n3)O(n^{3}) complexity when the number of detections is nn. However, our framework calculates the final triplet scores without any additional neural network inference after Union and Instance Branches. The calculation time of Eq. (6) is negligible (<1ms)(<1ms). The end-to-end inference time of our model is marginally increased (∼9ms\sim 9ms) compared to the vanilla object detector (RetinaNet with ResNet50-FPN) thanks to the parallel architecture.

Experiments

In this section, we demonstrate the effectiveness of UnionDet in HOI detection. We first describe the two public datasets that we use as our benchmark: V-COCO and HICO-DET. Next, we perform various qualitative and quantitative analysis to show that our union-level detector successfully addresses the proposed technical challenges and captures quality union regions, leading to a fast and accurate one-stage HOI detector.

To validate the performance of our model, we evaluate our model on two public benchmark datasets: the V-COCO (Verbs in COCO) dataset and HICO-DET dataset. V-COCO is a subset of COCO and has 5,400 trainval images and 4,946 test images. For V-COCO dataset, we report the AProle\text{AP}_{\text{role}} over T=29T=29 interactions. Including the four interaction types that do not involve target objects, V-COCO has Ts=26T_{s}=26 active actions and To=25T_{o}=25 passive actions. As previous works, we exclude the interaction point during inference time, because only 31 instances appear in the test set. HICO-DET is a subset of HICO dataset and has more than 150K annotated instances of human-object pairs in 47,051 images (37,536 training and 9,515 testing) and is annotated with 600 ⟨verb,object⟩\langle verb,object\rangle interaction types. There are 80 unique object types, identical to the COCO object categories, and T=117T=117 unique verbs. In the HICO-DET dataset, we separate the 117 action classes into as,ao{a_{s},a_{o}}, thus leading into a total action number of Ts+To=234T_{s}+T_{o}=234. For HICO-DET dataset, we follow the previous settings and report the mAP over three different category sets: (1) all 600 HOI categories in HICO (Full), (2) 138 HOI categories with less than 10 training instances (Rare), and (3) 462 HOI categories with 10 or more training instances (Non-Rare).

0.2 Union-level detection.

Our union-level detector (Union Branch in UnionDet) directly detects union regions of HOI, see Fig. 3. Interestingly, the union-level detections are useful to disambiguate the confusing pairs with the same action and target object types (e.g., horse, or motorcycle) in an image. For example, when multiple people ride the same target objects as in Fig. 3, instance-level appearances are not sufficient to associate the correct pairs. Union-level detections successfully group them using the context in the union-region.

0.3 Interactions in various distances.

We discussed in Sec. 3.1 that a vanilla object detector is not directly applicable to union-level detection due to the bias toward human regions. This bias gets severer especially when a human interacts with remote target objects. Fig. 4 shows that the bias is addressed by our pre-text task ‘Target Object Classification’ and UnionDet is able to detect target objects for various distances. We show four cases: included (bh⊃bob_{h}\supset b_{o}), adjacent (IoU(bh,bo)>0\text{IoU}(b_{h},b_{o})>0), distant (IoU(bh,bo)=0\text{IoU}(b_{h},b_{o})=0) and remote (IoU(bh,bo)=0\text{IoU}(b_{h},b_{o})=0 and large distance), where bhb_{h}, and bob_{o} are human and object bounding boxes. Especially the fourth column in Fig. 4 shows that UnionDet successfully captures the remote relation with small remote target objects (e.g., tennis ball and frisbee). Our ablation study in Table. 3 provides that the ‘Target Object Classification’ improves HOI detection. Qualitative results of a vanilla union-level detector without the Target Object Classification sub-branch are provided in the supplement.

0.4 HOI detection results.

In Figure 5, we highlight the detected humans and objects by object detection with the blue and yellow boxes and the union region predicted by UnionDet with red boxes. The detected human-object interactions are visualized and given a pair of objects, the ⟨human,verb,object⟩\langle human,verb,object\rangle triplet with the highest HOI score is listed below each image. Note that our model can detect various types of interactions including one-to-one, many-to-one (multiple persons interacting with a single object), one-to-many (one person interacting with multiple objects), and many-to-many (multiple persons interacting with multiple objects) relationships.

0.5 Performance Analysis.

We quantitatively evaluate our model on two datasets, followed by the ablation study of our proposed methods. We use the official evaluation code for computing the performance of both V-COCO and HICO-DET. In V-COCO, there are two versions of evaluation but most previous works have not explicitly stated which version was used for evaluation. We have specified the evaluation scenario if it has been referred in either the literature , authors’ code or the reproduced code. We report our performance in both scenarios for a fair comparison with heterogeneous baselines. In both scenarios, our model outperforms state-of-the-art methods . Further, our model shows competitive performance compared to the baselines that leverage heavy external features such as linguistic priors or human pose features . On HICO-DET, our model achieves state-of-the-art performance for both the official ‘Default’ setting and ‘Known Object’ setting. For a more comprehensive evaluation of HOI detectors, we also provide the performance of recent works leverage external knowledge , although the models with External Knowledge are beyond the scope of this paper. Note that our main focus is to build a fast single-stage HOI detector from visual features.

Our ablation study in Table 3 shows that each component (foreground focal loss, target object classification loss Lijcls(θ˘)\mathcal{L}_{ij}^{cls}(\breve{\theta}), union matching function μu\mu_{u}) in our approach improves the overall performance of HOI detection.

0.6 Interaction Prediction Time.

We measured inference time on a single Nvidia GTX1080Ti GPU. Our model achieved the fastest ‘end-to-end’ inference time (77.6 ms). However, the end-to-end inference time is not suitable for fair comparison since the end-to-end computation time of one approach may largely vary depending on the base networks or the backbone object detector. Therefore, we here compare the additional time for interaction prediction, excluding the time for object detection. The detailed analysis of end-to-end time will also be provided in the supplement. Our approach increases the minimal inference time on top of a standard object detector by eliminating the additional pair-wise neural network inference on detected object pairs, which is commonly required in previous works. Table 1 compares the inference time of the HOI interaction prediction excluding the time of the object detection. Note that compared to other multi-stage pipelines that have heavy network structures after the object detection phase, our model additionally requires significantly less time 9.06 ms (11.7%) compared to the base object detector. Our approach achieves 4X∼\sim14X speed-up compared to the baseline HOI detection models which require 40ms∼130ms40ms\sim 130ms per image after the object detection phase. Since most multi-stage pipelines have extra overhead for switching heavy models between different stages and saving/loading intermediate results. In a real-world application on a single GPU, the gain from our approach is much bigger.

Conclusions

In this paper, we present a novel one-stage human-object interaction detector. By performing action classification and union region detection in parallel with object detection, we achieved the fastest inference time while maintaining comparable performance with state-of-the-art methods. Also, our architecture is generally compatible with existing one-stage object detectors and end-to-end trainable. Our model enables a unified HOI detection that performs object detection and human-object interaction prediction at near real-time frame rates. Compared to heavy multi-stage HOI detectors, our model does not need to switch models across different stages and save/load intermediate results. In the real-world scenario, our model will more beneficial.

Acknowledgement

This work was supported by the National Research Council of Science & Technology (NST) grant by the Korea government (MSIT)(No.CAP-18-03-ETRI), National Research Foundation of Korea (NRF-2017M3C4A7065887), and Samsung Electronics, Co. Ltd.

References