Detecting and Recognizing Human-Object Interactions

Georgia Gkioxari, Ross Girshick, Piotr Dollár, Kaiming He

Introduction

Visual recognition of individual instances, e.g., detecting objects and estimating human actions/poses , has witnessed significant improvements thanks to deep learning visual representations . However, recognizing individual objects is just a first step for machines to comprehend the visual world. To understand what is happening in images, it is necessary to also recognize relationships between individual instances. In this work, we focus on human-object interactions.

The task of recognizing human-object interactions can be represented as detecting ⟨\langlehuman, verb, object⟩\rangle triplets and is of particular interest in applications and in research. From a practical perspective, photos containing people contribute a considerable portion of daily uploads to internet and social networking sites, and thus human-centric understanding has significant demand in practice. From a research perspective, the person category involves a rich set of actions/verbs, most of which are rarely taken by other subjects (e.g., to talk, throw, work). The fine granularity of human actions and their interactions with a wide array of object types presents a new challenge compared to recognition of entry-level object categories.

In this paper, we present a human-centric model for recognizing human-object interaction. Our central observation is that a person’s appearance, which reveals their action and pose, is highly informative for inferring where the target object of the interaction may be located (Figure 1(b)). The search space for the target object can thus be narrowed by conditioning on this estimation. Although there are often many objects detected (Figure 1(a)), the inferred target location can help the model to quickly pick the correct object associated with a specific action (Figure 1(c)).

We implement this idea as a human-centric recognition branch in the Faster R-CNN framework . Specifically, on a region of interest (RoI) associated with a person, this branch performs action classification and density estimation for the action’s target object location. The density estimator predicts a 4-d Gaussian distribution, for each action type, that models the likely relative position of the target object to the person. The prediction is based purely on the human appearance. This human-centric recognition branch, along with a standard object detection branch and a simple pairwise interaction branch (described later), form a multi-task learning system that can be jointly optimized.

We evaluate our method, InteractNet, on the challenging V-COCO (Verbs in COCO) dataset for detecting human-object interactions. Our human-centric model improves accuracy by 26% (relative) from 31.8 to 40.0 AP (evaluated by Average Precision on a triplet, called ‘role AP’ ), with the gain mainly due to inferring the target object’s relative position from the human appearance. In addition, we prove the effectiveness of InteractNet by reporting a 27% relative improvement on the newly released HICO-DET dataset . Finally, our method can run at about 135ms / image for this complex task, showing good potential for practical usage.

Related Work

Bounding-box based object detectors have improved steadily in the past few years. R-CNN, a particularly successful family of methods , is a two-stage approach in which the first stage proposes candidate RoIs and the second stage performs object classification. Region-wise features can be rapidly extracted from shared feature maps by an RoI pooling operation. Feature sharing speeds up instance-level detection and enables recognizing higher-order interactions, which would be computationally infeasible otherwise. Our method is based on the Fast/Faster R-CNN frameworks .

Human Action & Pose Recognition.

The action and pose of humans is indicative of their interactions with objects or other people in the scene. There has been great progress in understanding human actions and poses from images. These methods focus on the human instances and do not predict interactions with other objects. We rely on action and pose appearance cues in order to predict the interactions with objects in the scene.

Visual Relationships.

Research on visual relationship modeling has attracted increasing attention. Recently, Lu et al. proposed to recognize visual relationships derived from an open-world vocabulary. The set of relationships include verbs (e.g., wear), spatial (e.g., next to), actions (e.g., ride) or a preposition phrase (e.g., drive on). Our focus is related, but different. First, we aim to understand human-centric interactions, which take place in particularly diverse and interesting ways. These relationships involve direct interaction with objects (e.g., person cutting cake), unlike spatial or prepositional phrases (e.g., dog next to dog). Second, we aim to build detectors that recognize interactions in images with high precision, which is a requirement for practical applications. In contrast, in an open-world recognition setting, evaluating precision is not feasible, resulting in recall-based evaluation, as in .

Human-Object Interactions.

Human-object interactions are related to visual relationships, but present different challenges. Human actions are more fine-grained (e.g., walking, running, surfing, snowboarding) than the actions of general subjects, and an individual person can simultaneously take multiple actions (e.g., drinking tea and reading a newspaper while sitting in a chair). These issues require a deeper understanding of human actions and the objects around them and in much richer ways than just the presence of the objects in the vicinity of a person in an image. Accurate recognition of human-object interaction can benefit numerous tasks in computer vision, such as action-specific image retrieval , caption generation , and question answering .

Method

We now describe our method for detecting human-object interactions. Our goal is to detect and recognize triplets of the form ⟨\langlehuman, verb, object⟩\rangle. To detect an interaction triplet, we have to accurately localize the box containing a human and the box for the associated object of interaction (denoted by bhb_{h} and bob_{o}, respectively), as well as identify the action aa being performed (selected from among AA actions).

Our proposed solution decomposes this complex and multifaceted problem into a simple and manageable form. We extend the Fast R-CNN object detection framework with an additional human-centric branch that classifies actions and estimates a probability density over the target object location for each action. The human-centric branch reuses features extracted by Fast R-CNN for object detection so its marginal computation is lightweight.

Specifically, given a set of candidate boxes, Fast R-CNN outputs a set of object boxes and a class label for each box. Our model extends this by assigning a triplet score Sh,oaS_{h,o}^{a} to pairs of candidate human/object boxes bhb_{h}, bob_{o} and an action aa. To do so, we decompose the triplet score into four terms:

While the model has multiple components, the basic idea is straightforward. shs_{h} and sos_{o} are the class scores from Fast R-CNN of bhb_{h} and bob_{o} containing a human and object. Our human-centric branch outputs two extra terms. First, shas_{h}^{a} is the score assigned to action aa for the person at bhb_{h}. Second, μha\mu_{h}^{a} is the predicted location of the target of interaction for a given human/action pair, computed based on the appearance of the human. This, in turn, is used to compute gh,oag_{h,o}^{a}, the likelihood that an object with box bob_{o} is the actual target of interaction. We give details shortly and show that this target localization term is key for obtaining good results.

We discuss each component next, followed by an extension that replaces the action classification output shas_{h}^{a} with a dedicated interaction branch that outputs a score sh,oas_{h,o}^{a} for an action aa based on both the human and object appearances. Finally we give details for training and inference. Figure 3 illustrates each component in our full framework.

The object detection branch of our network, shown in Figure 3(a), is identical to that of Faster R-CNN . First, a Region Proposal Network (RPN) is used to generate object proposals . Then, for each proposal box bb, we extract features with RoiAlign , and perform object classification and bounding-box regression to obtain a new set of boxes, each of which has an associated score sos_{o} (or shs_{h} if the box is assigned to the person category). These new boxes are only used during inference; during training all branches are trained with RPN proposal boxes.

Action Classification.

The first role of the human-centric branch is to assign an action classification score shas_{h}^{a} to each human box bhb_{h} and action aa. Just like in the object classification branch, we extract features from bhb_{h} with RoiAlign and predict a score for each action aa. Since a human can simultaneously perform multiple actions (e.g., sit and drink), our output layer consists of binary sigmoid classifiers for multi-label action classification (i.e. the predicted action classes do not compete). The training objective is to minimize the binary cross entropy losses between the ground-truth action labels and the scores shas_{h}^{a} predicted by the model.

Target Localization.

The second role of the human-centric branch is to predict the target object location based on a person’s appearance (again represented as features pooled from bhb_{h}). However, predicting the precise target object location based only on features from bhb_{h} is challenging. Instead, our approach is to predict a density over possible locations, and use this output together with the location of actual detected objects to precisely localize the target.

We model the density over the target object’s location as a Gaussian function whose mean is predicted based on the human appearance and action being performed. Formally, the human-centric branch predicts μha\mu_{h}^{a}, the target object’s 4-d mean location given the human box bhb_{h} and action aa. We then write our target localization term as:

We can use gg to test the compatibility of an object box bob_{o} and the predicted target location μha\mu_{h}^{a}. In the above, bo∣hb_{o|h} is the encoding of bob_{o} in coordinates relative to bhb_{h}, that is:

This is a similar encoding as used in Fast R-CNN for bounding box regression. However, in our case bhb_{h} and bob_{o} are two different objects and moreover bob_{o} is not necessarily near or of the same size as bhb_{h}. The training objective is to minimize the smooth L1L_{1} loss between μha\mu_{h}^{a} and bo∣hb_{o|h}, where bob_{o} is the location of the ground truth object for the interaction. We treat σ\sigma as a hyperparameter that we empirically set to σ=0.3\sigma=0.3 using the validation set.

Figure 4 visualizes the predicted distribution over the target object’s location for example human/action pairs. As we can see, a carrying appearance suggests an object in the person’s hand, a throwing appearance suggests an object in front of the person, and a sitting appearance implies an object below the person. We note that the yellow dashed boxes depicting μha\mu_{h}^{a} shown in Figure 4 are inferred from bhb_{h} and aa and did not have direct access to the objects.

Intuitively, our formulation is predicated on the hypothesis that the features computed from bhb_{h} contain a strong signal pointing to the target of an action, even if that target object is outside of bhb_{h}. We argue that such ‘outside-the-box’ regression is possible because the person’s appearance provides a strong clue for the target location. Moreover, as this prediction is action-specific and instance-specific, our formulation is effective even though we model the target location using a uni-modal distribution. In Section 5 we discuss a variant of our approach which allows us to handle conditionally multi-modal distributions and predict multiple targets for a single action.

Interaction Recognition.

Our human-centric model scores actions based on the human appearance. While effective, this does not take into account the appearance of the target object. To improve the discriminative power of our model, and to demonstrate the flexibility of our framework, we can replace shas_{h}^{a} in (1) with an interaction branch that scores an action based on the the appearance of both the human and target object. We use sh,oas_{h,o}^{a} to denote this alternative term.

The computation of sh,oas_{h,o}^{a} reuses the computation from shas_{h}^{a} and additionally in parallel performs a similar computation based on features extracted from bob_{o}. The outputs from the two action classification heads, which are AA-dimensional vectors of logits, are summed and passed through a sigmoid activation to yield AA scores. This process is illustrated in Figure 3(c). As before, the training objective is to minimize the binary cross entropy losses between the ground-truth action labels and the predicted action scores sh,oas_{h,o}^{a}.

2 Multi-task Training

We approach learning human-object interaction as a multi-task learning problem: all three branches shown in Figure 3 are trained jointly. Our overall loss is the sum of all losses in our model including: (1) the classification and regression loss for the object detection branch, (2) the action classification and target localization loss for the human-centric branch, and (3) the action classification loss of the interaction branch. This is in contrast to our cascaded inference described in §3.3, where the output of the object detection branch is used as input for the human-centric branch.

We adopt image-centric training . All losses are computed over both RPN proposal and ground truth boxes as in Faster R-CNN . As in , we sample at most 64 boxes from each image for the object detection branch, with a ratio of 1:3 of positive to negative boxes. The human-centric branch is computed over at most 16 boxes bhb_{h} that are associated with the human category (i.e., their IoU overlap with a ground-truth person box is ≥0.5\geq 0.5). The loss for the interaction branch is only computed on positive example triplets (i.e., ⟨bh,a,bo⟩\langle b_{h},a,b_{o}\rangle must be associated with a ground truth interaction triplet). All loss terms have a weight of one, except the action classification term in the human-centric branch has a weight of two, which we found performs better.

3 Cascaded Inference

At inference, our goal is to find high-scoring triplets according to Sh,oaS_{h,o}^{a} in (1). While in principle this has O(n2)O(n^{2}) complexity as it requires scoring every pair of candidate boxes, we present a simple cascaded inference algorithm whose dominant computation has O(n)O(n) complexity.

Object Detection Branch: We first detect all objects (including the person class) in the image. We apply non-maximum suppression (NMS) with an IoU threshold of 0.3 on boxes with scores higher than 0.05 (set conservatively to retain most objects). This step yields a new smaller set of nn boxes bb with scores shs_{h} and sos_{o}. Unlike in training, these new boxes are used as input to the remaining two branches.

Human-Centric Branch: Next, we apply the human-centric branch to all detected objects that were classified as human. For each action aa and detected human box bhb_{h}, we compute shas_{h}^{a}, the score assigned to aa, as well as μha\mu_{h}^{a}, the predicted mean offset of the target object location relative to bhb_{h}. This step has a complexity of O(n)O(n).

Interaction Branch: If using the optional interaction branch, we must compute sh,oas_{h,o}^{a} for each action aa and pair of boxes bhb_{h} and bob_{o}. To do so we first compute the logits for the two action classification heads independently for each box bhb_{h} and bob_{o}, which is O(n)O(n). Then, to get scores sh,oas_{h,o}^{a}, these logits are summed and passed through a sigmoid for each pair. Although this last step is O(n2)O(n^{2}), in practice its computational time is negligible.

Once all individual terms have been computed, the computation of (1) is fast. However, rather than scoring every potential triplet, for each human/action pair we find the object box that maximizes Sh,oaS_{h,o}^{a}. That is we compute:

Recall that gh,oag_{h,o}^{a} is computed according to (2) and measures the compatibility between bob_{o} and the expected target location μha\mu_{h}^{a}. Intuitively, (4) encourages selecting a high-confidence object near the predicted target location of a high-scoring action. With bob_{o} selected for each bhb_{h} and action aa, we have a triplet of ⟨\langlehuman, verb, object⟩\rangle = ⟨bh\langle b_{h}, aa, bo⟩b_{o}\rangle. These triplets, along with the scores Sh,oaS_{h,o}^{a}, are the final outputs of our model. For actions that that do not interact with any object (e.g., smile, run), we rely on shas_{h}^{a} and the interaction output sh,oas_{h,o}^{a} is not used, even if present. The score of such a predicted ⟨\langlehuman, verb⟩\rangle pair is simply sh⋅shas_{h}\cdot s_{h}^{a}.

The above cascaded inference has a dominant complexity of O(n)O(n), which involves extracting features for each of the nn boxes and forwarding through a small network. The pairwise O(n2)O(n^{2}) operations require negligible computation. In addition, for the entire system, a portion of computation is spent on computing the full-image shared convolutional feature maps. Altogether, our system takes ∼\scriptstyle\sim135ms on a typical image running on a single Nvidia M40 GPU.

Datasets and Metrics

There exist a number of datasets for human-object interactions . The most relevant for this work are V-COCO (Verbs in COCO) and HICO-DET . V-COCO serves as the primary testbed on which we demonstrate the effectiveness of InteractNet and analyze its various components. The newly released HICO-DET contains ∼\scriptstyle\sim48k images and 600 types of interactions and serves to further demonstrate the efficacy of our approach. The older TUHOI and HICO datasets only have image-level labels and thus do not allow for grounding interactions in a detection setting, while COCO-a is promising but only a small beta-version is currently available.

V-COCO is a subset of COCO and has ∼\scriptstyle\sim5k images in the trainval set and ∼\scriptstyle\sim5k images in the test set.V-COCO’s trainval set is a subset of COCO’s train set, and its test set is a subset of COCO’s val set. See for more details. In this work, COCO’s val images are not used during training in any way. The trainval set includes ∼\scriptstyle\sim8k person instances and on average 2.9 actions/person. V-COCO is annotated with 26 common action classes (listed in Table 2). Of note, there are three actions (cut, hit, eat) that are annotated with two types of targets: instrument and direct object. For example, cut + knife involves the instrument (meaning ‘cut with a knife’), and cut + cake involves the direct object (meaning ‘cut a cake’). In , accuracy is evaluated separately for the two types of targets. To address this, for the target estimation, we train and infer two types of targets for these three actions (i.e., they are treated like six actions for target estimation).

Following , we evaluate two Average Precision (AP) metrics. We note that this is a detection task, and both AP metrics measure both recall and precision. This is in contrast to metrics of Recall@NN that ignore precision.

The AP of central interest in the human-object interaction task is the AP of the triplet ⟨\langlehuman, verb, object⟩\rangle, called ‘role AP’ (AProle{}_{\text{role}}) in . Formally, a triplet is considered as a true positive if: (i) the predicted human box bhb_{h} has IoU of 0.5 or higher with the ground-truth human box, (ii) the predicted object box bob_{o} has IoU of 0.5 or higher with the ground-truth target object, and (iii) the predicted and ground-truth actions match. With this definition of a true positive, the computation of AP is analogous to standard object detection (e.g., PASCAL ). Note that this metric does not consider the correctness of the target object category (but only the target object box location). Nevertheless, our method can predict the object categories, as shown in the visualized results (Figure 2 and Figure 5).

We also evaluate the AP of the pair ⟨\langlehuman, verb⟩\rangle, called ‘agent AP’ (APagent{}_{\text{agent}}) in , computed using the above criteria of (i) and (iii). APagent{}_{\text{agent}} is applicable when the action has no object. We note that APagent{}_{\text{agent}} does note require localizing the target, and is thus of secondary interest.

Experiments

Our implementation is based on Faster R-CNN with a Feature Pyramid Network (FPN) backbone built on ResNet-50 ; we also evaluate a non-FPN version in ablation experiments. We train the Region Proposal Network (RPN) of Faster R-CNN following . For convenient ablation, RPN is frozen and does not share features with our network (we note that feature sharing is possible ). We extract 7×\times7 features from regions by RoiAlign , and each of the three model branches (see Figure 3) consist of two 1024-d fully-connected layers (with ReLU ) followed by specific output layers for each output type (box, class, action, target).

Given a model pre-trained on ImageNet , we first train the object detection branch on the COCO train set (excluding the V-COCO val images). This model, which is in essence Faster R-CNN, has 33.8 object detection AP on the COCO val set. Our full model is initialized by this object detection network. We prototype our human-object interaction models on the V-COCO train split and perform hyperparameter selection on the V-COCO val split. After fixing these parameters, we train on V-COCO trainval (5k images) and report results on the 5k V-COCO test set.

We fine-tune our human-object interaction models for 10k iterations on the V-COCO trainval set with a learning rate of 0.001 and an additional 3k iterations with a rate of 0.0001. We use a weight decay of 0.0001 and a momentum of 0.9. We use synchronized SGD on 8 GPUs, with each GPU hosting 2 images (so the effective mini-batch size per iteration is 16 images). The fine-tuning time is ∼\scriptstyle\sim2.5 hours on the V-COCO trainval set on 8 GPUs.

Baselines.

To have a fair comparison with Gupta & Malik , which used VGG-16 , we reimplement their best-performing model (‘model C’ in ) using the same ResNet-50-FPN backbone as ours. In addition, only reported AProle{}_{\text{role}} on a subset of 19 actions, but we are interested in all actions (listed in Table 2). We therefore report comparisons in both the 19-action and all-action cases.

The baselines from are shown in Table 1. Our reimplementation of is solid: it has 37.5 AProle{}_{\text{role}} on the 19 action classes tested on the val set, 11 points higher than the 26.4 reported in . We believe that this is mainly due to ResNet-50 and FPN. This baseline model, when trained on the trainval set, has 31.8 AProle{}_{\text{role}} on all action classes tested on the test set. This is a strong baseline (31.8 AProle{}_{\text{role}}) to which we will compare our method.

Our method, InteractNet, has an AProle{}_{\text{role}} of 40.0 evaluated on all action classes on the V-COCO test set. This is an absolute gain of 8.2 points over the strong baseline’s 31.8, which is a relative improvement of 26%. This result quantitatively shows the effectiveness of our approach.

Qualitative Results.

We show our human-object interaction detection results in Figure 2 and Figure 5. Each subplot illustrates one detected ⟨\langlehuman, verb, object⟩\rangle triplet, showing the location of the detected person, the action taken by this person, and the location (and category) of the detected target object for this person/action. Our method can successfully detect the object outside of the person bounding box and associate it to the person and action.

Figure 7 shows our correctly detected triplets of one person taking multiple actions on multiple objects. We note that in this task, one person can take multiple actions and affect multiple objects. This is part of the ground-truth and evaluation and is unlike traditional object detection tasks in which one object has only one ground-truth class.

Moreover, InteractNet can detect multiple interaction instances in an image. Figure 6 shows two test images with all detected triplets shown. Our method detects multiple persons taking different actions on different target objects.

The multi-instance, multi-action, and multi-target results in Figure 6 and Figure 7 are all detected by one forward pass in our method, running at about 135ms per image on a GPU.

Ablation Studies.

In Table 3–5 we evaluate the contributions of different factors in our system to the results.

With vs. without target localization. Target localization, performed by the human-centric branch, is the key component of our system. To evaluate its impact, we implement a variant without target localization. Specifically, for each type of action, we perform k-means clustering on the offsets between the target RoIs and person RoIs (via cross-validation we found k=2k=2 clusters performs best). This plays a role similar to density estimation, but is not aware of the person appearance and thus is not instance-dependent. Aside from this, the variant is the same as our full approach.

Table 3 (a) vs. (c) shows that our target localization contributes significantly to AProle{}_{\text{role}}. Removing it shows a degradation of 5.6 points from 37.5 to 31.9. This result shows the effectiveness of our target localization (see Figure 4). The per-category results are in Table 2.

With vs. without the interaction branch. We also evaluate a variant of our method when removing the interaction branch. We can instead use the action prediction from the human-centric branch (see Figure 3). Table 3 (b) vs. (c) shows that removing the interaction branch reduces AProle{}_{\text{role}} just slightly by 0.7 point. This again shows the main effectiveness of our system is from the target localization.

With vs. without FPN. Our model is a generic human-object detection framework and can support various network backbones. We recommend using the FPN backbone, because it performs well for small objects that are more common in human-object detection.

Table 4 shows a comparison between ResNet-50-FPN and a vanilla ResNet-50 backbone. The vanilla version follows the ResNet-based Faster R-CNN presented in . Specifically, the full-image convolutional feature maps are from the last residual block of the 4-th stage (res4), on which the RoI features are pooled. On the RoI features, each of the region-wise branches consists of the residual blocks of the 5-th stage (res5). Table 4 shows a degradation of 1.6 points in AProle{}_{\text{role}} when not using FPN. We argue that this is mainly caused by the degradation of the small objects’ detection AP, as shown in . Moreover, the vanilla ResNet-50 backbone is much slower, 225ms versus 135ms for FPN, due to use of res5 in the region-wise branches.

Pairwise Sum vs. MLP. In our interaction branch, the pairwise outputs from two RoIs are added (Figure 3). Although simple, we have found that more complex variants do not improve results. We compare with a more complex transform in Table 5. We concatenate the two 1024-d features from the final fully-connected layers of the interaction branch for the two RoIs and feed it into an 2-layer MLP (512-d with ReLU for its hidden layer), followed by action classification. This variant is slightly worse (Table 5), indicating that it is not necessary to perform a complex pairwise transform (or there is insufficient data to learn this).

Per-action accuracy. Table 2 shows the AP for each action category defined in V-COCO, for the baseline, InteractNet without target localization, and our full system. We observe leading performance of AProle{}_{\text{role}} consistently. The actions with largest improvement are those with high variance in the spatial location of the object such as hold, look, carry, and cut. On the other hand, actions such as ride, kick, and read show small or no improvement.

Failure Cases.

Figure 8 shows some false positive detections. Our method can be incorrect because of false interaction inferences (e.g., top left), target objects of another person (e.g., top middle), irrelevant target objects (e.g., top right), or confusing actions (e.g., bottom left, ski vs. surf). Some of them are caused by a failure of reasoning, which is an interesting open problem for future research.

Mixture Density Networks.

To improve target localization prediction, we tried to substitute the uni-modal regression network with a Mixture Density Network (MDN) . The MDN predicts the mean and variance of MM relative locations for the objects of interaction conditioned on the human appearance. Note that MDN with M=1M=1 is an extension of our original approach that also learns the variance in (2). However, we found that the MDN layer does not improve accuracy. More details and discussion regarding the MDN experiments can be found in Appendix A.

HICO-DET Dataset.

We additionally evaluate InteractNet on HICO-DET which contains 600 types of interactions, composed of 117 unique verbs and 80 object types (identical to COCO objects). We train InteractNet on the train set, as specified by the authors, and evaluate performance on the test set using released evaluation code. Results are shown in Table 6 and discussed more in Appendix B.

Appendix A: Mixture Density Networks

The target localization module in InteractNet learns to predict the location of the object of interaction conditioned on the appearance of the human hypothesis hh. As an alternative, we can allow the relative location of the object to follow a multi-modal conditional distribution. For this, we replace our target localization module with a Mixture Density Network (MDN) , which parametrizes the mean, variance and mixing coefficients of MM components of the conditional normal distribution (MM is a hyperparameter). This instantiation of InteractNet is flexible can can capture different modes for the location of the objects of interaction.

The localization term for scoring bo∣hb_{o|h} is defined as:

The mixing coefficients are required to have whm∈w^{m}_{h}\in and ∑mwha,m=1\sum_{m}w^{a,m}_{h}=1. We parametrize μ\mu as a 4-D vector for each component and action type. We assume a diagonal covariance matrix and thus parametrize σ\sigma as a 4-D vector for each component and action type. Compare the localization term in (5) with the term in (2). Inference is unchanged, except we use the new form of gh,oag_{h,o}^{a} when evaluating (4).

To instantiate MDN, the human-centric branch of our network must predict ww, μ\mu, and σ\sigma. We train the network to minimize −log⁡(gh,oa)-\log(g_{h,o}^{a}) given bob_{o} (the location of the ground truth object for each interaction). We use a fully connected layer followed by a softmax operation to predict the mixing coefficients wha,mw^{a,m}_{h}. For μ\mu and σ\sigma, we also use fully connected layers. However, in the case of σ\sigma, we use a softplus operation (f(x)=log⁡(ex+1)f(x)=\log(e^{x}+1)) to enforce positive values for the covariance coefficients. This leads to stable training, compared to a variant which parametrizes log⁡σ2\log\sigma^{2} and becomes unstable due to an exponential term in the gradients. In addition, we found that enforcing a lower bound on the covariance coefficients was necessary to avoid overfitting. We set this value to 0.3 throughout our experiments.

Table 7 shows the performance of InteractNet with MDN and compare it to our original model. Note that the MDN with M=1M=1 is an extension of our original approach, with the only difference that it learns the variance in (2). The performance of MDN with M=1M=1 is similar to our original model, which suggests that a learned variance does not lead to an improvement. With M=2M=2, we do not see any gains in performance, possibly because the human appearance is a strong cue for predicting the object’s location and possibly due to the limited number of objects per action type in the dataset. Based on these results, along with the relative complexity of MDN, we chose to use our simpler proposed target localization model instead of MDN. Nonetheless, our experiments demonstrate that MDN can be trained within the InteractNet framework, which may prove useful.

Appendix B: HICO-DET Dataset

As discussed, we also test InteractNet on the HICO-DET dataset . HICO-DET contains approximately 48k images and is annotated with 600 interaction types. The annotations include boxes for the humans and the objects of interactions. There are 80 unique object types, identical to the COCO object categories, and 117 unique verbs.

Objects are not exhaustively annotated on HICO-DET. To address this, we first detect objects using a ResNet50-FPN object detector trained on COCO as described in . These detections are kept frozen during training, by setting a zero-valued weight on the object detection loss of InteractNet. To train the human-centric and interaction branch, we assign ground truth labels from the HICO-DET annotations to each person and object hypothesis based on box overlap. To diversify training, we jitter the object detections to simulate a set of region proposals. We define 117 action types and use the object detector to identify the type for the object of interaction, e.g. orange vs. apple.

We train InteractNet on the train set defined in , for 80k iterations and with a learning rate of 0.001 (a 10x step decay is applied after 60k iterations). In the interaction branch, we use dropout with a ratio of 0.5. We evaluate performance on the test set using released evaluation code.

Results are shown in Table 6. We show a 27% relative performance gain compared to the published results in and a 9% relative gain compared to our baseline approach. Figure 9 shows example predictions made by InteractNet.

References