Reformulating HOI Detection as Adaptive Set Prediction
Mingfei Chen, Yue Liao, Si Liu, Zhiyuan Chen, Fei Wang, Chen Qian
Introduction
Human-Object Interaction (HOI) detection aims to identify HOI triplets human, verb, object from a given image, it is an important step toward the high-level semantic understanding . Conventional HOI methods can be divided into two-stage methods and one-stage methods . Most two-stage methods detect instances (humans and objects), and match the detected humans and objects one by one to form pair-wise proposals in the first stage. Next, in the second stage, such methods infer the interactions based on the features of cropped human-object pair-wise proposals. Two-stage methods have made great progress in HOI detection, however, their efficiency and effectiveness are limited by their serial architectures. With the development of one-stage object detectors, one-stage HOI detectors have raised a new fashion. Existing one-stage HOI detectors formulate HOI detection as a parallel detection problem, which detects the HOI triplets from an image directly. One-stage methods have delivered great improvements in both efficiency and effectiveness.
Determining which regions to concentrate on is critical and challenging for HOI detectors. To obtain essential features for interaction prediction, conventional two-stage methods usually involve extra features, \eg, human pose and language . However, even with extra features, two-stage methods still focus on the detected instances that might be inaccurate, which are less adaptive and limited by the detected instances. One-stage methods partially alleviate these issues by inferring interactions directly from the whole image. Such methods intuitively define a location-relative medium to predict interactions, and can be mainly divided into anchor-based methods and point-based methods. Anchor-based methods predict the interactions based on the union box of each pair-wise human and object instances. While point-based methods infer the interaction midpoint of each corresponding human-object pair. However, we argue that it is sub-optimal to predict the interaction through a pre-defined interaction location. Figure 1 illustrates an example. The interaction “direct” (in yellow) and “drive” (in purple) are quite different and thus require different visual features for interaction prediction. However, their union boxes are considerably overlapped (Figure 1 (a)), and their interaction midpoints are very close (Figure 1 (b)). Therefore, these one-stage methods concentrate on similar visual features for the two different interactions.
To further address the limitation of interaction location in one-stage methods, we reformulate interaction detection as a set-based prediction problem. We define an interaction query set with several learnable embeddings, and an interaction prediction set. Each embedding in the query set is mapped by a transformer based interaction decoder to an interaction prediction set. By feeding the interaction query set into a multi-head co-attention module, we are able to adaptively aggregate features from global contexts. Our proposed method matches each ground-truth with the resembling interaction prediction for adaptive supervision. Therefore, our proposed method adaptively concentrates on the most suitable features for each prediction, free from the location limitation of conventional one-stage methods. As demonstrated in Figure 1 (c), our method aggregates arm features of the left person and pose features of the right person to make two different interaction predictions. The predictions are then matched with the ground-truth interaction “direct” and “drive” respectively.
To this end, we propose a novel Adaptive Set-based one-stage framework, namely AS-Net. Our AS-Net consists of two parallel branches: an instance branch and an interaction branch. Both branches leverage a transformer encoder-decoder structure, which utilize global features to perform set predictions. The instance branch predicts location and category for each instance, while the interaction branch predicts interaction vectors and their corresponding categories. The interaction vectors point from the centers of the human instances to the centers of the object instances. We obtain the predicted interaction triplets by matching each interaction vector from the interaction branch with the detected instances from the instance branch. Besides, we exploit an instance-aware attention module in a co-attention manner to perform branch aggregation. Specifically, this module aggregates information in the instance branch and introduces the aggregated features into the interaction branch. We also utilize semantic embeddings to perform more accurate human-object matching.
We test our proposed AS-Net on three datasets, \ie, HICO-Det , V-COCO , and HOI-A . Our proposed AS-Net outperforms all the other algorithms among all datasets. In specific, our proposed AS-Net has gained relative improvement comparing to the previous state-of-the-art one-stage method on HICO-DET.
Our contributions can be concluded in the following three aspects:
We formulate HOI detection as a set prediction problem, which breaks the instance-centric limitation and location limitation of the existing methods. Thereby, our method can adaptively concentrate on the most suitable features to improve the predicting accuracy.
We propose a novel one-stage transformer-based HOI detection framework, namely AS-Net. We also design an instance-aware attention module to introduce the information in the instance branch into the interaction branch.
Without introducing any extra features, our method outperforms all the previous state-of-the-art methods, achieving relative improvement over the second best one-stage method on the HICO-DET dataset.
Related Work
Two-stage Methods. Most conventional HOI detectors are in a two-stage manner. In the first stage, an object detector is applied to detect the instances. In the second stage, the cropped instance features are classified to obtain the interaction categories. In addition to the cropped instance features, previous methods leverage combined spatial features , union box features , or context features to improve the accuracy of HOI detection. In order to concentrate on more interaction-relevant features, some methods utilize extra features, such as human pose , human parts and language features . However, the serial architectures of such two-stage methods impair the efficiency of HOI detection. Moreover, the prediction accuracy is usually limited by the results of instance detection.
One-stage Methods. Recently, one-stage HOI detection methods with higher efficiency have attracted increasing attention. Most one-stage methods extract features with a bottom-up structure, and detect the HOI triplets in parallel from an image directly. Specifically, the one-stage methods can be divided into anchor-based methods and point-based methods according to the manners of their interaction prediction. The anchor-based methods predict the interactions based on each union box. The point-based methods perform inference at each interaction key point, such as the midpoint of each corresponding human-object pair. Though breaking the limitation from instance detection, such methods which pre-assign each ground-truth interaction to the predictions, are still non-adaptive and limited by the interaction locations.
Methods
HOI detection aims to predict the triplet of human, verb, object, which contains a pair of bounding-boxes for a human and an object, and a corresponding verb category. In this paper, we reformulate HOI detection as a set prediction problem, and propose an Adaptive Set-based one-stage Network (AS-Net).
Our AS-Net builds on a transformer encoder-decoder architecture and makes parallel set-based predictions for the HOI triplets. As illustrated in Figure 2, our proposed AS-Net consists of four parts. We first utilize a backbone (Section 3.1) to extract the visual feature sequence with global contexts. The instance (Section 3.2) and interaction branches (Section 3.3) following the backbone parallelly detect an instance and interaction prediction set from the feature sequence respectively. In order to intensify the instance features that are valuable for interaction inference, we design an instance-aware attention module (Section 3.4) to perform branch aggregation. Specifically, we introduce semantic embeddings (Section 3.5) in instance and interaction branches for more accurate triplet prediction. At the end, we match detected instances and interactions to obtain the final HOI triplets (Section 3.6).
2 Instance Branch
Training. For the set-based training process, we first find a one-to-one bipartite matching between the detected instance set and the ground-truth (padded with no-instance to a set of size ). To this end, we deploy a matching loss, which is the summation of bounding-box loss and category semantic distance between instance and all ground-truth bounding-boxes. Following , the bounding-box loss is composed of a loss and a GIoU loss . The category semantic distance is the negative of summation of the predicted scores for each ground-truth category.
The universal index permutation set of predictions is denoted as . We consider that minimizes the summation of all the matching cost as the optimal index permutation of the detected instance set, which we adopt the Hungarian algorithm to calculate. The -th element of the index permutation is defined as , and the is formulated as:
For the instance prediction with index permutation , the predicted bounding-box and category are represented as and respectively. We follow the DETR detector to construct the set-based instance detection loss :
where and denotes the bounding-box and category of the matched ground-truth instance respectively, is the confidence score for category .
3 Interaction Branch
Training. We denote the ground-truth interaction as , where is the interaction vector of , and indicates the ground-truth interaction categories of . We compute the matching loss between and each predicted interaction , where refers to the predicted interaction vector and indicates the confidence scores of the interaction categories. The matching cost can be computed by:
where refers to the score for the -th ground-truth interaction category of . Similar to the set-based training process for the instance branch, we utilize the Hungarian algorithm to find the optimal index assignment for the predicted interaction set w.r.t. the ground-truth.
For the interaction prediction with index , we define the predicted interaction vector and categories as and respectively, and the matched target interaction vector and categories are and respectively. To balance the ratio between the positive and negative samples for each classifier, we apply Focal loss , denoted as , for the training of interaction classification. Besides, we adopt loss, denoted as , for the regression of interaction vectors. The interaction loss is calculated as:
where and are the weight coefficients of and respectively.
Analysis. Adaptation is involved in the interaction prediction from two aspects. First, for each interaction query, we apply multi-head co-attention to aggregate information from each element in the feature sequence. Hence, each query can adaptively aggregate the interaction-relevant visual features. Second, instead of pre-assigning each ground-truth to the corresponding prediction, we consider both the predicted interaction vectors and categories to match each ground-truth interaction with the resembling prediction. Therefore, each interaction prediction can be supervised by the most suitable ground-truth more adaptively.
4 Instance-aware Attention
We construct an instance-aware attention module between each instance and interaction layer to emphasize relevant instance features for interaction prediction.
We then apply Softmax to obtain the instance-aware attention weight matrix :
5 Semantic Embedding
The interaction vectors are not pointing to a instance directly, instead, they point to a region. Instead of matching which only employs location indication from the interaction vectors, we introduce semantic embeddings inferred by an MLP block in our matching strategy. We infer semantic embeddings from for each detected instance in instance branch. And in the interaction branch, two semantic embeddings and are inferred from , one for the human instance and another for the object instance for each prediction.
In the training process, the semantic embeddings of different instances are pushed away from each other. The push procedure can be described as:
where refers to the total number of the ground-truth instances, and refers to the semantic embedding of the predicted instance matched to the -th target instance. If the distance between two semantic embeddings are more than a threshold , we consider two embeddings are separate enough and set to .
We pull the semantic embeddings that refer to the same instance towards each other:
where we denote the predicted human semantic embedding as and the object embedding as for the interaction prediction with index . The semantic embedding and refer to the same human instance in the instance and interaction branches respectively. Similarly, and refer to the same object instance. refers to the total number of the ground-truth interactions.
6 Training Loss and Post-processing
The target loss is the weighted sum of the losses mentioned above:
where is a hyper-parameter to balance different loss.
During the post-processing, we first match the detected human instances with object instances based on our predicted interaction vectors and semantic embeddings. A good human-object interaction match should meet the following three requirements: 1) the normalized center of the matched human/object instances is close to the start and the end point of the interaction vector respectively; 2) the matched instances have high confidence scores on their predicted categories; 3) the semantic embedding referring to the same matched instances are similar to each other.
We consider all detected instances as object instances. For each predicted interaction vector , the matching distance can be calculated as:
When the semantic embeddings are introduced for matching, given the human () and object () semantic embedding of the instance branch, the embedding matching distance for the predicted human () and object () semantic embedding of the interaction branch can be defined as:
The final matching cost is calculated as . We match the detected instances with the minimum matching cost to each interaction prediction. The HOI confidence score for each predicted triplet is the product of the interaction category score, and the matched instance scores and . Triplets with top confidence scores are preserved as the final HOI triplet predictions.
Experiments
Datasets. To verify the effectiveness of our model, we conduct experiments on three HOI detection datasets HICO-DET , V-COCO and HOI-A . HICO-DET contains images for training and images for testing, contains the same object categories as MS-COCO and verb categories. The objects and verbs form classes of HOI triplets. V-COCO provides images for training, images for validating and images for testing. V-COCO is derived from MS-COCO dataset, annotated with action categories. HOI-A dataset consists of annotated images, kinds of objects and action categories.
Metrics. Following , the mean average precision (mAP) is adopted as evaluation metric. For one positive predicted HOI triplet human, verb, object, both the predicted human and object bounding-boxes have IoUs greater than w.r.t. the ground-truth boxes, with the correct predicted verb simultaneously.
2 Implementation Details
During training, we resize the shortest side of the input image to the range $1,333\lambda_{\operatorname{cls}}\lambda_{\operatorname{reg}}\lambda_{\operatorname{emb}}120.1907510^{-4}107010^{-5}10809.06432$ GPUs.
3 Comparing to State-of-the-art
We conduct experiments on three HOI detection benchmarks to verify the effectiveness of our AS-Net. It is shown in Table 1, Table 2 and Table 3 that our AS-Net has achieved state-of-the-art across all the three benchmarks. Specifically, on the HICO-DET dataset, comparing to the previous state-of-the-art one-stage method PPDM which adopts Hourglass-104 as backbone, our AS-Net has achieved a performance gain with a relatively light-weight backbone, i.e., ResNet-. Since the object detectors in two-stage methods are purely trained on MS-COCO, which does not fine-tune on HICO-DET, thus we also show the result when only training the interaction branch for fair comparison. In this setting, our AS-Net* has achieved mAP, which is superior to all existing two-stage methods, and has achieved above mAP improvements.
We compare our results on the V-COCO dataset with other state-of-the-art methods. Freezing the instance detection related parameters pretrained on the MS-COCO dataset, we only train the remaining parameters of our model. As shown in Table 2, our model achieves on , outperforms the previous works. Considering the relatively small scale of the V-COCO dataset may impair the representation capability of the trained semantic embeddings, we test the results using the matching strategy without the semantic embeddings.
The Table 3 also illustrates our effectiveness on the HOI-A test set. We reach a mAP of , better than all the previous methods, including the method which adopts a relatively heavy-weight Hourglass-104 as backbone.
4 Ablation Study
Matching Strategy. Two variants of inference matching methods are implemented. As shown in Table 4a, when only using the vector matching distance , or only using the semantic embedding distance in Section 3.6, the effectiveness are both compromised.
Semantic Embedding Settings. To explore the suitable semantic embedding setting, we evaluate the models with different embedding dimension and weight coefficient of the training losses and . As shown in Table 4b, the effectiveness of our model is not sensitive to the embedding dimension. As changes from to , the changing of the mAP result is only point. The embedding dimension is set to regarding the trade-off for both effectiveness and computational cost. As illustrated in Table 4c, the model performs best when training with , while the effectiveness will be impaired when is increased or decreased.
Single Branch Variant. We implemented a single branch variant to detect instances along their interactions while keeping all hyper-parameters. As shown in Table 4d, the variant achieves mAP on the HICO-DET dataset, which is lower than our AS-Net. Especially, the Rare mAP is , which is lower than ours. We consider it is because detection and interaction rely on some different features. Lacking the interaction-related features such as human postures, the single branch variant is more likely to infer actions that frequently appear in the presence of the detected objects.
Basic Model. To verify the effectiveness of the basic framework, we implement a variant consists of one -layer instance detection branch and one -layer interaction detection branch, without the instance-aware interaction attention module and semantic embedding. Table 4d articulates that our basic model (Basic Model, Int) achieves mAP on the HICO-DET dataset, which outperforms the previous methods by a large margin.
Instance-aware Attention. Two other variants are evaluated by utilizing the instance-aware attention module to verify the contribution of branch aggregation. As presented in Table 4d, the instance-aware attention module on our basic model (+ IA Attn, Int w/o emb) improves mAP by point. For the basic model with the semantic embeddings (+ Int w/ emb), the improvements are points using the instance-aware attention module (+ IA Attn, Int w/ emb). Therefore, we conclude that the instance-aware attention features from the instance branch are valuable for the interaction prediction.
Semantic Embedding & Instance-aware Attention. The basic model with the semantic embeddings (+ Int w/ emb) improves slightly comparing to the basic model (Basic Model, Int) without the instance-aware attention as shown in Table 4d. As the bridge connecting predicted instances and interaction vectors, the semantic embeddings also contribute to the training. However, the semantic embedding is less powerful than the instance-aware interaction attention module from the results. Based on the basic model with the semantic embedding, several variants are implemented, which consist of different interaction decoder layers or attention modules additionally:1) instance-aware attention modules with -layer interaction decoder (+ IA Attn, Int w/ emb), performs attention every other layer; 2) instance-aware attention modules with -layer interaction decoder (+ IA Attn, Int w/ emb). From Table 4d, the performance is improved by about point utilizing the attention module and the semantic embedding jointly. Besides, it’s better to use the instance-aware modules and the decoder layers with the same number of times. The effectiveness reduces slightly when we utilize the two modules with less times, while the amount of the model parameters is reduced significantly.
5 Qualitative Results
As shown in the first three rows of the Figure 3, we visualize the interaction decoder attention for some interaction pairs in our basic model, the basic model with the semantic embeddings (+ Int w/ emb) and the model with both the instance-aware attention module as well as the semantic embeddings (+ IA Attn, Int w/ emb), respectively. We also visualize the instance-aware attention in the last row for each example interaction pairs to present how the attention module contributes to the interaction prediction.
From the Figure 3 (a), the basic model without any branch aggregation focuses on some scattered redundant feature regions and leave out some interaction-relevant features. From the Figure 3 (b), the model with semantic embeddings only partially alleviates the problem. For example, there is a girl holding an umbrella in the figures in the first column. To predict such interaction, the basic model concentrates on the head of the girl and the body of an irrelevant person. Correspondingly, the model with semantic embeddings pays attention to the edge of the umbrella and the body of the girl, while still concentrates on an irrelevant person. When we involve the instance-aware attention module, as shown in Figure 3 (c) and Figure 3 (d), the interaction branch concentrates on the whole umbrella and some body parts which are close to the umbrella, and the instance-aware attention module focuses on the body and head of the girl. In such a separated focus mechanism, our model can concentrate on the features more accurately.
Conclusion and Future Work
In this paper, we reformulate HOI detection as an adaptive set prediction problem and propose a novel one-stage HOI detection framework, namely AS-Net. By aggregating interaction-relevant features from global contexts, and matching each ground-truth with the interaction prediction, our method demonstrates adaptive ability on both feature aggregation and supervision. Moreover, the designed instance-aware attention module contributes to intensify the instructive instance features, and we also introduce semantic embeddings to improve performance. The ablation studies verify the effectiveness of each key component of our model. Our AS-Net outperforms all existing methods on three HOI detection datasets. In the future, we plan to extend AS-Net to handle more general association problems, \eg, visual relationship detection and multi-object tracking.
This work was partially supported by Sensetime Ltd. Group, National Natural Science Foundation of China under Grant 61876177, Zhejiang Lab (No. 2019KD0AB04), Beijing Natural Science Foundation 4202034 and Fundamental Research Funds for the Central Universities.