Cascaded Human-Object Interaction Recognition
Tianfei Zhou, Wenguan Wang, Siyuan Qi, Haibin Ling, Jianbing Shen
Introduction
Human-object interaction (HOI) recognition aims to identify meaningful triplets from images, such as in Fig. 1. It plays a crucial role in many vision tasks, e.g., visual question answering norcliffe2018learning; li2019relation; zheng2019reasoning, human-centric understanding wang2018attentive; Wang_2019_ICCV; zhoucvpr2020, image generation johnson2018image, and activity recognition shao2018find; wu2019learning; fan2019understanding; pang2018deep; ma2018attend, to name a few representative ones.
Though great advances have been made recently, the task is still far from being solved. One of the main challenges comes from its intrinsic complexity: a successful HOI recognition model must accurately 1) localize and recognize each interacted entity (human, object), and 2) predict the interaction classes (verb). Both subtasks are difficult, leading to HOI recognition itself a highly complex problem. With a broader view of other computer vision and machine learning related fields, coarse-to-fine and cascade inference have been shown to deal well with complex problems kirkpatrick1983optimization; felzenszwalb2010cascade; felzenszwalb2006efficient; wang2019iterative. The central idea is to leverage sequences of increasingly fine approximations to control the complexity of learning and inference. This motivates us to propose a cascade HOI recognition model, which builds up multiple stages of neural network inference in an annealing-style. For the two subtasks of instance localization and interaction recognition, this model arranges them in a successive manner within each single stage, and carries out cascade, cross-stage inference for each. Above designs result in a multi-task, coarse-to-fine inference framework, which enables asymptotically improved HOI representation learning. This also distinctively differentiates our method from previous efforts, which rely on single-stage architectures.
As shown in Fig. 1, our model consists of an instance localization network and an interaction recognition network, both working in a cascade manner. Through the instance localization network, the model step-by-step increases the selectiveness of the instance proposals. With such progressively refined HOI candidates, as well as the useful relation representation from the preceding stage, better action predictions can be achieved by current-stage interaction recognition network. Moreover, in the interaction recognition network, both human semantics and facial patterns are mined to boost relation reasoning, as these cues are tied to underlying purposes of human actions. With such human-centric features, a relation ranking module (RRM) is proposed to rank all the possible human-object pairs. Only the top-ranked, high-quality candidates are fed into a relation classification module (RCM) for final verb prediction.
More essentially, previous HOI literature mainly address relation detection, i.e., recognizing HOIs at a bounding-box level. In addition to addressing this classic setting, we take a further step towards more fine-grained HOI understanding, i.e., identifying the relations between interacted entities at the pixel level (see Fig. 1). Studying such relation segmentation setting not only further demonstrates the efficacy and flexibility of our cascade framework, but allows us to explore more powerful relation representations. This is because bounding box based representations only encode coarse object information with noisy backgrounds, while pixel-wise mask based features may capture more detailed and precise cues. We empirically study the effectiveness of bounding box and pixel-wise mask based relation representations as well as their hybrids. Our results suggest that the pixel-mask representation is indeed more powerful.
Our model reached the place in ICCV-2019 Person in Context Challenge http://picdataset.com/challenge/leaderboard/pic2019 (PIC19 Challenge), on both Human-Object Interaction in the Wild (HOIW) and Person in Context (PIC) tracks, where HOIW addresses relation detection, while PIC focuses on relation segmentation. Besides, it also obtains promising results on V-COCO gupta2015visual.
This paper makes three major contributions. First, we formulate HOI recognition as a coarse-to-fine inference procedure with a novel cascade architecture. Second, we introduce several techniques to learn rich features that represent the semantics of HOIs. Third, for the first time, we study the feature representations of HOI and find pixel-mask to be more powerful than the traditional bounding-box representation. We expect such a study could inspire more future efforts towards pixel-level HOI understanding.
Related Work
Human-object interaction recognition has a rich study history in computer vision. Early methods yao2010modeling; yao2010grouplet; delaitre2011learning mainly exploited human-object contextual information in structured models, such as Bayesian inference gupta2007objects; gupta2009observing, and compositional framework desai2012detecting.
With the recent renaissance of neural networks in computer vision, deep learning based solutions are now dominant in this field. For instance, in gkioxari2018detecting, a multi-branch architecture was explored to address human, object, and relation representation learning. Some researchers revisited the classic graph model and solved this task in a neural message passing framework qi2018learning. For learning more effective human feature representations, pose cues have been widely adopted in recent leading approaches li2019transferable; gupta2019no; wan2019pose; fang2018pairwise; Zhou_2019_ICCV. Some other efforts addressed long-tail distribution and zero-shot problems with external knowledge gu2019scene; kato2018compositional; zhuang2018hcvrd; shen2018scaling. All these models use single-stage pipelines for inference (Fig. 2 (a)), and they can potentially benefit from the general architecture we propose here: a multi-stage pipeline that performs coarse-to-fine inference as shown in Fig. 2 (b).
Object detection has gained remarkable progress recently, benefiting from the availability of large-scale datasets (e.g., MS-COCO lin2014microsoft) and strong representation power of deep neural networks. Mainstream methods are often categorized into two-stage ren2015faster; he2017mask; chen2019hybrid; cai2019cascade or single-stage redmon2016you; liu2016ssd; li2017fully; law2018cornernet paradigms. Recently, some multi-stage pipelines have been explored for coarse-to-fine object detection cai2018cascade; chen2019hybrid. Similarly, we revisit the general idea of cascade inference in HOI recognition, where both instance localization and relation recognition are coupled for step-by-step HOI reasoning.
Our Algorithm
To identify triplets in images, our method carries out progressive refinement on instance localization and relation recognition at multiple stages (see Fig. 2 (b)). At each stage , the multi-tasking is achieved by two networks: an instance localization network generates human and object proposals, and an instance recognition network identifies the action (i.e., verb) for each human-object pair sampled from the proposals, as shown in Fig. 3 (a). Our cascade network is organized as follows:
At stage , takes the detection results from as inputs and outputs refined results . Then, a human-object pair is sampled from . Finally, uses the relation features and of at current and previous stages to estimate a verb score vector . More details about the relation feature are given in §3.3.1. Notably, the instance localization and interaction recognition networks work closely at each stage, and can benefit from the improved localization results of and give better interaction predictions.
Next we will describe in detail our instance localization network in §3.2 and interaction recognition network in §3.3.
2 Instance Localization Network
The instance localization network outputs a set of human and object regions, from which human-object pair candidates are sampled and fed into the interaction recognition network for relation classification. It is built on a cascade of detectors, i.e., at stage , refines an object region detected from the preceding stage by:
Similar to previous cascade object detectors cai2018cascade; chen2019hybrid, at each stage, is trained with a certain interaction over union (IoU) threshold, and its output is re-sampled to train the next detector with a higher IoU threshold. In this way, we gradually increase the quality of training data for deeper stages in the cascade, thus boosting the selectiveness against hard negative examples. At each stage, the instance localization loss is the same as Faster R-CNN ren2015faster.
3 Interaction Recognition Network
As shown in Fig. 3 (a), the interaction recognition network comprises a relation ranking module (RRM, §3.3.2) and a relation classification module (RCM, §3.3.3). Both RRM and RCM rely on our elaborately designed human-centric relationship representation (§3.3.1).
Here H, O and U are specific instances of the RoIAlign feature Y in Eqs. (1, 2), which are renamed to make it clear that they come from different regions.
To better capture the underlying semantics in HOI, we introduce two feature-enhancement mechanisms: implicit human semantic mining to improve the human feature H and explicit facial region attending to enhance the object feature O. Then we have the visual feature as:
where and denote the enhanced human and object features, respectively, and is the concatenation operation. Next we detail our two feature-enhancement mechanisms.
1) Implicit Human Semantic Mining. To reason about human-object interactions, it is essential to understand how humans interact with the world, i.e., which human parts are involved for an action. Different from current leading methods resorting to expensive human pose annotations wan2019pose; fang2018pairwise; li2019transferable, we propose to implicitly learn human parts and their mutual interactions.
For each pixel (position) inside the human region (feature) H, we define its semantic context as the pixels that belong to the same semantic human part category of . We use such semantic context to enhance our human representation, as it captures the relations within and among parts. Such enhancement would require a human part label map. Here, we compute a semantic similarity map as a surrogate to expedite computation. Specifically, for each pixel we compute a semantic similarity map , where each element stores the ‘relation’ between the latent part categories of pixel and :
Then for a pixel , we collect information from its semantic context according to :
2) Explicit Facial Region Attending. Human face is vital for HOI understanding, as it conveys rich information closely tied to underlying attention and intention kleinke1986gaze of humans. There are many interactions that directly involve human face. For example, humans use eyes to watch TV, use mouth to eat food, and so on. Besides, face-related interactions are typically fine-grained and combined with heavy occlusions on the interacted objects, e.g., call a phone, play a phone, posing great difficulties for HOI models. To address the above issues, we propose another feature-enhancement mechanism, called explicit facial region attending. This mechanism enriches the object representation O via two attention mechanisms:
where is the sigmoid function, and stands for two stacked FC layers.
Considering Eqs. (8, 9), the object feature O is enhanced by:
We do not update semantic and geometric features.
3.2 Relation Ranking Module
Once obtaining the features of a human-object pair, we can directly predict its action label. However, a big issue here is how to sample human-object pairs. Given the proposals detected from the localization network, previous HOI methods typically pair all humans and objects, leading to large computational overhead. As a matter of fact, human beings interact with the world following some regularity rather than in a pure chaotic way baldassano2017human. By leveraging such regularity, we propose a human-object relation ranking module (RRM) to select high-quality HOI candidates for further relation recognition. This also helps decrease the difficulty in relation classification and erase the serious class imbalance, as the samples for ‘non-interaction’ class are much more than the ones of any other interaction classes.
Here means has a higher ranking than . gives the ranking score of :
In RRM, the learning of g is achieved by minimizing the following pairwise ranking hinge loss:
where the margin is empirically set as 0.2. This loss penalizes the situation that assigning an un-annotated pair with a higher ranking score, compared to a labeled pair .
3.3 Relation Classification Module
Through RRM, only a few top-ranked, high-quality human-object pairs are preserved and fed into a triple-stream zhang2019graphical, relation classification module (RCM) for final HOI recognition. For a HOI candidate (), the semantic , geometric and visual features, are separately fed into a corresponding stream in RCM for estimating a HOI action score vector independently:
where , and are the score vectors from semantic, geometric and visual streams, respectively, and is the number of pre-defined actions in HOI. Note that here follows a multi-label classification setting.
During training, for each stream, the binary cross-entropy loss is used to evaluate the discrepancy between the output score and truth target. The total loss is the sum of the ones from streams. During inference, the final prediction is obtained by:
where denotes the Hadamard product.
4 Relation Segmentation
So far, we strictly follow the classic relation detection setting in HOI recognition gkioxari2018detecting; li2019transferable; fang2018pairwise; xu2019learning, i.e., identify the interaction entities by bounding boxes. Now we focus on how to adapt our cascade framework to relation segmentation, which addresses more fine-grained HOI understanding by representing each entity at the pixel level.
Inspired by cai2018cascade, for the instance localization network at each stage , an instance segmentation head is added and the whole workflow (Eqs. (1, 2)) is changed as:
where indicates a generated object instance mask. Then, in our relation recognition network (§3.3), the human-object pair is sampled from the object masks and associated with finer features: H, O and U by pixel-wise RoI. In addition, the generation of geometric feature is based on pixel-level masks. The binary cross-entropy loss is used for training .
5 Implementation Details
Training Loss. Since all the modules mentioned above are differentiable, our cascade architecture can be trained in an end-to-end manner. In the relation detection setting, the entire loss is computed as:
Here, is the localization loss at stage (§3.3). and are the losses of RRM (§3.3.2) and RCM (§3.3.3), respectively. The coefficients and are used to balance the contributions of different stages and tasks. There are three stages used in our method (), and we set . In the relation segmentation setting, the instance segmentation head is injected into the network (§3.4). The corresponding instance segmentation loss is further added in Eq. (18), with coefficients .
Cascade Inference. During inference, the object proposals generated by the instance localization network in different stages are merged together. We remove the ones whose confidence scores are smaller than 0.3. Then, all the possible human-object pairs, generated from the remaining proposals, are fed into RRM for relation ranking. After that, we only select the top 64 pairs as candidates and feed them into RCM for final relation classification. The last-stage output of RCM is used as the final action score.
Experiments
Experiments are conducted on three datasets, i.e., HOIW, PIC and V-COCO gupta2015visual. The former two are from the PIC19 Challenge, and the last one is a gold standard benchmark.
Training Settings: Unless specially noted, we adopt the following training settings for all the experiments. We use ResNet-50 he2016deep as the backbone. The training includes two phases: 1) training the instance localization network; and then 2) jointly training the instance localization and interaction recognition networks. In the first phase, the network is initialized using the weights pre-trained on COCO lin2014microsoft. The three stages are trained using gradually increased IoU thresholds cai2018cascade; chen2019hybrid. Training images are resized to a maximum scale of , without changing the aspect ratio. We apply horizontal flipping for data augmentation and train the network for 12 epochs with batch size 16 and initial learning rate 0.02, which is reduced by 10 at epoch 8 and 11. In the second phase, we adopt the image-centric training strategy girshick2015fast, i.e., using pairwise samples from one image to make up a mini-batch. For each mini-batch, we sample at most 128 HOI proposals with a ratio of 1:3 of positive to negative samples to jointly train RRM and RCM. At each stage, the same IoU threshold is used to determine positive HOI proposals so that the training data for the interaction recognition network closely match the detection quality. Besides, ground-truth HOIs are also used at each stage for training. The second phase is trained with learning rate 0.02 and batch-size 8 for 7 epochs.
Reproducibility: Our model is implemented on PyTorch and trained on 8 NVIDIA Tesla V100 GPUs with a 32GB memory per-card. Testing is conducted on a single NVIDIA TITAN Xp GPU with 12 GB memory.
Dataset: The PIC19 Challenge includes two tracks, i.e., HOIW and PIC tracks, each with a standalone dataset:
HOIW liao2019ppdm is for human-object relation detection. It has 29,842 training and 8,794 testing images, with bounding box annotations for 11 object and 10 action categories. Since it does not provide train/val splits, in our ablation study, we randomly choose 9,999 images for val and the other 19,843 for train; for the challenge result, we use train+val for training.
PIC is for human-object relation segmentation. It has 17,606 images (12,654 for train, 1,977 for val and 2,975 for test) with pixel-level annotations for 143 objects. It covers 30 relationships, including 6 geometric (e.g., next-to) and 24 non-geometric (e.g., look, talk).
Evaluation Metrics: Standard evaluation metrics in the challenges are adopted. For HOIW, the performance is evaluated by . A detected triplet is considered as a true positive if the predicted verb is correct and both the human and object boxes have IoUs at least 0.5 with the corresponding ground-truths. For PIC, we use Recall@100 (R@100), which is averaged over two relationship categories (i.e., geometric and non-geometric) and three IoU thresholds (i.e., 0.25, 0.5 and 0.75). In our ablation study, we also consider R@50 and R@20 to measure the performance under stricter conditions.
Performance on the HOIW Track: Our approach reaches the place for relation detection on the HOIW track. As reported in Table 3.4, our result is substantially better than other teams. In particular, it is 5.78% absolutely better than the (GMVM) and 9.11% better than the (FINet). Our approach also significantly outperforms one published state-of-the-art, i.e., TIN li2019transferable . Fig. 4 presents some visual results on HOIW test. Our model shows robust to various challenges, e.g., occlusions, subtle relationships, etc.
Performance on the PIC Track: Our approach also reaches the place for relation segmentation on the PIC track. As reported in Table 3.4, our overall score (52.52%) outperforms the place by 3.85% and the by 7.56%. Fig. 5 depicts visual results of two complex scenes on PIC test. Our method shows outstanding performance in terms of instance segmentation as well as interaction recognition. It can identify both geometric and non-geometric relationships, and is capable of recognizing many fine-grained interactions, e.g., look human, hold tableware. In this track, the instance localization network is instantiate as Eq. (17).
2 Results on V-COCO
Dataset: V-COCO gupta2015visual provides verb annotations for MS-COCO lin2014microsoft. Proposed in 2015, it is the first large-scale dataset for HOI understanding and remains the most popular one today. It contains 10,346 images in total ( for train/val/test splits). 16,199 human instances are annotated with 26 action labels, wherein three actions (i.e., cut, hit, eat) are annotated with two types of targets (i.e., instrument and direct object).
Evaluation Metrics: We use the original role mean AP (), which is exactly same with in HOIW.
Performance: Since V-COCO has both bounding box and mask annotations, we provide two variants of our methods, i.e., and , where is trained with box annotations while uses groundtruth masks. For fairness, during evaluation, the mask outputs of are transformed to boxes. Table 3 summarizes the results in comparison with 8 state-of-the-arts. outperforms TIN li2019transferable by 0.5% and RPNN Zhou_2019_ICCV by 0.8%. further improves by 0.6%, which suggests the superiority of the mask-level representation over the box-level. We would like to note that wan2019pose reported a on V-COCO. However, It relies on an expensive pose estimator, thus it is unfair to directly compare with our method. Without the pose estimator, wan2019pose obtains a score of , slightly worse than .
In Fig. 6, we illustrate HOI segmentation results of on V-COCO test set. It precisely recognizes many fine-grained interactions, such as look computer, read book, etc.
Overall, our model consistently achieves promising results over different datasets as well as two different settings (i.e., relation detection and segmentation), which clearly reveals its remarkable performance and strong generalization.
3 Ablation Study
Key Component Analysis. First, we investigate the influence of essential components in our framework, i.e., implicit human semantic mining (IHSM), explicit facial region attending (EFRA), relation ranking module (RRM) and cascade network architecture (CAS). We first build a baseline model without any of these components, and then gradually add each into the baseline for investigation. As reported in Table 4, all these components can improve the performance in both PIC and HOIW datasets. 1) IHSM and EFRA help to learn more discriminative visual features and further boost the performance (e.g., 0.5% and 2.8% performance improvements on HOIW). 2) Fig. 7 shows the per-category performance improvement of IHSM and EFRA on HOIW val set. Obviously, EFRA improves the performance on face-related interactions (e.g., eat, drink, smoking, call) and discriminates these categories from some similar ones, e.g., play phone. In contrast, IHSM is more effective for the actions with specific poses, e.g., ride, kick ball. 3) RRM plays a key role in pruning negative human-object pairs, as proved by Table 4. Moreover, RRM improves the average inference speed by about ms on HOIW. 4) Our cascade architecture substantially boosts the performance, i.e., 8.8% absolute improvement in PIC and 5.1% in HOIW.
Cascade Architecture Analysis. We study the impact of the number of stages used in our cascade network by varying it from to . The IoU thresholds used for these five stages are . The results in Table 5 show that the performance is significantly improved by adding a second stage, i.e., 6.5% in terms of R@20 in PIC and 3.5% in terms of in HOIW. When further adding more than 3 stages, the performance gain is marginal. Table 5 also reports the average inference time for these variants on HOIW val set. The test speed decreases with adding more stages and drops quickly after using or stages. Considering the model complexity and performance, we choose as our default setting. Table 6 reports the performance comparison of our approach with () or without () cascade under different backbones, i.e., ResNet-50, ResNet-101 and ResNeXt-101. The results reveal that our cascade network consistently improves the performance on various backbones.
Efficacy of Our Relation Representation and Score Fusion Strategy. In our method, three kinds of features, , and , are used for capture semantic, geometric and visual information for relation modeling. Table 7 reports the performance with only considering one single feature. As seen, the visual feature is more important than the other two. In addition, we further investigate different ways to fuse the action scores from the three features, we find that the one used in Eq. (16) is the best.
Exploring Better Relation Representation. Existing HOI methods typically use coarse bounding boxes to represent the entities, however, is it the best choice? To answer this, we perform experiments to explore more powerful relation representation. We evaluate the performance of our model on PIC val set using four different representations: a) BBox; b) Mask; c) BBox+Mask (max); and d) BBox+Mask (sum). Here, a) and b) means that we extract the features H, O, U by applying RoIAlign over bbox and mask regions, respectively. c) and d) are the fusion of bbox and mask features with element-wise max and sum operations, respectively. Note that the detected entities are the same for all the baselines. The results in Table 8 show that mask is superior to bbox, especially under the strictest metric R@20. The two hybrid representations are better than solely using bbox, but slightly worse than the purely mask-based. In summary, mask-based representation indeed benefits HOI recognition as it provides more precise information.
Conclusion
This paper introduces a cascade network architecture for coarse-to-fine HOI recognition. It consists of an instance localization network and an interaction recognition network, which are densely connected at each stage to fully exploit the superiority of multi-tasking. The interaction recognition network leverages human-centric features to learn better semantics of actions, and comprises two crucial modules for relation ranking and classification. Our model achieves the place on both relation detection and relation segmentation tasks in PIC19 Challenge, and also outperforms prior methods on a gold standard benchmark, V-COCO. Besides, we empirically demonstrate the advantages of mask over bounding box for more precise relation representation, and will go deep into this in our future research.