HOI Analysis: Integrating and Decomposing Human-Object Interaction

Yong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li, Cewu Lu

Introduction

Human-Object Interaction (HOI) takes up most of the human activities. As a composition, HOI consists of three parts: . To detect HOI, machines need to simultaneously locate human and object and classify the verb . Except for the direct thinking that maps pixels to semantics, in this work we rethink HOI and explore two questions in a novel perspective (Fig. 1): First, as for the inner structure of HOI, how do isolated human and object compose HOI? Second, what is the relationship between two human-object pairs with the same HOI?

For the first question, we may find some clues from psychology. The view of Gestalt psychology is usually summarized as one simple sentence: “The whole is more than the sum of its parts” . This is also in line with human perception. Baldassano et al.et~{}al. studied the mechanism of how the brain builds HOI representation and concluded that the encoding of HOI is not the simple sum of human and object: a higher-level neural representation exists. Specific brain regions, e.g., posterior superior temporal sulcus (pSTS), are responsible for integrating isolated human and object into coherent HOI . Hence, to encode HOI, we may need complex nonlinear transformation (integration) to combine isolated human and object. Also, we argue that a reverse process is also essential to decompose HOI into isolated human and object (decomposition). Here we use TI(⋅)T_{I}(\cdot) and TD(⋅)T_{D}(\cdot) to indicate integration and decomposition functions. According to , isolated human and object are different from coherent HOI pair. Therefore, TI(⋅)T_{I}(\cdot) should be able to add interactive relationship to isolated elements. On the contrary, TD(⋅)T_{D}(\cdot) should eliminate this interactive information. Through the semantic change before and after transformations, we can reveal the “eigen” structure of HOI carrying the semantics. Considering that verb is hard to represent explicitly in image space, our transformations are conducted in latent space. For the second question, directly transforming one human-object pair to another (inter-pair transformation) is difficult. We need to consider not only isolated element differences but also interaction pattern change. However, with TI(⋅)T_{I}(\cdot) and TD(⋅)T_{D}(\cdot), things are different. We can first decompose HOI pair-ii into isolated person-ii and object-ii and eliminate the interaction semantics. Next, we transform human-ii (object-ii) to human-jj (object-jj) (j≠ij\neq i). The last step is to integrate human-jj and object-jj into pair-jj and add the interaction simultaneously.

Interestingly, we find the above process is kind of like Harmonic Analysis: to process the signal, we usually use Fourier Transform (FT) to decompose it into the integration of basic exponential functions; then we can modulate the exponential functions via very simple transformations like scalar-multiplication; finally, inverse FT can help us integrate the modulated elements and map them back to the input space. This elegant property brings a lot of convenience for signal processing. Therefore, we mimic this insight and design our methodology, i.e., HOI Analysis. To implement HOI Analysis, we propose an Integration-Decomposition Network (IDN). In detail, after extracting the features from human/object and human-object tight union boxes, we perform the integration TI(⋅)T_{I}(\cdot) to integrate the isolated human and object into the union in latent space. Moreover, decomposition TD(⋅)T_{D}(\cdot) is then performed to decompose the union into isolated human and object instances again. Through the transformations, IDN can learn to represent the interaction/verb with TI(⋅)T_{I}(\cdot) and TD(⋅)T_{D}(\cdot). That said, we first embed verbs in transformation function space, then learn to add and eliminate interaction semantics and classify interactions during transformations. For the inter-pair transformation, we adopt a simple instance exchange policy. For each human/object, we beforehand find its similar instances as candidates and randomly exchange the original instance with candidates in training. This policy can avoid complex transformation like motion transfer . Hence, we can focus on the learning of TI(⋅)T_{I}(\cdot) and TD(⋅)T_{D}(\cdot). Moreover, the lack of samples for rare HOIs can also be alleviated. To train IDN, we adopt the objectives derived from transformation principles, such as integration validity, decomposition validity and interactiveness validity (detailed in Sec. 3.4). With them, IDN can effectively model the interaction/verb in transformation function space. Subsequently, IDN can be applied to the HOI detection task by comparing the above validities and greatly advance it.

Our contributions are threefold: (1) Inspired by Harmonic Analysis, we thereon devise HOI Analysis to model the HOI inner structure. (2) A concise Integration-Decomposition Network (IDN) is proposed to conduct the transformations in HOI Analysis. (3) By learning verb representation in transformation function space, IDN achieves state-of-the-art performance on HOI detection.

Related Work

Human-Object Interaction (HOI) detection is crucial for deeper scene understanding and can facilitate behavior and activity learning . Recently, huge progress has been made in this field with the promotion of large-scale datasets and deep learning. HOI has been studied for a long history. Previously, most methods adopted hand-crafted features. With the renaissance of neural networks, recent works start to leverage learning-based features with end-to-end paradigm. HO-RCNN utilized a multi-stream model to leverage human, object and spatial patterns respectively, which is widely followed by subsequent works . Differently, GPNN adopted a graph model to address HOI learning for both images and videos. Instead of directly processing all human-object pairs generated from detection, TIN utilized interactiveness estimation to filter out non-interactive pairs in advance. In terms of modality, Peyre et al.et~{}al. explored to learn a joint space via aligning the visual and linguistic features and used word analogy to address unseen HOIs. DJ-RN recovered 3D human and object (location and size) and learned a 2D-3D joint representation. Finally, some works also explore to encode HOI with the help of a knowledge base. Based on human part-level semantics, HAKE built a large-scale part state knowledge base and Activity2Vec for finer-grained action encoding. Xu et al.et~{}al. constructed a knowledge graph from HOI annotations and the external source to advance the learning.

Besides the computer vision community, HOI is also studied in human perception and cognition researches. In , Baldassano et al.et~{}al. studied how human brain models HOI given HOI images. Interestingly, besides the brain regions responsible for encoding isolated human or object, certain regions can integrate isolated human and object into a higher-level joint representation. For example, pSTS can coherently model the HOI, instead of simply summing isolated human and object information. This phenomenon inspires us to rethink the nature of HOI representation. Thus we propose a novel HOI Analysis method to encode HOI by integration and decomposition.

On the other hand, HOI learning is similar to another compositional problem: attribute-object learning . Attribute-object compositions have many interesting properties such as contextuality, compositionality and symmetry . To learn the attribute-object, attributes are seen as primitives equal with objects or linear/non-linear transformations . Different from attributes expressed on object appearance, verbs in HOIs are more implicit and hard to locate in images. They are a kind of holistic representation of composed human and object instances. Thus, we propose several transformation validities to embed and capture the verbs in transformation function space, instead of utilizing an explicit classifier to classify them or using language priors .

Method

In an image, human and object can be explicitly seen. However, we can hardly depict which region is the verb. For “hold cup”, “hold” may be obvious and center on the hand and cup. But for “ride bicycle”, most parts of the person and bicycle all represent “ride”. Hence, vision systems may struggle given diverse interactions as it is hard to capture the appropriate visual regions. Though attention mechanism may help, the long-tail distribution of HOI data usually makes it unstable. In this work, instead of directly finding the interaction region and mapping it to semantics , we propose a novel learning paradigm, i.e., learning the verb representation via HOI Analysis.

Inspired by the perception study , we propose the integration TI(⋅)T_{I}(\cdot) and decomposition TD(⋅)T_{D}(\cdot) functions. HOI naturally consists of human, object and implicit verb. Thus, we can decompose HOI into basic elements and integrate them again like Harmonic Analysis. The overview of HOI Analysis is depicted in Fig. 2. As HOI is not the simple sum of isolated human and object , different from FT, our transformations are nonequivalent. The key difference lies in the addition and elimination of implicit interactions. We use binary interactiveness , which indicates whether human and object are interactive, to monitor these semantic changes. Hence, the interactiveness of isolated human/object is False, and joint human-object has True interactiveness. From the above, TI(⋅)T_{I}(\cdot) should have the ability to “add” interaction to isolated instances and make the integrated human-object has True interactiveness. On the contrary, TD(⋅)T_{D}(\cdot) can “eliminate” the interaction between coherent human-object and force their interactiveness to be False. At last, to encode the implicit verbs, we represent them in the transformation function space. A pair of decomposition and integration functions are constructed for each verb and forced to operate the appropriate transformations.

We introduce the feature preparation as follows. First, given an image, we use an object detector to obtain the human/object boxes bh,bob_{h},b_{o}. Then, we adopt a COCO pre-trained ResNet-50 to extract human/object RoI pooling features fha,foaf^{a}_{h},f^{a}_{o} from the third ResNet Block, where aa indicates visual appearance. For simplicity, we use tight union box of human and object to represent the coherent HOI pair (union). Notably, coherent HOI carries the interaction semantics and is more than the sum of isolated human and object , i.e., the incoherent ones. With bh,bob_{h},b_{o}, the union box bub_{u} can be easily obtained. The RoI pooling feature of bub_{u} is thus adopted from the fourth ResNet Block as the appearance representation of coherent HOI (fuaf^{a}_{u}). Note that fuaf^{a}_{u} is twice the size of fha,foaf^{a}_{h},f^{a}_{o}, for passing through one more ResNet Block. Second, to encode the box location, we generate location features fhb,fob,fubf^{b}_{h},f^{b}_{o},f^{b}_{u}, where bb indicates box location. We follow the box coordinate normalization method , getting the normalized box bh^,bo^\hat{b_{h}},\hat{b_{o}}. Next, for union box, we concatenate bh^\hat{b_{h}} and bo^\hat{b_{o}} and feed them to an MLP to get fubf^{b}_{u}. For human/object box, bh^\hat{b_{h}} or bo^\hat{b_{o}} is also fed to an MLP to get fhbf^{b}_{h} or fobf^{b}_{o}. The size of fhbf^{b}_{h} or fobf^{b}_{o} is half the size of fubf^{b}_{u}. Third, the location features fub,fhb,fobf^{b}_{u},f^{b}_{h},f^{b}_{o} are concatenated respectively to their corresponding appearance features fua,fha,foaf^{a}_{u},f^{a}_{h},f^{a}_{o}, getting fu^,fh^,fo^\hat{f_{u}},\hat{f_{h}},\hat{f_{o}}. The size of fh^\hat{f_{h}} and fo^\hat{f_{o}} are also half the size of fu^\hat{f_{u}}. For convenience, we concatenate fh^\hat{f_{h}} and fo^\hat{f_{o}} as fh^⊕fo^\hat{f_{h}}\oplus\hat{f_{o}}.

Before transformations, we compress these features to reduce the computational burden via an auto-encoder (AE). This AE is given fu^\hat{f_{u}} as input and pre-trained with an input-output reconstruction loss and a verb classification loss (Sec. 3.4). The classification score is denoted as SvAES_{v}^{AE}. After pre-training, we use AE to compress fu^\hat{f_{u}} and fh^⊕fo^\hat{f_{h}}\oplus\hat{f_{o}} to 1024 sized fuf_{u} (coherent) and fh⊕fof_{h}\oplus f_{o} (isolated) respectively, Finally, we have fu,fh⊕fof_{u},f_{h}\oplus f_{o} for integration and decomposition. The ideal transformations are:

where TD(⋅),TI(⋅)T_{D}(\cdot),T_{I}(\cdot) indicates the decomposition and integration functions, ⊕\oplus indicates the linear operation between isolated human and object features such as element-wise summation or concatenation. In most cases, concatenation performs better. As for the inter-pair transformation, we use

where fhi,foif^{i}_{h},f^{i}_{o} indicate the features of human/object instances, and gh(⋅),go(⋅)g_{h}(\cdot),g_{o}(\cdot) are the inter-human/object transformation functions. Because the strict inter-pair transformation like motion transfer is complex and not our main goal, we implement gh(⋅)g_{h}(\cdot) and go(⋅)g_{o}(\cdot) as simple feature replacement for simplicity. For human instances, we find their substitutional persons with the same HOI according to the pose similarity. As to object instances, we use the objects of the same category and similar sizes as the substitutions. All substitutional candidates come from the same dataset (train set) and are randomly sampled during training. From the experiment (Sec. 4.5), we find that this policy performs well and effectively improves the interaction representation learning.

We propose a concise Integration-Decomposition Network (IDN) as shown in Fig. 3. IDN mainly consists of two parts: the first one is the integration and decomposition transformations (Sec. 3.2) which construct a loop between the union and human/object features; the second one is the inter-pair transformation (Sec. 3.3) that exchanges the human/object instances between pairs with same HOI. In Sec 3.4, we introduce the training objectives derived from the transformation principles. With them, IDN would learn more effective interaction representations and advance HOI detection in Sec. 3.5.

2 Integration and Decomposition

As shown in Fig. 3, IDN constructs a loop consists of two inverse transformations: integration and decomposition implemented with MLPs. That is, we represent the verb/interaction in MLP weight space or transformation function space. For each verb, we adopt a pair of appropriative MLPs as integration and decomposition functions, e.g., TIvi(⋅)T^{v_{i}}_{I}(\cdot) and TDvi(⋅)T^{v_{i}}_{D}(\cdot) for verb viv_{i}. For integration, when inputting a pair of isolated fhf_{h} and fof_{o}, {TIvi(⋅)}i=1n\{T^{v_{i}}_{I}(\cdot)\}^{n}_{i=1} integrates them into nn outputs for nn verbs:

where ii = 1,2,3,....n1,2,3,....n and nn is the number of verbs, fuvif^{v_{i}}_{u} is the integrated union feature for the ii-thth verb. ⊕\oplus indicates concatenation. Through the integration function set {TIvi(⋅)}i=1n\{T^{v_{i}}_{I}(\cdot)\}^{n}_{i=1}, we get a set of 1024 sized integrated union features {fuvi}i=1n\{f^{v_{i}}_{u}\}^{n}_{i=1}. If the original fuf_{u} contains the semantics of the ii-thth verb, it should be close to fuvif^{v_{i}}_{u} and far away from the other integrated union features. Second, the subsequent decomposition is depicted as follows. Given the integrated union feature set {fuvi}i=1n\{f^{v_{i}}_{u}\}^{n}_{i=1}, we also use nn decomposition functions {TDvi(⋅)}i=1n\{T^{v_{i}}_{D}(\cdot)\}^{n}_{i=1} to decompose them respectively:

The decomposition output is also a set of features {fhvi⊕fovi}i=1n\{f^{v_{i}}_{h}\oplus f^{v_{i}}_{o}\}^{n}_{i=1}, where fhvif^{v_{i}}_{h} and fovif^{v_{i}}_{o} all have the same size with fhf_{h} and fof_{o}. Similarly, if this human-object pair is performing the ii-thth interaction, the original input fh⊕fof_{h}\oplus f_{o} should be close to the fhvi⊕fovif^{v_{i}}_{h}\oplus f^{v_{i}}_{o} and far away from the other {fhvj⊕fovj}j≠in\{f^{v_{j}}_{h}\oplus f^{v_{j}}_{o}\}^{n}_{j\neq i}.

3 Inter-Pair Transformation

Inter-Pair Transformation (IPT) is proposed to reveal the inherent nature of implicit verb, i.e., the shared information between different pairs with the same HOI. Here, we adopt a simple implementation: instance exchange policy. For humans, we first use pose estimation to obtain poses and then operate alignment and normalization. In detail, the pelvis keypoints of all persons are aligned and all the distances between head and pelvis are scaled to one. Hence, we can find similar persons according to the pose similarity, which is calculated as the sum of Euclidean distances between the corresponding keypoints of two persons. To keep the semantics, similar persons should have at least one same HOI. Selecting similar objects is simpler, we directly choose the objects of the same category. An extra criterion is that we choose objects with similar sizes. We use the area ratio between the object box and the paired human box as the criteria. Finally, mm similar candidates are selected for each human/object. This whole selection is operated within one dataset. Formally, with instance exchange, Eq. 3 can be rewritten as:

where k1,k2k_{1},k_{2} = 1,2,3,...m1,2,3,...m and mm is the number of selected similar candidates, here mm = 55. And {TIvi(⋅)}i=1n\{T^{v_{i}}_{I}(\cdot)\}^{n}_{i=1} and {TDvi(⋅)}i=1n\{T^{v_{i}}_{D}(\cdot)\}^{n}_{i=1} should be equally effective before and after instance exchange. During training, we first use Eq. 3 for a certain number of epochs and then replace Eq. 3 with Eq. 5 (Sec. 4.2). When using Eq. 5, we put the original instance and its exchanging candidates together and randomly sample them. Notably, we focus on the transformations between pairs with the same HOI. The transformations between different HOIs which need to manipulate the corresponding human posture, human-object spatial configuration and interactive pattern are beyond the scope of this paper. For IPT, more sophisticate approaches are also possible, e.g., using motion transfer to adjust 2D human posture according to another person with the same HOI but different posture (eating while sitting/standing), recovering 3D HOI and adjusting 3D pose to generate new images/features, using language priors to change the classes of interacted objects or HOI compositions , etc. But these are beyond the scope of our main insight, so we leave these to the future work.

4 Transformation Principles as Objectives

Before training, we first pre-train AE to compress the inputs. We first feed fu^\hat{f_{u}} to the encoder and obtain the compressed fuf_{u}. Then an MLP takes fuf_{u} as input to classify the verbs with Sigmoids (one pair can have multiple HOIs simultaneously) with cross-entropy loss LclsAEL^{AE}_{cls}. Meanwhile, fuf_{u} is decoded and generates fureconf^{recon}_{u}. We construct MSE reconstruction loss LreconAEL^{AE}_{recon} between fureconf^{recon}_{u} and fu^\hat{f_{u}}. The overall loss of AE is LAEL^{AE} = LclsAE+LreconAEL^{AE}_{cls}+L^{AE}_{recon}. After pre-training, AE will be fine-tuned together with transformation modules. Next, we detail the objectives derived from transformation principles.

Integration Validity. As aforementioned, we integrate fhf_{h} and fof_{o} into the union feature set {fuvi}i=1n\{f^{v_{i}}_{u}\}^{n}_{i=1} for all verbs (Eq. 3 or 5). If integration is able to “add” the verb semantics, the corresponding fuvif^{v_{i}}_{u} that belongs to the ongoing verb classes should be close to the real fuf_{u}. For example, if coherent fuf_{u} contains the semantics of verb vpv_{p} and vqv_{q}, then fuvpf^{v_{p}}_{u} and fuvqf^{v_{q}}_{u} should be close to fuf_{u}. Meanwhile, {fuvi}i≠p,qn\{f^{v_{i}}_{u}\}^{n}_{i\neq p,q} should be far away from fuf_{u}. Hence, we can construct the distance:

For nn verb classes, we can get distance set {duvi}i=1n\{d^{v_{i}}_{u}\}^{n}_{i=1}. Considering above principle, if fuf_{u} carries the pp-thth verb semantics, dupd^{p}_{u} should be small, and vice versa. Therefore, we can directly use the negative distances as the score of verb classification, i.e. SvuS^{u}_{v} = {−duvi}i=1n\{-d^{v_{i}}_{u}\}^{n}_{i=1}. Naturally, SvuS^{u}_{v} is then used to generate verb classification loss LclsuL^{u}_{cls} = Lentu+LhingeuL^{u}_{ent}+L^{u}_{hinge}, where LentuL^{u}_{ent} is cross-entropy loss, LhingeuL^{u}_{hinge} = ∑i=1n[yimax⁡(0,dvi−t1vi)+(1−yi)max⁡(0,t0vi−dvi)]\sum_{i=1}^{n}[y_{i}\max(0,d^{v_{i}}-t_{1}^{v_{i}})+(1-y_{i})\max(0,t_{0}^{v_{i}}-d^{v_{i}})]. yiy_{i} = 11 indicates this pair has verb viv_{i} and otherwise yiy_{i} = . t0t_{0} and t1t_{1} are chosen following semi-hard mining strategy: t1vit_{1}^{v_{i}} = min⁡B−vi(dvi)\min_{B_{-}^{v_{i}}}(d^{v_{i}}) and t0vit_{0}^{v_{i}} = max⁡B+vi(dvi)\max_{B_{+}^{v_{i}}}(d^{v_{i}}), where B−viB_{-}^{v_{i}} denotes all the pairs without verb viv_{i} in the current mini-batch, and B+viB_{+}^{v_{i}} denotes all the pairs with verb viv_{i} in the current mini-batch.

Decomposition Validity. This validity is proposed to constrain the decomposed {fhvi⊕fovi}i=1n\{f^{v_{i}}_{h}\oplus f^{v_{i}}_{o}\}^{n}_{i=1} (Eq. 4). Similar to Eq. 6, we also construct nn distances between {fhvi⊕fovi}i=1n\{f^{v_{i}}_{h}\oplus f^{v_{i}}_{o}\}^{n}_{i=1} and fh⊕fof_{h}\oplus f_{o} as

and obtain {dhovi}i=1n\{d^{v_{i}}_{ho}\}^{n}_{i=1}. Again {dhovi}i=1n\{d^{v_{i}}_{ho}\}^{n}_{i=1} should obey the same principle according to ongoing verbs. Thus, we get the second verb score SvhoS^{ho}_{v} = {−dhovi}i=1n\{-d^{v_{i}}_{ho}\}^{n}_{i=1} and verb classification loss LclshoL^{ho}_{cls}.

Interactiveness Validity. Interactiveness depicts whether a person and an object are interactive. Thus, it is False if and only if human-object do not have any interactions. As the “1+1>2” property , the interactiveness of isolated human fhf_{h} or object fof_{o} should be False, so does fh⊕fof_{h}\oplus f_{o}. But after we integrate fh⊕fof_{h}\oplus f_{o} into {fuvi}i=1n\{f^{v_{i}}_{u}\}^{n}_{i=1}, its interactiveness should be True. Meanwhile, the original union fuf_{u} should have True interactiveness. We adopt one shared FC-Sigmoid as the binary classifier for fuf_{u}, fh⊕fof_{h}\oplus f_{o} and {fuvi}i=1n\{f^{v_{i}}_{u}\}^{n}_{i=1}. The binary label converted from HOI label is zero if and only if a pair does not have any interactions. Notably, we also adopt the interactiveness validity upon decomposed {fhvi⊕fovi}i=1n\{f^{v_{i}}_{h}\oplus f^{v_{i}}_{o}\}^{n}_{i=1} but achieve limited improvement. To keep the model concise, we just adopt the other three effective interactiveness validities hereinafter. Thus, we obtain three binary classification cross entropy losses: Lbinu,Lbinho,LbinIL^{u}_{bin},L^{ho}_{bin},L^{I}_{bin}. For clarity, we use a unified LbinL_{bin} = Lbinu+Lbinho+LbinIL^{u}_{bin}+L^{ho}_{bin}+L^{I}_{bin}.

The overall loss of IDN is LL = Lclsu+Lclsho+LbinL^{u}_{cls}+L^{ho}_{cls}+L_{bin}. With the guidance of these principles, IDN can well capture the interaction changes during the transformations. Different from previous methods that aim at encoding the entire HOI representations statically, IDN focuses on dynamically inferring whether an interaction exists within human-object through the integration and decomposition. So IDN can alleviate the learning difficulty of complex and various HOI patterns.

5 Application: HOI Detection

We further apply IDN to HOI detection, which needs to simultaneously locate human-object and classify the ongoing interactions. For locations, we adopt the detected boxes from a COCO pre-trained Faster R-CNN , so does the object class probability PoP_{o}. Then, verb scores can be obtained from Eq. 6 and 7. SvuS^{u}_{v} = {−duvi}i=1n\{-d^{v_{i}}_{u}\}^{n}_{i=1}, SvhoS^{ho}_{v} = {−dhovi}i=1n\{-d^{v_{i}}_{ho}\}^{n}_{i=1} and SvAES^{AE}_{v} obtained from AE are then fed to exponential functions or Sigmoids to generate PvuP^{u}_{v} = exp⁡(Svu)\exp(S^{u}_{v}), PvhoP^{ho}_{v} = exp⁡(Svho)\exp(S^{ho}_{v}) and PvAEP^{AE}_{v} = Sigmoid(SvAE)Sigmoid(S^{AE}_{v}). Since the validity losses would pull the features that meet the labels together and push away the others, thus here we directly use three kinds of distances to classify verbs. For example, if fuf_{u} contains the ii-thth verb, duvi=∣∣fu−fuvi∣∣2d^{v_{i}}_{u}=||f_{u}-f^{v_{i}}_{u}||_{2} should be small (probability should be large); if not, duvid^{v_{i}}_{u} should be large (probability should be small). The final verb probabilities is acquired via PvP_{v} = α(Pvu+Pvho+PvAE)\alpha(P^{u}_{v}+P^{ho}_{v}+P^{AE}_{v}), here α\alpha = 13\frac{1}{3}. For HOI triplets, we get their HOI probabilities using PHOIP_{HOI} = Pv∗PoP_{v}*P_{o} for all possible compositions according to the benchmark setting.

Experiment

In this section, we first introduce the adopted datasets, metrics (Sec. 4.1) and implementation (Sec. 4.2). Next, we compare IDN with the state-of-the-art on HICO-DET and V-COCO in Sec. 4.3. As HOI detection metrics expect both accurate human/object locations and verb classification, the performance strongly relies on object detection. Hence, we conduct experiments to evaluate IDN with different object detectors. At last, ablation studies are conducted (Sec. 4.5).

We adopt the widely-used HICO-DET and V-COCO . HICO-DET consists of 47,776 images (38,118 for training and 9,658 for testing) and 600 HOI categories (80 COCO objects and 117 verbs). V-COCO contains 10,346 images (2,533 and 2,867 in train and validation sets, 4,946 in test set). Its annotations include 29 verb categories (25 HOIs and 4 body motions) and same 80 objects with HICO-DET. For HICO-DET, we use mAP following : true positive needs to contain accurate human and object locations (box IoU with reference to GT box is larger than 0.5) and accurate verb classification. The role means average precision is used for V-COCO.

2 Implementation Details

The encoder of the adopted AE compresses the input feature dimension from 4608 to 4096, then to 1024. The decoder is structured symmetrical to the encoder. For HICO-DET , AE is pre-trained for 4 epochs using SGD with a learning rate of 0.1, momentum of 0.9, while each batch contains 45 positive and 360 negative pairs. The whole IDN (AE and transformation modules) is first trained without inter-pair transformation (IPT) for 20 epochs using SGD with a learning rate of 2e-2, momentum of 0.9. Then we finetune IDN with IPT for 30 epochs using SGD, with a learning rate of 1e-3, momentum of 0.9. Each batch for the whole IDN contains 15 positive and 120 negative pairs. For V-COCO , AE is first pre-trained for 60 epochs. The whole IDN is trained without IPT for 45 epochs using SGD, then fine-tuned with IPT for 20 epochs. The other training parameters are the same as those for HICO-DET. In testing, LIS is adopted with TT = 8.3,k8.3,k=12.0,ω12.0,\omega=10.010.0. Following , we use NIS in all testings with the default threshold and the interactiveness estimation of fuf_{u}. All experiments are conducted on one single NVIDIA Titan Xp GPU.

3 Results

Setting. We compare IDN with state-of-the-art on two benchmarks in Tab. 1 and Tab. 3. For HICO-DET, we follow the settings in : Full (600 HOIs), Rare (138 HOIs), Non-Rare (462 HOIs) in Default and Known Object sets. For V-COCO, we evaluate AProleAP_{role} (24 actions with roles) on Scenario 1 (S1) and Scenario 2 (S2). To purely illustrate the HOI recognition ability without the influence of object detection, we conduct evaluations with three kinds of detectors: COCO pre-trained (COCO), pre-trained on COCO and then finetuned on HICO-DET train set (HICO-DET), GT boxes (GT) in Tab. 1.

Comparison. With TI(⋅)T_{I}(\cdot) and TD(⋅)T_{D}(\cdot), IDN outperforms previous methods significantly and achieves 23.36 mAP on the Default Full set of HICO-DET with COCO detector. Moreover, IDN is the first to achieve more than 20 mAP on all three Default sets without additional information used. Moreover, the improvement on the Rare set proves that the dynamically learned interaction representation can greatly alleviate the data deficiency of rare HOIs. With the HICO-DET finetuned detector, IDN also shows great improvements and achieves more than 26 mAP and further proves the affect from detections (23.36 to 26.29 mAP). Given GT boxes, the gaps among the other three methods are marginal. But IDN achieves more than 9 mAP improvement on HOI recognition solely. All these greatly verify the efficacy of our integration and decomposition. On V-COCO , IDN achieves 53.3 mAP on S1 and 60.3 mAP on S2, both significantly outperforming previous methods. Moreover, we also apply our IDN to existing HOI methods since its flexibility as a plug-in. In detail, we apply integration and decomposition to iCAN as a proxy task to enhance its feature learning. The performance improves from 14.84 mAP to 18.98 mAP (HICO-DET Full).

Efficiency and Scalability. In IDN, each verb is represented by a pair of MLPs (TIvi(⋅)T^{v_{i}}_{I}(\cdot) and TDvi(⋅)T^{v_{i}}_{D}(\cdot)). To ensure the efficiency, we carefully designed the data flow to make IDN is able to run on a single GPU. All transformations are operated in parallel and the inference speed is 10.04 FPS (iCAN : 4.90 FPS, TIN : 1.95 FPS, PPDM : 14.08 FPS, PMFNet : 3.95 FPS). We also considered an implementation which utilizes a single MLP for all verbs for scalability, i.e., conditioned MLP functions fuvi=TI(fh⊕fo,fviID)f^{v_{i}}_{u}=T_{I}(f_{h}\oplus f_{o},f^{ID}_{v_{i}}) and fhvi⊕fovi=TD(fuvi,fviID)f^{v_{i}}_{h}\oplus f^{v_{i}}_{o}=T_{D}(f^{v_{i}}_{u},f^{ID}_{v_{i}}), where fviIDf^{ID}_{v_{i}} is the verb indicator (one-hot/Word2Vec /Glove ). For new verbs, we just change the verb indicator instead of increasing MLPs. It works similar to zero-shot learning like TAFE-Net and Nan et al.et~{}al. , but performs worse (20.86 mAP, HICO-DET Full) than the reported version (23.36 mAP).

4 Visualization

To verify the effectiveness of transformations, we use t-SNE to visualize fu,fh⊕fo,TIv(fh⊕fo)f_{u},f_{h}\oplus f_{o},T_{I}^{v}(f_{h}\oplus f_{o}) for different vv in Fig. 5. We can find integrated TIv(fh⊕fo)T_{I}^{v}(f_{h}\oplus f_{o}) obviously closer to the real union fuf_{u}, while the simple linear combination fh⊕fof_{h}\oplus f_{o} cannot represent the interaction information. We also analyze the IPT. In detail, we randomly select a pair with verb vv and denote its features as fh,fo,fuf_{h},f_{o},f_{u}. Assume there are mm other pairs with verb vv, whose features are {fhi,foi,fui}i=1m\{f_{h}^{i},f_{o}^{i},f_{u}^{i}\}_{i=1}^{m}. Then we calculate D1vD_{1}^{v} = ∑i=1m∥fu−fui∥2m\frac{\sum_{i=1}^{m}\|f_{u}-f_{u}^{i}\|^{2}}{m} and D2vD_{2}^{v} = ∑i=1m(∥TIv(fh⊕foi)−fui∥2+∥TIv(fhi⊕fo)−fui∥2)2m\frac{\sum_{i=1}^{m}(\|T_{I}^{v}(f_{h}\oplus f_{o}^{i})-f_{u}^{i}\|_{2}+\|T_{I}^{v}(f_{h}^{i}\oplus f_{o})-f_{u}^{i}\|_{2})}{2m}. Here, D1vD_{1}^{v} is the mean distance from fuf_{u} to {fui}i=1m\{f_{u}^{i}\}_{i=1}^{m}. D2vD_{2}^{v} is the mean distance from TIv(fh⊕foi)T_{I}^{v}(f_{h}\oplus f_{o}^{i}) and TIv(fhi⊕fo)T_{I}^{v}(f_{h}^{i}\oplus f_{o}) to fuif_{u}^{i} (i=1,2,...,mi=1,2,...,m). If IPT can effectively transform one pair to another by exchanging the human/object, there should be D1v>D2vD_{1}^{v}>D_{2}^{v}. We compare D1vD_{1}^{v} and D2vD_{2}^{v} of 20 different verbs in Fig. 5. As shown, in most cases D1vD_{1}^{v} is much larger than D2vD_{2}^{v}, indicating the effectiveness of IPT.

5 Ablation Study

We conduct ablation studies on HICO-DET with COCO detector. The results are shown in Tab. 3. (1) Modules: The performance of each module is evaluated. TIT_{I}, TDT_{D} and AE achieve 21.26, 21.05, 17.27 mAP respectively and show complementary property. (2) Objectives: During training, we drop one of the three validity objectives respectively. Without anyone of them, IDN shows obvious degradation, especially integration validity. (3) Inter-Pair Transformation (IPT): IDN without IPT achieves 22.63 mAP, showing the importance of instance exchange policy. (4) AE: AE is pre-trained with: reconstruction loss LreconAEL^{AE}_{recon} and verb classification loss LclsAEL^{AE}_{cls}. The removal of LclsAEL^{AE}_{cls} hurts the performance more severely, especially on the Rare set, while LreconAEL^{AE}_{recon} also plays an important auxiliary role in boosting the performance. (5) Transformation Order: In practice, we construct a loop (fh⊕fof_{h}\oplus f_{o} to {fuvi}i=1n\{f^{v_{i}}_{u}\}^{n}_{i=1} to {fhvi⊕fovi}i=1n\{f^{v_{i}}_{h}\oplus f^{v_{i}}_{o}\}^{n}_{i=1}) to train IDN with the consistency. Using fuf_{u} instead of {fuvi}i=1n\{f^{v_{i}}_{u}\}^{n}_{i=1}, i.e., fh⊕fof_{h}\oplus f_{o} to {fuvi}i=1n\{f^{v_{i}}_{u}\}^{n}_{i=1} and fuf_{u} to {fhvi⊕fovi}i=1n\{f^{v_{i}}_{h}\oplus f^{v_{i}}_{o}\}^{n}_{i=1}, performs worse (21.77 mAP).

Conclusion

In this paper, we propose a novel HOI learning paradigm named HOI Analysis, which is inspired by Harmonic Analysis. And an Integration-Decomposition Network (IDN) is introduced to implement it. With the integration and decomposition between the coherent HOI and isolated human and object, IDN can effectively learn the interaction representation in transformation function space and outperform the state-of-the-art on HOI detection with significant improvements.

Broader Impact

In this work, we propose a novel paradigm for Human-Object Interaction detection, which would promote human activity understanding. Our work could be useful for vision applications, such as the health care system in an intelligent hospital. Current activity understanding systems are usually computationally expensive and require high computational resources, and could cost many financial and environmental resources. Considering this, we will release our code and trained models to the community, as part of efforts to alleviate the repeated training of future works.

Acknowledgments and Disclosure of Funding

This work is supported in part by the National Key R&D Program of China, No. 2017YFA0700800, National Natural Science Foundation of China under Grants 61772332 and Shanghai Qi Zhi Institute, SHEITC (2018-RGZN-02046).

References

Appendix A Visualized Results

We visualize some HOI detection results of our IDN on HICO-DET in Fig. 6. As shown, IDN is able to decompose and integrate various HOIs in diverse scenes and accurately detect them.

Appendix B Result Analysis

We illustrate the detailed comparison between our method, Peyre et al.et~{}al. and DJ-RN on Rare set on HICO-DET in Fig. 7. We can find that our IDN outperforms Peyre et al.et~{}al. and DJ-RN on various rare HOIs. The effectiveness of our IDN on Rare set proves that the dynamically learned interaction representation can greatly alleviate the data deficiency of the rare HOIs.

Appendix C Code

We provide our source code in https://github.com/DirtyHarryLYL/HAKE-Action-Torch/tree/IDN-(Integrating-Decomposing-Network) under our project HAKE-Action-Torch (https://github.com/DirtyHarryLYL/HAKE-Action-Torch).