HOI Analysis: Integrating and Decomposing Human-Object Interaction
Yong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li, Cewu Lu
Introduction
Human-Object Interaction (HOI) takes up most of the human activities. As a composition, HOI consists of three parts:
For the first question, we may find some clues from psychology. The view of Gestalt psychology is usually summarized as one simple sentence: “The whole is more than the sum of its parts” . This is also in line with human perception. Baldassano studied the mechanism of how the brain builds HOI representation and concluded that the encoding of HOI is not the simple sum of human and object: a higher-level neural representation exists. Specific brain regions, e.g., posterior superior temporal sulcus (pSTS), are responsible for integrating isolated human and object into coherent HOI . Hence, to encode HOI, we may need complex nonlinear transformation (integration) to combine isolated human and object. Also, we argue that a reverse process is also essential to decompose HOI into isolated human and object (decomposition). Here we use and to indicate integration and decomposition functions. According to , isolated human and object are different from coherent HOI pair. Therefore, should be able to add interactive relationship to isolated elements. On the contrary, should eliminate this interactive information. Through the semantic change before and after transformations, we can reveal the “eigen” structure of HOI carrying the semantics. Considering that verb is hard to represent explicitly in image space, our transformations are conducted in latent space. For the second question, directly transforming one human-object pair to another (inter-pair transformation) is difficult. We need to consider not only isolated element differences but also interaction pattern change. However, with and , things are different. We can first decompose HOI pair- into isolated person- and object- and eliminate the interaction semantics. Next, we transform human- (object-) to human- (object-) (). The last step is to integrate human- and object- into pair- and add the interaction simultaneously.
Interestingly, we find the above process is kind of like Harmonic Analysis: to process the signal, we usually use Fourier Transform (FT) to decompose it into the integration of basic exponential functions; then we can modulate the exponential functions via very simple transformations like scalar-multiplication; finally, inverse FT can help us integrate the modulated elements and map them back to the input space. This elegant property brings a lot of convenience for signal processing. Therefore, we mimic this insight and design our methodology, i.e., HOI Analysis. To implement HOI Analysis, we propose an Integration-Decomposition Network (IDN). In detail, after extracting the features from human/object and human-object tight union boxes, we perform the integration to integrate the isolated human and object into the union in latent space. Moreover, decomposition is then performed to decompose the union into isolated human and object instances again. Through the transformations, IDN can learn to represent the interaction/verb with and . That said, we first embed verbs in transformation function space, then learn to add and eliminate interaction semantics and classify interactions during transformations. For the inter-pair transformation, we adopt a simple instance exchange policy. For each human/object, we beforehand find its similar instances as candidates and randomly exchange the original instance with candidates in training. This policy can avoid complex transformation like motion transfer . Hence, we can focus on the learning of and . Moreover, the lack of samples for rare HOIs can also be alleviated. To train IDN, we adopt the objectives derived from transformation principles, such as integration validity, decomposition validity and interactiveness validity (detailed in Sec. 3.4). With them, IDN can effectively model the interaction/verb in transformation function space. Subsequently, IDN can be applied to the HOI detection task by comparing the above validities and greatly advance it.
Our contributions are threefold: (1) Inspired by Harmonic Analysis, we thereon devise HOI Analysis to model the HOI inner structure. (2) A concise Integration-Decomposition Network (IDN) is proposed to conduct the transformations in HOI Analysis. (3) By learning verb representation in transformation function space, IDN achieves state-of-the-art performance on HOI detection.
Related Work
Human-Object Interaction (HOI) detection is crucial for deeper scene understanding and can facilitate behavior and activity learning . Recently, huge progress has been made in this field with the promotion of large-scale datasets and deep learning. HOI has been studied for a long history. Previously, most methods adopted hand-crafted features. With the renaissance of neural networks, recent works start to leverage learning-based features with end-to-end paradigm. HO-RCNN utilized a multi-stream model to leverage human, object and spatial patterns respectively, which is widely followed by subsequent works . Differently, GPNN adopted a graph model to address HOI learning for both images and videos. Instead of directly processing all human-object pairs generated from detection, TIN utilized interactiveness estimation to filter out non-interactive pairs in advance. In terms of modality, Peyre explored to learn a joint space via aligning the visual and linguistic features and used word analogy to address unseen HOIs. DJ-RN recovered 3D human and object (location and size) and learned a 2D-3D joint representation. Finally, some works also explore to encode HOI with the help of a knowledge base. Based on human part-level semantics, HAKE built a large-scale part state knowledge base and Activity2Vec for finer-grained action encoding. Xu constructed a knowledge graph from HOI annotations and the external source to advance the learning.
Besides the computer vision community, HOI is also studied in human perception and cognition researches. In , Baldassano studied how human brain models HOI given HOI images. Interestingly, besides the brain regions responsible for encoding isolated human or object, certain regions can integrate isolated human and object into a higher-level joint representation. For example, pSTS can coherently model the HOI, instead of simply summing isolated human and object information. This phenomenon inspires us to rethink the nature of HOI representation. Thus we propose a novel HOI Analysis method to encode HOI by integration and decomposition.
On the other hand, HOI learning is similar to another compositional problem: attribute-object learning . Attribute-object compositions have many interesting properties such as contextuality, compositionality and symmetry . To learn the attribute-object, attributes are seen as primitives equal with objects or linear/non-linear transformations . Different from attributes expressed on object appearance, verbs in HOIs are more implicit and hard to locate in images. They are a kind of holistic representation of composed human and object instances. Thus, we propose several transformation validities to embed and capture the verbs in transformation function space, instead of utilizing an explicit classifier to classify them or using language priors .
Method
In an image, human and object can be explicitly seen. However, we can hardly depict which region is the verb. For “hold cup”, “hold” may be obvious and center on the hand and cup. But for “ride bicycle”, most parts of the person and bicycle all represent “ride”. Hence, vision systems may struggle given diverse interactions as it is hard to capture the appropriate visual regions. Though attention mechanism may help, the long-tail distribution of HOI data usually makes it unstable. In this work, instead of directly finding the interaction region and mapping it to semantics , we propose a novel learning paradigm, i.e., learning the verb representation via HOI Analysis.
Inspired by the perception study , we propose the integration and decomposition functions. HOI naturally consists of human, object and implicit verb. Thus, we can decompose HOI into basic elements and integrate them again like Harmonic Analysis. The overview of HOI Analysis is depicted in Fig. 2. As HOI is not the simple sum of isolated human and object , different from FT, our transformations are nonequivalent. The key difference lies in the addition and elimination of implicit interactions. We use binary interactiveness , which indicates whether human and object are interactive, to monitor these semantic changes. Hence, the interactiveness of isolated human/object is False, and joint human-object has True interactiveness. From the above, should have the ability to “add” interaction to isolated instances and make the integrated human-object has True interactiveness. On the contrary, can “eliminate” the interaction between coherent human-object and force their interactiveness to be False. At last, to encode the implicit verbs, we represent them in the transformation function space. A pair of decomposition and integration functions are constructed for each verb and forced to operate the appropriate transformations.
We introduce the feature preparation as follows. First, given an image, we use an object detector to obtain the human/object boxes . Then, we adopt a COCO pre-trained ResNet-50 to extract human/object RoI pooling features from the third ResNet Block, where indicates visual appearance. For simplicity, we use tight union box of human and object to represent the coherent HOI pair (union). Notably, coherent HOI carries the interaction semantics and is more than the sum of isolated human and object , i.e., the incoherent ones. With , the union box can be easily obtained. The RoI pooling feature of is thus adopted from the fourth ResNet Block as the appearance representation of coherent HOI (). Note that is twice the size of , for passing through one more ResNet Block. Second, to encode the box location, we generate location features , where indicates box location. We follow the box coordinate normalization method , getting the normalized box . Next, for union box, we concatenate and and feed them to an MLP to get . For human/object box, or is also fed to an MLP to get or . The size of or is half the size of . Third, the location features are concatenated respectively to their corresponding appearance features , getting . The size of and are also half the size of . For convenience, we concatenate and as .
Before transformations, we compress these features to reduce the computational burden via an auto-encoder (AE). This AE is given as input and pre-trained with an input-output reconstruction loss and a verb classification loss (Sec. 3.4). The classification score is denoted as . After pre-training, we use AE to compress and to 1024 sized (coherent) and (isolated) respectively, Finally, we have for integration and decomposition. The ideal transformations are:
where indicates the decomposition and integration functions, indicates the linear operation between isolated human and object features such as element-wise summation or concatenation. In most cases, concatenation performs better. As for the inter-pair transformation, we use
where indicate the features of human/object instances, and are the inter-human/object transformation functions. Because the strict inter-pair transformation like motion transfer is complex and not our main goal, we implement and as simple feature replacement for simplicity. For human instances, we find their substitutional persons with the same HOI according to the pose similarity. As to object instances, we use the objects of the same category and similar sizes as the substitutions. All substitutional candidates come from the same dataset (train set) and are randomly sampled during training. From the experiment (Sec. 4.5), we find that this policy performs well and effectively improves the interaction representation learning.
We propose a concise Integration-Decomposition Network (IDN) as shown in Fig. 3. IDN mainly consists of two parts: the first one is the integration and decomposition transformations (Sec. 3.2) which construct a loop between the union and human/object features; the second one is the inter-pair transformation (Sec. 3.3) that exchanges the human/object instances between pairs with same HOI. In Sec 3.4, we introduce the training objectives derived from the transformation principles. With them, IDN would learn more effective interaction representations and advance HOI detection in Sec. 3.5.
2 Integration and Decomposition
As shown in Fig. 3, IDN constructs a loop consists of two inverse transformations: integration and decomposition implemented with MLPs. That is, we represent the verb/interaction in MLP weight space or transformation function space. For each verb, we adopt a pair of appropriative MLPs as integration and decomposition functions, e.g., and for verb . For integration, when inputting a pair of isolated and , integrates them into outputs for verbs:
where = and is the number of verbs, is the integrated union feature for the - verb. indicates concatenation. Through the integration function set , we get a set of 1024 sized integrated union features . If the original contains the semantics of the - verb, it should be close to and far away from the other integrated union features. Second, the subsequent decomposition is depicted as follows. Given the integrated union feature set , we also use decomposition functions to decompose them respectively:
The decomposition output is also a set of features , where and all have the same size with and . Similarly, if this human-object pair is performing the - interaction, the original input should be close to the and far away from the other .
3 Inter-Pair Transformation
Inter-Pair Transformation (IPT) is proposed to reveal the inherent nature of implicit verb, i.e., the shared information between different pairs with the same HOI. Here, we adopt a simple implementation: instance exchange policy. For humans, we first use pose estimation to obtain poses and then operate alignment and normalization. In detail, the pelvis keypoints of all persons are aligned and all the distances between head and pelvis are scaled to one. Hence, we can find similar persons according to the pose similarity, which is calculated as the sum of Euclidean distances between the corresponding keypoints of two persons. To keep the semantics, similar persons should have at least one same HOI. Selecting similar objects is simpler, we directly choose the objects of the same category. An extra criterion is that we choose objects with similar sizes. We use the area ratio between the object box and the paired human box as the criteria. Finally, similar candidates are selected for each human/object. This whole selection is operated within one dataset. Formally, with instance exchange, Eq. 3 can be rewritten as:
where = and is the number of selected similar candidates, here = . And and should be equally effective before and after instance exchange. During training, we first use Eq. 3 for a certain number of epochs and then replace Eq. 3 with Eq. 5 (Sec. 4.2). When using Eq. 5, we put the original instance and its exchanging candidates together and randomly sample them. Notably, we focus on the transformations between pairs with the same HOI. The transformations between different HOIs which need to manipulate the corresponding human posture, human-object spatial configuration and interactive pattern are beyond the scope of this paper. For IPT, more sophisticate approaches are also possible, e.g., using motion transfer to adjust 2D human posture according to another person with the same HOI but different posture (eating while sitting/standing), recovering 3D HOI and adjusting 3D pose to generate new images/features, using language priors to change the classes of interacted objects or HOI compositions , etc. But these are beyond the scope of our main insight, so we leave these to the future work.
4 Transformation Principles as Objectives
Before training, we first pre-train AE to compress the inputs. We first feed to the encoder and obtain the compressed . Then an MLP takes as input to classify the verbs with Sigmoids (one pair can have multiple HOIs simultaneously) with cross-entropy loss . Meanwhile, is decoded and generates . We construct MSE reconstruction loss between and . The overall loss of AE is = . After pre-training, AE will be fine-tuned together with transformation modules. Next, we detail the objectives derived from transformation principles.
Integration Validity. As aforementioned, we integrate and into the union feature set for all verbs (Eq. 3 or 5). If integration is able to “add” the verb semantics, the corresponding that belongs to the ongoing verb classes should be close to the real . For example, if coherent contains the semantics of verb and , then and should be close to . Meanwhile, should be far away from . Hence, we can construct the distance:
For verb classes, we can get distance set . Considering above principle, if carries the - verb semantics, should be small, and vice versa. Therefore, we can directly use the negative distances as the score of verb classification, i.e. = . Naturally, is then used to generate verb classification loss = , where is cross-entropy loss, = . = indicates this pair has verb and otherwise = . and are chosen following semi-hard mining strategy: = and = , where denotes all the pairs without verb in the current mini-batch, and denotes all the pairs with verb in the current mini-batch.
Decomposition Validity. This validity is proposed to constrain the decomposed (Eq. 4). Similar to Eq. 6, we also construct distances between and as
and obtain . Again should obey the same principle according to ongoing verbs. Thus, we get the second verb score = and verb classification loss .
Interactiveness Validity. Interactiveness depicts whether a person and an object are interactive. Thus, it is False if and only if human-object do not have any interactions. As the “1+1>2” property , the interactiveness of isolated human or object should be False, so does . But after we integrate into , its interactiveness should be True. Meanwhile, the original union should have True interactiveness. We adopt one shared FC-Sigmoid as the binary classifier for , and . The binary label converted from HOI label is zero if and only if a pair does not have any interactions. Notably, we also adopt the interactiveness validity upon decomposed but achieve limited improvement. To keep the model concise, we just adopt the other three effective interactiveness validities hereinafter. Thus, we obtain three binary classification cross entropy losses: . For clarity, we use a unified = .
The overall loss of IDN is = . With the guidance of these principles, IDN can well capture the interaction changes during the transformations. Different from previous methods that aim at encoding the entire HOI representations statically, IDN focuses on dynamically inferring whether an interaction exists within human-object through the integration and decomposition. So IDN can alleviate the learning difficulty of complex and various HOI patterns.
5 Application: HOI Detection
We further apply IDN to HOI detection, which needs to simultaneously locate human-object and classify the ongoing interactions. For locations, we adopt the detected boxes from a COCO pre-trained Faster R-CNN , so does the object class probability . Then, verb scores can be obtained from Eq. 6 and 7. = , = and obtained from AE are then fed to exponential functions or Sigmoids to generate = , = and = . Since the validity losses would pull the features that meet the labels together and push away the others, thus here we directly use three kinds of distances to classify verbs. For example, if contains the - verb, should be small (probability should be large); if not, should be large (probability should be small). The final verb probabilities is acquired via = , here = . For HOI triplets, we get their HOI probabilities using = for all possible compositions according to the benchmark setting.
Experiment
In this section, we first introduce the adopted datasets, metrics (Sec. 4.1) and implementation (Sec. 4.2). Next, we compare IDN with the state-of-the-art on HICO-DET and V-COCO in Sec. 4.3. As HOI detection metrics expect both accurate human/object locations and verb classification, the performance strongly relies on object detection. Hence, we conduct experiments to evaluate IDN with different object detectors. At last, ablation studies are conducted (Sec. 4.5).
We adopt the widely-used HICO-DET and V-COCO . HICO-DET consists of 47,776 images (38,118 for training and 9,658 for testing) and 600 HOI categories (80 COCO objects and 117 verbs). V-COCO contains 10,346 images (2,533 and 2,867 in train and validation sets, 4,946 in test set). Its annotations include 29 verb categories (25 HOIs and 4 body motions) and same 80 objects with HICO-DET. For HICO-DET, we use mAP following : true positive needs to contain accurate human and object locations (box IoU with reference to GT box is larger than 0.5) and accurate verb classification. The role means average precision is used for V-COCO.
2 Implementation Details
The encoder of the adopted AE compresses the input feature dimension from 4608 to 4096, then to 1024. The decoder is structured symmetrical to the encoder. For HICO-DET , AE is pre-trained for 4 epochs using SGD with a learning rate of 0.1, momentum of 0.9, while each batch contains 45 positive and 360 negative pairs. The whole IDN (AE and transformation modules) is first trained without inter-pair transformation (IPT) for 20 epochs using SGD with a learning rate of 2e-2, momentum of 0.9. Then we finetune IDN with IPT for 30 epochs using SGD, with a learning rate of 1e-3, momentum of 0.9. Each batch for the whole IDN contains 15 positive and 120 negative pairs. For V-COCO , AE is first pre-trained for 60 epochs. The whole IDN is trained without IPT for 45 epochs using SGD, then fine-tuned with IPT for 20 epochs. The other training parameters are the same as those for HICO-DET. In testing, LIS is adopted with = ==. Following , we use NIS in all testings with the default threshold and the interactiveness estimation of . All experiments are conducted on one single NVIDIA Titan Xp GPU.
3 Results
Setting. We compare IDN with state-of-the-art on two benchmarks in Tab. 1 and Tab. 3. For HICO-DET, we follow the settings in : Full (600 HOIs), Rare (138 HOIs), Non-Rare (462 HOIs) in Default and Known Object sets. For V-COCO, we evaluate (24 actions with roles) on Scenario 1 (S1) and Scenario 2 (S2). To purely illustrate the HOI recognition ability without the influence of object detection, we conduct evaluations with three kinds of detectors: COCO pre-trained (COCO), pre-trained on COCO and then finetuned on HICO-DET train set (HICO-DET), GT boxes (GT) in Tab. 1.
Comparison. With and , IDN outperforms previous methods significantly and achieves 23.36 mAP on the Default Full set of HICO-DET with COCO detector. Moreover, IDN is the first to achieve more than 20 mAP on all three Default sets without additional information used. Moreover, the improvement on the Rare set proves that the dynamically learned interaction representation can greatly alleviate the data deficiency of rare HOIs. With the HICO-DET finetuned detector, IDN also shows great improvements and achieves more than 26 mAP and further proves the affect from detections (23.36 to 26.29 mAP). Given GT boxes, the gaps among the other three methods are marginal. But IDN achieves more than 9 mAP improvement on HOI recognition solely. All these greatly verify the efficacy of our integration and decomposition. On V-COCO , IDN achieves 53.3 mAP on S1 and 60.3 mAP on S2, both significantly outperforming previous methods. Moreover, we also apply our IDN to existing HOI methods since its flexibility as a plug-in. In detail, we apply integration and decomposition to iCAN as a proxy task to enhance its feature learning. The performance improves from 14.84 mAP to 18.98 mAP (HICO-DET Full).
Efficiency and Scalability. In IDN, each verb is represented by a pair of MLPs ( and ). To ensure the efficiency, we carefully designed the data flow to make IDN is able to run on a single GPU. All transformations are operated in parallel and the inference speed is 10.04 FPS (iCAN : 4.90 FPS, TIN : 1.95 FPS, PPDM : 14.08 FPS, PMFNet : 3.95 FPS). We also considered an implementation which utilizes a single MLP for all verbs for scalability, i.e., conditioned MLP functions and , where is the verb indicator (one-hot/Word2Vec /Glove ). For new verbs, we just change the verb indicator instead of increasing MLPs. It works similar to zero-shot learning like TAFE-Net and Nan , but performs worse (20.86 mAP, HICO-DET Full) than the reported version (23.36 mAP).
4 Visualization
To verify the effectiveness of transformations, we use t-SNE to visualize for different in Fig. 5. We can find integrated obviously closer to the real union , while the simple linear combination cannot represent the interaction information. We also analyze the IPT. In detail, we randomly select a pair with verb and denote its features as . Assume there are other pairs with verb , whose features are . Then we calculate = and = . Here, is the mean distance from to . is the mean distance from and to (). If IPT can effectively transform one pair to another by exchanging the human/object, there should be . We compare and of 20 different verbs in Fig. 5. As shown, in most cases is much larger than , indicating the effectiveness of IPT.
5 Ablation Study
We conduct ablation studies on HICO-DET with COCO detector. The results are shown in Tab. 3. (1) Modules: The performance of each module is evaluated. , and AE achieve 21.26, 21.05, 17.27 mAP respectively and show complementary property. (2) Objectives: During training, we drop one of the three validity objectives respectively. Without anyone of them, IDN shows obvious degradation, especially integration validity. (3) Inter-Pair Transformation (IPT): IDN without IPT achieves 22.63 mAP, showing the importance of instance exchange policy. (4) AE: AE is pre-trained with: reconstruction loss and verb classification loss . The removal of hurts the performance more severely, especially on the Rare set, while also plays an important auxiliary role in boosting the performance. (5) Transformation Order: In practice, we construct a loop ( to to ) to train IDN with the consistency. Using instead of , i.e., to and to , performs worse (21.77 mAP).
Conclusion
In this paper, we propose a novel HOI learning paradigm named HOI Analysis, which is inspired by Harmonic Analysis. And an Integration-Decomposition Network (IDN) is introduced to implement it. With the integration and decomposition between the coherent HOI and isolated human and object, IDN can effectively learn the interaction representation in transformation function space and outperform the state-of-the-art on HOI detection with significant improvements.
Broader Impact
In this work, we propose a novel paradigm for Human-Object Interaction detection, which would promote human activity understanding. Our work could be useful for vision applications, such as the health care system in an intelligent hospital. Current activity understanding systems are usually computationally expensive and require high computational resources, and could cost many financial and environmental resources. Considering this, we will release our code and trained models to the community, as part of efforts to alleviate the repeated training of future works.
Acknowledgments and Disclosure of Funding
This work is supported in part by the National Key R&D Program of China, No. 2017YFA0700800, National Natural Science Foundation of China under Grants 61772332 and Shanghai Qi Zhi Institute, SHEITC (2018-RGZN-02046).
References
Appendix A Visualized Results
We visualize some HOI detection results of our IDN on HICO-DET in Fig. 6. As shown, IDN is able to decompose and integrate various HOIs in diverse scenes and accurately detect them.
Appendix B Result Analysis
We illustrate the detailed comparison between our method, Peyre and DJ-RN on Rare set on HICO-DET in Fig. 7. We can find that our IDN outperforms Peyre and DJ-RN on various rare HOIs. The effectiveness of our IDN on Rare set proves that the dynamically learned interaction representation can greatly alleviate the data deficiency of the rare HOIs.
Appendix C Code
We provide our source code in https://github.com/DirtyHarryLYL/HAKE-Action-Torch/tree/IDN-(Integrating-Decomposing-Network) under our project HAKE-Action-Torch (https://github.com/DirtyHarryLYL/HAKE-Action-Torch).