PaStaNet: Toward Human Activity Knowledge Engine
Yong-Lu Li, Liang Xu, Xinpeng Liu, Xijie Huang, Yue Xu, Shiyi Wang, Hao-Shu Fang, Ze Ma, Mingyang Chen, Cewu Lu
Introduction
Understanding activity from images is crucial for building an intelligent system. Facilitated by deep learning, great advancements have been made in this field. Recent works mainly address this high-level cognition task in one-stage, i.e. from pixels to activity concept directly based on instance-level semantics (Fig. 1(a)). This strategy faces performance bottleneck on large-scale benchmarks . Understanding activities is difficult for reasons, e.g. long-tail data distribution, complex visual patterns, etc. Moreover, action understanding expects a knowledge engine that can generally support activity related tasks. Thus, for data from another domain and unseen activities, much smaller effort is required for knowledge transfer and adaptation. Additionally, for most cases, we find that only a few key human parts are relevant to the existing actions, the other parts usually carry very few useful clues.
Consider the example in Fig. 1, we argue that perception in human part-level semantics is a promising path but previously ignored. Our core idea is that human instance actions are composed of fine-grained atomic body part states. This lies in strong relationships with reductionism . Moreover, the part-level path can help us to pick up discriminative parts and disregard irrelevant ones. Therefore, encoding knowledge from human parts is a crucial step toward human activity knowledge engine. The generic object part states reveal that the semantic state of an object part is limited. For example, after exhaustively checking on manually labeled body part state samples, we find that there are only about 12 states for “head” in daily life activities, such as “listen to”, “eat”, “talk to”, “inspect”, etc. Therefore, in this paper, we exhaustively collect and annotate the possible semantic meanings of human parts in activities to build a large-scale human part knowledge base PaStaNet (PaSta is the abbreviation of Body Part State). Now PaStaNet includes 118 K+ images, 285 K+ persons, 250 K+ interacted objects, 724 K+ activities and 7 M+ human part states. Extensive analysis verifies that PaStaNet can cover most of the the part-level knowledge in general. Using learned PaSta knowledge in transfer learning, we can achieve 3.2, 4.2 and 3.2 improvements on V-COCO , images-based AVA and HICO-DET (Sec.5.4).
Given PaStaNet, we propose two powerful tools to promote the image-based activity understanding: 1) Activity2Vec: With PaStaNet, we convert a human instance into a vector consisting of PaSta representations. Activity2Vec extracts part-level semantic representation via PaSta recognition and combines its language representation. Since PaSta encodes common knowledge of activities, Activity2Vec works as a general feature extractor for both seen and unseen activities. 2) PaSta-R: A Part State Based Reasoning method (PaSta-R) is further presented. We construct a Hierarchical Activity Graph consisting of human instance and part semantic representations, and infer the activities by combining both instance and part level sub-graph states.
The advantages of our method are two-fold: 1) Reusability and Transferability: PaSta are basic components of actions, their relationship can be in analogy with the amino acid and protein, letter and word, etc. Hence, PaSta are reusable, e.g., is shared by various actions like “hold horse” and “eat apple”. Therefore, we get the capacity to describe and differentiate plenty of activities with a much smaller set of PaSta, i.e. one-time labeling and transferability. For few-shot learning, reusability can greatly alleviate its learning difficulty. Thus our approach shows significant improvements, e.g. we boost 13.9 mAP on one-shot sets of HICO . 2) Interpretability: we obtain not only more powerful activity representations, but also better interpretation. When the model predicts what a person is doing, we can easily know the reasons: what the body parts are doing.
In conclusion, we believe PaStaNet will function as a human activity knowledge engine. Our main contributions are: 1) We construct PaStaNet, the first large-scale activity knowledge base with fine-grained PaSta annotations. 2) We propose a novel method to extract part-level activity representation named Activity2Vec and a PaSta-based reasoning method. 3) In supervised and transfer learning, our method achieves significant improvements on large-scale activity benchmarks, e.g. 6.4 (16%), 5.6 (33%) mAP improvements on HICO and HICO-DET respectively.
Related Works
Activity Understanding. Benefited by deep learning and large-scale datasets, image-based or video-based activity understanding has achieved huge improvements recently. Human activities have a hierarchical structure and include diverse verbs, so it is hard to define an explicit organization for their categories. Existing datasets often have a large difference in definition, thus transferring knowledge from one dataset to another is ineffective. Meanwhile, plenty of works have been proposed to address the activity understanding . There are holistic body-level approaches , body part-based methods , and skeleton-based methods , etc. But compared with other tasks such as object detection or pose estimation , its performance is still limited.
Human-Object Interaction. Human-Object Interaction (HOI) occupies the most of daily human activities. In terms of tasks, some works focus on image-based HOI recognition . Furthermore, instance-based HOI detection needs to detect accurate positions of the humans and objects and classify interaction simultaneously. In terms of the information utilization, some works utilized holistic human body and pose , and global context is also proved to be effective . According to the learning paradigm, earlier works were often based on hand-crafted features . Benefited from large scale HOI datasets, recent approaches started to use deep neural networks to extract features and achieved great improvements.
Body Part based Methods. Besides the instance pattern, some approaches studied to utilize part pattern . Gkioxari et al. detects both the instance and parts and input them all into a classifier. Fang et al. defines part pairs and encodes pair features to improve HOI recognition. Yao et al. builds a graphical model and embed parts appearance as nodes, and use them with object feature and pose to predict the HOIs. Previous work mainly utilized the part appearance and location, but few studies tried to divide the instance actions into discrete part-level semantic tokens, and refer them as the basic components of activity concepts. In comparison, we aim at building human part semantics as reusable and transferable knowledge.
Part States. Part state is proposed in . By tokenizing the semantic space as a discrete set of part states, constructs a sort of basic descriptors based on segmentation . To exploit this cue, we divide the human body into natural parts and utilize their states as discretized part semantics to represent activities. In this paper, we focus on the part states of humans instead of daily objects.
Constructing PaStaNet
In this section, we introduce the construction of PaStaNet. PaStaNet seeks to explore the common knowledge of human PaSta as atomic elements to infer activities.
PaSta Definition. We decompose human body into ten parts, namely head, two upper arms, two hands, hip, two thighs, two feet. Part states (PaSta) will be assigned to these parts. Each PaSta represents a description of the target part. For example, the PaSta of “hand” can be “hold something” or “push something”, the PaSta of “head” can be “watch something”, “eat something”. After exhaustively reviewing collected 200K+ images, we found the descriptions of any human parts can be concluded into limited categories. That is, the PaSta category number of each part is limited. Especially, a person may have more than one action simultaneously, thus each part can have multiple PaSta, too.
Data Collection. For generality, we collect human-centric activity images by crowdsourcing (30K images paired with rough activity label) as well as existing well-designed datasets (185K images), which are structured around a rich semantic ontology, diversity, and variability of activities. All their annotated persons and objects are extracted for our construction. Finally, we collect more than 200K images of diverse activity categories.
Activity Labeling. Activity categories of PaStaNet are chosen according to the most common human daily activities, interactions with object and person. Referred to the hierarchical activity structure , common activities in existing datasets and crowdsourcing labels, we select 156 activities including human-object interactions and body motions from 118K images. According to them, we first clean and reorganize the annotated human and objects from existing datasets and crowdsourcing. Then, we annotate the active persons and the interacted objects in the rest of the images. Thus, PaStaNet includes all active human and object bounding boxes of 156 activities.
Body Part Box. To locate the human parts, we use pose estimation to obtain the joints of all annotated persons. Then we generate ten body part boxes following . Estimation errors are addressed manually to ensure high-quality annotation. Each part box is centered with a joint, and the box size is pre-defined by scaling the distance between the joints of the neck and pelvis. A joint with confidence higher than 0.7 will be seen as visible. When not all joints can be detected, we use body knowledge-based rules. That is, if the neck or pelvis is invisible, we configure the part boxes according to other visible joint groups (head, main body, arms, legs), e.g., if only the upper body is visible, we set the size of the hand box to twice the pupil distance.
PaSta Annotation. We carry out the annotation by crowdsourcing and receive 224,159 annotation uploads. The process is as follows: 1) First, we choose the PaSta categories considering the generalization. Based on the verbs of 156 activities, we choose 200 verbs from WordNet as the PaSta candidates, e.g., “hold”, “pick” for hands, “eat”, “talk to” for head, etc. If a part does not have any active states, we depict it as “no_action”. 2) Second, to find the most common PaSta that can work as the transferable activity knowledge, we invite 150 annotators from different backgrounds to annotate 10K images of 156 activities with PaSta candidates (Fig. 2). For example, given an activity “ride bicycle”, they may describe it as , , , etc. 3) Based on their annotations, we use the Normalized Point-wise Mutual Information (NPMI) to calculate the co-occurrence between activities and PaSta candidates. Finally, we choose 76 candidates with the highest NPMI values as the final PaSta. 4) Using the annotations of 10K images as seeds, we automatically generate the initial PaSta labels for all of the rest images. Thus the other 210 annotators only need to revise the annotations. 5) Considering that a person may have multiple actions, for each action, we annotate its corresponding ten PaSta respectively. Then we combine all sets of PaSta from all actions. Thus, a part can also have multiple states, e.g., in “eating while talking”, the head has PaSta , and simultaneously. 6) To ensure quality, each image will be annotated twice and checked by automatic procedures and supervisors. We cluster all labels and discard the outliers to obtain robust agreements.
Activity Parsing Tree. To illustrate the relationships between PaSta and activities, we use their statistical correlations to construct a graph (Fig. 2): activities are root nodes, PaSta are son nodes and edges are co-occurrence.
Finally, PaStaNet includes 118K+ images, 285K+ persons, 250K+ interacted objects, 724K+ instance activities and 7M+ PaSta. Referred to well-designed datasets and WordNet , PaSta can cover most part situations with good generalization. To verify that PaSta have encoded common part-level activity knowledge and can adapt to various activities, we adopt two experiments:
Coverage Experiment. To verify that PaSta can cover most of the activities, we collect other 50K images out of PaStaNet. Those images contain diverse activities and many of them are unseen in PaStaNet. Another 100 volunteers from different backgrounds are invited to find human parts that can not be well described by our PaSta set. We found that only cases cannot find appropriate descriptions. This verifies that PaStaNet is general to activities.
Recognition Experiment. First, we find that PaSta can be well learned. A shallow model trained with a part of PaStaNet can easily achieve about 55 mAP on PaSta recognition. Meanwhile, a deeper model can only achieve about 40 mAP on activity recognition with the same data and metric (Sec. 5.2). Second, we argue that PaSta can be well transferred. To verify this, we conduct transfer learning experiments (Sec. 5.4), i.e. first trains a model to learn the knowledge from PaStaNet, then use it to infer the activities of unseen datasets, even unseen activities. Results show that PaSta can be well transferred and boost the performance (4.2 mAP on image-based AVA). Thus it can be considered as the general part-level activity knowledge.
Activity Representation by PaStaNet
In this section, we discuss the activity representation by PaStaNet.
Conventional Paradigm Given an image , conventional methods mainly use a direct mapping (Fig. 1(a)):
to infer the action score with instance-level semantic representations . is the human box and are the interacted object boxes of this person.
PaStaNet Paradigm. We propose a novel paradigm to utilize general part knowledge: 1) PaSta recognition and feature extraction for a person and an interacted object :
where are part boxes generated from the pose estimation automatically following (head, upper arms, hands, hip, thighs, feet). indicates the Activity2Vec, which extracts ten PaSta representations . 2) PaSta-based Reasoning (PaSta-R), i.e., from PaSta to activity semantics:
where indicates the PaSta-R, is the object feature. is the action score of the part-level path. If the person does not interact with any objects, we use the ROI pooling feature of the whole image as . For multiple object case, i.e., a person interacts with several objects, we process each human-object pair respectively and generate its Activity2Vec embedding.
Following, we introduce the PaSta recognition in Sec. 4.1. Then, we discuss how to map human instance to semantic vector via Activity2Vec in Sec. 4.2. We believe it can be a general activity representation extractor. In Sec. 4.3, a hierarchical activity graph is proposed to largely advance activity related tasks by leveraging PaStaNet.
With the object and body part boxes , we operate the PaSta recognition as shown in Fig. 3. In detail, a COCO pre-trained Faster R-CNN is used as the feature extractor. For each part, we concatenate the part feature from and object features from as inputs. For body only motion, we input the whole image feature as . All features will be first input to a Part Relevance Predictor. Part relevance represents how important a body part is to the action. For example, feet usually have weak correlations with “drink with cup”. And in “eat apple”, only hands and head are essential. These relevance/attention labels can be converted from PaSta labels directly, i.e. the attention label will be one, unless its PaSta label is “no_action”, which means this part contributes nothing to the action inference. With the part attention labels as supervision, we use part relevance predictor consisting of FC layers and Sigmoids to infer the attentions of each part. Formally, for a person and an interacted object:
where is the part attention predictor. We compute cross-entropy loss for each part and multiply with its scalar attention, i.e. .
Second, we operate the PaSta recognition. For each part, we concatenate the re-weighted with , and input them into a max pooling layer and two subsequent 512 sized FC layers, thus obtain the PaSta score for the -th part. Because a part can have multiple states, e.g. head performs “eat” and “watch” simultaneously. Hence we use multiple Sigmoids to do this multi-label classification. With PaSta labels, we construct cross-entropy loss . The total loss of PaSta recognition is:
2 Activity2Vec
In Sec. 3, we define the PaSta according to the most common activities. That is, choosing the part-level verbs which are most often used to compose and describe the activities by a large number of annotators. Therefore PaSta can be seen as the fundamental components of instance activities. Meanwhile, PaSta recognition can be well learned. Thus, we can operate PaSta recognition on PaStaNet to learn the powerful PaSta representations, which have good transferability. They can be used to reason out the instance actions in both supervised and transfer learning. Under such circumstance, PaStaNet works like the ImageNet . And PaStaNet pre-trained Activity2Vec functions as a knowledge engine and transfers the knowledge to other tasks.
Language PaSta feature. Our goal is to bridge the gap between PaSta and activity semantics. Language priors are useful in visual concept understanding . Thus the combination of visual and language knowledge is a good choice for establishing this mapping. To further enhance the representation ability, we utilize the uncased BERT-Base pre-trained model as the language representation extractor. Bert is a language understanding model that considers the context of words and uses a deep bidirectional transformer to extract contextual representations. It is trained with large-scale corpus databases such as Wikipedia, hence the generated embedding contains helpful implicit semantic knowledge about the activity and PaSta. For example, the description of the entry “basketball” in Wikipedia: “drag one’s foot without dribbling the ball, to carry it, or to hold the ball with both hands…placing his hand on the bottom of the ball;..known as carrying the ball”.
3 PaSta-based Activity Reasoning
With part-level , we construct a Hierarchical Activity Graph (HAG) to model the activities. Then we can extract the graph state to reason out the activities.
Hierarchical Activity Graph. Hierarchical activity graph is depicted in Fig. 4. For human-object interactions, . For body only motions, . In instance level, a person is a node with instance representation from previous instance-level methods as a node feature. Object node and has as node feature. In part level, each body part can be seen as a node with PaSta representation as node feature. Edge between body parts and object is , and edge within parts is .
Our goal is to parse HAG and reason out the graph state, i.e. activities. In part-level, we use PaSta-based Activity Reasoning (PaSta-R) to infer the activities. That is, with the PaSta representation from Activity2Vec, we use (Eq. 3) to infer the activity scores . For body motion only activities e.g. “dance”, Eq. 3 is , is the feature of image. We adopt different implementations of .
Linear Combination. The simplest implementation is to directly combine the part node features linearly. We concatenate the output of Activity2Vec with and input them to a FC layer with Sigmoids.
MLP. We can also operate nonlinear transformation on Activity2Vec output. We use two 1024 sized FC layers and an action category sized FC with Sigmoids.
Graph Convolution Network. With part-level graph, we use Graph Convolution Network (GCN) to extract the global graph feature and use an MLP subsequently.
Sequential Model. When watching an image in this way: watch body part and object patches with language description one by one, human can easily guess the actions. Inspired by this, we adopt an LSTM to take the part node features gradually, and use the output of the last time step to classify actions. We adopt two input orders: random and fixed (from head to foot), and fixed order is better.
Tree-Structured Passing. Human body has a natural hierarchy. Thus we use a tree-structured graph passing. Specifically, we first combine the hand and upper arm nodes into an “arm” node, its feature is obtained by concatenating the features of three son nodes and passed a 512 sized FC layer. Similarly, we combine the foot and thigh nodes to an “leg” node. Head, arms, legs and feet nodes together form the second level. The third level contains the “upper body“ (head, arms) and “lower-body” (hip, legs). Finally, the body node is generated. We input it and the object node into an MLP.
The instance-level graph inference can be operated by instance-based methods using Eq. 1: . To get the final result upon the whole graph, we can use either early or late fusion. In early fusion, we concatenate with and input them to PaSta-R. In late fusion, we fuse the predictions of two levels, i.e. . In our test, late fusion outperforms early fusion in most cases. If not specified, we use late fusion in Sec. 5. We use and to indicate the cross-entropy losses of two levels. The total loss is:
Experiments
We design a simplified experiment to give an intuition (Fig. 5). We randomly sample MNIST digits from 0 to 9 () and generate images consists of 3 to 5 digits. Each image is given a label to indicate the sum of the two largest numbers within it (0 to 18). We assume that “PaSta-Activity” resembles the “Digits-Sum”. Body parts can be seen as digits, thus human is the union box of all digits. To imitate the complex body movements, digits are randomly distributed, and Gaussian noise is added to the images. For comparison, we adopt two simple networks. For instance-level model, we input the ROI pooling feature of the digit union box into an MLP. For hierarchical model, we operate single-digit recognition, then concatenate the union box and digit features and input them to an MLP (early fusion), or use late fusion to combine scores of two levels. Early fusion achieves 43.7 accuracy and shows significant superiority over instance-level method (10.0). And late fusion achieves a preferable accuracy of 44.2. Moreover, the part-level method only without fusion also obtains an accuracy of 41.4. This supports our assumption about the effectiveness of part-level representation.
2 Image-based Activity Recognition
Usually, Human-Object Interactions (HOIs) often take up most of the activities, e.g., more than 70% activities in large-scale datasets are HOIs. To evaluate PaStaNet, we perform image-based HOI recognition on HICO . HICO has 38,116 and 9,658 images in train and test sets and 600 HOIs composed of 117 verbs and 80 COCO objects . Each image has an image-level label which is the aggregation over all HOIs in an image and does not contain any instance boxes.
Modes. We first pre-train Activity2Vec with PaSta labels, then fine-tune Activity2Vec and PaSta-R together on HICO train set. In pre-training and finetuning, we exclude the HICO testing data in PaStaNet to avoid data pollution. We
adopt different data mode to pre-train Activity2Vec: 1) “PaStaNet*” mode (38K images): we use the images in HICO train set and their PaSta labels. The only additional supervision here is the PaSta annotations compared to conventional way. 2) “GT-PaStaNet*” mode (38K images): the data used is same with “PaStaNet*”. To verify the upper bound of our method, we use the ground truth PaSta (binary labels) as the predicted PaSta probabilities in Activity2Vec. This means we can recognize PaSta perfectly and reason out the activities from the best starting point. 3) “PaStaNet” mode (118K images): we use all PaStaNet images with PaSta labels except the HICO testing data.
Settings. We use image-level PaSta labels to train Activity2Vec. Each image-level PaSta label is the aggregation over all existing PaSta of all active persons in an image. For PaSta recognition, i.e., we compute the mAP for the PaSta categories of each part, and compute the mean mAP of all parts. To be fair, we use the person, body part and object boxes from and VGG-16 as the backbone. The batch size is 16 and the initial learning rate is 1e-5. We use SGD optimizer with momentum (0.9) and cosine decay restarts (the first decay step is 5000). The pre-training costs 80K iterations and fine-tuning costs 20K iterations. Image-level PaSta and HOI predictions are all generated via Multiple Instance Learning (MIL) of 3 persons and 4 objects. We choose previous methods as the instance-level path in the hierarchical model, and uses late fusion. Particularly, uses part-pair appearance and location but not part-level semantics, thus we still consider it as a baseline to get a more abundant comparison.
Results. Results are reported in Tab. 1. PaStaNet* mode methods all outperform the instance-level method. The part-level method solely achieves 44.5 mAP and shows good complementarity to the instance-level. Their fusion can boost the performance to 45.9 mAP (6 mAP improvement). And the gap between and is largely narrowed from 3.8 to 0.9 mAP. Activity2Vec achieves 55.9 mAP on PaSta recognition in PaStaNet* mode: 46.3 (head), 66.8 (arms), 32.0 (hands), 68.6 (hip), 56.2 (thighs), 65.8 (feet). This verifies that PaSta can be better learned than activities, thus they can be learned ahead as the basis for reasoning. In GT-PaStaNet* mode, hierarchical paradigm achieves 65.6 mAP. This is a powerful proof of the effectiveness of PaSta knowledge. Thus what remains to do is to improve the PaSta recognition and further promote the activity task performance. Moreover, in PaStaNet mode, we achieve relative 16% improvement. On few-shot sets, our best result significantly improves 13.9 mAP, which strongly proves the reusability and transferability of PaSta.
3 Instance-based Activity Detection
We further conduct instance-based activity detection on HICO-DET , which needs to locate human and object and classify the actions simultaneously. HICO-DET is a benchmark built on HICO and add human and object bounding boxes. We choose several state-of-the-arts to compare and cooperate.
Settings. We use instance-level PaSta labels, i.e. each annotated person with the corresponding PaSta labels, to train Acitivty2Vec, and fine-tune Activity2Vec and PaSta-R together on HICO-DET. All testing data are excluded from pre-training and fine-tining. We follow the mAP metric of , i.e. true positive contains accurate human and object boxes ( with reference to ground truth) and accurate action prediction. The metric for PaSta detection is similar, i.e., estimated part box and PaSta action prediction all have to be accurate. The mAP of each part and the mean mAP are calculated. For a fair comparison, we use the object detection from and ResNet-50 as backbone. We use SGD with momentum (0.9) and cosine decay restart (the first decay step is 80K). The pre-training and fine-tuning take 1M and 2M iterations respectively. The learning rate is 1e-3 and the ratio of positive and negative samples is 1:4. A late fusion strategy is adopted. Three modes in Sec. 5.2 and different PaSta-R are also evaluated.
Results. Results are shown in Tab. 2. All PaStaNet* mode methods significantly outperform the instance-level methods, which strongly prove the improvement from the learned PaSta information. In PaStaNet* mode, the PaSta detection performance are 30.2 mAP: 25.8 (head), 44.2 (arms), 17.5 (hands), 41.8 (hip), 22.2 (thighs), 29.9 (feet). This again verifies that PaSta can be well learned. And GT-PaStaNet* (upper bound) and PaStaNet (more PaSta labels) modes both greatly boosts the performance. On Rare sets, our method obtains 7.7 mAP improvement.
4 Transfer Learning with Activity2Vec
To verify the transferability of PaStaNet, we design transfer learning experiments on large-scale benchmarks: V-COCO , HICO-DET and AVA . We first use PaStaNet to pre-train Activity2Vec and PaSta-R with 156 activities and PaSta labels. Then we change the last FC in PaSta-R to fit the activity categories of the target benchmark. Finally, we freeze Activity2Vec and fine-tune PaSta-R on the train set of the target dataset. Here, PaStaNet works like the ImageNet and Activity2Vec is used as a pre-trained knowledge engine to promote other tasks.
V-COCO. V-COCO contains 10,346 images and instance boxes. It has 29 action categories, COCO 80 objects . For a fair comparison, we exclude the images of V-COCO and corresponding PaSta labels in PaStaNet, and use remaining data (109K images) for pre-training. We use SGD with 0.9 momenta and cosine decay restarts (the first decay is 80K). The pre-training costs 300K iterations with the learning rate as 1e-3. The fine-tuning costs 80K iterations with the learning rate as 7e-4. We select state-of-the-arts as baselines and adopt the metric (requires accurate human and object boxes and action prediction). Late fusion strategy is adopted. With the domain gap, PaStaNet still improves the performance by 3.2 mAP (Tab. 3.).
Image-based AVA. AVA contains 430 video clips with spatio-temporal labels. It includes 80 atomic actions consists of body motions and HOIs. We utilize all PaStaNet data (118K images) for pre-training. Considering that PaStaNet is built upon still images, we use the frames per second as still images for image-based instance activity detection. We adopt ResNet-50 as backbone and SGD with momentum of 0.9. The initial learning rate is 1e-2 and the first decay of cosine decay restarts is 350K. For a fair comparison, we use the human box from . The pre-training costs 1.1M iterations and fine-tuning costs 710K iterations. We adopt the metric from , i.e. mAP of the top 60 most common action classes, using IoU threshold of 0.5 between detected human box and the ground truth and accurate action prediction. For comparison, we adopt a image-based baseline: Faster R-CNN detector with
ResNet-101 provided by the AVA website . Recent works mainly use a spatial-temporal model such as I3D . Although unfair, we still employ two video-based baselines as instance-level models to cooperate with the part-level method via late fusion. Results are listed in Tab. 4. Both image and video based methods cooperated with PaStaNet achieve impressive improvements, even our model is trained without temporal information. Considering the huge domain gap (films) and unseen activities, this result strongly proves its great generalization ability.
HICO-DET. We exclude the images of HICO-DET and the corresponding PaSta labels, and use left data (71K images) for pre-training. The test setting in same with Sec. 5.3. The pre-training and fine-tuning cost 300K and 1.3M iterations. PaStaNet shows good transferability and achieve 3.25 mAP improvement on Default Full set (20.28 mAP).
5 Ablation Study
We design ablation studies on HICO-DET with TIN +PaSta*-Linear (22.12 mAP). 1) w/o Part Attention degrades the performance with 0.21 mAP. 2) Language Feature: We replace the PaSta Bert feature in Activity2Vec with: Gaussian noise, Word2Vec and GloVe . The results are all worse (20.80, 21.95, 22.01 mAP). If we change the PaSta triplet into a sentence and convert it to Bert vector, this vector performs sightly better (22.26 mAP). This is probably because the sentence carries more contextual information.
Conclusion
In this paper, to make a step toward human activity knowledge engine, we construct PaStaNet to provide novel body part-level activity representation (PaSta). Meanwhile, a knowledge transformer Activity2Vec and a part-based reasoning method PaSta-R are proposed. PaStaNet brings in interpretability and new possibility for activity understanding. It can effectively bridge the semantic gap between pixels and activities. With PaStaNet, we significantly boost the performance in supervised and transfer learning tasks, especially under few-shot circumstances. In the future, we plan to enrich our PaStaNet with spatio-temporal PaSta.
This work is supported in part by the National Key R&D Program of China, No. 2017YFA0700800, National Natural Science Foundation of China under Grants 61772332 and Shanghai Qi Zhi Institute.
References
Appendix A Dataset Details
In this section, we give a more detailed introduction of our knowledge base PaStaNet, covering characteristics and the annotator backgrounds.
Tab. 5 shows some characteristics of PaStaNet. Now PaStaNet contains about 118K+ images and the corresponding instance-level and human part-level annotations. To give an intuitive presentation of our PaSta, we use t-SNE to visualize the part-level features of some typical PaSta samples in Fig. 6. We use the human body part patches with different colored borders to replace the embedding points in Fig. 6, e.g. the embeddings of “hand holds something” and “foot kicks something”.
A.2 Annotator Backgrounds
There are about 360 annotators have participated in the construction of PastaNet. They have various backgrounds, thus we can ensure annotation diversity and reduce the bias. The specific information is shown in Fig. 7.
Appendix B Selected Activities and PaSta
We list all selected 156 activities and 76 PaSta in Tab. 6 and Tab. 7. PaStaNet contains both person-object/person interactions and body only motions, which cover the vast majority of human daily life activities. And all the annotated interacted objects in interactions all belong to COCO 80 object categories .
PaStaNet can provide abundant activity knowledge for both instance and part levels and help construct a large-scale activity parsing tree, as seen in Fig. 8. Moreover, we can represent the parsing tree as a co-occurrence matrix of the activities and PaSta. A part of the matrix is depicted in Fig. 9.
Appendix C Additional Details of PaSta-R
Fig. 10 gives more details about the different PaSta-R implementations, i.e. directly input the Activity2Vec output to the Linear Combination, MLP or GCN , sequential LSTM-based model and the tree-structured passing model.
Appendix D Additional Details of MNIST-Action
In this section, we provide some details of the MNIST-Action experiment. The instance-based and hierarchical models are shown in Fig. 11. The train and test set sizes are 5,000 and 800. In the instance-based model, we directly use the ROI pooling feature of the digit union box to predict the target (summation of the largest and second largest numbers). We use optimizer RMSProp and the initial learning rate is 0.0001. The batch size is 16 and we train the model 1K epochs. In the hierarchical model, we first operate the single digit recognition and then use single digit features (part features) together with the instance feature to infer the sum (early fusion). If using late fusion, we directly fuse the scores of instance branch and part branch. We also use optimizer RMSProp and the initial learning rate is 0.001. The batch size is 32 and the training also costs 1K epochs. Two models are all implemented with 4 convolution layers with subsequent fully-connect layers.
Results are shown in Fig. 12. We can find that the hierarchical method largely outperforms the instance-based method.
Appendix E Effectiveness on Few-shot Problems
In this section, we show more detailed results on HICO to illustrate the effectiveness of Pasta on few-shot problems. We divide the 600 HOIs into different sets according to their training sample numbers. On HICO , there is an obvious positive correlation between performance and the number of training samples. From Fig. 13 we can find that our hierarchical method outperforms the previous state-of-the-art on all sets, especially on the few-shot sets.
Appendix F Additional Activity Detection Results
We report visualized PaSta and activity predictions of our method in Figure 14. The with the highest scores are visualized in blue, green and red boxes, and their corresponding PaSta descriptions are demonstrated under each image with colors consisted with boxes. The final activity predictions with the highest scores are also represented. We can find that our model is capable to detect various kinds of activities covering interactions with various objects.
Appendix G Data usage
The data usages of pre-training and finetuning are clarified in Tab. 8. We have carefully excluded the testing data in all pre-training and finetuning to avoid data pollution.