Hierarchical Human Parsing with Typed Part-Relation Reasoning

Wenguan Wang, Hailong Zhu, Jifeng Dai, Yanwei Pang, Jianbing Shen, Ling Shao

Introduction

Human parsing involves segmenting human bodies into semantic parts, e.g., head, arm, leg, etc. It has attracted tremendous attention in the literature, as it enables fine-grained human understanding and finds a wide spectrum of human-centric applications, such as human behavior analysis , human-robot interaction , and many others.

Human bodies present a highly structured hierarchy and body parts inherently interact with each other. As shown in Fig.​ 1(b), there are different relations between parts : decompositional and compositional relations (full line:) between constituent and entire parts (e.g., {upper body, lower body} and full body), and dependency relations (dashed line:) between kinematically connected parts (e.g., hand and arm). Thus the central problem in human parsing is how to model such relations. Recently, numerous structured human parsers have been proposed . Their notable successes indeed demonstrate the benefit of exploiting the structure in this problem. However, three major limitations in human structure modeling are still observed. (1) The structural information utilized is typically weak and relation types studied are incomplete. Most efforts directly encode human pose information into the parsing model, causing them to suffer from trivial structural information, not to mention the need of extra pose annotations. In addition, previous structured parsers focus on only one or two of the aforementioned part relations, not all of them. For example, only considers dependency relations, and relies on decompositional relations. (2) Only a single relation model is learnt to reason different kinds of relations, without considering their essential and distinct geometric constraints. Such a relation modeling strategy is over-general and simple; do not seem to characterize well the diverse part relations. (3) According to graph theory, as the human body yields a complex, cyclic topology, an iterative inference is desirable for optimal result approximation. However, current arts are primarily built upon an immediate, feed-forward prediction scheme.

To respond to the above challenges and enable a deeper understanding of human structures, we develop a unified, structured human parser that precisely describes a more complete set of part relations, and efficiently reasons structures with the prism of a message-passing, feed-back inference scheme. To address the first two issues, we start with an in-depth and comprehensive analysis on three essential relations, namely decomposition, composition, and dependency. Three distinct relation networks (, , and in Fig.​ 1(c)) are elaborately designed and imposed to explicitly satisfy the specific, intrinsic relation constraints. Then, we construct our parser as a tree-like, end-to-end trainable graph model, where the nodes represent the human parts, and edges are built upon the relation networks. For the third issue, a modified, relation-typed convolutional message passing procedure ( in Fig.​ 1(c)) is performed over the human hierarchy, enabling our method to obtain better parsing results from a global view. All components, i.e., the part nodes, edge (relation) functions, and message passing modules, are fully differentiable, enabling our whole framework to be end-to-end trainable and, in turn, facilitating learning about parts, relations, and inference algorithms.

More crucially, our structured human parser can be viewed as an essential variant of message passing neural networks (MPNNs) , yet significantly differentiated in two aspects. (1) Most previous MPNNs are edge-type-agnostic, while ours addresses relation-typed structure reasoning with a higher expressive capability. (2) By replacing the Multilayer Perceptron (MLP) based MPNN units with convolutional counterparts, our parser gains a spatial information preserving property, which is desirable for such a pixel-wise prediction task.

We extensively evaluate our approach on five standard human parsing datasets , achieving state-of-the-art performance on all of them (§4.2). In addition, with ablation studies for each essential component in our parser (§4.3), three key insights are found: (1) Exploring different relations reside on human bodies is valuable for human parsing. (2) Distinctly and explicitly modeling different types of relations can better support human structure reasoning. (3) Message passing based feed-back inference is able to reinforce parsing results.

Related Work

Human parsing: Over the past decade, active research has been devoted towards pixel-level human semantic understanding. Early approaches tended to leverage image regions , hand-crafted features , part templates and human keypoints , and typically explored certain heuristics over human body configurations in a CRF , structured model , grammar model , or generative model framework. Recent advance has been driven by the streamlined designs of deep learning architectures. Some pioneering efforts revisit classic template matching strategy , address local and global cues , or use tree-LSTMs to gather structure information . However, due to the use of superpixel or HOG feature , they are fragmentary and time-consuming. Consequent attempts thus follow a more elegant FCN architecture, addressing multi-level cues , feature aggregation , adversarial learning , or cross-domain knowledge . To further explore inherent structures, numerous approaches choose to straightforward encode pose information into the parsers, however, relying on off-the-shelf pose estimators or additional annotations. Some others consider top-down or multi-source semantic information over hierarchical human layouts. Though impressive, they ignore iterative inference and seldom address explicit relation modeling, easily suffering from weak expressive ability and risk of sub-optimal results.

With the general success of these works, we make a further step towards more precisely describing the different relations residing on human bodies, i.e., decomposition, composition, and dependency, and addressing iterative, spatial-information preserving inference over human hierarchy.

Graph neural networks (GNNs):GNNs have a rich history (dating back to ) and became a veritable explosion in research community over the last few years . GNNs effectively learn graph representations in an end-to-end manner, and can generally be divided into two broad classes: Graph Convolutional Networks (GCNs) and Message Passing Graph Networks (MPGNs). The former directly extend classical CNNs to non-Euclidean data. Their simple architecture promotes their popularity, while limits their modeling capability for complex structures . MPGNs parameterize all the nodes, edges, and information fusion steps in graph learning, leading to more complicated yet flexible architectures.

Our structured human parser, which falls in the second category, can be viewed as an early attempt to explore GNNs in the area of human parsing. In contrast to conventional MPGNs, which are mainly MLP-based and edge-type-agnostic, we provide a spatial information preserving and relation-type aware graph learning scheme.

Our Approach

Formally, we represent the human semantic structure as a directed, hierarchical graph G ⁣= ⁣(V,E,Y)\mathcal{G}\!=\!(\mathcal{V},\mathcal{E},\mathcal{Y}). As show in Fig.​ 2(a), the node set V ⁣= ⁣∪l=13 ⁣Vl\mathcal{V}\!=\!\cup_{l=1}^{3}\!\mathcal{V}_{l} represents human parts in three different semantic levels, including the leaf nodes V1\mathcal{V}_{1} (i.e., the most fine-grained parts: head, arm, hand, etc.) which are typically considered in common human parsers, two middle-level nodes V2 ⁣ ⁣=\mathcal{V}_{2\!}\!={upper-body, lower-body} and one root V3 ⁣=\mathcal{V}_{3}\!={full-body}As the classic settings of graph models, there is also a ‘dummy’ node in V\mathcal{V}, used for interpreting the background class. As it does not interact with other semantic human parts (nodes), we omit this node for concept clarity.. The edge set E ⁣ ⁣∈ ⁣ ⁣(V2)\mathcal{E}\!\!\in\!\!\binom{\mathcal{V}}{2} represents the relations between human parts (nodes), i.e., the directed edge e ⁣= ⁣(u,v) ⁣∈ ⁣Ee\!=\!(u,v)\!\in\!\mathcal{E} links node uu to v ⁣ ⁣: ⁣u ⁣ ⁣→ ⁣ ⁣vv_{\!}\!:\!u\!\!\rightarrow\!\!v. Each node vv and each edge (u,v)(u,v) are associated with feature vectors: hv\textbf{{h}}_{v} and hu,v\textbf{{h}}_{u,v}, respectively. yv ⁣ ⁣∈ ⁣Yy_{v\!}\!\in\!\mathcal{Y} indicates the groundtruth segmentation map of part (node) vv and the groundtruth maps Y\mathcal{Y} are also organized in a hierarchical manner: Y ⁣= ⁣∪l=13Yl\mathcal{Y}\!=\!\cup_{l=1}^{3}\mathcal{Y}_{l}.

Our human parser is trained in a graph learning scheme, using the full supervision from existing human parsing datasets. For a test sample, it is able to effectively infer the node and edge representations by reasoning human structures at the levels of individual parts and their relations, and iteratively fusing the information over the human structures.

2 Structured Human Parsing Network

where each node embedding hv\textbf{{h}}_{v} is a (WW​, HH​, cc)-dimensional tenor that encodes full spatial details ( in Fig.​ 2(c)).

where r ⁣ ⁣ ⁣∈ ⁣ ⁣ ⁣{dec,com,dep}r_{\!}\!\!\in_{\!}\!\!\{\text{dec},\text{com},\text{dep}\}. Fr ⁣(⋅)F^{r\!}(\cdot) is an attention-based relation-adaption operation, which is used to enhance the original node embedding hu ⁣\textbf{{h}}_{u\!} by addressing geometric characteristics in relation rr. The attention mechanism is favored here as it allows trainable and flexible feature enhancement and explicitly encodes specific relation constraints. From the view of information diffusion mechanism in the graph theory , if there exists an edge (u,vu,v) that links a starting node uu to a destination vv, this indicates vv should receive incoming information (i.e., hu,v\textbf{{h}}_{u,v}) from uu. Thus, we use Fr ⁣(⋅)F^{r\!}(\cdot) to make hu\textbf{{h}}_{u} better accommodate the target vv. RrR^{r} is edge-type specific, employing the more tractable feature Fr(hu)F^{r}(\textbf{{h}}_{u}) in place of hu\textbf{{h}}_{u}, so more expressive relation feature hu,v\textbf{{h}}_{u,v} for vv can be obtained and further benefit the final parsing results. In this way, we learn more sophisticated and impressive relation patterns within human bodies.

1) Decompositional relation modeling: Decompositional relations (full line: in Fig.​ 2(a)) are represented by those vertical edges starting from parent nodes to corresponding child nodes in the human hierarchy G\mathcal{G}. For example, a parent node full-body can be separated into {upper-body, lower-body}, and upper-body can be decomposed into {head, torso, upper-arm, lower-arm}. Formally, for a node uu, let us denote its child node set as Cu\mathcal{C}_{u}. Our decompositional relation network RdecR^{\text{dec}} aims to learn the rule for ‘breaking down’ uu into its constituent parts Cu\mathcal{C}_{u} (Fig.​ 3):

‘⊙\odot’ indicates the attention-based feature enhancement operation, and attu,v ⁣dec(hu ⁣) ⁣ ⁣∈ ⁣ ⁣W ⁣× ⁣H ⁣\verb"att"_{u,v\!}^{\text{dec}}(\textbf{{h}}_{u\!})_{\!}\!\in_{\!}\!^{W\!\times\!H\!} produces an attention map. ⁣{}_{\!} For ⁣{}_{\!} each ⁣{}_{\!} sub-node ⁣{}_{\!} v ⁣ ⁣∈ ⁣ ⁣Cu ⁣v_{\!}\!\in_{\!}\!\mathcal{C}_{u\!} of uu, attu,v ⁣dec(hu ⁣) ⁣\verb"att"_{u,v\!}^{\text{dec}}(\textbf{{h}}_{u\!})_{\!} is defined as:

where LCE\mathcal{L}_{\text{CE}} represents the standard cross-entropy loss.

2) Compositional relation modeling: In the human hierarchy G\mathcal{G}, compositional relations are represented by vertical, downward edges. To address this type of relations, we design a compositional relation network RcomR^{\text{com}} as (Fig.​ 4):

For each parent node v ⁣∈ ⁣V2∪V3v\!\in\!\mathcal{V}_{2}\cup\mathcal{V}_{3}, with its groundtruth map yv ⁣∈ ⁣{0,1}W ⁣× ⁣Hy_{v}\!\in\!\{0,1\}^{W\!\times\!H}, the compositional attention for all its child nodes Cv\mathcal{C}_{v} is trained by minimizing the following loss:

3) Dependency relation modeling: In G\mathcal{G}, dependency relations are represented as horizontal edges (dashed line: in Fig.​ 2(a)), describing pairwise, kinematic connections between human parts, such as (head, torso), (upper-leg, lower-leg), etc. Two kinematically connected human parts are spatially adjacent, and their dependency relation essentially addresses the context information. For a node uu, with its kinematically connected siblings Ku\mathcal{K}_{u}, a dependency relation network RdepR^{\text{dep}} is designed as (Fig. 5):

For each sibling node v ⁣∈ ⁣Kuv\!\in\!\mathcal{K}_{u} of uu, attu,vdep\verb"att"^{\text{dep}}_{u,v} is defined as:

where YKu ⁣ ⁣ ⁣∈ ⁣W ⁣× ⁣H ⁣× ⁣∣Ku∣\mathcal{Y}_{\mathcal{K}_{u\!\!}}\!\in\!^{W\!\times\!H\!\times\!|\mathcal{K}_{u}|} stands for the groundtruth maps {yv}v∈Ku\{y_{v}\}_{v\in\mathcal{K}_{u}} of all the sibling nodes Ku\mathcal{K}_{u} of uu.

Iterative inference over human hierarchy: Human bodies present a hierarchical structure. According to graph theory, approximate inference algorithms should be used for such a loopy structure G\mathcal{G}. However, previous structured human parsers directly produce the final node representation hv\textbf{{h}}_{v} by either simply accounting for the information from the parent node u ⁣u_{\!} : hv ⁣ ⁣← ⁣ ⁣R(hu,hv)\textbf{{h}}_{v}\!\!\leftarrow\!\!R(\textbf{{h}}_{u},\textbf{{h}}_{v}), where v ⁣∈ ⁣Cuv\!\in\!\mathcal{C}_{u}; or from its neighbors Nv ⁣\mathcal{N}_{v\!} :  ⁣hv ⁣ ⁣← ⁣ ⁣∑u∈Nv ⁣R(hu,hv)\!\textbf{{h}}_{v}\!\!\leftarrow\!\!\sum_{u\in\mathcal{N}_{v}\!}R(\textbf{{h}}_{u},\textbf{{h}}_{v}). They ignore the fact that, in such a structured setting, information is organized in a complex system. Iterative algorithms offer a more favorable solution, i.e., the node representation should be updated iteratively by aggregating the messages from its neighbors; after several iterations, the representation can approximate the optimal results . In graph theory parlance, the iterative algorithm can be achieved by a parametric message passing process, which is defined in terms of a message function MM and node update function UU, and runs TT steps. For each node vv, the message passing process recursively collects information (messages) mv\textbf{{m}}_{v} from the neighbors Nv\mathcal{N}_{v} to enrich the node embedding hv\textbf{{h}}_{v}:

where hv(t)\textbf{{h}}_{v}^{(t)} stands for vv’s state in the tt-th iteration. Recurrent neural networks are typically used to address the iterative nature of the update function UU.

Inspired by previous message passing algorithms, our iterative algorithm is designed as (Fig.​ 2(e)-(f)):

where the initial state hv(0)\textbf{{h}}_{v}^{(0)} is obtained by Eq.​ 1. Here, the message aggregation step (Eq.​ 13) is achieved by per-edge relation function terms, i.e., node vv updates its state hv\textbf{{h}}_{v} by absorbing all the incoming information along different relations. As for the update function U ⁣U_{\!} in Eq.​ 14, we use a convGRU , which replaces the fully-connected units in the original MLP-based GRU with convolution operations, to describe its repeated activation behavior and address the pixel-wise nature of human parsing, simultaneously. Compared to previous parsers, which are typically based on feed-forward architectures, our massage-passing inference essentially provides a feed-back mechanism, encouraging effective reasoning over the cyclic human hierarchy G\mathcal{G}.

Given the hierarchical human parsing results {Y^l(t)}l=13\{\hat{\mathcal{Y}}^{(t)}_{l}\}^{3}_{l=1} and corresponding groundtruths {Yl}l=13\{\mathcal{Y}_{l}\}^{3}_{l=1}, the learning task in the iterative inference can be posed as the minimization of the following loss (Fig.​ 2(h)):

Considering ⁣{}_{\!} Eqs.​ 5,​ 7,​ 11,​ and​ 16, ⁣{}_{\!} the ⁣{}_{\!} overall ⁣{}_{\!} loss ⁣{}_{\!} is ⁣{}_{\!} defined ⁣{}_{\!} as:

where the coefficient α\alpha is empirically set as 0.10.1. We set the total inference time T ⁣= ⁣2T\!=\!2 and study how the performance changes with the number of inference iterations in §4.3.

3 Implementation Details

Iterative inference: In Eq. ​14, the update function UconvGRUU_{\text{convGRU}} is implemented by a convolutional GRU with 3 ⁣× ⁣33\!\times\!3 convolution kernels. The readout function OO in Eq. ​15 applies a 1 ⁣× ⁣11\!\times\!1 convolution operation on the feature-prediction projection. In addition, before sending a node feature hv(t)\textbf{{h}}_{v}^{(t)} into OO, we use a light-weight decoder (built using a principle of upsampling the node feature and merging it with the low-level feature of the backbone network) that outputs the segmentation mask with 1/4 the spatial resolution of the input image.

As seen, all the units of our parser are built on convolution operations, leading to spatial information preservation.

Experiments

Datasets:As the datasets provide different human part labels, we make proper modifications of our human hierarchy. For some labels that do not deliver human structures, such as hat, sun-glasses, we treat them as isolate nodes. Five standard benchmark datasets are used for performance evaluation. LIP contains 50,462 single-person images, which are collected from realistic scenarios and divided into 30,462 images for training, 10,000 for validation and 10,000 for test. The pixel-wise annotations cover 19 human part categories (e.g., face, left-/right-arms, left-/right-legs, etc.). PASCAL-Person-Part includes 3,533 multi-person images with challenging poses and viewpoints. Each image is pixel-wise annotated with six classes (i.e., head, torso, upper-/lower-arms, and upper-/lower-legs). It is split into 1,716 and 1,817 images for training and test. ATR is a challenging human parsing dataset, which has 7,700 single-person images with dense annotations over 17 categories (e.g., face, upper-clothes, left-/right-arms, left-/right-legs, etc.). There are 6,000, 700 and 1,000 images for training, validation, and test, respectively. PPSS is a collection of 3,673 single-pedestrian images from 171 surveillance videos and provides pixel-wise annotations for hair, face, upper-/lower-clothes, arm, and leg. It presents diverse real-word challenges, e.g., pose variations, illumination changes, and occlusions. There are 1,781 and 1,892 images for training and testing, respectively. Fashion Clothing has 4,371 images gathered from Colorful Fashion Parsing , Fashionista , and Clothing Co-Parsing . It has 17 clothing categories (e.g., hair, pants, shoes, upper-clothes, etc.) and the data split follows 3,934 for training and 437 for test.

Training: ResNet101 , pre-trained on ImageNet , is used to initialize our DeepLabV3​ backbone. The remaining layers are randomly initialized. We train our model on the five aforementioned datasets with their respective training samples, separately. Following the common practice , we randomly augment each training sample with a scaling factor in [0.5, 2.0], crop size of 473 ⁣× ⁣473473\!\times\!473, and horizontal flip. For optimization, we use the standard SGD solver, with a momentum of 0.9 and weight_decay of 0.0005. To schedule the learning rate, we employ the polynomial annealing procedure , where the learning rate is multiplied by (1 ⁣− ⁣itertotal_iter)power(1\!-\!\frac{iter}{total\_iter})^{power} with power as 0.90.9.

Testing: For each test sample, we set the long side of the image to 473 pixels and maintain the original aspect ratio. As in , we average the parsing results over five-scale image pyramids of different scales with flipping, i.e., the scaling factor is 0.5 to 1.5 (with intervals of 0.25).

Reproducibility: Our method is implemented on PyTorch and trained on four NVIDIA Tesla V100 GPUs (32GB memory per-card). All the experiments are performed on one NVIDIA TITAN Xp 12GB GPU. To provide full details of our approach, our code will be made publicly available.

Evaluation: For fair comparison, we follow the official evaluation protocols of each dataset. For LIP, following , we report pixel accuracy, mean accuracy and mean Intersection-over-Union (mIoU). For PASCAL-Person-Part and PPSS, following , the performance is evaluated in terms of mIoU. For ATR and Fashion Clothing, as in , we report pixel accuracy, foreground accuracy, average precision, average recall, and average F1-score.

2 Quantitative and Qualitative Results

LIP : ⁣{}_{\!} LIP is a gold standard benchmark for human parsing. Table​ 1 reports the comparison results with 16 state-of-the-arts on LIP val. We first find that general semantic segmentation methods tend to perform worse than human parsers. This indicates the importance of reasoning human structures in this problem. In addition, though recent human parsers gain impressive results, our model still outperforms all the competitors by a large margin. For instance, in terms of pixAcc., mean Acc., and mean IoU, our parser dramatically surpasses the best performing method, CNIF , by 1.02%, 1.78% and 1.51%, respectively. We would also like to mention that our parser does not use additional pose or edge information.

PASCAL-Person-Part : In Table​ 2, we compare our method against 18 recent methods on PASCAL-Person-Part test using IoU score. From the results, we can again see that our approach achieves better performance compared to all other methods; specially, 73.12% vs 70.76% of CNIF and 68.40% of PGN , in terms of mIoU. Such a performance gain is particularly impressive considering that improvement on this dataset is very challenging.

ATR : Table​ 3 presents comparisons with 14 previous methods on ATR test. Our approach sets new state-of-the-arts for all five metrics, outperforming all other methods by a large margin. For example, our parser provides a considerable performance gain in F-1 score, i.e., 1.74% and 5.49% higher than the current top-two performing methods, CNIF and TGPNet , respectively.

Fashion Clothing : The quantitative comparison results with six competitors on Fashion Clothing test are summarized in Table 4. Our model yields an F-1 score of 60.19%, while those for Attention , TGPNet , and CNIF are 48.68%, 51.92%, and 58.12%, respectively. This again demonstrates our superior performance.

PPSS : Table ​5 compares our method against six famous methods on PPSS test set. The evaluation results demonstrate that our human parser achieves 65.3% mIoU, with substantial gains over the second best, CNIF , and third best, LCPC , of 4.8% and 11.8%, respectively.

Runtime ⁣{}_{\!} comparison: ⁣{}_{\!} As ⁣{}_{\!} our ⁣{}_{\!} parser ⁣{}_{\!} does ⁣{}_{\!} not ⁣{}_{\!} require ⁣{}_{\!} extra pre-/post-processing steps (e.g., human pose used in , over-segmentation in , and CRF in ), it achieves a high speed of 12fps (on PASCAL-Person-Part), faster than most of the counterparts, such as Joint (0.1fps), Attention+SSL (2.0fps), MMAN (3.5fps), SS-NAN (2.0fps), and LG-LSTM (3.0fps).

Qualitative results: Some qualitative comparison results on PASCAL-Person-Part test are depicted in Fig. ​6. We can see that our approach outputs more precise parsing results than other competitors , despite the existence of rare pose (2nd row) and occlusion (3rd row). In addition, with its better understanding of human structures, our parser gets more robust results and eliminates the interference from the background (1st row). The last row gives a challenging case, where our parser still correctly recognizes the confusing parts of the person in the middle.

Overall, our human parser attains strong performance across all the five datasets. We believe this is due to our typed relation modeling and iterative algorithm, which enable more trackable part features and better approximations.

3 Diagnostic Experiments

To demonstrate how each component in our parser contributes to the performance, a series of ablation experiments are conducted on PASCAL-Person-Part test.

Type-specific relation modeling: We first investigate the necessity of comprehensively exploring different relations, and discuss the effective of our type-specific relation modeling strategy. Concretely, we studied six variant models, as ⁣{}_{\!} listed ⁣{}_{\!} in ⁣{}_{\!} Table ​6: ⁣{}_{\!} (1) ⁣{}_{\!} ‘Baseline’ ⁣{}_{\!} denotes ⁣{}_{\!} the ⁣{}_{\!} approach ⁣{}_{\!} only ⁣{}_{\!} using ⁣{}_{\!} the ⁣{}_{\!} initial ⁣{}_{\!} node ⁣{}_{\!} embeddings ⁣{}_{\!} {hv(0) ⁣}v∈V ⁣\{\textbf{{h}}^{(0)\!}_{v}\}_{v\in\mathcal{V}\!} without ⁣{}_{\!} any relation information; (2) ‘Type-agnostic’ shows the performance when modeling different human part relations in a type-agnostic manner:  ⁣hu,v ⁣ ⁣ ⁣= ⁣ ⁣ ⁣R([hu,hv])\!\textbf{{h}}_{u,v\!}\!\!=_{\!}\!\!R([\textbf{{h}}_{u},\textbf{{h}}_{v}]); (3) ‘Type-specific w/o FrF^{r}’ gives the performance without the relation-adaption operation Fr ⁣F^{r\!} in Eq. ​2: ⁣{}_{\!} hu,v ⁣ ⁣ ⁣= ⁣ ⁣ ⁣Rr([hu,hv])\textbf{{h}}_{u,v\!}\!\!=_{\!}\!\!R^{r}([\textbf{{h}}_{u},\textbf{{h}}_{v}]); (4-6) ‘Decomposition relation’, ‘Composition relation’ and ‘Dependency relation’ are three variants that only consider the corresponding single one of the three kinds of relation categories, using our type-specific relation modeling strategy (Eq. ​2). Four main conclusions can be drawn: (1) Structural information are essential for human parsing, as all the structured models outperforms ‘Baseline’. (2) Typed relation modeling leads to more effective human structure learning, as ‘Type-specific w/o FrF^{r}’ improves ‘Type-agnostic’ by 1.28%. (3) Exploring different kinds of relations are meaningful, as the variants using individual relation types outperform ‘Baseline’ and our full model considering all the three kinds of relations achieves the best performance. (4) Encoding relation-specific constrains helps with relation pattern learning as our full model is better than the one without relation-adaption, ‘Type-specific w/o FrF^{r}’.

Iterative inference: Table 6 shows the performance of our parser with regard to the iteration step tt as denoted in Eq. ​13 and Eq. ​14. Note that, when t ⁣= ⁣0t\!=\!0, only the initial node feature is used. It can be observed that setting T ⁣= ⁣2T\!=\!2 or T ⁣= ⁣3T\!=\!3 provided a consistent boost in accuracy of 4∼\sim5%, on average, compared to T ⁣= ⁣0T\!=\!0; however, increasing TT beyond 3 gave marginal returns in performance (around 0.1%). Accordingly, we choose T ⁣= ⁣2T\!=\!2 for a better trade-off between accuracy and computation time.

Conclusion

In the human semantic parsing task, structure modeling is an essential, albeit inherently difficult, avenue to explore. This work proposed a hierarchical human parser that addresses this issue in two aspects. First, three distinct relation networks are designed to precisely describe the compositional/decompositional relations between constituent and entire parts and help with the dependency learning over kinetically connected parts. Second, to address the inference over the loopy human structure, our parser relies on a convolutional, message passing based approximation algorithm, which enjoys the advantages of iterative optimization and spatial information preservation. The above designs enable strong performance across five widely adopted benchmark datasets, at times outperforming all other competitors.

References