Semantic Graph Convolutional Networks for 3D Human Pose Regression

Long Zhao, Xi Peng, Yu Tian, Mubbasir Kapadia, Dimitris N. Metaxas

Introduction

Convolutional Neural Networks (CNNs) have successfully tackled classic computer vision problems such as image classification elhoseiny2017link; krizhevsky2012imagenet; li2018joint; simonyan2015very, object detection he2017mask; ren2015faster; tang2018quantized; wang2015object; zhao2018pseudo; zhu2017multilayer and generation radford2016unsupervised; tian2018cr; han2017stackgan; zhao2018learning; zhu2018generative, where the input image has a grid-like structure. However, many real-world tasks, e.g., molecular structures, social networks and 3D meshes, can only be represented in the form of irregular structures, where CNNs have limited applications.

In order to address this limitation, Graph Convolutional Networks (GCNs) gori2005new; kipf2016semi; scarselli2009graph have been introduced recently as a generalization of CNNs that can directly deal with a general class of graphs. They have achieved state-of-the-art performance when applied to 3D mesh deformation ranjan2018generating; wang2018pixel, image captioning yao2018exploring, scene understanding yang2018graph, and video recognition wang2018videos; stgcn2018spatial. These works utilize GCNs to model relations of visual objects for classification. In this paper, we investigate using deep GCNs for regression, which is another core problem of computer vision with many real-world applications.

However, GCNs cannot be directly applied to regression problems due to the following limitations in baseline methods kipf2016semi; wang2018pixel; stgcn2018spatial. First, to handle the issue that graph nodes may have various numbers of neighborhoods, the convolution filter shares the same weight matrix for all nodes, which is not comparable with CNNs. Second, previous methods are simplified by restricting the filters to operate in a one-step neighborhood around each node according to the guidance of kipf2016semi. The receptive field of the convolution kernel is limited to one due to this formulation, which severely impairs the efficiency of information exchanging especially when the network goes deeper.

In this work, we propose a novel graph neural network architecture for regression called Semantic Graph Convolutional Networks (SemGCN) to address the above limitations. Specifically, we investigate learning semantic information encoded in a given graph, i.e., the local and global relations of nodes, which is not well-studied in previous works. SemGCN does not rely on hand-crafted constraints du2015hierarchical; fang2018grammar; shahroudy2016ntu to analyze the patterns for a specific application, and thus can be easily generalized to other tasks.

In particular, we study SemGCN for 2D to 3D human pose regression. Given a 2D human pose (and the optional relevant image) as input, we aim to predict the locations of its corresponding 3D joints in a certain coordinate space. Using SemGCN to formulate this problem is intuitive. Both 2D and 3D poses are able to be naturally represented by a canonical skeleton in the form of 2D or 3D coordinates, and SemGCN can explicitly exploit their spatial relations, which are crucial for understanding human actions stgcn2018spatial.

Our work makes the following contributions. First, we propose an improved graph convolution operation called Semantic Graph Convolution (SemGConv) which is derived from CNNs. The key idea is to learn channel-wise weights for edges as priors implied in the graph, and then combine them with kernel matrices. This significantly improves the power of graph convolutions. Second, we introduce SemGCN where SemGConv and non-local wang2018non layers are interleaved. This architecture captures both local and global relationships among nodes. Third, we present an end-to-end learning framework to show that SemGCN can also incorporate external information, such as image content, to further boost the performance for 3D human pose regression.

The effectiveness of our approach is validated by comprehensive evaluation with a rigorous ablation study and comparisons with state of the art on standard 3D benchmarks. Our approach matches the performance of state-of-the-art techniques on Human3.6M ionescu2014human3 using only 2D joint coordinates as inputs and 90% fewer parameters. Meanwhile, our approach outperforms state of the art when incorporating image features. Furthermore, we also show the visual results of SemGCN, which demonstrate the effectiveness of our approach qualitatively. Note that the proposed framework can be easily generalized to other regression tasks, and we leave this for future work.

Related Work

Graph convolutional networks. Generalizing CNNs to inputs with graph-like structures is an important topic in the field of deep learning. In the literature, there have been several attempts to use recursive neural networks to process data represented in graph domains as directed acyclic graphs frasconi1998general. GNNs were introduced in gori2005new; kipf2016semi; scarselli2009graph as a more common solution to handle arbitrary graph data. The principle of constructing GCNs on graph generally follows two streams: the spectral perspective and the spatial perspective. Our work belongs to the second stream kipf2016semi; niepert2016learning; velickovic2018graph, where the convolution filters are applied directly on the graph nodes and their neighbors.

Recent studies on computer vision have achieved state-of-the-art performance by leveraging GCNs to model the relations among visual objects yang2018graph; yao2018exploring or temporal sequences wang2018videos; stgcn2018spatial. This paper follows the spirit of them, while we explore applying GCNs for regression tasks, especially, 2D to 3D human pose regression.

3D pose estimation. Lee and Chen lee1985determination first investigated inferring 3D joints from their corresponding 2D projections. Later approaches either exploited nearest neighbors to refine the results of pose inference gupta20143d; jiang20103d or extracted hand-crafted features agarwal2006recovering; ionescu2011latent; rogez2008randomized for later regression. Other methods created over-complete bases which are suitable for representing human poses as sparse combinations akhter2015pose; bogo2016keep; ramakrishna2012reconstructing; wang2014robust; zhou2016sparseness. More and more studies focus on making use of deep neural networks to find the mapping between 2D and 3D joint locations. A couple of algorithms directly predicted 3D pose from the image zhou2017towards, while others combined 2D heatmaps with volumetric representation pavlakos2017coarse, pairwise distance matrix estimation moreno20173d or image cues tekin2017learning for 3D human pose regression.

Recently, it has been proven that 2D pose information is crucial for 3D pose estimation. Martinez et al. martinez2017simple introduced a simple yet effective method which predicted 3D key points purely based on 2D detections. Fang et al. fang2018grammar further extended this approach through pose grammar networks. These works focus on 2D to 3D pose regression, which are most relevant to the context of this paper.

Other methods use synthetic datasets which are generated from deforming a human template model with the ground truth chen2016synthesizing; peng2018jointly; rogez2016mocap or introduce loss functions involving high-level knowledge park20183d; sun2017compositional; yang20183d in addition to joints. They are complementary to the others. Remaining works target at exploiting temporal information du2016marker; gupta20143d; hossain2018exploiting; tekin2016direct for 3D pose regression. They are out of the scope of this paper, since we aim at handling the 2D pose from one single image. However, our method can be easily extended to sequence inputs, and we leave it for future work.

Semantic Graph Convolutional Networks

We propose a novel graph network architecture to handle general regression tasks involving data that can be represented in the form of graphs. We first provide the background of GCNs and related baseline method. Then we introduce the detailed design of SemGCN.

We assume that graph data share the same topological structure, such as human skeletons du2015hierarchical; ke2017new; vemulapalli2014human; stgcn2018spatial, 3D morphable models loper2015smpl; ranjan2018generating; zhao2019cartoonish and citation networks sen2008collective. Other problems which own different graph structures in the same domain, e.g., protein-protein interaction velickovic2018graph and quantum chemistry gilmer2017neural, are out of the scope of this paper. This assumption makes it possible to learn priors implied in the graph structure, which motivates SemGCN.

Wang et al. wang2018pixel rephrased a very deep graph network based on Eq. 1 with residual connections he2016deep to learn the mapping between image features and 3D vertexes. We adopt its network architecture and treat it as our baseline which is denoted as ResGCN.

There are two clear drawbacks in Eq. 1. First, in order to make the graph convolution work on nodes with arbitrary topologies, the learned kernel matrix W\mathbf{W} is shared for all edges. As a result, the relationships of neighboring nodes, or the internal structure in the graph, is not well exploited. Second, previous works only collect features from the first-order neighbors of each node. This is also limited because the receptive field is fixed to 1.

2 Semantic Graph Convolutions

We show that learning semantic relationships of neighboring nodes implied in edges of the graph is effective to address the limitation of the shared kernel matrix.

The proposed approach builds on concepts from CNNs. Fig. 1(a) shows a CNN with a convolution kernel of size 3×33\times 3. It learns nine transformation matrices which are different from each other to encode features inside the kernel in the spatial dimension. This makes the operation own expressive power to model feature patterns contained in images. We find that this formulation can be approximated by learning a weighting vector a→i\overrightarrow{\boldsymbol{a}}_{i} for each position, and then combining them with a shared transformation matrix W\mathbf{W}. If we represent the image feature map as a square grid graph whose nodes represent pixels, this approximated formulation can be directly extended to GCNs as shown in Fig. 1(c).

where ρi\rho_{i} is Softmax nonlinearity which normalizes the input matrix across all choices of node ii; ⊙\odot is an element-wise operation which returns mijm_{ij} if aij=1a_{ij}=1 or negatives with large exponents saturating to zero after ρi\rho_{i}; A\mathbf{A} serves as a mask which forces that for node ii in the graph, we only compute the weights of its neighboring nodes j∈N(i)j\in\mathcal{N}(i).

where ∥\| represents channel-wise concatenation, and w→d\overrightarrow{\boldsymbol{w}}_{d} is the dd-th row of the transformation matrix W\mathbf{W}.

Comparison to previous GCNs. Both aGCN yang2018graph and GAT velickovic2018graph follow a self-attention strategy vaswani2017attention to compute the hidden representations of each node in the graph by attending over its neighbors. They aim to estimate a weighting function depending on inputs for edges to modulate information flow throughout the graph. By contrast, we target at learning input-independent weights for edges which represent priors implied in the graph structures, e.g., how one joint influences other body parts in human pose estimation.

The edge importance weighting mask introduced in ST-GCN stgcn2018spatial is the most related work to ours but with following two sharp differences. First, no Softmax nonlinearity is leveraged after weighting by stgcn2018spatial, while we find it stabilizes the training and obtains better results, since the contributions of nodes to their neighbors are normalized by Softmax. Second, ST-GCN applies only one single learnable mask to all channels, but our Eq. 3 learns channel-wise different weights for edges. As a result, our model owns better capability to fit the data mapping.

3 Network Architecture

Capturing global and long-range relationships among nodes in the graph is able to efficiently address the problem of the limited receptive field. However, in order to maintain the behavior of GCNs, we restrict the feature updating mechanism by computing responses between nodes based on their representations other than learning new convolution filters. Therefore, we follow the non-local mean concept buades2005non; wang2018non and define the operation as:

where WxW_{x} is initialized as zero; ff is a pairwise function to compute the affinity between node ii and all other jj; gg computes the representation of the node jj. In practice, Eq. 4 can be implemented by the non-local layers proposed in wang2018non.

Based on Eq. 3 and 4, we propose a new network architecture for regression tasks called Semantic Graph Convolutional Networks, where SemGConv and non-local layers are interleaved to capture local and global semantic relations of nodes. Fig. 2 shows an example. In this work, SemGCN in all blocks has the same structure, which consists of one residual block he2016deep built by two SemGConv layers with 128 channels, and then followed by one non-local layer. This block is repeated several times to make the network deeper. At the beginning of the network, one SemGConv is used for mapping the inputs into the latent space; and we have an additional SemGConv which projects the encoded features back to the output space. All SemGConv layers are followed by batch normalization ioffe2015batch and a ReLU activation nair2010rectified except the last one. Note that if SemGConv layers are replaced with vanilla graph convolutions and all non-local layers are removed, SemGCN downgrades to ResGCN in Sect. 3.1.

Intuitively, SemGCN can be regarded as a form of neural message passing system gilmer2017neural where the forward pass has two phases: messages are updated locally and then refined by the global state of the system. These two phases take turns to process messages so that the efficiency of information exchanging is improved for the whole system.

3D Human Pose Regression

In this section, we present a novel end-to-end trainable framework which incorporates SemGCN in Sect. 3 with image features for 3D human pose regression.

We argue that image content is able to offer important cues for solving ambiguous cases, such as the classic turning ballerina optical illusion. Therefore, we extend Eq. 5 by treating image content as an additional constraint. The extended formulation can be denoted as:

where IiI_{i} is the image containing the aligned human pose of the 2D joints Pi\mathbf{P}_{i}. In practice, P\mathbf{P} may be obtained as 2D ground truth locations under known camera parameters or from a 2D joint detector. In the latter case, the 2D detector has already encoded the perceptual features of the input image during the training process. This observation motivates the design of our framework.

An overview of our framework is shown in Fig. 3. The whole framework consists of two neural networks. Given an image, one deep convolutional network is leveraged for 2D joints prediction; at the same time, it also serves as a backbone network and image features are pooled from its intermediate layers. Since 2D and 3D joint coordinates can be encoded in a human skeleton, the proposed SemGCN is used for automatically capturing the patterns embedded in the spatial configuration of the human joints. It predicts 3D coordinates according to the 2D pose as well as perceptual features from the backbone network.

Note that our framework effectively reduces to Eq. 5 when image features are not considered. As we demonstrate in experiments, SemGCN manages to effectively encode the mapping from 2D to 3D poses, and the performance can be further boosted when incorporating image content.

2 Perceptual Feature Pooling

ResNet he2016deep and HourGlass newell2016stacked are widely adopted in conventional human pose detection problems. Empirically, we employ ResNet as the backbone network since its intermediate layers provide hierarchical features from images which are useful in computer vision problems such as object detection and segmentation ren2015faster; zhao2018pseudo.

Given the coordinate of each 2D joint in the input image, we pool features from multiple layers in ResNet. In particular, we concatenate features extracted from layer conv_1 to conv_4 using RoIAlign he2017mask. These perceptual features are then concatenated with the 2D coordinates and fed into SemGCN. Note that since all joints in the input image share the same scale, we pool the features in a squared bounding box centered on each joint with a fixed size, i.e., the mean bone length of the skeleton. This is illustrated in Fig. 3.

3 Loss Function

Most previous regression-based methods directly minimize the mean squared errors (MSE) of the predicted and ground truth joint positions carreira2016human; martinez2017simple; tekin2016direct; zhou2016deep or bone vectors sun2017compositional. Following the spirit of them, we employ a simple combination of joint and bone constraints in human poses as our loss function, which is defined as:

Experiments

In this section, we first introduce settings and implementation details for evaluation, and then conduct an ablation study on components in our method, and finally report our results and comparisons with state-of-the-art methods.

As suggested in the previous works martinez2017simple; sun2017compositional; zhou2017towards, it is impossible to train an algorithm to infer the 3D joint locations in an arbitrary coordinate space system. Therefore, we choose to predict 3D pose in the camera coordinate system du2016marker; li2015maximum; pavlakos2017coarse; tekin2016direct, which makes the 2D to 3D regression problem similar across different cameras.

We make use of the ground truth 2D joint locations provided in the dataset to align the 3D and 2D poses following the setting of zhou2017towards. This implies that we implicitly use the camera calibration information. Then, we zero-center both the 2D and 3D poses around the predefined root joint, i.e., the pelvis joint, which is in line with previous works and the standard protocol. Moreover, we do not use data augmentation during the training process for simplicity.

Network training. We use ResNet50 in sun2017integral as our backbone network, which is compatible with the integral loss and pre-trained on ImageNet deng2009imagenet. During training, we employ Adam kingma2014adam for optimization with a initial learning rate of 0.001 and use mini-batches of size 64. The learning rate is dropped with a decay rate of 0.5 when the loss on the validation set saturates. We initialize weights of the graph network using the initialization described in glorot2010understanding.

In our preliminary experiments, we observe that the direct end-to-end training of the whole network from scratch cannot achieve the best performance. We argue that this is likely because of the highly non-linear dependency between the graph network and conventional deep convolutional module for 2D pose estimation. Therefore, we utilize a multi-stage training scheme which is more stable and effective in practice. We first train the backbone network for 2D pose estimation from images using 2D ground truth. As described in sun2017integral, the integral loss is used. Then we fix the 2D pose estimation module and train the graph network for 2D to 3D pose regression using the output of 2D estimation module and the 3D ground truth. In this stage, the loss function defined in Eq. 7 is employed. At last, the whole network is fine-tuned with all data. Both integral loss and Eq. 7 are activated. Note that the final stage is end-to-end.

2 Datasets and Evaluation Protocols

Our proposed approach is comprehensively evaluated on the most widely used dataset for 3D human pose estimation: Human3.6M ionescu2014human3, following the standard protocol.

Datasets. Human3.6M ionescu2014human3 is currently the largest publicly available dataset for 3D human pose estimation. This dataset contains 3.6 million of images captured by a MoCap System in an indoor environment, where 7 professional actors perform 15 everyday activities such as walking, eating, sitting, making a phone call and engaging in a discussion. Both 2D and 3D ground truth are available for supervised learning. Following the setting of zhou2017towards, the videos are down-sampled from 50fps to 10fps for both the training and testing sets to reduce redundancy. We also use MPII dataset andriluka20142d, the state-of-the-art benchmark for 2D human pose estimation, for pre-training the 2D pose detector and qualitatively evaluation in the experiment.

Evaluation protocols. For Human3.6M ionescu2014human3, there are two common evaluation protocols using different training and testing data split in the literature. One standard protocol uses all 4 camera views in subjects S1, S5, S6, S7 and S8 for training and the same 4 camera views in subjects S9 and S11 for testing. Errors are calculated after the ground truth and predictions are aligned with the root joint. We refer to this as Protocol #1. The other protocol makes use of six subjects S1, S5, S6, S7, S8 and S9 for training, and evaluation is performed on every 64th frame of S11. It also utilizes a rigid transformation to further align the predictions with the ground truth. This protocol is referred as Protocol #2. In this work, we use Protocol #1 in all the experiments for evaluation, since it is more challenging and matches the settings of our method.

The evaluation metric is the Mean Per Joint Position Error (MPJPE) in millimeter between the ground truth and the predicted 3D coordinates across all cameras and joints after aligning the pre-defined root joints (the pelvis joint). We use this metric for evaluation in the following sections.

Our network predicts the normalized locations of 3D joints. During testing, to calibrate the scale of the outputs, we require that the sum of length of all 3D bones is equal to that of a canonical skeleton as shown in pavlakos2017coarse; zhou2017towards; zhou2018monocap. Therefore, we follow the method in zhou2017towards for calibration.

Configurations. Our method is evaluated with the following two different configurations for 3D human pose estimation on Human3.6M.

Configuration #1. We only leverage 2D joints of the human pose as inputs. SemGCN in Sect. 3 is trained for regression and the SemGConv layer defined in Eq. 2 is utilized. 2D ground truth (GT) or outputs from pre-trained 2D pose detectors are used for training and testing. In order to be in line with the setting of previous works fang2018grammar; martinez2017simple, we employ HourGlass newell2016stacked (HG) as the 2D detector. It is first pre-trained on MPII and then fine-tuned on Human3.6M. Only the joint loss in Eq. 7 is employed.

Configuration #2. We use 2D images as inputs, and the proposed framework in Sect. 4 is trained for regression. The channel-wise weighted SemGConv defined in Eq. 3 is employed. ResNet50 he2016deep is utilized as the backbone network for 2D pose estimation and feature pooling (RN w/ FP).

3 Ablation Study

We conduct the ablation study on the proposed method in Sect. 3. Configuration #1 is employed. Our SemGCN consists of two main components: SemGConv and non-local layers. To verify them, we train two variants of SemGCN: one only uses SemGConv and the other only uses non-local layers. Then we evaluate them together with the baseline method in Sect. 3.1 (ResGCN) and our full model in Sect. 3.3 on Human3.6M. Note that in order to get rid of the influence from the 2D pose detector, we report the results using 2D ground truth for training and testing.

All models are trained based on the architecture shown in Fig. 2 after 200 epochs. Results are shown in Table 1. We also show their curves of training losses and testing errors in Fig. 4. We can see that our model with more components performs better than those with fewer components, which indicates the efficacy of each part of our algorithm. Moreover, our networks with SemGConv have much smoother training curves which demonstrates that learning local relations among nodes stabilizes the training process as well.

4 Evaluation on 3D Human Pose Regression

2D to 3D pose regression. We first evaluate our method for 2D to 3D pose regression and only Configuration #1 is leveraged. We compared ours with three GCN-based methods: aGCN yang2018graph, GAT velickovic2018graph and ST-GCN stgcn2018spatial, and two state-of-the-art approaches: FC martinez2017simple and PG fang2018grammar. As ST-GCN stgcn2018spatial is designed for videos, we set its temporal dimension to one for images. PG proposed a framework to refine the 3D pose, which is complementary to FC and ours. Therefore, we also report our results refined by PG.

The results are reported in Table 3. Our approach outperforms other GCN-based approaches by a large margin (about 20%). More importantly, our method achieves the state-of-the-art performance with around 90% fewer parameters than martinez2017simple. Meanwhile, the runtime of SemGCN reduces 10% compared with martinez2017simple, which is around 1.8ms for a forward pass on a Titan Xp GPU. After we refined our results by PG, our approach obtains the best performance.

Comparison with the state of the art. We show evaluation results under Configuration #1 and #2. Note that many leading methods have sophisticated frameworks or learning strategies. Some of them aim at in-the-wild images sun2017integral; yang20183d; zhou2017towards or exploit temporal information du2016marker; gupta20143d; hossain2018exploiting; tekin2016direct, while some other approaches use complex loss functions sun2017compositional; yang20183d. These methods are with different research targets compared to ours. Therefore, we include some of them during evaluation for completeness. Table 2 reports the results.

We find that our method using only 2D joints as inputs is able to match the state-of-the-art performance. After incorporating image features, our network sets the new state of the art. Especially, we improve previous methods by a large margin for the action of directions, taking photo, posing, sitting down, walking dog and walking together. We hypothesize that this is due to the severe self-occlusions in these actions, while they can be effectively encoded by our SemGCN using relations within graphs. The result of our method trained and tested with ground truth 2D joint locations shows our upper bound.

Qualitative results. In Fig. 5, we show the visual results of our method on Human3.6M and the test set of MPII. MPII contains in-the-wild images with novel human poses which are not similar to the examples in Human3.6M. As seen, our method is able to accurately predict 3D pose for both indoor and most in-the-wild images. It indicates that SemGCN can effectively encode relationships among joints and further generalize them to some novel cases.

The bottom row of Fig. 5 also shows typical failure cases of our method. These images include extreme poses which are largely different from those in Human3.6M. Our method failed to handle them but still yields reasonable 3D poses.

Conclusions

We present a novel model for 3D human pose regression, the Semantic Graph Convolutional Networks (SemGCN). Our method has addressed the key challenges of GCNs by learning local and global semantic relations among nodes in the graph. The combination of SemGCN and features pooled from image content further improves the performance in 3D human pose estimation. Comprehensive evaluation results show that our network obtains state-of-the-art performance with 90% fewer parameters compared with the closest work. The proposed SemGCN also opens up many possible directions for future works. For example, how to incorporate temporal information, such as videos, into SemGCN becomes a natural question.

Acknowledgments. This work was funded partly by grant BAAAFOSR-2013-0001 to Dimitris Metaxas. This work was also partly supported by NSF 1763523, 1747778, 1733843 and 1703883 Awards. Mubbasir Kapadia was funded partly by NSF IIS-1703883, NSF S&AS-1723869, and DARPA SocialSim-W911NF-17-C-0098.

References

Appendix A Supplementary Material

This supplementary material provides additional results supporting the claims of the main paper. First, we provide more details about Semantic Graph Convolutional Networks (SemGCN), including the skeleton representation for building the graph (Sect. A.1) and the implementation of graph convolutions (Sect. A.2) and non-local layers (Sect. A.3). Additionally, to better understand the proposed Semantic Graph Convolutions, we provide the visualization results of the learned weights implied in the graph after training (Sect. A.4).

Following the setting of previous works fang2018grammar; martinez2017simple; sun2017compositional; sun2017integral; zhou2017towards, we utilize a common human skeleton representation for Human3.6M ionescu2014human3 and MPII andriluka20142d to build the graph of SemGCN. This skeleton is visualized in Fig. 6(left). It consists of 16 joints and we define the pelvis joint as the root joint. Note that the skeleton is initialized as an undirected graph in SemGCN before training. After we finish training the network, it will transform to a weighted directed graph represented by ρi(M⊙A)\rho_{i}\big(\mathbf{M}\odot\mathbf{A}\big) in Eq. 2 and 3.

In Fig. 6(left), we also show the bone vectors we employed in Eq. 7 to compute the bone loss. Let the bone Bk\mathbf{B}_{k} be directed from the joint Jparent(k)\mathbf{J}_{parent(k)} to the target joint Jk\mathbf{J}_{k}, and we define the bone vector as:

This formulation is consistency with sun2017compositional. However, in order to be in line with the setting of previous works fang2018grammar; martinez2017simple for fair comparison, the bone loss is not employed in Configuration #1 of the experiments.

A.2 Implementation of Graph Convolutions

Some previous approaches wang2018pixel; stgcn2018spatial proposed to leverage two different transformation matrixes other than one in the graph convolutions. To be specific, when the graph convolutional filter is applied to node ii in the graph, one matrix W0\mathbf{W}_{0} is employed to transform the representation of node ii while the other matrix W1\mathbf{W}_{1} is learned for all its neighbors. According to this formulation, we rewrite Eq. 1 to:

where ⊗\otimes denotes element-wise multiplication and I\mathbf{I} is the identity matrix. We also implement the proposed SemGConv defined by Eq. 2 and 3 in the similar manner.

A.3 Non-local Layers

We follow the guidance of Wang et al. wang2018non to implement the non-local layers in SemGCN. For computational efficiency, we down-sample both the feature dimension and number of nodes in the graph when calculating the embedding of f(x→i(l),x→j(l))f(\overrightarrow{\boldsymbol{x}}^{(l)}_{i},\overrightarrow{\boldsymbol{x}}^{(l)}_{j}) in Eq. 4.

Feature embedding. We use “concatenation” for the implementation of ff. Two mapping functions θ(⋅)\theta(\cdot) and ϕ(⋅)\phi(\cdot) are employed to down-sample the feature of each node from 128 to 64 channels. They are implemented as convolutions with the kernel size of 1. Then we define ff as:

where [⋅∥⋅][\cdot\|\cdot] denotes concatenation, and wf\mathbf{w}_{f} is the parameters to be learned to project the concatenated vector to a scalar.

Node grouping. We also use max pooling to down-sample the number of nodes in the graph. Fig 6(right) illustrates the grouping strategy we employed. The number of nodes contained in the graph reduces from 16 to 8 after the max pooling operation. This strategy is used for all non-local layers in SemGCN. In the experiments, we find that this pooling operation can speed up the runtime, while it does not influence the final accuracy of the regression.

A.4 Visualization of Weights in SemGCN

To better understand the proposed SemGCN, we visualize the learned weighting matrix M\mathbf{M} of each SemGConv layer in the network. For simplicity, we utilize a simplified version of SemGCN, where Eq. 2 is employed so that all feature channels share the same M\mathbf{M}. We use the network architecture as illustrated in Fig. 2 and train it according to Configuration #1 of the experiments.

The trained network consists of 4 residual blocks where each block contains 2 SemGConv layers. Therefore, we visualize the weighting matrixes of these 8 SemGConv layers respectively. The matrixes are shown in Fig. 7. We have made two important observations. First, although all SemGConv layers share the same graph structure in the network, each of them has learned a different weighting matrix. Second, we can find that SemGConv layers have learned higher weights for nodes which are farer from the gravity center of the human skeleton on average.

To further illustrate the second observation, we compute the average learned weight of each joint among all 8 SemGConv layers. The quantitative results are shown in Fig. 8(left). We can see that the left wrist, right wrist, left ankle, right ankle and head own the top highest weights which are greater than 0.4; while the neck, thorax and pelvis have the lowest weights less than 0.3. Other joints have quite similar weights around 0.3. This result can be better visualized by representing the human skeleton with a regional map where joints are grouped into three regions according to their weights. Fig. 8(right) shows the result.

This result is intuitive since joints farer from the center always encode more information of the pose while central joints determine the position of the skeleton. This observation is also consistency with sun2017compositional; stgcn2018spatial. This demonstrates that the proposed SemGCN is able to effectively encode spatial relationships of nodes in the graph. However, we only rely on the ground truth for supervision, and no additional hand-crafted constraints or rules are employed.