Graph Stacked Hourglass Networks for 3D Human Pose Estimation

Tianhan Xu, Wataru Takano

Introduction

In recent years, with the application of deep learning methods, the performance of 2D human pose estimation has been greatly improved. Recent works show that using such detected 2D joints positions, the 3D human pose can also be efficiently and accurately regressed . Due to the graph structure formed by the topology of the human skeleton, many attempts have been made to use the generalized form of CNN: Graph Convolutional Networks (GCN), to perform the regression task of 2D-to-3D human pose estimation . Since the graph convolution has a good feature extraction capability for the graph-structured data, the GCN-based approaches work well and some of them achieve the state-of-the-art results in the 2D-to-3D human pose estimation task.

However, the existing GCN-based approaches have the following limitations: First, the graph convolutions exploit all node information, which can be seen as that all features are processed at only “one scale”. Thus it is difficult to extract features that can represent spatial local and global information and limits the representation capabilities. Second, most existing approaches use a straightforward architecture of sequentially connecting the graph convolution layers (Fig. 4 (Top)). Such a model architecture does not take advantage of the benefits of model depth, such as intermediate features at each depth, and therefore limits its performance.

The core issue mentioned above is that the features extracted by existing architecture are oversimple, which limits the expressiveness of the model. Features with greater representation capabilities, such as multi-scale and multi-level features, are commonly used in image-related tasks. For multi-scale features, they denote the information from small to large resolutions of the image features, thus bringing rich image understanding from local to global , while multi-level features denote the latent representations in different depths of the latent space, bringing important semantic information at all levels from shallow to deep. The introduction of the above multi-scale and multi-level features can enrich the performance of the model. However, because the upsampling and downsampling operations required for multi-scale features are defined on the image, and the graph has an irregular structure, such methods cannot be directly applied to the graph-structured data.

To address these issues, we propose a novel architecture for 2D-to-3D human pose estimation: Graph Stacked Hourglass Networks. Specifically, we do not focus on specific graph convolution operations, but rather consider how to integrate them in the architecture which gives the best performance improvement. Given the advantages of the 2D human pose estimation approaches, proposed architecture adopts the repeated encoder-decoder applicable to the graph-structured data for multi-scale feature extraction, as well as the intermediate features at each depth of the model for multi-level feature extraction. Such multi-scale and multi-level feature information makes the model more expressive and enables the model to achieve high-precision 3D human pose estimation.

Our work makes the following contributions. First, we propose Graph Hourglass modules suitable for extracting multi-scale human skeletal features, which includes novel pooling and unpooling operations considering human skeletal structure, called Skeletal Pool and Skeletal Unpool. Second, we introduce Graph Stacked Hourglass Networks (GraphSH) consisting of the proposed Graph Hourglass module , which incorporates multi-level feature representations at various depths of the architecture. Our architecture incorporates multi-scale, multi-level features and a priori knowledge of the human skeleton, achieving impressive performance improvement for 2D-to-3D human pose estimation tasks.

Related Work

3D Human Pose Estimation. Predicting the 3D human pose from images or videos has been an essential topic in computer vision for a long time. In the early days, handcrafted features, perspective relationships, and geometric constraints were used to predict 3D human pose . In recent years, with the development of deep learning, there has been an increase in using deep neural networks for image-to-3D human pose estimation .

Some methods regress 3D human pose directly from images. Tekin et al. propose a method that first train an autoencoder to learn the latent representation of the 3D human pose, then use CNN to regress the image with the latent representation, and finally connect the trained CNN with the decoder to achieve the prediction from image to 3D human pose. Pavlakos et al. exploit voxel to discretize representations of the space around the human body and use 3D heatmaps to estimate 3D human pose.

There are also methods that break the problem down into two steps: first predicting 2D human joints from the image, and then using the 2D joints information to predict 3D human pose. Our approach falls into this category. Martinez et al. propose a simple yet effective baseline for 3D human pose estimation that uses only 2D joints information but get highly accurate results, showing the importance of 2D joints information for 3D human pose estimation. Since the human skeleton’s topology can be viewed as a graph structure, there has been increasing use of Graph Convolutional Networks (GCN) for 2D-to-3D human pose estimation tasks .

Graph Convolutional Networks. Graph Convolutional Networks (GCN) are used to perform convolution operations on graph-structured data, such as human skeleton, thus enabling effective feature extraction. The early simple GCN is the ‘vanilla’ GCN proposed by Kipf and Welling , which consists of a simple graph convolution operation that performs the transformation and aggregation of graph-structured data, and it becomes the basic model for various graph convolution later on. The following GCNs are based on this model with some improvements and are applied to 2D-to-3D human pose estimation. Zhao et al. propose Semantic Graph Convolution (SemGConv), which has learnable adjacency matrix parameters, enabling the model to learn the semantic relationships between the human joints. In the two graph convolutions just introduced, each node information is transformed using the same weight matrix and then aggregated. Liu et al. point out that sharing the same weight by all nodes limits the representation capabilities of graph convolution, and propose a new method that, first transforming each node information using different weights and then aggregating them together, called Pre-Aggregation Graph Convolution (PreAggr). They also introduce an approach that decouples the self-connections in the graph and use separate weight to compute the self-information transformation.

The GCN using the above graph convolution for 3D human pose estimation, however, only use a straightforward overall architecture, as in Fig. 4 (Top). Such a simple architecture prevents the model from using the multi-scale, multi-level features common in image-based tasks, limiting the performance of the model.

Multi-scale and Multi-level Learning. Multi-scale and multi-level features learning drives advances in a wide variety of image-based tasks . Multi-scale feature learning refers to integrating features in different resolutions to provide a better understanding within the spatial domain. Feature Pyramid Network (FPN) is a powerful multi-scale feature extractor and achieves encouraging results in object detection tasks. In the task of 2D human pose estimation from image, Hourglass structure for extracting multi-scale features allows the model to learn both local and global features, which are essential for human pose understanding. (\eg, spatial configuration relationships between human joints) . Multi-level feature learning, on the other hand, represents the use of features at various depths of the network. Some methods use a skip layer to incorporate features from the intermediate layer of the network and then combine them into the output layer . Zhao et al. incorporate multi-scale and multi-level features, combining image pyramids of different depths to extract higher feature representations.

These image-based methods take advantage of the fact that images can be easily scaled up and down and the richness of intermediate features, thus enabling multi-scale and multi-level feature extraction. However, due to the graph structure’s irregularity, it is not trivial to scale it up or down like an image, so such approaches have not been applied much to tasks with graph-structured data.

In the next section, we present our proposed novel graph convolutional network architecture that integrates multi-scale and multi-level features of the graph-structured data.

Graph Stacked Hourglass Networks

Our approach is inspired by Stacked Hourglass Networks proposed by Newell et al. for estimating 2D human pose from images, which exploits repeated hourglass-like encoder-decoder architecture. We aim to extend such an hourglass structure to the graph for extracting multi-scale features of the graph-structured data.

Using deep neural networks for computer vision tasks, multi-scale features of the image are essential for image understanding. Since images have large amounts of information at high resolution, the model can extract much detailed information from them. Alternatively, while the image is at low resolution, the model can better extract globalized information. The hourglass structure, accompanied by downsampling and upsampling operations, enables the image features to go through all resolutions so that crucial information can be extracted at all scales. Previous works have shown that this hourglass structure has strong feature extraction capabilities, especially for tasks requiring both local and global information, such as human pose estimation . Moreover, stacking such a structure enables repeated feature extraction and enhances model performance.

Our motivation is to extend such a structure with powerful multi-scale feature extraction capabilities to graph-structured data to achieve highly accurate 2D-to-3D human pose estimation. In the hourglass structure, downsampling and upsampling of the data are implemented by pooling and unpooling operations, respectively. Such operations are easy to define on the images because of their regularized structure, but there is no consistent way to define them on the graph. On this point, there are some previous works regarding pooling and unpooling operations on the graph . However, these pooling and unpooling operations are defined on the more generalized arbitrary-shaped graph structure, while the human skeletal graph used in this study has a fixed structure as shown in Fig. 1(a). Here, we exploit this property to propose pooling and unpooling methods applicable to the human skeletal graph, called Skeletal Pooling and Skeletal Unpooling.

Skeletal pooling. According to the property of the human body structure, we group the human body nodes in pairs, where the corresponding two nodes’ features are fused into one node in the lower-scale skeleton structure, using max pooling operation. With repeated pooling operation, we can obtain three different scales of skeletal structures containing 16, 8, and 4 nodes, respectively, and each of them corresponds to a different graph structure. For relatively lower-scale graph representations, we use more channels to encode information to prevent information degradation. The illustration of skeletal pooling and the three-scales skeleton graph structures are shown in Fig. 1(b).

Skeletal unpooling. We use unpooling operation to restore lower-scale skeletal structures to their original size, enabling them to fuse higher-scale skeletal information and pass it to subsequent processing. We adopt a very simple approach: since lower-scale nodes are generated by two grouped higher-scale nodes, in unpooling operation, we duplicate the feature representations of the lower-scale node and assign them to corresponding two nodes to recover the higher-scale skeletal representations. The illustration of skeletal unpooling is shown in Fig. 1(c).

Graph hourglass design. Using the above skeletal pooling and unpooling operations, we propose a novel graph hourglass module applicable to human skeleton representation.The detailed hourglass structure is shown in Fig. 2. For the same scale of the skeletal structure, we applied the residual connections to pass information and prevent the vanishing gradient problem. Note that our hourglass structure does not depend on the specific graph convolution layer, so arbitrary graph convolution operation can be implemented on our model, such as the three introduced in Sect. 2.

Graph U-Nets proposed by Gao et al. is the closest work to our architecture, but differs in two ways. First, uses an input-related dynamic pooling operation so that different pooled skeletal structures are obtained depending on the input. Our method utilizes a priori knowledge of human skeletal structure, and this approach is more suitable for feature extraction of specific skeletal structures while ensuring the stability of pooling. Second, we use more channels at relatively low scales of skeleton representation to reduce information loss due to scale changes and make our architecture more expressive, while Graph U-Nets uses the same number of channels at all scales.

2 Multi-scale and Multi-level Features

Features are extracted at multi-scales as the hourglass module processes the information across three skeletal structures. As multi-scale features that can represent information on the spatial aspect of the graph, we believe that multi-level features in terms of the depth of latent space can also bring valuable information to the final prediction. Specifically, we integrate the intermediate features at each depth level of the network for the final 3D human pose estimation. As shown in Fig. 3, we use the spatial 1x1 convolution to reduce the channels of intermediate features and concatenate them into an overall feature representation fcat\mathbf{f}_{cat}. For the architecture that stacks nn hourglass modules, the overall feature can be represented as:

3 Network Architecture

As in previous works , our backbone network consists of the proposed graph hourglass module stacked.

The input 2D joint information is first mapped to the latent feature space via a pre-processing graph convolution layer. The features that through the hourglass module are fed into the 1x1 convolution layer to be transformed into intermediate features, and also passed to the next hourglass module, except the last one. All the intermediate features are concatenated into the final feature, and the SE block adjusts its channel-wise weights, and which is then fed into the output convolution layer and mapped to the output space. Our overall network structure is shown in Fig. 3.

In the following experiments, our model uses PreAggr as the graph convolution layer, with 4-stacking hourglass approaches and latent space of 64 channels.

Experiments

In this section, we first describe the experimental setup for 2D-to-3D human pose estimation tasks. Next, we introduce the dataset used and its evaluation protocols. Then several ablation studies are conducted regarding the proposed architecture. Finally, we show our experimental results and comparisons with state-of-the-art methods.

In this study, Mean Squared Error (MSE) is used as a loss function L\mathcal{L} for training.

2 Datasets and Evaluation Protocols

Datasets. The Human3.6M dataset is the most widely used dataset in the 3D human pose estimation tasks. It uses motion captures to obtain the 3D pose information of the subjects and 4 cameras with different orientations to record the corresponding video image information. The provided camera parameters allow us to obtain the ground truth of the corresponding 2D joint coordinates in each image frame. The dataset provides 3.6 million images by recording 11 professional actors performing 15 different actions, such as eating, walking, etc. In the following experiments, we mainly use the Human3.6M for training and testing. The MPI-INF-3DHP test set provides images in three different scenarios: studio with a green screen (GS), studio without green screen (noGS) and outdoor scene (Outdoor). We use this dataset to test the generalization capabilities of our proposed architecture.

Evaluation protocols. For the Human3.6M, There are two evaluation protocols used in previous works . Protocol #1 uses the Mean Per Joint Position Error (MPJPE) in millimeter as evaluation metric, which calculates the Euclidean distance error between the prediction and the ground truth after the origin (pelvis) alignment. Protocol #2 aligns the prediction with the ground truth by rigid transformation and then calculates the error. In this study, we use Protocol #1 to evaluate our approach since performance under both protocols is usually consistent and Protocol #1 is more appropriate for our experimental setup. For the MPI-INF-3DHP test set, we follow previous works and use 3D-PCK and AUC as evaluation metrics.

3 Implementation Details

Our implementation follows the settings of previous works . As introduced in , we normalize the coordinates of the 2d and 3d joints and align the root joint (pelvis) to the origin.

In our experiments, we use a 4-stacked hourglass architecture. Previous works typically use 128 as the number of channels , but due to the relative complexity of our model architecture, we use 64 channels to keep the number of parameters at the same scale as previous works. We use Adam as the optimizer with an initial learning rate set to 0.0001 and decay by 0.92 per 20,000 iterations. We use a mini-batch size of 256. Since the learning of graph-structured data is very prone to overfitting, we apply Dropout with a dropout probability of 0.25 to all graph convolutional layers within the hourglass module. The entire training follows an end-to-end fashion.

4 Ablation Study

Pooling and Unpooling. Pooling and unpooling layers play an important role in the hourglass module.

Due to the graph structure’s irregularity, there is no consistent way to pool graph-structured data, so we compare the impact of different pooling methods on model performance. Here we compare the performance of three graph pooling operations: gPool , SAGPool , and our proposed Skeletal Pool.

The first two pooling methods take the same idea: calculate the scores of each node by some operation, and keep the part of the node with the higher score. Such pooling methods are initially designed for more general graph pooling situations, with the benefit that pooling operations can be defined for different graph-structured inputs. However, since the human skeleton used in this study has a fixed graph structure, it does not take advantage of the benefits of these pooling approaches. Moreover, since these pooling methods are input-dependent, different subgraphs are generated based on different inputs, which not only introduces computational complexity but also makes it difficult for the model to learn valuable features stably.

Compared to the above pooling methods that require computation, our pooling approach is more like methods that focus on geometric information of graph structure (\eg, edges, nodes), such as mesh sampling or edge contraction in mesh convolution. Specifically, We follow the node grouping method in to perform pooling operations on paired nodes. In their work, the features lose their graph structure after pooling. We extend their grouping concept by designing three subgraph structures of the human body consisting of 16, 8, and 4 nodes, respectively. Such a pooling approach not only exploits the topology of the human skeleton, which is more interpretable relative to other pooling approaches, but also, due to its simplicity, greatly reduces the computational complexity of the pooling layer and improves the speed of training and inference. Moreover, the two pooling methods mentioned above completely discard some nodes’ information, resulting in some degree of information loss. Instead, our proposed Skeletal Pool refers to the concept of image pooling, which performs a maximum pooling operation between two nodes, thus can summarize the information of both nodes.

Regarding unpooling, in Graph U-Net , the gUnpool operation assigns zero vectors to nodes that are not selected at the time of pooling, which loses much valuable information and results in more sparse features. Experimental results show that such an unpooling operation is not suitable for understanding structures with few nodes like the human skeleton. On the contrary, our proposed Skeletal Unpool operation copies and assigns the node features in the lower-scale graph representation to the corresponding two nodes in the higher-scale graph, passes them into the following graph convolutional layer, makes our model have better representation capability.

We conduct experiments based on three pool/unpool settings, and the results are shown in Table 1. We also add a comparison that removes all the pool/unpool layers to validate the importance of the proposed skeletal pool/unpool approaches. Note that we use ground truth 2D joints as input in this and the following ablation experiments to eliminate the influence of the 2D human pose detector.

Multi-scale and Multi-level Features. Multi-scale feature extraction is achieved by pooling and unpooling in the hourglass module that transforms features across three different scales. For comparison, we remove all pooling and unpooling layers in our architecture, which means that the features are always processed at the highest scale. The results in Table 1 show that multi-scale features derived from skeletal pooling and unpooling can improve model performance. For multi-level features, our architecture concatenates each level’s intermediate features and feeds them into the SE block to obtain the final overall feature. We compare three different settings: (1) No multi-level features are used. At this point the model simply connect all the hourglass modules sequentially, and the last hourglass modules is connected to the final output layer of 1x1 convolution. (2) Remove the SE block that calculates the weights of each intermediate feature. (3) Our proposed GraphSH architecture. Results are shown in Table 2.

Stacked Hourglass Architecture. To verify that our architecture has better performance than the simple Sequential Residual blocks (denoted as SeqRes) in previous works , we use Vanilla Graph Convolution, Semantic Graph Convolution (SemGConv), and Pre-aggregation Graph Convolution (PreAggr) introduced in Sect. 2 as convolution layers in our GraphSH architecture, respectively, and compare them to the corresponding SeqRes models. To make the comparison fair, we reduce the number of channels in the convolution to 64, so that the overall number of parameters is at the same scale as the SeqRes model. Results are shown in Table 3.

The results show that the model using our architecture performs better even though we use fewer channels than other GCN-based approaches. Our architecture does not rely on specific graph convolution layers, which suggests that any graph convolution layers for 3D human pose estimation can be applied to our architecture and improve its performance compared to the SeqRes architecture.

Moreover, our hourglass module can be seen as a well-integrated, high-performance graph convolution module, indicating that our hourglass module is general and can be easily extended to other tasks using graph convolution, such as action recognition , motion prediction , etc.

5 Comparison with the State-of-the-Art

We use two types of 2D joint detection data for evaluation: Cascaded Pyramid Network (CPN) detections and ground truth 2D keypoints, the results are shown in Table 4 and Table 5, respectively. Among the other methods, some use temporal information , some use additional data for training , and some use 3D pose scale in both training and testing . These results suggest that our approach outperforms the state-of-the-art.

Table 5 show that when given precise 2D joint information, the performance improvement of our model is significant, outperforming other GCN-based methods by a large margin. Therefore, we believe that in combination with methods that refine detected 2D joint information, such as deep kinematic analysis , the performance of our method on noisy 2D joints can be enhanced even more. Compared to the second-place method (4.38M), our model uses fewer parameters (3.70M), showing that our architecture can extract more important features, which provides a better understanding of the human pose.

To evaluate the generalization capabilities of our approach to domain shift, we apply the model trained on the Human3.6M to the MPI-INF-3DHP test set. Results are shown in Table 6. Although we train the model using only the Human3.6M, our approach outperforms the others, indicating that our architecture has strong generalization capabilities to unseen datasets.

The qualitative results of our method are shown in Fig. 5.

Conclusions

We present a novel architecture for 2D-to-3D human pose estimation, the Graph Stacked Hourglass Networks (GraphSH). With our unique skeletal pooling and skeletal unpooling scheme, together with the proposed architecture which has powerful multi-scale and multi-level feature extraction capabilities on graph-structured data, our method achieves accurate 2D-to-3D human pose estimation outperforming the state-of-the-art. As future work, we hope to introduce temporal multi-frame features in our architecture for further improvement.

References