3D Hand Shape and Pose Estimation from a Single RGB Image

Liuhao Ge, Zhou Ren, Yuncheng Li, Zehao Xue, Yingying Wang, Jianfei Cai, Junsong Yuan

Introduction

Vision-based 3D hand analysis is a very important topic because it has many applications in virtual reality (VR) and augmented reality (AR). However, despite years of studies , it remains an open problem due to the diversity and complexity of hand shape, pose, gesture, occlusion, etc. In the past decade, we have witnessed a rapid advance in 3D hand pose estimation from depth images . Considering RGB cameras are more widely available than depth cameras, some recent works start looking into 3D hand analysis from monocular RGB images, and mainly focus on estimating sparse 3D hand joint locations but ignore dense 3D hand shape . However, many immersive VR and AR applications often require accurate estimation of both 3D hand pose and 3D hand shape.

This motivates us to bring out a more challenging task: how to jointly estimate not only the 3D hand joint locations, but also the full 3D mesh of hand surface from a single RGB image? In this work, we develop a sound solution to this task, as illustrated in Fig. 1.

The task of single-view 3D hand shape estimation has been studied previously, but mostly in controlled settings, where a depth sensor is available. The basic idea is to fit a generative 3D hand model to the input depth image with iterative optimization . In contrast, here we consider to estimate 3D hand shape from a monocular RGB image, which has not been extensively studied yet. The absence of explicit depth cues in RGB images makes this task difficult to be solved by iterative optimization approaches. In this work, we apply deep neural networks that are trained in an end-to-end manner to recover 3D hand mesh directly from a single RGB image. Specifically, we predefine the topology of a triangle mesh representing the hand surface, and aim at estimating the 3D coordinates of all the vertices in the mesh using deep neural networks. To achieve this goal, there are several challenges.

The first challenge is the high dimensionality of the output space for 3D hand mesh generation. Compared with estimating sparse 3D joint locations of the hand skeleton (e.g., 21 joints), it is much more difficult to estimate 3D coordinates of dense mesh vertices (e.g., 1280 vertices) using conventional CNNs. One straightforward solution is to follow the common approach used in human body shape estimation , namely to regress low-dimensional parameters of a predefined deformable hand model, e.g., MANO .

In this paper we argue that the output 3D hand mesh vertices in essence are graph-structured data, since a 3D mesh can be easily represented as a graph. To output such graph-structured data and better exploit the topological relationship among mesh vertices in the graph, motivated by recent works on Graph CNNs , we propose a novel Graph CNN-based approach. Specifically, we adopt graph convolutions hierarchically with upsampling and nonlinear activations to generate 3D hand mesh vertices in a graph from image features which are extracted by backbone networks. With such an end-to-end trainable framework, our Graph CNN-based method can better represent the highly variable 3D hand shapes, and can better express the local details of 3D hand shapes.

Besides the computational model, an additional challenge is the lack of ground truth 3D hand mesh training data for real-world images. Manually annotating the ground truth 3D hand meshes on real-world RGB images is extremely laborious and time-consuming. We thus choose to create a large-scale synthetic dataset containing the ground truth of both 3D hand mesh and 3D hand pose for training. However, models trained on the synthetic dataset usually produce unsatisfactory estimation results on real-world datasets due to the domain gap between them. To address this issue, inspired by , we propose a novel weakly-supervised method by leveraging depth map as a weak supervision for 3D mesh generation, since depth map can be easily captured by an RGB-D camera when collecting real-world training data. More specifically, when fine-tuning on real-world datasets, we render the generated 3D hand mesh to a depth map on the image plane and minimize the depth map loss against the reference depth map, as shown in Fig. 3. Note that, during testing, we only need an RGB image as input to estimate full 3D hand shape and pose.

To the best of our knowledge, we are the first to handle the problem of estimating not only 3D hand pose but also full 3D hand shape from a single RGB image. Our main contributions are summarized as follows:

We propose a novel end-to-end trainable hand mesh generation approach based on Graph CNN . Experiments show that our method can well represent hand shape variations and capture local details. Furthermore, we observe that by estimating full 3D hand mesh, our method boost the accuracy performance of 3D hand pose estimation, as validated in Sec. 5.4.

We propose a weakly-supervised training pipeline on real-world dataset, by rendering the generated 3D mesh to a depth map on the image plane and leveraging the reference depth map as a weak supervision, without requiring any annotations of 3D hand mesh or 3D hand pose for real-world images.

We introduce the first large-scale synthetic RGB-based 3D hand shape and pose dataset as well as a small-scale real-world dataset, which contain the annotation of both 3D hand joint locations and the full 3D meshes of hand surface. We will share our datasets publicly upon the acceptance of this work.

We conduct comprehensive experiments on our proposed synthetic and real-world datasets as well as two public datasets . Experimental results show that our proposed method can produce accurate and reasonable 3D hand mesh with real-time speed on GPU, and can achieve superior accuracy performance on 3D hand pose estimation when compared with state-of-the-art methods.

Related Work

3D hand shape and pose estimation from depth images: Most previous methods estimate 3D hand shape and pose from depth images by fitting a deformable hand model to the input depth map with iterative optimization . A recent method was proposed to estimate pose and shape parameters from the depth image using CNNs, and recover 3D hand meshes using LBS. The CNNs are trained in an end-to-end manner with mesh and pose losses. However, the quality of their recovered hand meshes is restricted by their simple LBS model.

3D hand pose estimation from RGB images: Pioneering works estimate hand pose from RGB image sequences. Gorce et al. proposed estimating 3D hand pose, the hand texture and the illuminant dynamically through minimization of an objective function. Sridhar et al. adopted multi-view RGB images and depth data to estimate the 3D hand pose by combining a discriminative method with local optimization. With the advance of deep learning and the wide applications of monocular RGB cameras, many recent works estimate 3D hand pose from a single RGB image using deep neural networks . However, few works focus on 3D hand shape estimation from RGB images. Panteleris et al. proposed to fit a 3D hand model to the estimated 2D joint locations. But the hand model is controlled by 27 hand pose parameters, thus it cannot well adapt to various hand shapes. In addition, this method is not an end-to-end framework for generating 3D hand mesh.

3D human body shape and pose estimation from a single RGB image: Most recent methods rely on SMPL, a body shape and pose model . Some methods fit the SMPL model to the detected 2D keypoints . Some methods regress SMPL parameters using CNNs with supervisions of silhouette and/or 2D keypoints . A more recent method predicts a volumetric representation of human body. Different from these methods, we propose to estimate 3D mesh vertices using Graph CNNs in order to learn nonlinear hand shape variations and better utilize the relationship among vertices in the mesh topology. In addition, instead of using 2D silhouette or 2D keypoints to weakly supervise the network training, we propose to leverage the depth map as a weak 3D supervision when training on real-world datasets without 3D mesh or 3D pose annotations.

3D Hand Shape and Pose Dataset Creation

Manually annotating the ground truth of 3D hand meshes and 3D hand joint locations for real-world RGB images is extremely laborious and time-consuming. To overcome the difficulties in real-world data annotation, some works have adopted synthetically generated hand RGB images for training. However, existing hand RGB image datasets only provide the annotations of 2D/3D hand joint locations, and they do not contain any 3D hand shape annotations. Thus, these datasets are not suitable for the training of the 3D hand shape estimation task.

In this work, we create a large-scale synthetic hand shape and pose dataset that provides the annotations of both 3D hand joint locations and full 3D hand meshes. In particular, we use Maya to create a 3D hand model and rig it with joints, and then apply photorealistic textures on it as well as natural lighting using High-Dynamic-Range (HDR) images. We model hand variations by creating blend shapes with different shapes and ratios, then applying random weights on the blend shapes. To fully explore the pose space, we create hand poses from 500 common hand gestures and 1000 unique camera viewpoints. To simulate real-world diversity, we use 30 lightings and five skin colors. We render the hand using global illumination with off-the-shelf Arnold renderer . The rendering tasks are distributed onto a cloud render farm for maximum efficiency. In total, our synthetic dataset contains 375,000 hand RGB images with large variations. We use 315,000 images for training and 60,000 images for validation. During training, we randomly sample and crop background images from COCO , LSUN , and Flickr datasets, and blend them with the rendered hand images, as shown in Fig. 2.

In addition, to quantitatively evaluate the performance of hand mesh estimation on real-world image, we create a real-world dataset containing 583 hand RGB images with the annotations of 3D hand mesh and 3D hand joint locations. To facilitate the 3D annotation, we capture the corresponding depth images using an Intel RealSense RGB-D camera and manually adjust the 3D hand model in Maya with the reference of both RGB images and depth points. In this work, this real-world dataset is only used for evaluation.

Methodology

We propose to generate a full 3D mesh of the hand surface and the 3D hand joint locations directly from a single monocular RGB image, as illustrated in Fig. 3. Specifically, the input is a single RGB image centered on a hand, which is passed through a two-stacked hourglass network to infer 2D heat-maps. The estimated 2D heat-maps, combined with the image feature maps, are encoded as a latent feature vector by using a residual network that contains eight residual layers and four max pooling layers. The encoded latent feature vector is then input to a Graph CNN to infer the 3D coordinates of N{N} vertices V={vi}i=1N{\mathcal{V}=\left\{{{{\bm{v}}_{i}}}\right\}_{i=1}^{N}} in the 3D hand mesh. The 3D hand joint locations Φ={ϕj}j=1J{{\bf{\Phi}}=\left\{{{{\bm{\phi}}_{j}}}\right\}_{j=1}^{J}} are linearly regressed from the reconstructed 3D hand mesh vertices by using a simplified linear Graph CNN.

In this work, we first train the network models on a synthetic dataset and then fine-tune them on real-world datasets. On the synthetic dataset that contains the ground truth of 3D hand meshes and 3D hand joint locations, we train the networks end-to-end in a fully-supervised manner by using 2D heat-map loss, 3D mesh loss, and 3D pose loss. More details will be presented in Section 4.3. On the real-world dataset, the networks can be fine-tuned in a weakly-supervised manner without requiring the ground truth of 3D hand meshes or 3D hand joint locations. To achieve this target, we leverage the reference depth map available in training, which can be easily captured from a depth camera, as a weak supervision during the fine-tuning, and employ a differentiable renderer to render the generated 3D mesh to a depth map from the camera viewpoint. To guarantee the mesh quality, we generate the pseudo-ground truth mesh from the pretrained model as an additional supervision. More details will be presented in Section 4.4.

2 Graph CNNs for Mesh and Pose Estimation

Graph CNNs have been successfully applied in modeling graph structured data . As 3D hand mesh is of graph structure by nature, in this work we adopt the Chebyshev Spectral Graph CNN to generate 3D coordinates of vertices in the hand mesh and estimate 3D hand pose from the generated mesh.

A 3D mesh can be represented by an undirected graph M=(V,E,W){\mathcal{M}=\left({\mathcal{V},\mathcal{E},{W}}\right)}, where V={vi}i=1N{\mathcal{V}=\left\{{{{\bm{v}}_{i}}}\right\}_{i=1}^{N}} is a set of N{N} vertices in the mesh, E={ei}i=1E{\mathcal{E}=\left\{{{{\bm{e}}_{i}}}\right\}_{i=1}^{E}} is a set of E{E} edges in the mesh, W=(wij)N×N{{{W}}={\left({{w_{ij}}}\right)_{N\times N}}} is the adjacency matrix, where wij=0{w_{ij}=0} if (i,j)∉E{\left({i,j}\right)\notin\mathcal{E}}, and wij=1{w_{ij}=1} if (i,j)∈E{\left({i,j}\right)\in\mathcal{E}}. The normalized graph Laplacian is computed as L=IN−D−1/122WD−1/122{L={I_{N}}-{D^{-{1\mathord{\left/{\vphantom{12}}\right.\kern-1.2pt}2}}}W{D^{-{1\mathord{\left/{\vphantom{12}}\right.\kern-1.2pt}2}}}}, where D=diag(∑jwij){D={\rm{diag}}\left({\sum\nolimits_{j}{{w_{ij}}}}\right)} is the diagonal degree matrix, IN{I_{N}} is the identity matrix. Here, we assume that the topology of the triangular mesh is fixed and is predefined by the hand mesh model, i.e., the adjacency matrix W{W} and the graph Laplacian L{L} of the graph M{\mathcal{M}} are fixed during training and testing.

In this work, we design a hierarchical architecture for mesh generation by performing graph convolution on graphs from coarse to fine, as shown in Fig. 4. The topologies of coarse graphs are precomputed by graph coarsening, as shown in Fig. 5 (a), and are fixed during training and testing. Following Defferrard et al. , we use the Graclus multilevel clustering algorithm to coarsen the graph, and create a tree structure to store correspondences of vertices in graphs at adjacent coarsening levels. During the forward propagation, we upsample features of vertices in the coarse graph to corresponding children vertices in the fine graph, as shown in Fig. 5 (b). Then, we perform the graph convolution to update features in the graph. All the graph convolutional filters have the same support of K=3{K=3}. To make the network output irrelevant to the camera intrinsic parameters, we design the network to output UV coordinates on input image and depth of vertices in the mesh, which can be converted to 3D coordinates in the camera coordinate system using the camera intrinsic matrix. Similar to , we estimate scale-invariant and root-relative depth of mesh vertices.

Considering that 3D joint locations can be estimated directly from the 3D mesh vertices using a linear regressor , we adopt a simplified Graph CNN with two pooling layers and without nonlinear activation to linearly regress the scale-invariant and root-relative 3D hand joint locations from 3D coordinates of hand mesh vertices.

3 Fully-supervised Training on Synthetic Dataset

We first train the networks on our synthetic hand shape and pose dataset in a fully-supervised manner. As shown in Fig. 3 (a), the networks are supervised by heat-map loss LH{{\mathcal{L}}_{\mathcal{H}}}, mesh loss LM{{\mathcal{L}}_{\mathcal{M}}}, and 3D pose loss LJ{{\mathcal{L}}_{\mathcal{J}}}.

Heat-map Loss. LH=∑j=1J∥Hj−H^j∥22{{{\mathcal{L}}_{\mathcal{H}}}=\sum\nolimits_{j=1}^{J}{\left\|{{\mathcal{H}}_{j}-{{\hat{\mathcal{H}}}_{j}}}\right\|_{2}^{2}}}, where Hj{{\mathcal{H}}_{j}} and H^j{\hat{\mathcal{H}}_{j}} are the ground truth and estimated heat-maps, respectively. We set the heat-map resolution as 64×\times64 px. The ground truth heat-map is defined as a 2D Gaussian with a standard deviation of 4 px centered on the ground truth 2D joint location.

Mesh Loss. Similar to , LM=λvLv+λnLn+λeLe+λlLl{{\mathcal{L}}_{\mathcal{M}}}={\lambda_{v}}{{\mathcal{L}}_{v}}+{\lambda_{n}}{{\mathcal{L}}_{n}}+{\lambda_{e}}{{\mathcal{L}}_{e}}+{\lambda_{l}}{{\mathcal{L}}_{l}} is composed of vertex loss Lv{{\mathcal{L}}_{v}}, normal loss Ln{{\mathcal{L}}_{n}}, edge loss Le{{\mathcal{L}}_{e}}, and Laplacian loss Ll{{\mathcal{L}}_{l}}. The vertex loss Lv{{\mathcal{L}}_{v}} is to constrain 2D and 3D locations of mesh vertices:

where vi{{\bm{v}}_{i}} and v^i{\hat{\bm{v}}_{i}} denote the ground truth and estimated 2D/3D locations of the mesh vertices, respectively. The normal loss Ln{{\mathcal{L}}_{n}} is to enforce surface normal consistency:

where t{t} is the index of triangle faces in the mesh; (i,j){\left({i,j}\right)} are the indices of vertices that compose one edge of triangle t{t}; and nt{{\bm{n}}_{t}} is the ground truth normal vector of triangle face t{t}, which is computed from ground truth vertices. The edge loss Le{{\mathcal{L}}_{e}} is introduced to enforce edge length consistency:

where ei{{\bm{e}}_{i}} and e^i{\hat{\bm{e}}_{i}} denote the ground truth and estimated edge vectors, respectively. The Laplacian loss Ll{{\mathcal{L}}_{l}} is introduced to preserve the local surface smoothness of mesh:

where δi=vi3D−v^i3D{{{\bm{\delta}}_{i}}={\bm{v}}_{i}^{3D}-\hat{\bm{v}}_{i}^{3D}} is the offset from the estimation to the ground truth, N(vi){{\mathcal{N}}\left({{{\bm{v}}_{i}}}\right)} is the set of neighboring vertices of vi{{{\bm{v}}_{i}}}, and Bi{B_{i}} is the number of vertices in the set N(vi){{\mathcal{N}}\left({{{\bm{v}}_{i}}}\right)}. This loss function prevents the neighboring vertices from having opposite offsets, thus making the estimated 3D hand surface mesh smoother. For the hyperparameters, we set λv=1{\lambda_{v}=1}, λn=1{\lambda_{n}=1}, λe=1{\lambda_{e}=1}, λl=50{\lambda_{l}=50} in our implementation.

3D Pose Loss. LJ=∑j=1J∥ϕj3D−ϕ^j3D∥22{{{\mathcal{L}}_{\mathcal{J}}}=\sum\nolimits_{j=1}^{J}{\left\|{{\bm{\phi}}_{j}^{3D}-\hat{\bm{\phi}}_{j}^{3D}}\right\|_{2}^{2}}}, where ϕj3D{{\bm{\phi}}_{j}^{3D}} and ϕ^j3D{\hat{\bm{\phi}}_{j}^{3D}} are the ground truth and estimated 3D joint locations, respectively.

In our implementation, we first train the stacked hourglass network and the 3D pose regressor separately with the heat-map loss and the 3D pose loss, respectively. Then, we train the stacked hourglass network, the residual network and the Graph CNN for mesh generation with the combined loss Lfully{{\mathcal{L}}_{fully}}:

where λH=0.5{\lambda_{\mathcal{H}}=0.5}, λM=1{\lambda_{\mathcal{M}}=1}, λJ=1{\lambda_{\mathcal{J}}=1}.

4 Weakly-supervised Fine-tuning

On the real-world dataset, i.e., the Stereo Hand Pose Tracking Benchmark , there is no ground truth of 3D hand mesh. Thus, we fine-tune the networks in a weakly-supervised manner. Moreover, our model also supports the fine-tuning without the ground truth of 3D joint locations, which can further removes the burden of annotating 3D joint locations on training data and make it more applicable for large-scale real-world dataset.

Depth Map Loss. As shown in Fig. 3 (b), we leverage the reference depth map, which can be easily captured by a depth camera, as a weak supervision, and employ a differentiable renderer, similar to , to render the estimated 3D hand mesh to a depth map from the camera viewpoint. We use smooth L1 loss for the depth map loss:

where D{\mathcal{D}} and D^{\hat{\mathcal{D}}} denote the ground truth and rendered depth maps, respectively; R(⋅){\mathcal{R}\left(\cdot\right)} is the depth rendering function; M^{\hat{\mathcal{M}}} is the estimated 3D hand mesh. We set the resolution of a depth map as 32×\times32 px.

In our implementation, we first fine-tune the stacked hourglass network with the heat-map loss, and then end-to-end fine-tune all networks with the combined loss Lweakly{{\mathcal{L}}_{weakly}}:

where λH=0.1{\lambda_{\mathcal{H}}=0.1}, λD=0.1{\lambda_{\mathcal{D}}=0.1}, λpM=1{\lambda_{p{\mathcal{M}}}=1}. Note that Eq. 8 is the loss function for fine-tuning on the dataset without 3D pose supervision. When the ground truth of 3D joint locations is provided during training, we add the 3D pose loss LJ{{\mathcal{L}}_{\mathcal{J}}} in the loss function and set the weight λJ=10{\lambda_{\mathcal{J}}=10}.

Experiments

In this work, we evaluate our method on two aspects: 3D hand mesh reconstruction and 3D hand pose estimation.

For 3D hand mesh reconstruction, we evaluate the generated 3D hand meshes on our proposed synthetic and real-world datasets, which are introduced in Section 3, since no other hand RGB image dataset contains the ground truth of 3D hand meshes. We measure the average error in Euclidean space between the corresponding vertices in each generated 3D mesh and its ground truth 3D mesh. This metric is denoted as “mesh error” in the following experiments.

For 3D hand pose estimation, we evaluate our proposed methods on two publicly available datasets: Stereo Hand Pose Tracking Benchmark (STB) and the Rendered Hand Pose Dataset (RHD) . STB is a real-world dataset containing 18,000 images with the ground truth of 21 3D hand joint locations and corresponding depth images. Following , we split the dataset into 15,000 training samples and 3,000 test samples. To make the joint definition consistent with our settings and RHD dataset, following , we move the root joint location from palm center to wrist. RHD is a synthetic dataset containing 41,258 training images and 2,728 testing images. This dataset is challenging due to the large variations in viewpoints and the low image resolution. We evaluate the performance of 3D hand pose estimation with three metrics: (i) Pose error: the average error in Euclidean space between the estimated 3D joints and the ground truth joints; (ii) 3D PCK: the percentage of correct keypoints of which the Euclidean error distance is below a threshold; (iii) AUC: the area under the curve on PCK for different error thresholds.

We implement our method within the PyTorch framework. The networks are trained using the RMSprop optimizer with mini-batches of size 32. The learning rate is set as 10−3{10^{-3}} when pretraining on our synthetic dataset, and is set as 10−4{10^{-4}} when fine-tuning on RHD and STB . The input image is resized to 256×{\times}256 px. Following the same condition used in , we assume that the global hand scale and the absolute depth of root joint are provided at test time. The global hand scale is set as the length of the bone between MCP and PIP joints of the middle finger.

2 Ablation Study of Loss Terms

We first evaluate the impact of different losses used in the fully-supervised training (Eq. 6) on the performance of mesh reconstruction and pose estimation. We conduct this experiment on our synthetic dataset. As presented in Table 1, the model trained with the full loss achieves the best performance in both mesh reconstruction and pose estimation, which indicates that all the losses have contributions to producing accurate 3D hand mesh as well as 3D hand joint locations.

3 Evaluation of 3D Hand Mesh Reconstruction

We demonstrate the advantages of our proposed Graph CNN-based 3D hand mesh reconstruction method by comparing it with two baseline methods: direct Linear Blend Skinning (LBS) method and MANO-based method.

Direct LBS. We train the network to directly regress 3D hand joint locations from the heat-maps and the image features, which is similar to the network architecture proposed in . We generate the 3D hand mesh from only the estimated 3D hand joint locations by applying inverse kinematics and LBS with the predefined mesh model and skinning weights (see the supplementary for details). As shown in Table 2, the average mesh error of direct LBS method is worse than our method on both our synthetic dataset and our real-world dataset, since the LBS model for mesh generation is predefined and cannot be adapt to hands with different shapes. As can be seen in Fig. 7, the hand meshes generated by direct LBS method have unrealistic deformation at joints and suffer from serious inherent artifacts.

MANO-based Method. We also implement a MANO based method that regresses hand shape and pose parameters from the latent image features using three fully-connected layers. Then, the 3D hand mesh is generated from the estimated shape and pose parameters using MANO hand model (see the supplementary for details). The networks are trained in fully-supervised manner using the same loss functions as Eq. 6 on our synthetic dataset. For fair comparison, we align our hand mesh with the MANO hand mesh, and compute mesh error on the aligned mesh. As shown in Table 2 and Fig. 7, the MANO-based method exhibits inferior performance on mesh reconstruction compared with our method. Note that direct supervising MANO parameters on synthetic dataset may obtain better performance . But it is infeasible on our synthetic dataset since our dataset does not contain MANO parameters.

4 Evaluation of 3D Hand Pose Estimation

We also evaluate our approach on the task of 3D hand pose estimation.

Self-comparisons. We conduct self-comparisons on STB dataset by fine-tuning the networks pretrained on our synthetic dataset in a weakly-supervised manner, as described in Section 4.4. In Table 3, we compare our proposed weakly-supervised method (Full model) with two baselines: (i) Baseline 1: directly regressing 3D hand joint locations from the heat-maps and the feature maps without using the depth map loss during training; (ii) Baseline 2: regressing 3D hand joint locations from the estimated 3D hand mesh without using the depth map loss during training. As presented in Fig. 8, the estimation accuracy of Baseline 2 is superior to that of Baseline 1, which indicates that our proposed 3D hand mesh reconstruction network is beneficial to 3D hand pose estimation. Furthermore, the estimation accuracy of our full model is superior to that of Baseline 2, especially when fine-tuning without 3D hand pose supervision, which validates the effectiveness of introducing the depth map loss as a weak supervision.

In addition, to explore a more efficient way for 3D hand pose estimation without mesh generation, we directly regress the 3D hand joint locations from the latent feature extracted by our full model instead of regressing them from the 3D hand mesh (see the supplementary for details). This task transfer method is denoted as “Full model, task transfer” in Fig. 8. Although this method has the same pipeline as that of Baseline 1, the estimation accuracy of this task transfer method is better than that of Baseline 1 and is only a little bit worse than that of our full model, which indicates that the latent feature extracted by our full model is more discriminative and is easier to regress accurate 3D hand pose than the latent feature extracted by Baseline 1.

Comparisons with State-of-the-arts. We compare our method with state-of-the-art 3D hand pose estimation methods on RHD and STB datasets. The PCK curves over different error thresholds are presented in Fig. 9. On RHD dataset, as shown in Fig. 9 (left), our method outperforms the three state-of-the-art methods over all the error thresholds on this dataset. On STB dataset, when the 3D hand pose ground truth is given during training, we compare our methods with seven state-of-the-art methods , and our method outperforms these methods over most of the error thresholds, as shown in Fig. 9 (middle). We also experiment with the situation when 3D hand pose ground truth is unknown during training on STB dataset, and compare our method with the weakly-supervised method proposed by Cai et al. , both of which adopt reference depth maps as a weak supervision. As shown in Fig. 9 (right), our 3D mesh-based method outperforms Cai et al. by a large margin.

5 Runtime and Qualitative Results

Runtime. We evaluate the runtime of our method on one Nvidia GTX 1080 GPU. The runtime of our full model outputting both 3D hand mesh and 3D hand pose is 19.9ms on average, including 12.6ms for the stacked hourglass network forward propagation, 4.7ms for the residual network and Graph CNN forward propagation, and 2.6ms for the forward propagation of the pose regressor. Thus, our method can run in real-time on GPU at over 50fps.

Qualitative Results. Some qualitative results of 3D hand mesh reconstruction and 3D hand pose estimation for our synthetic dataset, our real-world dataset, RHD , and STB datasets are shown in Fig. 10. More qualitative results are presented in the supplementary.

Conclusion

In this paper we have tackled the challenging task of 3D hand shape and pose estimation from a single RGB image. We have developed a Graph CNN-based model to reconstruct a full 3D mesh of hand surface from an input RGB image. To train the model, we have created a large-scale synthetic RGB image dataset with ground truth annotations of both 3D joint locations and 3D hand meshes, on which we train our model in a fully-supervised manner. To fine-tune our model on real-world datasets without 3D ground truth, we render the generated 3D mesh to a depth map and leverage the observed depth map as a weak supervision. Experiments on our proposed new datasets and two public datasets show that our method can recover accurate 3D hand mesh and 3D joint locations in real-time.

In future work, we will use Mocap data to create a larger 3D hand pose and shape dataset. We will also consider the cases of hand-object and hand-hand interactions in order to make the hand pose and shape estimation more robust.

Acknowledgment: This work is in part supported by MoE Tier-2 Grant (2016-T2-2-065) of Singapore. This work is also supported in part by start-up grants from University at Buffalo and a gift grant from Snap Inc.

References

Supplementary

Appendix A Qualitative Results

We present more qualitative results of 3D hand mesh reconstruction and 3D hand pose estimation for our synthetic dataset, our real-world dataset, STB dataset , RHD dataset , and Dexter+Object dataset , as shown in Fig. 11. Please see the supplementary video for more qualitative results on continuous sequences.

Appendix B Details of Baseline Methods for 3D Hand Mesh Reconstruction

In Section 5.3 of our main paper, we compare our proposed method with two baseline methods for 3D hand mesh reconstruction: direct Linear Blend Skinning (LBS) method and MANO-based method. Here, we describe more details of these two baseline methods, as illustrated in Fig. 12.

In the direct LBS method, we train the network to regress 3D hand joint locations from the heat-maps and the image features with heat-map loss and 3D pose loss. As illustrated in Fig. 12 (b), the latent feature extracted from the input image is mapped to 3D hand joint locations through a multi-layer perceptron (MLP) network with three fully-connected layers. Then, we apply inverse kinematics (IK) to compute the transformation matrix of each hand joint from the the estimated 3D hand joint locations. The 3D hand mesh is generated by applying LBS with the predefined hand model and skinning weights. In this method, the 3D hand mesh is only determined by the estimated 3D hand joint locations, thus it cannot be adapted to various hand shapes. In addition, the IK often suffers from singularity and multiple solutions, which makes the solutions to transformation matrices unreliable. Experimental results in Figure 7 and Table 2 of our main paper have shown the limitations of this direct LBS method.

In the MANO-based method, we train the network to regress hand shape and pose parameters of the MANO hand model . As illustrated in Fig. 12 (c), the latent feature extracted from the input image is mapped to hand shape and pose parameters θ{\theta}, β{\beta} through an MLP network with three fully-connected layers. Then, the 3D hand mesh is generated from the regressed parameters θ{\theta}, β{\beta} using the MANO hand model . Note that the MANO mesh generation module is differentiable and is involved in the network training. The networks are trained with heat-map loss, mesh loss and 3D pose loss, which are the same as our method. Since the MANO hand model is fixed during training and is essentially LBS with blend shapes , the representation power of this method is limited. Experimental results in Figure 7 and Table 2 of our main paper have shown the limitations of this MANO-based method.

Appendix C Details of the Task Transfer Method

In Section 5.4 of our main paper, we implement an alternative method (“full model, task transfer”) for 3D hand pose estimation by transferring our full model trained for 3D hand mesh reconstruction to the task of 3D hand pose estimation. Here, we describe more details of our task transfer method. As illustrated in Fig. 13, we directly regress the 3D hand joint locations from the latent feature extracted by our full model using an MLP network with three fully-connected layers. We first train the MLP network with 3D pose loss on our synthetic dataset. When experimenting on STB dataset with 3D pose supervision, we fine-tune the MLP network with 3D pose loss. When experimenting on STB dataset without 3D pose supervision, we directly use the MLP network pretrained on our synthetic dataset. Experimental results in Figure 8 of our main paper show that our task transfer method is better than the baseline method which is only trained for 3D hand pose estimation, even though these two methods have the same pipeline. This indicates that the latent feature extracted by our full model is more discriminative and is easier to regress accurate 3D hand pose since our full model is trained with the dense supervision of the 3D hand mesh that contains richer information than the 3D hand pose. In addition, although the estimation accuracy of our task transfer method is a little bit worse than that of our full model, our task transfer method is faster than our full model, since it does not generate 3D hand mesh. The runtime of our task transfer method is 15.1ms, while the runtime of our full model which estimate 3D hand pose from hand mesh is 19.9ms. Thus, in applications that only require 3D hand pose estimation but not 3D hand shape estimation, we can choose to use this task transfer method, which can maintain a comparable accuracy as our full model while runs at faster speed.