Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals

Shanxin Yuan, Guillermo Garcia-Hernando, Bjorn Stenger, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee, Pavlo Molchanov, Jan Kautz, Sina Honari, Liuhao Ge, Junsong Yuan, Xinghao Chen, Guijin Wang, Fan Yang, Kai Akiyama, Yang Wu, Qingfu Wan, Meysam Madadi, Sergio Escalera, Shile Li, Dongheui Lee, Iason Oikonomidis, Antonis Argyros, Tae-Kyun Kim

Introduction

The field of 3D hand pose estimation has advanced rapidly, both in terms of accuracy and dataset quality . Most successful methods treat the estimation task as a learning problem, using random forests or convolutional neural networks (CNNs). However, a review from 2015 surprisingly concluded that a simple nearest-neighbor baseline outperforms most existing systems. It concluded that most systems do not generalize beyond their training sets , highlighting the need for more and better data. Manually labeled datasets such as contain just a few thousand examples, making them unsuitable for large-scale training. Semi-automatic annotation methods, which combine manual annotation with tracking, help scaling the dataset size , but in the case of the annotation errors are close to the lowest estimation errors. Synthetic data generation solves the scaling issue, but has not yet closed the realism gap, leading to some kinematically implausible poses .

A recent study confirmed that cross-benchmark testing is poor due to different capture set-ups and annotation methods . It showed that training a standard CNN on a million-scale dataset achieves state-of-the-art results. However, the estimation accuracy is not uniform, highlighting the well-known challenges of the task: variations in view point and hand shape, self-occlusion, and occlusion caused by objects being handled.

Public benchmarks and challenges in other areas such as ImageNet for scene classification and object detection, PASCAL for semantic and object segmentation, and the VOT challenge for object tracking, have been instrumental in driving progress in their respective field. In the area of hand tracking, the review from 2007 by Erol et al. proposed a taxonomy of approaches. Learning-based approaches have been found effective for solving single-frame pose estimation, optionally in combination with hand model fitting for higher precision, e.g., . The review by Supancic et al. compared 13 methods on a new dataset and concluded that deep models are well-suited to pose estimation . It also highlighted the need for large-scale training sets in order to train models that generalize well. In this paper we extend the scope of previous analyses by comparing deep learning methods on a large-scale dataset, carrying out a fine-grained analysis of error sources and different design choices.

Evaluation tasks

We evaluate three different tasks on a dataset containing over a million annotated images using standardized evaluation protocols. Benchmark images are sampled from two datasets: BigHand2.2M and First-Person Hand Action dataset (FHAD) . Images from BigHand2.2M cover a large range of hand view points (including third-person and first-person views), articulated poses, and hand shapes. Sequences from the FHAD dataset are used to evaluate pose estimation during hand-object interaction. Both datasets contain 640×480640\times 480-pixel depth maps with 21 joint annotations, obtained from magnetic sensors and inverse kinematics. The 2D bounding boxes have an average diagonal length of 162.4 pixels with a standard deviation of 40.7 pixels. The evaluation tasks are 3D single hand pose estimation, i.e., estimating the 3D locations of 21 joints, from (1) individual frames, (2) video sequences, given the pose in the first frame, and (3) frames with object interaction, e.g., with a juice bottle, a salt shaker, or a milk carton. See Figure 1 for an overview. Bounding boxes are provided as input for tasks (1) and (3). The training data is sampled from the BigHand2.2M dataset and only the interaction task uses test data from the FHAD dataset. See Table 1 for dataset sizes and the number of total and unseen subjects for each task.

Evaluated methods

We evaluate the top 10 among 17 participating methods . Table 2 lists the methods with some of their key properties. We also indirectly evaluate DeepPrior and REN , which are components of rvhand , as well as DeepModel , which is the backbone of LSL . We group methods based on different design choices.

2D CNN vs. 3D CNN. 2D CNNs have been popular for 3D hand pose estimation . Common pre-processing steps include cropping and resizing the hand volume by normalizing the depth values to . Recently, several methods have used a 3D CNN , where the input can be a 3D voxel grid , or a projective D-TSDF volume . Ge et al. project the depth image onto three orthogonal planes and train a 2D CNN for each projection, then fusing the results. In they propose a 3D CNN by replacing 2D projections with a 3D volumetric representation (projective D-TSDF volumes ). In the HIM2017 challenge , they apply a 3D deep learning method , where the inputs are 3D points and surface normals. Moon et al. propose a 3D CNN to estimate per-voxel likelihoods for each hand joint. NAIST_RV proposes a 3D CNN with a hierarchical branch structure, where the input is a 50350^{3}-voxel grid.

Detection-based vs. Regression-based. Detection-based methods produce a probability density map for each joint. The method of RCN-3D is an RCN+ network , based on Recombinator Networks (RCN) with 17 layers and 64 output feature maps for all layers except the last one, which outputs a probability density map for each of the 21 joints. V2V-PoseNet uses a 3D CNN to estimate per-voxel likelihood of each joint, and a CNN to estimate the center of mass from the cropped depth map. For training, 3D likelihood volumes are generated by placing normal distributions at the locations of hand joints. Regression-based methods directly map the depth image to the joint locations or the joint angles of a hand model . rvhand combines ResNet , Region Ensemble Network (REN) , and DeepPrior to directly estimate the joint locations. LSL uses one network to estimate a global scale factor and a second network to estimate all joint angles, which are fed into a forward kinematic layer to estimate the hand joints.

Hierarchical models divide the pose estimation problem into sub-tasks . The evaluated methods divide the hand joints either by finger , or by joint type . mmadadi designs a hierarchically structured CNN, dividing the convolution+ReLU+pooling blocks into six branches (one per finger with palm and one for palm orientation), each of which is then followed by a fully connected layer. The final layers of all branches are concatenated into one layer to predict all joints. NAIST_RV chooses a similar hierarchical structure of a 3D CNN, but uses five branches, each to predict one finger and the palm. THU_VCLab , rvhand , and REN apply constraints per finger and joint-type (across fingers) in their multiple regions extraction step, each region containing a subset of joints. All regions are concatenated in the last fully connected layers to estimate the hand pose.

Structured methods embed physical hand motion constraints into the model . Structural constraints are included in the CNN model or in the loss function . DeepPrior learns a prior model and integrates it into the network by introducing a bottleneck in the last CNN layer. LSL uses prior knowledge in DeepModel by embedding a kinematic model layer into the CNN and using a fixed hand model. mmadadi includes the structure constraints in the loss function, which incorporates physical constraints about natural hand motion and deformation. strawberryfg applies a structure-aware regression approach, Compositional Pose Regression , and replaces the original ResNet-50 with ResNet-152. It uses phalanges instead of joints for representing pose, and defines a loss function that encodes long-range interaction between the phalanges.

Multi-stage methods propagate results from each stage to enhance the training of the subsequent stages . THU_VCLab uses REN to predict an initial hand pose. In the following stages, feature maps are computed with the guidance of the hand pose estimate in the previous stage. RCN-3D has five stages: (1) 2D landmark estimation using an RCN+ network , (2) estimation of corresponding depth values by multiplying probability density maps with the input depth image, (3) inverse perspective projection of the depth map to 3D, (4) error compensation for occlusions and depth errors (a 3-layer network of residual blocks) and, (5) error compensation for noise (another 3-layer network of residual blocks).

Residual networks. ResNet is adopted by several methods . V2V-PoseNet uses residual blocks as main building blocks. strawberryfg implements the Compositional Pose Regression method by using ResNet-152 as basic network. RCN-3D uses two small residual blocks in its fourth and fifth stage.

Results

The aim of this evaluation is to identify success cases and failure modes. We use both standard error metrics and new proposed metrics to provide further insights. We consider joint visibility, seen vs. unseen subjects, hand view point distribution, articulation distribution, and per-joint accuracy.

We evaluate ten state-of-the-art methods (Table 2) directly and three methods indirectly, which were used as components of others, DeepPrior , REN , and DeepModel . Figure 2 shows the results in terms of two metrics: (top) the proportion of frames in which all joint errors are below a threshold and (bottom) the total proportion of joints below an error threshold .

1.2 Analysis by occlusion and unknown subject

Average error for four cases: To analyze the results with respect to joint visibility and hand shape, we partition the joints into four groups, based on whether or not they are visible, and whether or not the subject was seen at training time. Different hand shapes and joint occlusions are responsible for a large proportion of errors, see Table 3. The error for unseen subjects is significantly larger than for seen subjects. Moreover, the error for visible joints is smaller than for occluded joints. Based on the first group (visible, seen), we carry out a best-case performance estimate for the current state-of-the-art. For each frame of seen subjects, we first choose the best result from all methods, and calculate the success rate based on the average error for each frame, see the black curve in Figure 4 (top-middle).

2D vs. 3D CNNs: We compare two hierarchical methods with similar structure but different representation. The bottom-left plot of Figure 4 shows mmadadi , which employs a 2D CNN, and NAIST_RV , using a 3D CNN. mmadadi and NAIST_RV have almost the same structure, but NAIST_RV uses a 3D CNN, while mmadadi uses a 2D CNN. NAIST_RV outperforms mmadadi in all four cases.

Detection-based vs. regression-based methods: We compare the average of the top two detection-based methods with the average of the top two regression-based methods. In all four cases, detection-based methods outperform regression-based ones, see the top-right plot of Figure 4. In the challenge, the top two methods are detection-based methods, see Table 2. Note that a similar trend can be seen in the field of full human pose estimation, where only one method in a recent challenge was regression-based .

Single- vs. multi-stage methods: Cascaded methods work better than single-stage methods, see the bottom-right plot of Figure 4. Compared to other methods, rvhand and THU_VCLab both embed structural constraints, employing REN as their basic structure. THU_VCLab takes a cascaded approach to iteratively update results from previous stages, outperforming rvhand .

1.3 Analysis by number of occluded joints

Most frames contain joint occlusions, see Figure 5 (top). We assume that a visible joint lies within a small range of the 3D point cloud. We therefore detect joint occlusion by thresholding the distance between the joint’s depth annotation value and its re-projected depth value. As shown in Figure 5 (bottom), the average error decreases nearly monotonously for increasing numbers of visible joints.

1.4 Analysis based on view point

1.5 Analysis based on articulation

1.6 Analysis by joint type

2 Hand pose tracking

In this task we evaluate three state-of-the-art methods, see Table 4 and Figure 8. Discriminative methods break tracking into two sub-tasks: detection and hand pose estimation, sometimes merging the sub-tasks . Based on the detection methods, 3D hand pose estimation can be grouped into pure tracking , tracking-by-detection , and a combination of tracking and re-initialization , see Table 4.

Pure tracking: RCN-3D_track estimate the bounding box location by scanning windows based on the result in the previous frame, including a motion estimate. Hand pose within the bounding box is estimated using RCN-3D .

Tracking-by-detection: NAIST_RV_track is a tracking-by-detection method with three components: hand detector, hand verifier, and pose estimator. The hand detector is built on U-net to predict a binary hand-mask, which, after verification, is passed to the pose estimator NAIST_RV . If verification fails, the result from the previous frame is chosen.

Hybrid tracking and detection: THU_VCLab_track makes use of the previous tracking result and the current frame’s scanning window. The hand pose of the previous frame is used as a guide to predict the hand pose in the current frame. The previous frame’s bounding box is used for the current frame. During fast hand motion, Faster R-CNN is used for re-initialization.

Detection accuracy: We first evaluate the detection accuracy by evaluating the bounding box overlap, i.e., the intersection over union (IoU) of the detection and ground truth bounding boxes, see Figure 8 (middle). Overall, RCN-3D_track is more accurate than THU_VCLab_track, which itself outperforms NAIST_RV_track. Pure detection methods have a larger number of false negatives, especially when multiple hands appear in the scene, see Figure 8 (left). There are 72 and 174 missed detections (IoU of zero), for NAIST_RV_track and THU_VCLab_track, respectively. By tracking and re-initializing, THU_VCLab_track achieves better detection accuracy overall. RCN-3D_track, using motion estimation and single-frame hand pose estimation, shows the lowest error.

Tracking accuracy is shown in Figure 8 (right). Even through THU_VCLab performs better than NAIST_RV in the Single frame pose estimation task, NAIST_RV_track performs better on the tracking tasks due to per-frame hand detection.

3 Hand object interaction

Current state-of-the-art methods have difficulty generalizing to the hand-object interaction scenario. However, NAIST_RV_obj and rvhand_obj show similar performance for visible joints and occluded joints, indicating that CNN-based segmentation can better preserve structure than image processing operations, see the middle plot of Figure 9.

Discussion and conclusions

The analysis of the top 10 among 17 participating methods from the HIM2017 challenge provides insights into the current state of 3D hand pose estimation.

(1) 3D volumetric representations used with a 3D CNN show high performance, possibly by better capturing the spatial structure of the input depth data.

(2) Detection-based methods tend to outperform regression-based methods, however, regression-based methods can achieve good performance using explicit spatial constraints. Making use of richer spatial models, e.g., bone structure , helps further. Regression-based methods perform better in extreme view point cases , where severe occlusion occurs.

(3) While joint occlusions pose a challenge for most methods, explicit modeling of structure constraints and spatial relation between joints can significantly narrow the gap between errors on visible and occluded joints .

(4) Discriminative methods still generalize poorly to unseen hand shapes. Data augmentation and scale estimation methods model only global shape changes, but not local variations. Integrating hand models with better generative capability may be a promising direction.

(6) In hand tracking, current discriminative methods divide the problem into two sub-tasks: detection and pose estimation, without using the hand shape provided in the first frame. Hybrid methods may work better by using the provided hand shape.

(7) Current methods perform well on single hand pose estimation when trained on a million-scale dataset, but have difficulty in generalizing to hand-object interaction. Two directions seem promising, (a) designing better hand segmentation methods, and (b) training the model with large datasets containing hand-object interaction.

Acknowledgement: This work was partially supported by Huawei Technologies.

References