FrankMocap: Fast Monocular 3D Hand and Body Motion Capture by Regression and Integration

Yu Rong, Takaaki Shiratori, Hanbyul Joo

Introduction

Billions of daily human activities are being recorded as videos and uploaded to public internet websites, capturing extremely diverse human behaviors in various real-world scenarios. A technology that can digitize the human motions from these videos has enormous potentials in various applications including human-computer interaction, social artificial intelligent, and robotics. A motion capture system with a commodity camera and reduced computation would enable people to make full use of such applications.

While the hands and body are equally important for human motion understanding toward these applications, the hands are physically small body parts, making it difficult to capture the motion of the hands and body jointly even with a professional motion capture system. The same is true for recent work on 3D body pose (i.e., torso and limbs) estimation from a single RGB image . Although the accuracy improvement of body pose estimation is significant, subtle finger gestures are ignored, losing the original nuance of the human motion. Similarly, there have been noticeable achievements in 3D hand pose estimation from a depth input or a single RGB image . However, these approaches are often demonstrated with hand-specific camera views, rather than the more challenging in-the-wild scenarios where a camera is capturing whole bodies of people and the hands tend to be in low resolutions with frequent motion blurs. A few recent approaches aim to capture 3D motions of the whole body (i.e., the body and hands) by leveraging the 3D parametric models that can express both hands and body . However, these approaches rely on optimization techniques to fit the parametric models to image measurements, which are relatively slow and not suitable for real-time applications.

In this paper, we present a fast and accurate motion capture method to estimate both 3D body and hand poses from monocular RGB images or videos, as shown in Figure LABEL:fig:teaser. Our method consists of two regression modules that predict 3D poses of the body and hands individually from a single RGB image input, followed by an integration module that produces the whole body pose from the outputs from the body and hand modules We use whole body to represent all the body parts including the fingers, while we use body to represent the torso and limbs excluding the fingers.. A main idea of our approach is to make the outputs from body module and hand module as compatible as possible, enabling us to efficiently integrate the outputs for whole body motion capture. To make this integration process tractable, we employ the SMPL-X model that has a unified skeleton and mesh representation for both the body and hands. Based on that, the body module and hand module contribute the different part of the same output structures. Inspired by the achievements of the recent work on 3D body motion estimation , the hand and body modules are designed based on deep neural network techniques, and directly regress 3D poses from an input single RGB image. Given the 3D body and hand pose estimations from the hand and body modules, the integration can be performed by direct copy-and-paste of the predicted body and hand poses to the SMPL-X model, achieving a near real-time performance for whole body 3D motion capture (∼\sim9.5 fps). A further improvement for better estimate of the whole body pose can be achieved via the optimization framework we also present. Follow the similar spirit of Joo et al. , we call our method as FrankMocap to represent this regression and integration manner.

We demonstrate the fast and accurate performance of FrankMocap on various real-world monocular videos, including a real-time demo. Notably, our 3D hand pose estimation outperforms previous approaches in public benchmarks. We also present thorough ablation studies to demonstrate the advantage of our method, and compare our method with previous hand motion capture methods and whole body motion capture methods.

Related Work

3D Parametric Human Body Models. 3D Parametric human body models are widely used for markerless motion capture, to model the deformation of 3D human (including body, face and hands) via low dimension parameters instead of the original vertices . The SCAPE is a pioneering work that accounts for shape variations and pose deformations . After that, Loper et al. introduce the SMPL that learns local pose-dependent blendshape on top of linear blend skinning for holistic mesh deformation as well as shape variations. Later, Romero et al. extend the SMPL to a hand and introduce the hand deformation model called MANO. They also develop a unified body and hand model called SMPL+H. Joo et al. build a unified model of body, hands and face, called Adam, and use it to achieve whole body motion capture of face, body and hands from a multiview setup . Pavlakos et al. also similarly develop a unified model of body, hands and face, called SMPL-X, in which all template body parts are designed by artists for consistent quality over the body parts, and learned deformation statistics.

Single Image 3D Body Pose Estimation. Many monocular 3D body pose estimation approaches consider to predict 3D body keypoint locations from single images . As the major limitations, the output of these methods cannot be directly used for graphics applications, since 3D joint angles are missing and the lengths of parts are not preserved. More recent monocular 3D body pose estimation approaches adopt parametric 3D human model such as SMPL or Adam model for 3D body representation. The use of 3D parametric models allows them to reconstruct a 3D body pose by fitting the 3D body model to 2D observations, such as 2D keypoints, via optimization framework . More recent work leverages the deep learning framework to directly regress parameters of a body model.

There are also approaches that use a hybrid framework by using a deep learning framework to produce an intermediate representation such as 2D and depth heat maps and fitting a skeletal model on these outputs to reconstruct joint angles . Various types of inputs are considered in these deep network approaches, including single RGB images , 2D keypoint heatmaps , body part segmentation or densepose maps . Due to the lack of training data with 3D annotations, these models are trained with mixed datasets including indoor datasets such as Human3.6M and in-the-wild datasets such as COCO or 3DPW . Most papers in this area such as HMR and SPIN use single images as input. There are some other works takes sequences as input. The representative ones are Zhang et al. and VIBE . In this paper, we mainly focus on processing single images and applying simple temporal smoothness to handle sequence inputs.

Single Image 3D Hand Pose Estimation. Previous works on 3D hand joint estimation takes depth images as input . Although achieving good performance, these methods cannot be easily applied to in-the-wild RGB images and videos. Recent work begins to use single RGB images as input . These approaches focus more on 3D hand joints location estimation instead of joint angles. Inspired by recent success in 3D body motion capture , there are several methods on single image 3D hand pose estimation. Boukhayma et al. uses images and 2D pose predicted from OpenPose as input and regress the parameters of the MANO model . Zhang et al. share a similar framework as Boukhayma. The key difference is that their model contains a 2D heatmap prediction module, instead of using predictions from OpenPose. Baek et al.’s method also has similar architecture. The difference lies in that they additionally adopt 2D masks as an intermediate representation. Different from these approaches, the work of Ge et al. use a self-created 3D hand model instead of MANO. The proposed framework takes single images as input and predicts 2D heatmaps as intermediate representation. After that, graph convolutional network is used to regress the vertices of the hand model.

Joint 3D Pose Estimation of Body and Hands. There are a few methods that pursue to estimate 3D poses of body and hands together. Due to the lack of annotated data for whole body capture, all these previous works resort to optimization methods. SMPLify-X uses the SMPL-X model to represent a whole body pose. The model parameters are optimized by fitting to 2D keypoints with additional constraints including body pose priors and collision penalizer. Monocular Total Capture (MTC) is based on the Adam model . It adopts deep neural networks to get 2.5D predictions first. Then the parameters of Adam are obtained through optimization. Both of these methods rely on optimization with relatively slow computation time (from 10 seconds to a few minutes). Besides, when 2D heatmap detection failed, their accuracy degrades significantly.

Method

Our method aims to estimate 3D body (the torso and limb parts) and 3D hands (both left and right) from monocular inputs (either monocular images or videos). Our method produces, as output, the parameters of the SMPL-X model to represent both 3D body and hand poses in a unified form. An important aspect of our method is to use separate expert modules for each body and hand pose estimation while both modules produce the compatible outputs as part of SMPL-X model. An overview of the framework is shown in Figure 1.

Notably, our hand pose estimator leverages the hand part of the SMPL-X model, by treating it as a stand-alone parametric hand model. While still showing the state-of-the-art monocular hand pose estimation performance, the output of our hand module can be directly merged to the body estimation output , to pose the whole body SMPL-X model. We also present an optimization framework to improve the hand and body pose estimation output with additional constraints for better accuracy. In the remaining part of this section, we assume the inputs to our model are single images. Details of processing sequence inputs are included in section 4.1.

Given a single image input cropped around a single person, our method produces whole body motion capture output as a form of shape and pose parameters of the SMPL-X model. As an extension of the SMPL model , the SMPL-X model can represent the shape variations and pose-dependent deformation of human bodies via a combination of low-dimensional shape and pose parameters. As a key difference from the SMPL model that only focuses on body parts, SMPL-X can also express finger motions and facial expressions, by including additional sets of parameters for them.

We formulate the SMPL-X model, denoted by WW, as:

Our hand models are defined by taking the hand parts of SMPL-X:

The major advantage of our representation is that the components of 3D hand model, including pose parameters, vertices, and 3D joints, are directly compatible with the whole body parameterization. This enables us to efficiently integrate outputs from the body module and the hand module.

2 3D Hand Estimation Module

We present a monocular 3D hand pose estimation module, denoted by MHM_{H}, estimating the parameters of the hand model HH. In particular, our hand module is inspired by the recently proposed monocular body pose estimation approaches , thus follows the similar model architecture, parameterizations, and training stages. Leveraging the achievement in body pose estimation area, we found that our hand pose estimation method can be robustly applicable for various in-the-wild situations, showing the state-of-the-art performance in public hand pose estimation benchmarks.

Hand Module Architecture. Our hand module MHM_{H} is built upon an end-to-end deep neural network architecture to regress the hand pose parameters defined in Eq. (3). Our hand module MHM_{H} is defined as:

where Π\Pi is an orthographic projection.

Following the body pose estimation approaches , the architecture of our hand module MHM_{H} is composed of an encoder and a decoder structure, where the encoder outputs the encoded features from input images, and the decoder regresses the hand pose parameters from the features. See Figure 2 for the overview of our hand module. We use the ResNet-50 for the encoder network. The decoder network is composed of a group of fully connected layers. Our hand module is trained with the data for the right hand. The images and annotations for the left hand are used after vertical flipping. During the testing time, the left hand images are flipped and processed as if they were a right hand, and their outputs are flipped back to the original left hand space.

Note that the shape parameter βh\boldsymbol{\beta}_{h} is originally defined for whole body model βw\boldsymbol{\beta}_{w}, but we only consider the deformation for the hand vertices defined in 3, ignoring the body part. We describe how this can be handled in our integration module.

We consider three different types of annotations: (1) 3D pose annotations (in angle-axis representation), (2) 3D keypoint (joint) annotations, and (3) 2D keypoint annotations. The losses for each of the annotations, namely LθL_{\boldsymbol{\theta}}, L3DL_{3D} and L2DL_{2D}, are defined as follows:

where θh^\hat{\boldsymbol{\theta}_{h}}, J^h3D\hat{\boldsymbol{J}}^{3D}_{h} and Jh2D^\hat{\boldsymbol{J}^{2D}_{h}} are the ground-truth annotations of angle-axis pose parameters, 3D keypoints, and 2D keypoints. In particular, the 2D keypoint loss L2DL_{2D} is necessary to estimate the camera projection parameters. We do not use the shape parameters provided by the 3D hand datasets such as FreiHAND , since these are defined for the MANO model and not compatible with our hand model from SMPL-X. Instead, an additional shape parameter regularization loss LregL_{reg} is applied:

The overall loss LL used to train our hand module is defined as follows:

In experiments, the balanced weights are set as λ1=10\lambda_{1}=10, λ2=100\lambda_{2}=100, λ3=10\lambda_{3}=10 and λ4=0.1\lambda_{4}=0.1.

Datasets Preprocessing. 3D hand pose datasets are often built by multi-view setups in controlled environments to obtain ground-truth annotations. A model trained with these datasets often suffer from overfitting, showing limited performance when applied to outdoor in-the-wild data. Notably, recent 3D body pose estimation approaches have shown that leveraging diverse datasets can greatly improve its generalization ability . Following this, we include as many publicly available datasets as possible towards in-the-wild 3D hand pose estimation. More details of the datasets are discussed in section 4.2. The major challenge in using diverse datasets is that their annotation types vary. For example, there exist available ground-truth joint angle parameters in FreiHand and HO-3D datasets, while others do not contain it. Furthermore, the details of hand annotations including skeleton hierarchy and scales are also different across datasets. To handle this, we perform several pre-processing steps to make them consistent and compatible with our hand model, including 1. Rescaling all 3D keypoint annotations to be compatible with our hand model, by using the middle finger’s knuckle lengththe skeleton between 44-th and 55-th joints efined in Figure 3. as a reference. 2. Re-ordering the 3D keypoints joints to be the same as our hand model’s skeleton hierarchy shown in Figure 3.

Training Data Augmentation. Performing data augmentations during training is a common practice to enable model with better generalization ability. Following previous approaches , we apply common data augmentation strategies including random scale, random translation, color jittering, and random rotation.

Importantly, we recognize that in-the-wild videos are often accompanied by severe motion blur. To achieve robustness to motion blur, we additional apply motion blur augmentation to the images. We first use the methods in previous papers to generate blur kernels and then use 2D filtering to add blurriness to images. The experiments show that our motion blur augmentation is beneficial to generalize our hand module for in-the-wild scenes. Examples of motion blur augmentation are shown in Figure 4.

3 3D Body Estimation Module

We leverage the state-of-the-art monocular 3D pose estimation methods with a few modifications for our body estimation module. The recent monocular body pose estimation methods are based on the SMPL model to capture torso and limb motion. Although our body module produces similar outputs, these previous methods cannot be directly applicable for our objective, since SMPL model’s shape parameters are not compatible with SMPL-X. Thus, we fine-tune the publicly available state-of-the-art pose estimator by replacing SMPL with SMPL-X in the training pipelines. For training, we use the publicly available indoor 3D pose datasets such as human3.6M . We also include the pseudo-ground truth annotations introduced in that provides SMPL fitting paired with the in-the-wild 2D keypoint datasets (e.g., COCO and MPII ). Since all these annotations are in SMPL format, we ignore the shape parameters of ground truth SMPL annotations, and only use the pose parameters and 2D keypoint annotations that are compatible to SMPL-X model. We follow the same neural network architecture with the similar training steps as in the work of , without using the SMPLify part.

Our body module MBM_{B} produces the torso and limb parameters defined in Eq. 1 from an single image:

Note that the existing body pose estimators including our fine-tuned version do not accurately estimate the the wrist and arm orientation due to inaccurate or insufficient annotations (e.g., only one keypoint is annotated for a wrist), as shown in Figure 5. Our integration module solves this issue.

4 Whole Body Integration Module

Our integration module combines the outputs from the 3D body and hand modules into a unified representation as a form of SMPL-X model. For the integration, we present two strategies: (1) a fast method by simple copy-and-paste composition, and (2) an optimization framework to include additional 2D keypoint cues for more accurate output.

where Γl\Gamma_{l} and Γr\Gamma_{r} are the functions to convert the global wrist orientation ϕh\boldsymbol{\phi}_{h} obtained from the hand module to the local wrist pose parameters w.r.t. its parent joint in the SMPL-X skeleton hierarchy. This can be implemented by comparing ϕh\boldsymbol{\phi}_{h} with the global orientation of the current wrist pose from θb\boldsymbol{\theta}_{b} that can be computed by following the forward kinematics of the body skeleton hierarchy. This strategy requires almost no extra computation, making our separate modules to contribute a common whole body model simultaneously. We found this simple integration produces convincing results, especially for the scenarios with computational bottlenecks as in our live demo.

Hand and Body Composition via Optimization. As an alternative integration method, we build an optimization framework to fit the whole body model parameters given the outputs from body and hand modules. This strategy is particularly helpful to reduce the artifact around the wrist parts over the copy-and-paste strategy, and also can take advantage from the 2D keypoint estimation output for better 2D localization quality. In particular, our optimization framework finds the whole body model parameters that minimize the following objective cost function:

where F2d\mathcal{F}^{2d} is the 2d reprojection cost term between the 2D keypoint estimation and the projection of 3D joints (body and both hands), and the prior term Fpri\mathcal{F}^{pri} is needed to keep the 3D pose and shape parameters in plausible space, as in SMPLify method. We first initialize all parameters by our copy-and-paste strategy except that we do not apply Γ\Gamma to transfer the global hand orientation to whole body model. Instead, the wrist orientations of the hands can be obtained by minimizing the Eq. (13) with other parameters. While a Gaussian mixture model learnt from motion capture dataset is often used for the body pose prior term as in SMPLify method , we use the the exemplar fine-tuning approach introduced in for the similar goal, by applying neural network fine-tuning of MBM_{B} for each frame independently, which does not require additional regularization term but still keep the 3D pose in plausible space. Note that our optimization framework requires only a few iteration (20 iterations in all our experiments), since outputs from the body and hand modules output is already close to the target status. See Figure 5 for the example of our optimization.

Experiments

In this section, we first describe the implementation details. Then we summarize the datasets used for our hand module training. After that, we quantitatively and qualitatively compare our methods with the state-of-the-art approaches. We also perform ablation studies to examine the key designs of our methods.

Bounding Boxes. For the online version, we use OpenPose to obtain body bounding boxes. After processing the body, the hand bounding boxes are obtained by projecting the hand part of the estimated 3D body to image space. For the offline processing of internet videos, we use OpenPose detections to localize both bodies and hands.

Video Processing. For copy-and-paste strategy (used in online demo and offline internet videos), the videos are processed frame-by-frame without any post-processing. For optimization-based strategy, after obtaining per-frame outputs, we apply a naive temporal smoothing for each separate dimension of parameters (shape, pose, and camera). We use a 5-frame-size smoothing kernel with the weight [0.1,0.2,0.5,0.2,0.1]\left[0.1,0.2,0.5,0.2,0.1\right]. It is noted that our copy-and-paste method can generate temporally-stable results even without smoothness. We believe it is due to the fact that recent CNN pose regressors tend to produce such output as demonstrated in recent papers (e.g. SPIN ), thanks to multiple augmentation tricks in training. The optimization-based method (SMPLify-X and MTC) suffers from temporal instability due to the complicated optimization procedures with multiple-stages (e.g. torso first and others later) and elaborated balancing issues between data term and prior term. The processing time of each method are compared in Table 1. The processing time of our copy and paste method is about 9.5 fps, where the code is implemented in python and runs in a single GeForce RTX 2080 GPU. In our supplementary video, we also show a live demo using a single webcam, which cannot be performed by alternative approaches.

2 Datasets

FreiHAND. FreiHAND is a dataset with ground truth 3D hand joints and MANO parameters for real human hand images. The 3D annotations are obtained by a multi-camera system and a semi-automated approach. The obtained data is further augmented with synthetic backgrounds. In our experiments, we randomly select 80% of samples from original training set as training data and use the remaining 20% of samples for validation.

HO-3D. HO-3D dataset is a dataset aiming to study the interaction between hands and objects. The dataset has 3D joints and MANO pose parameters for hands, and also has 3D bounding boxes for objects the hands interact with. In this paper, we only use 3D annotations of hands. The training set is composed of different sequences, each of which records one type of hand-object interaction. Following the similar practice in processing FreiHAND, we randomly choose 80% of sequences from the original training set as training data and use the remaining 20% of sequences for validation.

MTC. Monocular Total Capture is a dataset captured by Panoptic Studio in a multi-view setup with 30 HD cameras. It has 3D hand joints annotations for both body and hands. The sequences are mainly the range of motion data of multiple subjects. To polish the dataset, we filter out erroneous samples where hands are not visible or too small.

STB. Stereo Hand Pose Tracking Benchmark is composed of 15,000 training samples and 3,000 testing samples. The provided annotations include 3D joints and depth images. In our experiments, we use 3D joints only. We use training set of STB to train our model and compare with other state-of-the-art methods on the validation set. To unify definition of joints, following the practice of , we move the root joint from palm center to wrist.

RHD. Rendered Hand Dataset is a synthetic dataset that has 2D and 3D hand joint annotations. It is composed of 41,258 training samples and 2,728 testing samples. We train our model on the training set and compare with other state-of-the-art methods on the testing set.

MPII+NZSL. MPII+NZSL dataset is composed of in-the-wild images with manually annotated 2D hand joints. It includes challenging images with occlusion, blur, and low resolution. To show our models’ generalization ability, our models are no trained on the MPII+NZSL, we only use it for validation.

3 Hand Module Evaluation

Comparison with State-of-the-art Methods. We compare our hand module with the previous state-of-the-art hand approaches on three public hand benchmarks, STB , RHD and MPII+NZSL . For each validation dataset, we calculate the percentage of correct keypoints (PCK) under different thresholds and calculate the corresponding Area Under Curve (AUC) for PCK. For STB and RHD , we use 3D AUC and the threshold ranges from 20mm to 50mm. For MPII+NZSL , we use 2D AUC and the threshold ranges from 0px to 30px.

The results are listed in Table 2. For fair comparison, all the methods takes single RGB image as input. “Ours” refers to our best model trained with all the datasets and all data augmentation srategies. It outperforms previous methods on RHD and MPII+NZSL and shows a comparable performance in STB. Notably, our method shows significantly better 2D localization accuracy on challenging in-the-wild dataset MPII+NZSL, demonstrating its generalization ability to in-the-wild scenarios.

We also compare our own best model with variants of our method. “Ours-no-shape-params” differs from “Ours” in that the shape parameters β\beta are not used and fixed to zero. “Ours-no-data-augment” refers to the model trained without using any data augmentation. “Ours-less-datasets” refers to the model trained without using latest datasets, namely FreiHAND and HO-3D. This model is trained with MTC, STB and RHD. These datasets are also used by previous methods.

Comparison between the variants of our model verifies that including various datasets and applying data augmentation are important in achieving better results. Inferring shape variations by estimating shape parameters is also helpful to improve the accuracy. Note that the results of “Ours-less-datasets” show that our model still achieves comparable performance on STB and RHD with the previous methods by using only limited datasets, and shows better 2D localization accuracy on in-the-wild MPII+NZSL dataset. These results demonstrate that our method takes advantage from both network design (including the training strategy) and larger training datasets.

We also qualitatively compare our method with previous work, as shown in Figure 6 and our supplementary video. The results indicate that our hand model can generate more precise 3D hand poses under challenging in-the-wild scenarios with occlusion, blur and low resolution.

Ablation Study. We further examine two key designs used in training our hand module, the mixture of different datasets and data augmentation. The results for ablation study on the datasets are listed in Table 3 and Figure 7. As expected, the results in Table 3 shows that using more datasets will lead to better performance. We also show the examples of qualitative comparison in Figure 7. Similar to the conclusion from the quantitative study, the qualitative results show that incorporating more datasets can increase the models’ generalization ability and generate more precise results for in-the-wild images. In the figure, “Subset-01” means using the combination of datasets FreiHADN and HO-3D . “Subset-02” means using the combination of datasets: STB , RHD and MTC . “Full set” means using all the datasets.

The results for the ablation study on data augmentation are listed in Table 4 and Figure 8. The results in Table 4 demonstrate that applying data augmentation leads to better results. We also show qualitative results in Figure 8, where, by adopting data augmentation, our models can generalize better to challenging scenarios including blur, challenging poses and occlusion. In Figure 8, “No Augment” refers to model trained without any data augmentation, and “No Blur” refers to model trained with all data augmentation strategies except motion blur augmentation. “Full Augment” refers to model trained with all data augmentation strategies.

4 Integration Module Evaluation

We qualitatively compare our method with previous whole body motion capture approaches, MTC and SMPLif-X . We also compare between two version of our model. We refer to the copy-and-paste integration and the optimization-based integration as “Ours-CP” and “Ours-OP”, respectively The results are shown in Figure 9 and our supplementary videos, indicating that our method outperforms previous approaches in terms of both speed and accuracy. Notably, as shown in Table 1, our copy-and-paste method runs in two orders of magnitude faster speed than the alternative approaches, yet showing better 3D pose estimation quality.

Discussion

We present FrankMacop, a fast motion capture system to estimate both 3D hand and 3D body motion from monocular inputs in the wild. We design the body and hand expert modules to produce compatible outputs for whole body motion capture. We present two integration strategies, copy-and-paste for faster speed and an optimization framework for better quality. The performance of our method has been demonstrated in-the-wild monocular videos. In particular, we also demonstrate our whole body motion capture system in a live demo at near real-time speed (9.5 fps), which is orders of magnitude faster than alternative methods. Our 3D hand pose estimation module outperforms previous state-of-the art on hand only methods in public benchmarks, and ours also can be used as a stand-alone monocular 3D hand pose estimator.

Our method still suffered from a few limitations: 1. Hand pose estimation become erroneous if two hands are too close each other. 2. Bounding boxes are required to infer 3D body and hands. It would be an interesting future direction to mitigate these problems and extend the method to handle cases of multiple people interacting with each other, such as two people greeting with hand shaking.

Acknowledgements. We thank Yuting Ye for her helpful discussions and feedbacks. We also want to thank Xintao Wang for helping us in implementing motion blur augmentation.

References