HybrIK-X: Hybrid Analytical-Neural Inverse Kinematics for Whole-body Mesh Recovery
Jiefeng Li, Siyuan Bian, Chao Xu, Zhicun Chen, Lixin Yang, Cewu Lu
Introduction
Recovering the whole-body 3D surface from visual content has a broad spectrum of applications. Advancements in parametric statistical human body shape models have enabled the generation of realistic and animatable 3D meshes using only a small set of parameters. Despite the recent progress, recovering the 3D mesh by inferring the abstract pose and shape parameters is highly non-linear and still challenging.
Existing approaches can be divided into two categories: optimization-based and learning-based. Optimization-based approaches estimate the pose and shape of the human body through an iterative fitting process. The parameters of the statistical model are optimized to reduce the error between its 2D projection and 2D observations, e.g., 2D joint positions and silhouettes. However, this optimization problem is non-convex. Its solution can be time-consuming and its results are sensitive to the initialization. These challenges have shifted the research focus towards learning-based approaches. Learning-based approaches leverage parametric body models and employ neural networks to directly regress the pose parameters . Nevertheless, the pose parameters represent relative rotations that lie in the rotation group, which puts difficulties for neural networks to learn directly from RGB images. Consequently, the learned body mesh suffers from image-mesh misalignment.
Such a challenge prompts us to look into the field of 3D keypoint estimation. Previous 3D keypoint estimation approaches adopt volumetric heatmap as the target representation to learn 3D joint positions in the Cartesian coordinate system. The learned 3D joints can accurately align with the 2D RGB image. This inspires us to establish a collaboration between the 3D joints and the body mesh via forward kinematics (FK) and inverse kinematics (IK) (as illustrated in Fig. 1). On the one hand, the accurate 3D joints can improve image-mesh alignment for mesh recovery. Since the body mesh is recovered from 3D joints, the recovered mesh can obtain pixel-aligned accuracy as long as the 3D joints are well-aligned with the image. On the other hand, the shape prior inherent in the parametric body model can be utilized to mitigate the issue of the unrealistic body structure in 3D keypoint estimation approaches. Since existing 3D keypoint estimation approaches lack explicit modeling of body bone length, they may predict unrealistic body structures like left-right asymmetry and abnormal proportions of limbs. If we can leverage the parametric body model, the presented human shape can better conform to the actual human body.
In this work, we present a hybrid analytical-neural inverse kinematics solution (HybrIK) to establish a collaboration between 3D keypoint estimation and whole-body mesh estimation. Inverse kinematics (IK) is used to find the corresponding body-part rotations from 3D body joints. This is an ill-posed problem due to the lack of a unique solution. The core of our approach is an innovative solution for this problem using twist-and-swing decomposition. Specifically, the rotation of a skeleton part is decomposed into twist and swing, i.e., a longitudinal rotation and an in-plane rotation. The unique solution of the body-part rotations is composited iteratively along the kinematic tree by analytically calculating swing rotations from 3D joints and using a neural network to predict twist rotations from visual cues.
This IK framework is extended as HybrIK-X for articulated hand and expressive face reconstruction, improving the fine-grained image-mesh alignment with a new backward-updated solution. Unlike previous approaches , HybrIK-X gets rid of separated expert models and recovers the whole-body mesh with a one-stage network, resulting in improved efficiency and reduced computational resources. The robustness of 3D keypoints estimation against occlusions and truncations is enhanced by using a new regression approach to infer the truncated body parts. Furthermore, we exploit the body structure from the parametric human body model to alleviate the depth ambiguity in predicting camera parameters and estimate a more stable human motion.
A critical characteristic of our approach is that the estimated mesh is inherently aligned with the 3D skeleton, without the need for additional optimization procedures in the previous approaches . We conduct comprehensive experiments on various benchmarks for body-only, hand-only, and whole-body scenarios, including 3DPW , Human3.6M , and MPI-INF-3DHP , FreiHAND , HO3D , and AGORA . HybrIK and HybrIK-X show pixel-aligned accuracy and significantly outperform state-of-the-art approaches.
The contributions of our approach can be summarized as follows:
We propose a novel whole-body mesh recovery framework that uses a hybrid analytical-neural IK algorithm to convert accurate 3D joints to pixel-aligned body meshes.
Our approach closes the loop between the 3D skeleton and the parametric model. It improves the image-mesh alignment for body mesh recovery and addresses the unrealistic body structure problem of 3D keypoint estimation approaches at the same time.
Our approach achieves state-of-the-art performance across various body-only, hand-only, and whole-body benchmarks.
A preliminary version of this work was accepted in CVPR 2021 . This paper extends the previous work in the following ways. First, we extend the IK framework to whole-body mesh recovery with a backward-updated IK solution and an efficient one-stage model. Second, we enhance the robustness of our framework to occlusions and truncations by a new regression paradigm. Third, we propose a structure-aware cycle for the mitigation of depth ambiguity, thereby providing a more stable camera parameters estimation. Fourth, additional quantitative comparisons on various benchmarks for body-only, hand-only, and whole-body scenarios demonstrate the effectiveness and generalization of the proposed IK framework. Finally, we conduct further ablation studies to investigate and analyze our framework.
Related Work
Many studies formulates 3D human pose estimation as the problem of locating the 3D joints of the human body. Previous work can be divided into two categories: single-stage and two-stage approaches. Single-stage approaches directly estimate the 3D joint locations from the input image. Various representations are developed, including 3D heatmap , location-map , and 2D heatmap + regression . Two-stage approaches first estimate 2D pose and then lift them to 3D joint locations by a learned dictionary of 3D skeleton or regression . Two-stage approaches highly rely on accurate 2D pose estimators, which have achieved impressive performance through the combination of a powerful backbone network and the 2D heatmap.
These privileged forms of supervision contribute to the recent performance leaps of 3D keypoint estimation. However, the human structural information is modeled implicitly by the neural network, which can not ensure the output 3D skeletons are realistic. Our approach combines the advantages of both the 3D skeleton and parametric model to predict accurate and realistic human pose and shape.
2 Model-based 3D Body Pose and Shape Estimation
Prior work on the model-based 3D pose and shape estimation uses parameters of the statistical body model as the output target because they capture the statistics prior of body shape. Compared with the model-free methods , the model-based methods directly predict controllable body mesh, which can facilitate many downstream tasks for both computer graphics and computer vision. Bogo et al. propose SMPLify, a fully automatic approach, without manual user intervention . This optimization paradigm was further extended with silhouette cues , volumetric grids , multiple people , and whole-body parametric models .
With the advances in deep learning networks, there are increasing studies that focus on learning-based methods , using a deep network to estimate the pose and shape parameters. Since the mapping from RGB image to body shape and body-part rotation is hard to learn, many studies use intermediate representations to alleviate this problem, such as keypoints and silhouettes , semantic part segmentation , and 2D heatmap input . Kanazawa et al. use an adversarial prior and an iterative error feedback (IEF) loop to reduce the difficulty of regression. Arnab et al. and Kocabas et al. exploit temporal context, while Guler et al. use a part-voting expression and test-time post-processing to improve the regression network. Kolotouros et al. leverage the optimization paradigm to provide extra 3D supervision from unlabeled images. Zhang et al. propose to use saliency maps to infer occluded bodies. PARE uses part-based attention to improve body-part regression. PyMAF uses a mesh-aligned feedback loop to exploit locally aligned features.
In this work, we address this challenging learning problem by a transformation from the pixel-aligned 3D joints to the body-part rotations.
3 Whole-body Mesh Recovery
Numerous prior studies solve body, face , and hand meshes separately. Recent studies start using whole-body statistical models to jointly recover the 3D surface of the body, face, and hands.
Pavlakos et al. propose to estimate whole-body mesh by automatically fitting the SMPL-X model to the 2D body, face, and hand keypoints estimated by off-the-shelf whole-body keypoint estimators . Xiang et al. propose to learn 2D keypoints and part orientation fields for model fitting. Xu et al. fit GHUM with a reprojection error and a semantic body-part alignment error under anatomical joint angle limits. Similar to body-only pose estimation, optimization-based methods are slow and sensitive to initialization.
Learning-based methods tackle the above limitations by directly regressing the pose and shape parameters from the input image. ExPose uses three expert sub-networks to estimate body, face, and hands parameters separately. The body expert first estimates the body pose and the rough poses for the face and hands. Then the face and hand experts estimate the refined part poses from the cropped local image. Finally, the whole-body mesh is reconstructed by merging the results from three experts. Follow-up studies follow the same paradigm of ExPose . They use separate experts to handle body, face, and hand poses and merge them into a holistic whole-body pose. FrankMocap integrates the results from three experts through an integration network to approximate the optimization to better align the wrist and hand poses. Zhou et al. propose to regress the body and hand poses from the detected keypoints. PIXIE uses an attention-based moderator network to integrate body, face, and hand results from experts. Hand4Whole introduces to learn the wrist rotations from the hand keypoints. PyMAF-X proposes an adaptive integration module to better integrate the elbow and wrist poses.
Although using separate experts can benefit network training by using high-resolution input and multiple data sources, it costs higher computational complexity and lacks holistic information to integrate results from different experts. Unlike previous methods, we propose a one-stage paradigm that directly estimates the whole-body mesh. Our one-stage method gets rid of the time-consuming separated experts and the entire model is trained in an end-to-end manner.
4 Body-part Rotation in Pose Estimation
The core of our approach is to calculate the relative rotation of human body parts through a hybrid IK process. There are several studies that estimate the rotations in the 3D pose estimation literature. Zhou et al. use the network to predict the rotation angle of each body joint, followed by an FK layer to generate the 3D joint coordinates. Pavllo et al. switch to quaternions, while Yoshiyasu et al. directly predict the rotation matrices. Mehta et al. first estimate the 3D joints and then use a fitting procedure to find the rotation Euler angles. Previous approaches are either limited to a hard-to-learn problem or require an additional fitting procedure. Our approach recovers the body-part rotation from 3D joint locations in a direct, accurate and feed-forward manner.
5 Inverse Kinematics Process
The inverse kinematics (IK) problem has been extensively studied during recent decades. Numerical solutions are simple ways to implement the IK process, but they suffer from time-consuming iterative optimization. Heuristic methods are efficient solutions to the IK problem. For example, CDC , FABRIK , and IK-FA have a low computational cost for each heuristic iteration. In some special cases, there exist analytical solutions to the IK problem. Tolani et al. propose a reliable algorithm by the combination of analytical and numerical methods. Kallmann et al. solve the IK for arm linkage, i.e., a three-joint system. Recently, researchers have been interested in using neural networks to solve the IK problem in robotic control , motion retargeting , and hand pose estimation .
In this work, we combine the interpretable characteristic of the analytical solution and the flexibility of the neural network, introducing a feed-forward hybrid IK algorithm with twist-and-swing decomposition. Twist-and-swing decomposition is first introduced by Baerlocher et al. . The twist angles are limited based on the particular body joint. In our work, the twist angles are estimated by a neural network, which is more flexible and can be generalized to all body joints. Compared with previous analytical solutions designed for specific joint linkage, our algorithm can be applied to the entire body skeleton in a direct and differentiable manner.
Method
In this section, we present our hybrid analytical-neural inverse kinematics solution for whole-body mesh recovery (Fig. 2). First, in §3.1, we briefly describe the forward kinematics process, the inverse kinematics process, and the SMPL/SMPL-X model. In §3.2, we introduce the proposed inverse kinematics solution, HybrIK, for body-only mesh recovery, and HybrIK-X, for whole-body mesh recovery. Then, in §3.3, we present the overall learning framework to estimate the pixel-aligned whole-body mesh and realistic 3D skeleton. Finally, we provide the necessary implementation details in §3.4.
Forward Kinematics. Forward kinematics (FK) for human pose usually refers to the process of computing the reconstructed pose , with the rest pose template and the relative rotations as input:
For the root joint that has no parent, we have .
Inverse Kinematics. Inverse kinematics (IK) is the reverse process of FK, computing relative rotations that can generate the desired locations of input body joints . This process can be formulated as:
where denotes the -th joint of the input pose. Ideally, the resulting rotations should satisfy the following condition:
Similar to the FK process, we have for the root joint that has no parent. While the FK problem is well-posed, the IK problem is ill-posed because there are either no solutions or too many solutions to fulfill the target joint locations.
2 Hybrid Analytical-Neural Inverse Kinematics
Estimating the human body mesh by directly regressing the body part rotations is too difficult . In this paper, we propose a hybrid analytical-neural inverse kinematics solution that use 3D keypoints to recover 3D body mesh. Since the IK problem is ill-posed, we cannot uniquely determine the relative rotation just by the 3D joints. Here, we first decompose the original rotation into twist and swing. The 3D joints are utilized to calculate the swing rotation analytically, and we exploit the visual cues by a neural network to estimate the 1-DoF twist rotation. In our IK algorithm, the body part rotations are solved recursively along the kinematic tree.
where is the twist angle that estimated by a neural network, is a closed-form solution of the swing rotation, and transforms to the twist rotation. Here, should satisfy the condition in Eq. 5, i.e., .
Swing Rotation. The swing rotation has the axis that is perpendicular to and . Therefore, it can be formulated as:
Hence, the closed-form solution of the swing rotation can be derived by the Rodrigues formula:
where is the skew symmetric matrix of and is the identity matrix.
Twist Rotation. The twist rotation is rotating around itself. Thus, with itself the axis and the angle, we can determine twist rotation :
where is the skew symmetric matrix of .
Note that function and are fully differentiable, which allows us to integrate the twist-and-swing decomposition into the training process. Although we need a neural network to learn the twist angle, the difficulty of learning is significantly reduced. Compared with previous work that directly regresses the 3-DoF rotation, we regress the twist angle that is only a 1-DoF variable. Moreover, due to the physical limitation of the human body, the twist angle has a small range of variation. Therefore, it is much easier for the networks to learn the mapping function. We further analyze its variation in §4.5.
2.2 Body-only Inverse Kinematics
Naive HybrIK. Using the twist-and-swing decomposition, the IK process can be performed recursively along the kinematic tree like the FK process. First of all, we need to determine the global root rotation , which has a closed-form solution using the locations of , , and Singular Value Decomposition (SVD). Detailed mathematical proof is provided in appendix §A. Then, in each step, e.g., the -th step, we assume the rotation of the parent joint is known. Hence, we can reformulate Eq. 5 with Eq. 3 as:
Let and , we can solve the relative rotation via Eq. 6:
The whole process is named Naive HybrIK and summarized in Alg. 1. Note that we solve the relative rotation instead of the global rotation . The reason is that if we directly decompose the global rotation, the resulting twist angle will depend on all ancestors’ rotations along the kinematic tree, which increases the variation of the distal limb joints and makes it difficult for the network to learn.
Adaptive HybrIK. Although the Naive HybrIK process seems effective, it follows an unstated hypothesis: . Otherwise, there is no solution for Eq. 5. Unfortunately, in our case, the body-parts predicted by the 3D keypoint estimation method are not always consistent with the rest pose template. In Naive HybrIK, Eq. 6 can still be solved because the condition is turned into:
where denotes the error in the -th step, which has the same direction of and . To analyze the reconstruction error, we compare the difference between the input pose and the reconstructed pose :
where . Combining Eq. 2 and Eq. 13, we have:
where denotes the parent index of the -th joint, and denotes the set of ancestors of the -th joint. That means the difference between the input joint and the reconstructed joint will accumulate along the kinematic tree, which brings more uncertainty to the distal joint.
To address this error accumulation problem, we further propose Adaptive HybrIK. In Adaptive HybrIK, the target vector is adaptively updated with the newly reconstructed parent joints. Let and the same as the one in the naive solution. In this way, the condition in Adaptive HybrIK can be formulated as:
Compared to the naive solution (Eq. 15), the reconstructed error of the adaptive solution only depends on the current joint position. As illustrated in Fig. 4, in Naive HybrIK, once the parent joint is out of position, its children will continue this mistake. Instead, in Adaptive HybrIK, the solved relative rotation is always pointing towards the target joint and tries to reduce the error. We conduct empirical experiments in §4.5 to validate its robustness. The whole process of Adaptive HybrIK is summarized in Alg. 2.
2.3 Whole-body Inverse Kinematics
Adaptive HybrIK is accurate for body-only mesh recovery. It reduces the accumulated errors by adaptively updating the parent joints in a feedforward process. Nevertheless, extending it to whole-body mesh recovery is nontrivial. Although adaptive HybrIK tries to minimize the error in each step, the error won’t be totally eliminated since we cannot fix the erroneous positions of the ancestor joints. As long as the bone length calculated from keypoints is different from the SMPL-X model, misalignment is inevitable. This misalignment is particularly pronounced when considering fine-grained face and hand mesh recovery since the kinematics tree is much deeper and the distal joints (e.g., fingers and face) are further away from the root joint.
HybrIK-X. To achieve well-aligned whole-body mesh recovery, we further present HybrIK-X. As depicted in Fig. 5, the core of HybrIK-X is a divide-and-conquer IK process. Specifically, we first divide the whole-body kinematic tree into four sub-trees (namely the left/right hands, face, and body), which have shorter lengths compared to the whole-body tree. By doing so, the distal joints in each sub-tree are closer to their corresponding root joint. Thus the sub-trees are more robust to noisy bone lengths and have a better model-image alignment. Subsequently, we apply HybrIK independently in each sub-tree. The recovered mesh of each sub-tree will align with its corresponding root joint (, , and ). Finally, we utilize a novel backward-updated algorithm to merge the results of all sub-trees and resolve the conflict joints.
The conflict joints refer to those joints that appear in two sub-trees concurrently, such as the and . These joints are considered both the root joints of the sub-trees and the distal joints of the body-tree. If we directly combine the solved rotations of 4 sub-trees, the positions of the distal joints will be only determined by the results from the body-tree because the blend skinning function of SMPL-X depends on the feedforward FK process. In order to preserve the well-aligned results from the sub-trees, we propose a backward-updated algorithm for recalculating the rotations of the parents of the conflict joints.
Given the estimated position of the distal joint and the reconstructed position of its grandparent joint , our objective is to recalculate the rotation that satisfies:
By recalculating the rotation of the parent joint, the distal joint position will be consistent with the root position of the sub-tree. However, Eq. 18 involves the space and is not easy to solve. To make this problem solvable in a differentiable manner, we reformulate Eq. 18 to an equivalent problem:
The analytical solution of the position of the parent joint can be easily solved and computed differentiably. Detailed derivations are provided in appendix §B. Once we obtain , we follow the proposed twist-and-swing decomposition to calculate and .
Robust Jaw Pose Estimation. The rotation in the SMPL-X model relies on two joints, namely and . However, accurately determining the position of the joint can be challenging, particularly when people are wearing hats or looking upwards. Therefore, using Eq. 7 and Eq. 8 to calculate the swing rotation can result in an anomalous mouth-opening posture. To address this issue, we calculate rotation by measuring the angle of mouth opening. In particular, we replace the node in the kinematic tree with two nodes picked from the SMPL-X mesh ( and ). We then utilize these two joints to determine the swing axis and swing angle, and subsequently obtain the swing rotation for the joint using Eq.9.
3 Learning Framework
The overall framework of our approach is illustrated in Fig. 2. Firstly, a neural network is utilized to predict 2.5D joints , twist angles , shape parameters , expression parameters , and the initial camera parameter . The 2.5D joints and the initial camera parameter are sent to the iterative camera estimation module to obtain the final camera prediction and the 3D joints . Secondly, the shape and expression parameters are used to obtain the rest pose from the SMPL/SMPL-X model. Then, by combining , and , we employ HybrIK/HybrIK-X to solve the rotations of the 3D body, i.e., the pose parameters . Finally, with the function provided by the SMPL/SMPL-X model, the whole-body mesh is obtained. The reconstructed pose can be obtained from by FK or a regressor, which is guaranteed to be realistic. Since HybrIK and HybrIK-X are differentiable, the whole framework is trained in an end-to-end manner.
Regression-based 3D Keypoint Estimation. Previous approaches to 3D keypoint estimation use heatmaps to represent the likelihood of joint positions. However, this heatmap representation costs a significant computational burden. Besides, it limits the output range within the input bounding box, which fails in the truncated scenarios where part of the human body is outside the input image. In real-world challenging scenarios, object detection methods are not perfect and always generate truncated bounding boxes when people are occluded or partially visible.
where is the probability density function of the standard Laplace distribution, is the distribution learned by RLE , and and denotes the ground-truth joint position. In practice, the network output for each joint is a 2.5D coordinate, i.e., the two-dimensional pixel coordinate with a relative depth value. To obtain the 3D coordinates, we back-project the 2.5D point to the 3D space with the estimated camera parameters.
With the regression paradigm, we get rid of the heatmap representation, and the model can be trained to infer the invisible body joints in challenging scenarios. Besides, the predicted deviation can serve as the uncertainty index to evaluate the reliability of predicted keypoints, which is essential for downstream applications.
where is the projection function and denotes the back-projection function. Since the initial scale prediction is not correct, the projected 3D joints might be wrong, e.g., overly small or overly large. We then input to HybrIK-X to reconstruct the whole-body pose and retrieve reconstructed keypoints , which satisfy the constraints imposed by the human body structure. We then use the least square method to analytically calculate the updated that minimizes projection error between and the detected 2D keypoints:
where denotes the ground-truth scale factor.
where denotes the ground-truth twist angle for the -th joint.
Collaboration with SMPL-X. The SMPL-X model allows us to obtain the rest pose skeleton with the additive offsets according to the shape parameters and expression parameters :
where is the mesh vertices of mean rest pose, and are the blend shape functions provided by SMPL-X. Then the pose parameters are calculated by HybrIK-X in a differentiable manner. In the training phase, we supervise the shape parameters :
The overall loss of the learning framework is formulated as:
where , , and are weights of the loss items.
4 Implementation Details
Here we elaborate more implementation details. We use HRNet-W48 as the network backbone by default, initialized with ImageNet pre-trained weights. The HRNet output is fed to an average pooling layer, followed by the fully-connected layers to regress , , , and . The input image is resized to . The learning rate is set to at first and reduced by a factor of 10 at the 90th and 120th epoch. We use the Adam solver and train for 140 epochs, with a mini-batch size of per GPU and GPUs in total. In all experiments, and . Implementation is in PyTorch.
Empirical Evaluation
In this section, we first describe the datasets employed for training and quantitative evaluation. Next, we compare HybrIK and HybrIK with state-of-the-art approaches on body-only, hand-only, and whole-body mesh recovery benchmarks. Finally, ablation experiments are conducted to evaluate HybrIK and HybrIK-X.
Following previous work, we train the body-only HybrIK on the Human3.6M , 3DPW , MPI-INF-3DHP , MSCOCO , and AGORA datasets. For hand-only HybrIK, we train and evaluate on FreiHAND and HO3D-v2 datasets. For whole-body HybrIK-X, we use Human3.6M , 3DPW , MPI-INF-3DHP , MSCOCO , and AGORA datasets for training. Detailed descriptions of the datasets are provided appendix §C.
2 Evaluation on Body-only Mesh Recovery
To make a fair comparison with previous body-only mesh recovery methods, we use a regressor to obtain the LSP joints from the body mesh for the evaluation on 3DPW and Human3.6M datasets and joints for the MPI-INF-3DHP dataset. Procrustes aligned mean per joint position error (PA-MPJPE), mean per joint position error (MPJPE), percentage of correct keypoints (PCK), and area under curve (AUC) are reported to evaluate the 3D pose results. Per/Mean vertex error (PVE/MVE) is reported to evaluate the entire estimated body mesh. We further conduct experiments on the official AGORA test set. Normalized mean joint error (NMJE) and normalized mean vertex error (NMVE) are additionally reported.
In Tab. I, we compare our method with previous 3D human pose and shape estimation methods, including both model-based and model-free methods, on 3DPW, Human3.6M, and MPI-INF-3DHP datasets. Without bells and whistles, our method surpasses all previous state-of-the-art methods by a large margin on all three datasets. It is worth noting that our method improves mm PA-MPJPE (10.1% relative improvement) on the 3DPW dataset, which shows that it is accurate and reliable to recover body mesh through inverse kinematics.
In Tab. II, we futher compare our method on the SMPL track of the AGORA test set. Compared to other 3D pose datasets, AGORA contains more challenging scenarios with severe occlusions and truncations. HybrIK shows consistent improvements on in this dataset. Qualitative results are shown in Fig. 6.
3 Evaluation on Hand-only Mesh Recovery
To validate the generalization of our inverse kinematics solution, we conduct experiments on two widely-used hand pose benchmarks: FreiHAND and HO3D . MPJPE and PVE are reported for evaluation. F-Score is also reported to evaluate the harmonic mean between the predicted vertices and the ground-truth vertices.
The comparisons of HybrIK with previous state-of-the-art methods are reported in Tab. IV and Tab. V. It is noteworthy that, as observed in prior research , recent hand-only methods that adopt non-parametric representation exhibit a numerical advantage over parametric-based methods. HybrIK significantly outperforms the most accurate whole-body method by 1.5 mm MPJPE (19.5% relative improvement) on the FreiHAND dataset. Additionally, even when compared against the state-of-the-art hand-only methods, HybrIK exhibits a 7.5% relative improvement. Compared to the methods designed specifically for challenging interaction scenarios on the HO3D dataset, HybrIK also exhibits state-of-the-art performance.
4 Evaluation on Whole-body Mesh Recovery
To evaluate our hybrid inverse kinematics solution on whole-body mesh recovery, we conduct experiments on the SMPL-X track of the AGORA test set. The evaluation on AGORA is affected by the detection results since there are multiple persons on one image. We follow previous work and use the same detection results for a fair comparison.
The evaluation is depicted in Tab. III. Note that previous methods use separate expert networks to handle body, face, and hand estimation. They cost higher computational resources and longer running time. HybrIK-X is a one-stage approach and demonstrates a significant superiority over the state-of-the-art whole-body methods, achieving 24.3 mm and 20.7 mm improvement in full-body NMJE and NVME, respectively. For the fine-grained face and hand results, HybrIK-X obtains comparable performance against the most accurate method with a much smaller input resolution and less computation. Qualitative results on the MSCOCO validation set are shown in Fig. 7. Qualitative comparison with state-of-the-art approaches are provided in appendix §E. Detailed comparisons of computation complexity are reported in §4.5.
5 Ablation Study
In this study, we evaluate the effectiveness of the twist-and-swing decomposition and the proposed inverse kinematics algorithms. Evaluation is conducted on the AGORA validation set by default as it contains challenging in-the-wild scenarios. More experimental results are provided in appendix §D.
Analysis of the twist rotation. To demonstrate the effectiveness of twist-and-swing decomposition, we first count the distribution of the twist angle in the AGORA validation set. The distribution is illustrated in Fig. 8. As expected, due to the physical limitation, only , and have a wide range of variations. All other joints have a limited range of twist angle (less than 30∘). It indicates that the twist angle can be reliably estimated.
Besides, we develop an experiment to see how the twist angles affect the reconstructed pose and shape. We take the ground-truth 24 SMPL joints and shape parameters as the input of the HybrIK process. As for the twist angle, we compare random values in and the values estimated by the network. We evaluation the mean error of the reconstructed 24 SMPL joints, the 14 LSP joints, the body mesh and the twist angle. Here, following previous work , the 14 LSP joints are regressed from the body mesh by a pretrained regressor. Quantitative results are reported in Tab. VI. It shows that the regressed twist angles significantly reduce the error on the mesh vertices and the LSP joints that regressed from the mesh. Since most of the twist angles are close to zeros, the zero twist angles produce acceptable performance. Notice that the wrong twist angles do not affect the reconstructed SMPL joints. Only the swing rotations change the joint locations.
Robustness to Noisy Joint Positions. To demonstrate the superiority of HybrIK-X over Adaptive and Naive HybrIK, we compare the reconstruction errors with inputs in different noise levels on the AGORA validation set. We add jitters to the ground-truth joint positions and then feed them to the IK algorithms. Quantitative comparisons on whole-body recovery are reported in Tab. VII. We report the reconstruction errors of body and hand joints separately. It shows that when the input joints are correct, all three algorithms introduce negligible errors. As the noise level increases, HybrIK-X and Adaptive HybrIK are more robust than Naive HybrIK. For body reconstruction, HybrIK-X has similar performance to Adaptive HybrIK. For fine-grained hand reconstruction, HybrIK-X exhibits better robustness against the noisy joint positions.
Error correction capability. In this experiment, we examine the error correction capability of the HybrIK algorithm. The HybrIK algorithm is fed with the 3D joints, twist angles and shape parameters that predicted by the neural network. Additionally, we apply the SMPLify algorithm on the predicted pose and compare it to our method. As shown in Tab. VIII, the error of reconstructed joints after HybrIK is reduced the error to mm, while SMPLify raises the error to 100.1 mm. The error correction capability of HybrIK comes from the fact that the network may predict unrealistic body pose, e.g., left-right asymmetry and abnormal limbs proportions. In contrast, the rest pose is generated by the parametric statistical body model, which guarantees that the reconstructed pose is consistent with the realistic body shape distribution. Since our proposed framework is agnostic to the way we obtain 3D joints, we can improve the performance of any 3D keypoint estimation approaches.
Effectiveness of the Regression Paradigm. We further evaluate the effectiveness of the regression paradigm. Since the AGORA test set does not provide ground-truth bounding boxes, the model performance will be affected by the object detection results. Besides, the test set is quite challenging because the persons in it are severely occluded and truncated. Therefore, the AGORA test set can serve as the benchmark to evaluate the robustness of the mesh recovery algorithms. Moreover, we augment the current validation set with challenging truncated bounding boxes. We simulate truncations by cropping each frame with a truncated window. The areas of the cropped windows are of the original ones. Quantitative comparisons about different 3D keypoint estimation paradigms are reported in Tab. IX. Heatmap cannot represent the joints outside the input bounding box. Therefore, there is a significant performance degradation when using heatmaps to obtain joint positions. Additionally, the improved version of RLE shows better performance than the original RLE since it introduces fewer parameters to the distribution and avoids overfitting. It is demonstrated that our method is robust to challenging applications with occlusions and truncations. Qualitative comparisons between the heatmap-based paradigm and the regression-based paradigm are provided in appendix §E.
Effectiveness of Iterative Camera Estimation. To study the effectiveness of the proposed iterative camera estimation (ICE), we evaluate the performance with different iterative steps. We calculate the mean error of the estimated distance between the human body and the camera. Quantitative results are summarized in Tab. X. “0 step” is equivalent to direct regression without the iterative update. We can observe that ICE can improve the accuracy of the camera distance estimation. In practice, 3 steps are accurate enough.
Computation Complexity. The experimental results of computation complexity and model parameters are listed in Tab. XI. The proposed one-stage HybrIK-X achieves more accurate whole-body mesh recovery results with significantly lower computation complexity and fewer model parameters. Specifically, the total FLOPs are reduced by 26.8%, and the parameters are reduced by 60.5%. The efficiency and effectiveness of HybrIK-X are of great value in downstream applications.
Conclusion
In this paper, we present a hybrid analytical-neural inverse kinematics framework, HybrIK, for body-only mesh recovery. This framework is further extended to whole-body mesh recovery and named HybrIK-X. HybrIK and HybrIK-X transform the 3D joint locations to a pixel-aligned accurate human body mesh via inverse kinematics, and then obtains a more accurate and realistic 3D skeleton from the reconstructed 3D mesh with forward kinematics, closing the loop between the 3D skeleton and the parametric body model. To evaluate the effectiveness of our approach, we perform experiments on various benchmarks for body-only, hand-only, and whole-body scenarios. The results indicate that our approach outperforms existing state-of-the-art approaches by a significant margin. Moreover, the proposed approach is fully differentiable and uses a one-stage network to recover the whole-body mesh, making it considerably more efficient than existing approaches. Overall, we believe our approach can serve as a strong baseline for future research and offers a new perspective on whole-body mesh recovery.
Appendix A Rigid Registration of Global Rotation
In the SMPL/SMPL-X model , the pose parameters control the rotations of the rigid body parts. The three joints named , and constitute a rigid body part, which is controlled by the global root rotation. Consequently, the global rotation can be ascertained by registering the rest pose template of , and to the predicted locations of these three joints. Let , and denote their locations in the rest pose template, and , and denote the predicted locations. Our objective is to identify a rigid rotation that optimally aligns the two sets of joints. Here, we assume the root joint of the predicted pose and the rest pose are aligned. Hence, the problem is formulated as:
This formula can be written in matrix form:
where denotes the Frobenius norm, denotes , and denotes . Let us simplify the expression in Eq. 31 as:
Further, we can leverage the property of the matrix trace,
Then, we apply Singular Value Decomposition (SVD) to the joint locations:
Besides, is a diagonal matrix with non-negative values, i.e., , , . Therefore:
The trace is maximized if . That means , where is the identity matrix. Finally, the optimal rotation is:
Appendix B Backward-updated Algorithm
In the backward-updated algorithm of HybrIK-X, our aim is to analytically calculate the position of the parent joint, denoted as . Recall that should satisfy the following equation:
To simplify the notation, we define , , , and for the subsequent derivation. The aforementioned equation can be reformulated as:
The vector can be orthogonally decomposed with respect to as follows:
where is parallel to and is perpendicular to . We introduce a point , such that and . Since is parallel to , it can be expressed as , with representing the corresponding scalar. Consequently, by determining and , we can ascertain the position of .
According to the norm constrains in Eq. 42, we can determine as:
Once we determine , the original problem can be written as:
can be orthogonally decomposed according to as:
where and .
Thus, , where is a constant an irrelevant to . Intuitively, the optimal must be parallel to since they are both perpendicular to and sharing the same start point . Consequently, we have , where denotes the norm of . Following the norm constrains in Eq. 42, we can determine as:
Appendix C Datasets
AGORA : It is a synthetic dataset featuring precise SMPL and SMPL-X annotations fitted to 3D scans. We use this dataset for training only when conducting experiments on it. The evaluation is performed on the official platform with separated SMPL and SMPL-X tracks.
3DPW : It is a challenging outdoor benchmark for body-only 3D pose and shape estimation. It contains video sequences obtained from a hand-held moving camera. It employs IMU sensors to calculate ground-truth pose and shape data.
MPI-INF-3DHP : It is a body-only 3D keypoint dataset with both constrained indoor and complex outdoor scenes. It includes 8 actors performing 8 activities from 14 camera views. Following , we use its train set for training and evaluate on its test set.
Human3.6M : It is an indoor benchmark for body-only 3D pose and shape estimation. It includes 11 subjects performing 15 different activities in a laboratory environment. Following , we use 5 subjects (S1, S5, S6, S7, S8) for training and 2 subjects (S9, S11) for evaluation.
FreiHAND : It is a single-hand 3D pose dataset with MANO annotations. It contains over 130k training samples with the right hands. We employ HybrIK on the hand pose model and use this dataset for evaluation.
HO3D-v2 : It is a dataset with 3D pose annotations for hands and objects under severe occlusions from each other. It contains sequences of the right hand interacting with an object. We use this dataset to evaluate the performance of hand pose estimation under challenging scenarios.
MSCOCO : It is a large-scale in-the-wild 2D human pose dataset consisting of over 150k instances. We incorporate its train set for training.
Appendix D Ablation Experiments.
Robustness of HybrIK to Noisy Joint Positions. In the main paper, we evaluate the robustness of HybrIK in whole-body scenarios. Here, we further report the evaluation of robustness on body-only scenarios. We use the same evaluation protocol as in the main paper. We randomly add Gaussian noise to the 3D joint positions with different standard deviations. The results are shown in Tab. XII. We can see that adaptive HybrIK is more robust to the noise than naive HybrIK.
Effect of . In this experiment, we examine the impact of shape parameters on the AGORA validation set. As shown in Tab. XIII, using the ground-truth yields a mm improvement in MPJPE and PVE, while using zero results in a mm error. It shows that our model also gives an accurate body shape estimation.
Error correction capability of HybrIK. In this experiment, we investigate the error correction capability of body-only HybrIK on the 3DPW , Human3.6M , and AGORA datasets. Quantitative results are reported in Tab. XIV. They demonstrate that HybrIK can leverage the structural information embedded in the statistical body model to rectify unrealistic body joints derived from off-the-shelf 3D keypoint estimation approaches. The error correction capability of HybrIK is more pronounced on the AGORA dataset, which poses significant challenges for pose and joint estimation.
Appendix E Qualitative Results
Additional qualitative results for body-only, hand-only, and whole-body scenarios are presented in Fig. 9, 10, and 11, respectively. Qualitative comparisons between heatmap-based backbone and our regression-based backbone are displayed in Fig. 12. Qualitative comparisons with state-of-the-art approaches are presented in Fig. 13. More results can be found in our project page https://jeffli.site/HybrIK-X/.