AvatarPoser: Articulated Full-Body Pose Tracking from Sparse Motion Sensing
Jiaxi Jiang, Paul Streli, Huajian Qiu, Andreas Fender, Larissa Laich, Patrick Snape, Christian Holz
Introduction
Interaction in today’s Mixed Reality (MR) environments is driven by the user’s head pose and input from the hands. Cameras embedded in head-mounted displays (HMD) track the user’s position inside the world and estimate articulated hand poses during interaction, which finds frequent application in Augmented Reality (AR) scenarios. Virtual Reality (VR) systems commonly equip the user with two hand-held controllers for spatial input to render haptic feedback. In both cases, even this sparse amount of tracking information suffices for interacting with a large variety of immersive first-person experiences.
However, the lack of complete body tracking can break immersion and reduce the fidelity of the overall experience as soon as interactions exceed manual first-person tasks. This not just becomes evident as users see their own bodies during interaction in VR, but also in collaborative tasks in AR that necessarily limit the representation of other participants to their upper bodies, rendered to hover through space. Studies on avatar appearances have shown the importance of holistic avatar representations to achieve embodiment and to establish presence in the virtual environment . Applications such as telepresence or productivity meetings would greatly benefit from more holistic avatar representations that approach the fidelity of motion-capture animations.
This challenge will likely not be addressed by future hardware improvements, as MR systems increasingly optimize for mobile use outside controlled spaces that could accommodate comprehensive tracking. Therefore, we cannot expect future systems to expand much on the tracking information that is available today. While the headset’s cameras may partially capture the user’s feet in opportune moments with a wide field of view, head-mounted cameras are generally in a challenging location for capturing ego-centric poses .
Animating a complex full-body avatar based on the sparse input available on today’s platforms is a vastly underdetermined problem. To estimate the complete set of joint positions from the limited tracking sources, previous work has constrained the extend of motion diversity or used additional trackers on the user’s body, such as a 6D pelvis tracker or several body-worn inertial sensors . Dittadi et al.’s recent method estimates full-body poses from only the head and hand poses with promising results . However, since the method encodes all joints relative to the pelvis, it implicitly assumes knowledge of a fourth 3D input (i.e., the pelvis).
For practical application, existing methods for full-body avatar tracking come with three limitations: (1) Most general-purpose applications use Inverse Kinematics (IK) to estimate full-body poses. This often generates human motion that appears static and unnatural, especially for those joints that are far away from the known joint locations in the kinematic chain. (2) Despite the goal of using input from only the head and hands, existing deep learning-based methods implicitly assume knowledge of the pelvis pose. However, pelvis tracking may never be available in most portable MR systems, which increases the difficulty of full-body estimation. (3) Even with a tracked pelvis joint, animations from estimated lower-body joints sometimes contain jitter and sliding artifacts. These tend to arise from unintended movement of the pelvis tracker, which is attached to the abdomen and thus moves differently from the actual pelvis joint.
In this paper, we propose a novel Transformer-based method for full human pose estimation with only the sparse tracking information from the head and hand (or controller) poses as input. With AvatarPoser, we decouple the global motion from learned pose features and use it to guide our pose estimation. This provides robust results in the absence of other inputs, such as pelvis location or inertial trackers. To the best of our knowledge, our method is the first to recover the full-body motion from only the three inputs across a wide variety of motion classes. Because the predicted end effector poses of an avatar accumulate errors through the kinematic chain, we optimize our initial parameter estimations through inverse kinematics. This combination of our learning-based method with traditional model-based optimization strikes a good balance between full-body style realism and accurate hand control.
We demonstrate the effectiveness of AvatarPoser on the challenging AMASS dataset. Our proposed method achieves state-of-the-art accuracy on full-body avatar estimation from sparse inputs. For inference, our network reaches rates of up to 662 fps. In addition, we test our method on data we recorded with an HTC VIVE system and find good generalization of AvatarPoser to unseen user input. Taken together, our method provides a suitable solution for practical applications that operate based on the available tracking information on current MR headsets for application in both, Augmented Reality scenarios and Virtual Reality environments.
Related Work
Full-Body Pose Estimation from Sparse Inputs. Much prior work on full-body pose estimation from sparse inputs has used up to 6 body-worn inertial sensors . Because these 6 IMUs are distributed over head, arms, pelvis and legs, motion capture becomes inflexible and unwieldly. CoolMoves was first to use input from only the headset and hand-held controllers to estimate full-body poses. However, the proposed KNN-based method interpolates poses from a smaller dataset with only specific motion activities and it is unclear how well it scales to large datasets with diverse subjects and activities, also for inference. LoBSTr used a GRU network to predict the lower-body pose from the past sequence of tracking signals of the head, hands, and pelvis, while it computes the upper-body pose to match the tracked end-effector transformations via an IK solver. The authors also highlight the difficulty of developing a system for estimations from 3 sources only, especially when distinguishing a wide range of human poses due to the large amount of ambiguity. More recently, Dittadi et al. proposed a VAE-based method to generate plausible and diverse body poses from sparse input . However, their method implicitly uses knowledge of the pelvis as a fourth input location by encoding all joints relative to the pelvis, which leaves the highly ill-posed problem with only three inputs unsolved.
Vision Transformer. Transformers have achieved great success in their initial application in natural language processing . The use of Transformer-based models has also significantly improved the performance on various computer vision tasks such as image classification , image restoration , object detection , and object tracking . In the area of human pose estimation, METRO was first to apply Transformer models to vertex-vertex and vertex-joint interactions for 3D human pose and mesh reconstruction from a single image. PoseFormer and ST-Transformer used Transformers to capture both body joint correlations and temporal dependencies. MHFormer leveraged the spatio-temporal representations of multiple pose hypotheses to predict 3D human pose from monocular videos. In contrast to their offline setting where the complete time series of motions are available, our method focuses on the practical scenario where streaming data is processed by our Transformer in real-time without looking ahead.
Inverse Kinematics. Inverse kinematics (IK) is the process of calculating the variable joint parameters to produce a desired end-effector location. IK has been extensively studied in the past, with various applications in robotics and computer animation . Because no analytical solution usually exists for an IK problem, the most common way to solve the problem is through numerical methods via iterative optimization, which is costly. To speed up computation, several heuristic methods have been proposed to approximate the solution . Recently, learning-based IK solutions have attracted attention , because they can speed up inference. However, these methods are usually restricted to a scenario with a known data distribution and may not generalize well. To overcome this problem, recent works have combined IK with deep learning to make the prediction more robust and flexible . Our proposed method combines a deep neural network with IK optimization, where the IK component of our method refines the arm articulation to match the tracked hand positions from the original input (i.e., position of the hands or hand-held controllers).
Method
Although MR systems differ in the tracking technology they rely on, the global positions in Cartesian coordinates and orientations in axis-angle representation of the headset and the hand-held controllers or hands are generally available. From these, AvatarPoser reconstructs the position of the articulated joints of the user’s full body within the world . This mapping is described through the following equation,
where corresponds to the number of joints tracked by the MR system, is the number of joints of the full-body skeleton, matches the number of observed MR frames that are considered from the past, and is the body joint pose which is represented by .
Specifically, we use the SMPL model to represent and animate our human body pose. We use the first 22 joints defined in the kinematic tree of the SMPL human skeleton and ignore the pose of fingers similar to previous work .
2 Input and Output Representation
Since the 6D representation of rotations has proved effective for training neural networks due to its continuity , we convert the default axis-angle representation in the SMPL model to the rotation matrix and discard the last row to get the 6D rotation representation . During development, we observed that this 6D representation produces smooth and robust rotation predictions.
In addition to the accessible positions and orientations of the headset and hands, we also calculate the corresponding linear and angular velocities to obtain a signal of temporal smoothness. The linear velocity is given by backward finite difference at each time step :
Similar, the angular velocity can be calculated by:
followed by also converting to its 6D representation . As a result, the final input representation is a concatenated vector of position, linear velocity, rotation, and angular velocity from all given sparse inputs, which we write as:
Therefore, when the number of sparse trackers equals , the number of input features at each time step is 54.
The output of our rotation-based pose estimation network is the local rotation at each joint with respect to the parent joints . The rotation value at the pelvis, which is the root of the SMPL model, refers to the global orientation . As we use 22 joints to represent the full-body motion, the output dimension at each time step is 132.
3 Overall Framework for Avatar Full-Body Pose Estimation
Fig. 2 illustrates the overall framework of our proposed method AvatarPoser. AvatarPoser is a time series network that takes as input the 6D signals from the sparse trackers over the previous frames and the current frame and predicts global orientation of the human body as well as the local rotations at each joint with respect to its parent joint. Specifically, AvatarPoser consists of four components: a Transformer Encoder, a Stabilizer, a Forward-Kinematics (FK) Module, and an Forward-Kinematics (IK) Module. We designed the network such that each component solves a specific task.
Transformer Encoder. Our method builds on a Transformer model to extract the useful information from time-series data, following its benefits in efficiency, scalability, and long-term modeling capabilities. We particularly leverage the Transformer’s self-attention mechanism to distinctly capture global long-range dependencies in the data. Specifically, given the input signals, we apply a linear embedding to enrich the features to 256 dimensions. Next, our Transformer Encoder extracts deep pose features from previous time steps from the headset and hands, which are shared by the Stabilizer for global motion prediction, and a 2-layer multi-layer perceptron (MLP) for local pose estimation, respectively. We set the number of heads to 8 and the number of self-attention layers to 3.
Stabilizer. The Stabilizer is a 2-layer MLP that takes as input the 256-dimensional pose features from our Transformer Encoder. We set the number of nodes in the hidden layer to 256. The output of the network produces the estimated global orientation represented as the rotation of the pelvis; therefore, it is responsible for global motion navigation by decoupling global orientation from pose features and obtaining global translation from the head position through the body kinematic chain. Although it may be intuitive and possible to calculate the global orientation from a given head pose through the kinematic chain, the user’s head rotation is often independent of the motions of other joints. As a result, the global orientation at the pelvis is sensitive to the rotation of the head. Considering the scenario where a user stands still and only rotates their head, it is likely that the global orientation may have a large error, which often results in a floating avatar.
Forward-Kinematics Module. The Forward-Kinematics (FK) Module calculates all joint positions given a human skeleton model and predicted local rotations as input. While rotation-based methods provide robust results without the need to reproject onto skeleton constraints to avoid bone stretching and invalid configurations, they are prone to error accumulating along the kinematic chain. Training the network without FK could only minimize the rotation angles, but would not consider the actually resulting joint positions during optimization.
Inverse-Kinematics Module. A main problem of rotation-based pose estimation is that the prediction of end-effectors may deviate from their actual location—even if the end effector served as a known input, such as in the case of hands. This is because for end-effectors, the error accumulates along the kinematic chain. Accurately estimating the position of end-effectors is particularly important in MR, however, because hands typically often used for providing input and even small errors in position can significantly disturb interaction with virtual interface elements. To account for this, we integrate a separate IK algorithm that adjusts the arm limb positions according to the known hand positions.
Our methods performs IK-based optimization based on the estimated parameters output by our neural network. This combines the individual benefits of both approaches as explored in prior work (e.g., ). Specifically, after our network produces an output, our IK Module adjusts the estimated rotation angles of joints on the shoulder and elbow to reduce the error of hand positions as shown in Fig. 3. We thereby fix the position of the shoulder and do not optimize the other rotation angles, because we found the resulting overall body posture to appear more accurate than the output of the IK algorithm.
Given the initial rotation values estimated from our Transformer network, we calculate the positional error of the hand according to the input signals and estimated hand position through the FK Module by
where is the learning rate and is decided by the specific optimizer. To enable fast inference for real application, we stop the optimization after a fixed number of iterations.
There are several classical non-linear optimization algorithms that are suitable for optimizing inverse kinematics problems, such as Gauss-Newton method or the Levenberg-Marquardt method . In our experiment, we leverage the Adam optimizer due to its compatibility with Pytorch. We set the learning rate as .
Loss Function. The final loss function is composed of an L1 local rotational loss, an L1 global orientation loss, and an L1 positional loss, denoted by:
We set the weights , , and to 0.05, 1, and 1, respectively. For fast training, we do not include our IK Module into the training stage.
Experiments
We use the subsets CMU , BMLrub and HDM05 in AMASS dataset for training and testing. The AMASS dataset is a large human motion database that unifies different existing optical marker-based MoCap datasets by converting them into realistic 3D human meshes represented by SMPL model parameters. We split the three datasets into random training and test sets with 90% and 10% of the data, respectively. For use on VR devices, we unified the frame rate to 60 Hz.
To optimize the parameters of AvatarPoser, we adopt the Adam solver with batch size 256. We set the chunk size of input as 40 frames. The learning rate starts from and decays by a factor of 0.5 every iterations. We train our model with PyTorch on one NVIDIA GeForce GTX 3090 GPU. It takes about two hours to train AvatarPoser.
2 Evaluation Results
We use MPJRE (Mean Per Joint Rotation Error [°]), MPJPE (Mean Per Joint Position Error [cm]), and MPJVE (Mean Per Joint Velocity Error [cm/s]) as our evaluation metrics. We compare our proposed AvatarPoser with Final IK , CoolMoves , LoBSTr , and VAE-HMD , which are state-of-the-art methods working on the problems of avatar pose estimation from sparse inputs.
Since these state-of-the-art methods do not provide public source codes, we directly run Final IK in Unity and reproduce other methods to the best of our knowledge. For a fair comparison, we train all the methods on the same training and testing data. It should be noted that the original CoolMoves is a position-based method, we adapt it to rotation-based method for a fair comparison with other methods. We make all the methods work with both three (headset, controllers) and four inputs (headset, controllers, pelvis tracker). When only three inputs are provided, for input representation we do not use the pose of pelvis as a reference frame, and for the output we calculate the global orientation and translation of human body at pelvis through the kinematic chains from the given global pose of head.
The numerical results for the considered metrics (MRJRE, MPJPE, and MPJVE) for both four and three inputs are reported in Table 1. It can be seen that our proposed AvatarPoser achieves the best results on all three metrics and outperforms all other methods. VAE-HMD achieves the second best performance on MPJPE, which is followed by CoolMoves (KNN). Final IK gives the worst result on MPJPE and MPJRE because it optimizes the pose of the end-effectors without considering the smoothness of other body joints. As a result, the performance of LoBSTr, which uses Final IK for upper body pose estimation, is also low. We believe this shows the value in data-driven methods to learn motion from existing mocap datasets. However, it does not mean that traditional optimization methods are not useful. In our ablation studies, we show how inverse kinematics when combined with deep learning can improve the accuracy of hand positions.
To further evaluate the generalization ability of our proposed method, we perform a 3-fold cross-dataset evaluation among different methods. To do so, we train on two subsets and test on the other subset in a round robin fashion. Table 2 shows the experimental results of different methods tested on CMU, BMLrub, and HDM05 datasets. We achieve the best results over almost all evaluation metrics in all three datasets. Although Final IK performs slightly better than AvatarPoser in terms of MPJVE in CMU, which can only means the motions are a little bit smoother. However, the rotation error MPJRE and the position error MPJPE of Final IK, which represent the accuracy of predictions, are much larger than our method.
3 Ablation Studies
We perform an ablation study on the different submodules of our method and provide results in Table 3. The experiments are conducted on the same test set as HDM05 in Table 2. We use MPJRE [°], MPJPE [cm] as our evaluation metrics in the ablation studies to show the need for each component. In addition to the position error across the full-body joints, we specifically calculate the mean error on hands to show how the IK module helps improve the hand positions.
No Stabilizer. We remove the Stabilizer module, which predicts the global orientation, and calculate the global orientation through the body kinematic chain directly from the given orientation of the head. Table 3 shows that the MPJPE drops without Stabilizer. This is because the rotation of the head is relatively independent to the rest of the body. Therefore, the global orientation is highly sensitive to random rotations of the head. Learning the global orientation from richer information via the network is a superior way to solve the problem.
Predict Pelvis Position. In our final model, we calculate the global translation of the human body, which is located at the pelvis, from the input head position through the kinematic chain. We also try directly regressing to the global translation within the network, but the result is worse than computing via the kinematic chain according to our evaluation results.
No FK Module. We also remove the FK Module, which means the network is only trained to minimize the rotation angles without considering the positions of joints after forward kinematics calculation. When we remove the FK module, the MPJPE increases and the MPJRE decreases. This is intuitive as we only optimize the joint rotations without the IK module. While rotation-based methods provide robust results without the need to reproject onto skeleton constraints to avoid bone stretching and invalid configurations, they are prone to error accumulation along the kinematic chain.
No IK Module. We remove IK Module and only provide the results directly predicted by our neural network. Removing the IK module has little effect on the average position error of full-body joints. However, the average position error of the hands increases by almost 41%.
4 Running Time Analysis
We evaluated the run-time inference performance of our network AvatarPoser and compared it to the inference of VAE-HMD , LoBSTr , CoolMoves as shown in Fig. 6. Note that we did not include Final IK , because its integration into Unity makes accurate measurements difficult. To conduct our comparison, we modified LoBSTr to directly predict full-body motion via the GRU (denoted as LoBSTr-GRU) instead of combining Final IK and the GRU together. We measured the run time per frame (in milliseconds) on the evaluated test set on one NVIDIA 3090 GPU. For a fair comparison, we only calculated the network inference time of AvatarPoser here. Our AvatarPoser achieves a good trade-off between performance and inference speed.
Our method also requires executing an IK algorithm after the network forward pass. Each iteration costs approximately 6 ms, so we set the number of iterations to 5 to keep a balance between inference speed and the accuracy of the final hand position. Note that the speed could be accelerated by adopting a more standard non-linear optimization.
5 Test on a Commercial VR System
To qualitatively assess the robustness of our method, we executed our algorithm on live recordings from an actual VR system. We used an HTC VIVE HMD as well as two controllers, each providing real-time input with six degrees of freedom (rotation and translation). Fig. 7 shows a few examples of our method’s output based on sparse inputs.
Conclusions
We presented our novel Transformer-based method AvatarPoser to estimate realistic human poses from just the motion signals of a Mixed Reality headset and the user’s hands or hand-held controllers. By decoupling the global motion information from learned pose features and using it to guide pose estimation, we achieve robust estimation results in the absence of pelvis signals. By combing learning-based methods with traditional model-based optimization, we keep a balance between full-body style realism and accurate hand control. Our extensive experiments on the AMASS dataset demonstrated that AvatarPoser surpasses the performance of state-of-the-art methods and, thus, provides a useful learning-based IK solution for practical VR/AR applications.
Acknowledgments: We thank Christian Knieling for his early explorations of learning-based methods for pose estimation with us at ETH Zürich. We thank Zhi Li, Xianghui Xie, and Dengxin Dai from Max Planck Institute for Informatics for their helpful discussions. We also thank Olga Sorkine-Hornung and Alexander Sorkine-Hornung for early discussions.