One-Stage 3D Whole-Body Mesh Recovery with Component Aware Transformer

Jing Lin, Ailing Zeng, Haoqian Wang, Lei Zhang, Yu Li

Introduction

Expressive whole-body mesh recovery aims to jointly estimate the 3D human body poses, hand gestures, and facial expressions from monocular images. It is gaining increasing attention due to recent advancements in whole-body parametric models (e.g., SMPL-X ). This task is a key step in modeling human behaviors and has many applications, e.g., motion capture, human-computer interaction. Previous research focus on individual tasks of reconstructing human body , face , or hand . However, whole body mesh recovery is particularly challenging as it requires accurate estimation of each part and natural connections between them.

Existing learning-based works use multi-stage pipelines for body, hand, and face estimation to achieve the goal of this task. As depicted in Figure 1(a), these methods typically detect different body parts, crop and resize each region, and feed them into separate expert models to estimate the parameters of each part. The multi-stage pipeline with different estimators for body, hand, and face results in a complicated system with a large computational complexity. Moreover, the blocked communications among different components inevitably cause incompatible configurations, unnatural articulation of the mesh, and implausible 3D wrist rotations as they cannot obtain informative and consistent clues from other components. Some methods attempt to alleviate these issues by designing additional complicated integration schemes or elbow-twist compensation fusion among individual body parts. However, these approaches can be regarded as a late fusion strategy and thus have limited ability to enhance each other and correct implausible predictions.

In this work, we propose a one-stage framework named OSX for 3D whole-body mesh recovery, as shown in Figure 1(b), which does not require separate networks for each part. Inspired by recent advancements in Vision Transformers , which are effective in capturing spatial information in a plain architecture, we design our pipeline as a component-aware Transformer (CAT) composed of a global body encoder and a local component-specific decoder. The encoder equipped with body tokens as inputs captures the global correlation, predicts the body parameters, and simultaneously provides high-quality feature map for the decoder. The decoder utilizes a differentiable upsample-crop scheme to extract part-specific high-resolution features and adopt the keypoint-guided deformable attention to precisely locate and estimate hand and face parameters. The proposed pipeline is simple yet effective without any manual post-processing. To the best of our knowledge, this is the first one-stage pipeline for 3D whole-body estimation. We conduct comprehensive experiments to investigate the effects of the above designs and compare our method, with existing works on three benchmarks. Results show that OSX outperforms the state-of-the-art (SOTA) by 9.59.5% on AGORA, 7.87.8% on EHF, and 13.413.4% on the body-only 3DPW dataset.

In addition, existing popular benchmarks, as illustrated in the first row of Figure 2, are either indoor single-person scenes with limited images (e.g., EHF ) or outdoor synthetic scenes (e.g., AGORA ), where the people are often too far from the camera and the hands and faces are frequently obscured. In fact, human pose estimation and mesh recovery is a fundamental task that benefits many downstream applications, such as sign language recognition, gesture generation, and human-computer interaction. Many scenarios, such as talk shows and online classes, are of vital importance to our daily life yet under-explored. In such scenarios, the upper body is a major focus, whereas the hand and face are essential for analysis. To address this issue, we build a large-scale upper-body dataset with fifteen human-centric real-life scenes, as shown in Figure 2(f) to (t). This dataset contains many unseen poses, diverse appearances, heavy truncation, interaction, and abrupt shot changes, which are quite different from previous datasets. Accordingly, we design a systematical annotation pipeline and provide precise 2D whole-body keypoint and 3D whole-body mesh annotations. With this dataset, we perform a comprehensive benchmarking of existing whole-body estimators.

Our contributions can be summarized as follows.

We propose a one-stage pipeline, OSX, for 3D whole-body mesh recovery, which can regress the SMPL-X parameters in a simple yet effective manner.

Despite the conceptual simplicity of our one-stage framework, it achieves the new state of the art on three popular benchmarks.

We build a large-scale upper-body dataset, UBody, to bridge the gap between the basic task and downstream applications and provide precise annotations, with which we conduct benchmarking of existing methods. We hope UBody can inspire new research topics.

Related Work

Whole-body mesh recovery targets to localize mesh vertices of all human components, including body, hands, and face from monocular images. Most previous works focus only on individual hand , face , and body reconstruction. In contrast, the joint whole-body estimation methods are less addressed. Some optimization-based works reconstruct 3D bodies by fitting the detected 2D keypoints from images with additional constraints, but they are slow and prone to local optima . Thanks to the whole-body parametric model (e.g., SMPL-X ), learning-based models emerge to train networks to predict expressive body pose, shape, hand gesture, and facial expression. Due to the low resolution of hands and face, these whole-body methods crop and resize the hands and face images to higher resolutions and feed them into separate expert networks to conduct the corresponding parameter regression. Specifically, ExPose introduces body-driven attention for higher-resolution crops of the face and hand estimation, a dedicated refinement module, and part-specific knowledge from existing hand-only and face-only datasets. FrankMocap presents a regression-and-integration method to build a fast and accurate system. PIXIE produces animatable whole body with realistic facial details via a moderator to fuse body part features adaptively. Recently, Hand4Whole utilizes both body and hand joint features for accurate 3D wrist rotation and smooth connection between body and hands.

Nevertheless, these methods aim at high performance by using separate networks in a divide-and-conquer fashion for different components and a specific fusion module to paste them together. The multi-stage pipelines lead to high complexity and inevitably cause inconsistent and unnatural articulation of the mesh and implausible 3D wrist rotations, especially in occluded, truncated, and blurry contexts. Until now, one-stage methods in this task are unexplored.

2 Benchmarks of Expressive Body

Some datasets with parametric model annotations have been developed to advance the field. Table 1 summarizes these datasets from the annotation type, size, scene diversity, etc. To be specific, EHF is the first evaluation dataset for SMPL-X-based models, which is built by capturing 3D body shapes with a scanning system and then fitting the SMPL-X model to the scans. AGORA is a synthetic dataset with high realism and accurate ground truth, which is by far the most commonly used test data due to the diversity of subjects, environments, clothes, and occlusions. Notably, people in AGORA are often far from the camera, and their hands and face are obscured and have small resolutions, making existing methods focus more on body rather than hand and face estimation.

Since marker-based 3D mocap labels are hard to obtain, there are a few annotation methods for high-precision labeling for both monocular indoor and outdoor scenes. FBA emphasizes the severe failure cases of existing body recovery methods on consumer video data due to unusual camera viewpoints and aggressive truncations. They annotate pseudo 2D body keypoints and SMPL annotations via HMR on 13k frames across four action recognition datasets. Multi-shot-AVA also argues that data from edited media, like movies with rich appearances, interactions between humans, and various temporal contexts, is valuable. They apply the proposed multi-shot optimization on AVA to get pseudo 3D ground truth. Interestingly, a body recovery benchmark finds that simply using the 2D COCO dataset with pseudo-3D labels can surprisingly achieve a better performance and generalization ability. To complement these prior datasets and focus on expressive body recovery, we construct a new benchmark with high-quality 2D and 3D whole-body annotations.

Method

A one-stage framework is vital to simplify the cumbersome processes without hand-craft and complex integration designs. However, translating from multi-stage methods directly to a one-stage method is nontrivial. We take the present state-of-the-art method Hand4Whole as an example to perform some preliminary studies on bringing the gap between the multi-stage method and one-stage approach. On the one hand, we replace its separate backbones with a shared backbone for all human components. On the other hand, we explore different crop-and-resize image resolutions for the hands and face, as they usually have small image resolutions.

Table 2 shows that, when we transition from the original setup (Ori.) to a shared backbone (Share Backbone), all recovery errors are severely deteriorated on two datasets. Specifically, MPVPE increases from 183.8183.8mm to 202.3202.3mm (a 10.1% drop) on AGORA , and from 77.577.5mm to 84.784.7mm (a 9.3% drop) on EHF for all components (All). These results indicate that extracting the multi-component whole-body features with a shared backbone is difficult. Notably, the hand estimation performance deteriorates by 30.0% on EHF. Based on the results of different resolutions, we summarize some interesting observations as follows: (i) Overall, changing the resolution of the hand results in a larger performance drop than the face on EHF; (ii) When not sharing a backbone, the results are generally worse with smaller input resolutions of the hands and face.

2 Building Component Aware Transformer

3 Body Regression via Global Encoder

4 High-Resolution Decoder for Hand and Face

Loss Function. OSX is trained in an end-to-end manner by minimizing the following loss function:

The four items are calculated as the L1 distance between the ground truth values and the predicted ones. Specifically, LsmplxL_{smplx} provides the explicit supervision of the SMPL-X parameters. Lkpt3DL_{kpt3D}, Lkpt2DL_{kpt2D}, and Lbbox2DL_{bbox2D} are regression losses for 3D whole-body keypoints, projected 2D whole-body keypoints, and left/right hands and face 2D bounding boxes. More details are provided in the Appendix.

UBody–An Upper Body Dataset

3D whole-body mesh recovery from videos is a basic computer vision task, where it can provide comprehensive motion, gesture, and expression information to understand how humans perceive and act. However, existing datasets lack scenes of downstream tasks, such as sign language recognition, gesture generation, emotion recognition, and real-life scenarios recorded as VLOGs, making recent state-of-the-art methods hard to generalize well on these scenes. Interestingly, these scenarios are more concerned with the representations of upper bodies. We take this insight and present a novel large-scale benchmark for the expressive upper body mesh recovery as shown in Figure 2(f) to (t), named UBody. Our annotation pipeline is in Figure 4. Due to the page limit, we put the data collection, data annotation processes, and annotation visualization in Appendix.

Our annotation pipeline produces far better 3D pseudo-GT fits with a shorter running time than the previous optimization-based and learning-based methods . Figure 5(a) compares our 2D annotation results with the two wildly used annotation methods (OpenPose and MediaPipe ). The quality of our 2D annotations is much more accurate, especially in terms of hand details and the robustness of occlusion and blur. Figure 5(b) compares the 3D annotation of ours with the SOTA NeuralAnnot method on COCO. The quality of our approach is also better for the naked eye in terms of the fit of the body shape and the whole-body poses.

2 Data Characteristics

Compared to the popular datasets illustrated in Figure 2 (a) to (e) and the related human-centric datasets listed in Table 1, UBody possesses unique features that present new challenges for future research. Many videos are from edited media with highly diverse scenes and rich human actions and gestures. They have abrupt shot changes and dynamic camera viewpoints, leading to discontinuities between the frames. Close-up shots of humans cause severe truncation, making existing methods tend to fail. Meanwhile, they have varying degrees of interaction with objects and body components, subtitles, and special effects as occluded scenes. Also, there are high variations in background and light. Those conditions have not appeared in previous datasets. All scenes in UBody have rich hand gestures and facial expressions, making the recognition models pay more attention to these important body components. Lastly, all of these real-life videos provide audio as additional information to serve future multi-modality methods. We also provide statistical comparisons between the key features of UBody and the wildly used dataset AGORA in Figure 6. AGORA’s hand/face bounding box area is generally small, while UBody pays more attention to diverse hand and face scales as evidenced by its more dispersed area distribution. Meanwhile, UBody has more visible face/hand keypoints, underscoring the importance of recognizing hand gestures and facial expressions. Lastly, UBody’s inclusion of real-life videos provides new possibilities for subsequent spatio-temporal modeling that are not available in AGORA, which is an image-based dataset.

Experiment

Due to the page limit, we leave the detailed experiment setup, implementation, annotation visualization, qualitative comparison with SOTA methods, and more benchmark results and analyses in the appendix.

Datasets. We use COCO-Wholebody , MPII , and Human3.6M as the training set. Unlike previous multi-stage methods , we do not use additional hand-only and face-only datasets for training as a simple baseline for a one-stage method. The SMPL/SMPL-X pseudo-GTs are obtained from EFT and NeuralAnnot .

Evaluation metrics. For 3D whole-body mesh recovery, we utilize the mean per-vertex position error (MPVPE) as our primary metric. In addition, we apply Procrustes Analysis (PA) to the recovered mesh, and report the PA-MPVPE after rigid alignment. For AGORA, we also report normalized mean vertex error (N-PMVPE) to compensate for missing detection. Hand error is calculated as the mean of the left and right hands. For 3D body-only recovery on 3DPW, we follow previous works to report the mean per joint position error (MPJPE) and PA-MPJPE. All reported errors are in units of millimeters.

Implementation details. OSX is implemented in Pytorch and trained using the Adam optimizer with an initial learning rate of 1×10−41\times 10^{-4} for 14 epochs. Scaling, rotation, random horizontal flip, and color jittering are used as data augmentations during training. We set the number of body tokens Tb\mathbf{T}_{b} and component tokens Tc\mathbf{T}_{c} to 27 and 92, respectively.

2 Comparisons with Existing Methods

Table 3 provides a comprehensive comparison of OSX and existing whole-body mesh recovery methods. As the first one-stage method, OSX surpasses existing multi-stage models with complex designs in most cases. Notably, OSX has not been trained on hand-only and face-only datasets . Our All MPVPEs show a 9.59.5% improvement on AGORA test set and 7.87.8% improvement on EHF than SOTA . Since AGORA is a more complex and natural dataset than EHF, previous works claim it is more convincing and representative of real-world scenarios. We also visualize the misleading high-error cases on EHF in Figure 7. Besides, we obtain a SOTA performance on the body-only dataset, 3DPW, with a 13.413.4% error reduction compared to these whole-body methods. More qualitative results are available in the appendix.

3 Ablation Study

Impact of the component-aware decoder. Unlike body-only pose estimation, whole-body mesh recovery requires attention to both the body’s posture, which is on a larger spatial scale, and the gesture and expression of the hands and face, which are on a finer scale. To handle the resolution issue in a one-stage pipeline, we propose the component-aware decoder attached to the component-aware encoder. First, in the upper Table 5, we verify the effectiveness of the proposed decoder for both hand and face regression. We can observe a significant drop without the decoder (e.g., w/o H.D. and w/o F.D., indicating that simply regressing the low-resolution hand and face directly from the encoder is inferior. Moreover, the errors will also increase without the proposed keypoint-guided deformable attention scheme, as shown in the medium Table 5. In particular, the performance of the hand estimation is highly influenced, showing that hand pose estimation attends more to the sparsely deformable spatial information to obtain better queries.

Impact of the up-sampling strategy. To relieve the low-resolution problem of hand and facial features, we design the feature up-sampling strategy in the decoder to obtain multi-scale higher-resolution features. The lower Table 5 presents the impact of different up-sampling scale. As the up-sampling scale increases, the MPVPE decreases and then reaches a saturation point. Therefore, we use three scales (i.e., [×1,×2,×4][\times 1,\times 2,\times 4]) by default in our experiments.

4 Benchmark on UBody

As a new dataset, we provide both quantitative and qualitative results on UBody. Table 5 presents the performance comparisons of existing 3D whole-body methods. The general result ranking is similar to AGORA. Since the upper body is closer to the camera, their errors will be smaller than AGORA. However, the hand and face will play a more important role than previous data. Besides, we finetune Hand4Whole on AGORA and test again, and we find all errors are significantly enlarged. This observation can be attributed to the data distribution gap between AGORA and UBody, as shown in Figure 6. Moreover, we train OSX on our train set and find a 16.1% improvement compared to the original pretrained model, indicating that UBody can serve to improve the performance on downstream real-life scenes.

Conclusion

In this work, we propose the first one-stage pipeline for 3D whole-body mesh recovery that achieves SOTA performance on three benchmarks in a simple yet effective manner. Moreover, to bridge the gap between the basic task of full-body pose and shape estimation and their downstream tasks, we develop a large-scale dataset with comprehensive scenes covering our daily life. With our proposed annotation method, we show that training on UBody can effectively improve the performance of mesh recovery in upper-body scenes. We hope this work can contribute new insights to this area, both in terms of methodology and dataset.

Limitation and future work. Currently, our training does not use additional hand and face-specific datasets. It is worth studying how to make the best use of them in our pipeline to further improve performance. Also, we can validate the effectiveness of UBody on some downstream applications, e.g., gesture recognition, driving avatar.

Acknowledgements: This work was partially funded through the National Key Research and Development Program of China (Project No.2022YFB36066), in part by the Shenzhen Science and Technology Project under Grant (CJGJZD20200617102601004, JCYJ20220818101001004).

References

Overview

This supplementary material presents more details and additional results not included in the main paper due to page limitation. The list of items included are:

More experiment setup and details in Sec. A.

Efficiency comparison with SOTA in Sec. B.

Inter-scene benchmark on UBody dataset in Sec. E.

Qualitative comparisons with SOTA in Sec. F.

A Experiment Setup

Evaluation metrics. To quantitatively evaluate the performance of human mesh recovery, MPVPE, PA-MPVPE, MPJPE, and PA-MPJPE are used as evaluation metrics. Besides, we also report normalized mean vertex error (NMVE) and normalized mean joint error (NMJE) by the standard detection metric, F1 score (the harmonic mean of recall and precision) to penalize models for misses and false positives on AGORA test set with many multi-person scenes.

Implementation details. Our OSX model is implemented in Pytorch. It is trained with Adam optimizer (β1=0.1,β2=0.999\beta_{1}=0.1,\beta_{2}=0.999) using the Cosine Annealing scheme for 14 epochs. The learning rate is initially set to 1×10−41\times 10^{-4}. The batch size is set to 192. Random scaling, rotation, horizontal flip, and color jittering are used as data augmentations during training. The spatial size of the input image is 256×192256\times 192. The number of body tokens Tb\mathbf{T}_{b} and component tokens Tc\mathbf{T}_{c} are set to 27 and 92, respectively. During experiments on the AGORA-test set, we remove the decoder as we find that the decoder increases training time and does not significantly improve performance on AGORA-test set. This observation may be attributed to the fact that the main problem of AGORA is occlusion, while the decoder aims to estimate hands/face at a finer level.

B Efficiency comparison with SOTA methods

We report the complexity comparisons including average inference time, number of model parameters, FLOPs, and the NMJE-All on AGORA-test in Table S-1. The numbers are measured for single-person regression on the same resolution input using a machine with an NVIDIA A100 GPU. OSX has the shortest inference time and lowest error, indicating the advantages in practical applications.

C Experiment on AGORA Dataset

In this part, we report the complete result on the AGORA test set and the experiment result on the AGORA val set.

AGORA Test Set. Table S-2 depicts the complete result on the AGORA test set. All the results are taken from the official leaderboard. As shown, our OSX outperforms other competitors on most metrics, especially on the evaluation of the body and full-body recovery. More specifically, for full-body reconstruction, OSX even surpasses PyMAF-X by 10.6 mm, 9.1 mm, 2.9 mm, and 4.7 mm on NMVE, NMJE, MVE, and MPJPE, respectively. Since PyMAF-X has a lower detected person ratio, they have similar results on MVE and MPJPE metrics, which only calculate the matched person. The NMVE and NMJE will take the misses and false positives into account, and we have overall better multi-person estimation with more improvement under the metrics. Notably, although OSX does not use extra hand-only and face-only datasets, it can achieve competitive results on hand and face metrics, which demonstrates the effectiveness of our component-aware decoder.

AGORA Val Set. Table S-3 shows the result on the AGORA val set. All the results are taken from except OSX. Although we do not use extra hand/face specific datasets during training, OSX outperforms the SOAT method Hand4Whole by 8.3% on the MPVPE-all, demonstrating the effectiveness of our one-stage method.

D UBody: An Upper Body Dataset

To bridge the gap between the basic human mesh recovery task and its downstream applications, we design UBody with two rules. First, we research a wide range of human-related downstream tasks with upper-body scenes, including gesture recognition , sign language recognition, and translation , person clustering , emotion analysis, speaker verification , micro-gesture understanding , audio-visual generation and separation , human action recognition, and localization , and human video segmentation . We select the corresponding high-quality datasets from these existing tasks as a part of our data for the corresponding scenarios. In order to ensure a balanced amount of data for each scene, for datasets with many videos (e.g., lasting 20k minutes), we manually selected the videos in which the upper body appeared more frequently.

Second, with all kinds of athletic competitions, entertainment shows, we media, online conferences, and online classes being more and more indispensable, we carefully selected a large number of rich videos from YouTube to provide new opportunities and challenges for potential applications.

Since some untrimmed videos may have missing main characters, extraneous images such as opening and closing credits, and repetitive actions, we manually fine-cut the long videos. Each edited video is 10 seconds long, which ensures the high quality of the video.

In order to prevent infringement of ownership rights, we only provide download links to the corresponding videos and our labels without any personal information.

In summary, we collect fifteen real-life scenarios with more than 105,1k frames. We split the train/test sets from two protocols as follows.

Intra-scene: in each scene, the former 70% of the videos are the training set, and the last 30% are the test set. The benchmark was provided in the main paper.

Inter-scene: we use ten scenes of the videos as the training set and the other five scenes as the test set. Due to the page limit, we present the benchmark in Table S-4.

D.2 Data Annotation Processes

As shown in Figure S-1, we design a thorough whole-body annotation pipeline with high precision. It is divided into two stages: 2D whole-body keypoint annotation and 3D SMPLX annotations fitting. Since UBody scenes have a number of unpredictable transitions and cutscenes that make it difficult to use the temporal smoothing approaches , the annotation is conducted on a single frame.

2D whole-body keypoint annotation: We first detect all persons and their hands in an image via a specific human and hand detector BodyHands shown as Body Detector and Hand Detector in Figure 4. Leveraging the recent state-of-the-art 2D pose estimator ViT-Body-only , we use the pre-trained model trained on the COCO dataset to localize 17 body keypoints for each detected single person, named KBodyK_{Body}, which shows highly robust results on many scenes. Due to the diverse scales and motion blur for the fast-moving hands, we find that Hand Detector will output false positive samples or miss some hands. To enhance the performance of hand detection, we train a 2D whole-body estimator on COCO-wholeBody with 133133 2D keypoints, called ViT-WholeBody following the model design of ViTPose and masked autoencoder pre-trained scheme . ViT-WholeBody can provide high-recall hand keypoints KHandK_{Hand}, but the localization precision is low because of the fully one-stage pipeline and low-resolution of hands from the raw image. Accordingly, We can obtain coarse hand bounding boxes by calculating the maximum, and minimum values of the detected left and right-hand keypoints to correct the hand boxes from Hand Detector via an IoU matching strategy. Then, we use the fine hand boxes to crop the hand patches, resize them to a larger size, and put them into our specific pre-trained ViT-Hand-only model trained with the hand labels from the COCO-Whole dataset. In summary, ViT-WholeBody will output the body, hand, and face 2D keypoints. We use the body output from ViT-Body-only to replace the KBodyK_{Body}, and use the fine hand keypoints from ViT-Hand-only to change the KHandK_{Hand}. As the face of the current SMPL-X model does not require much detail, we simply use the 2D face keypoints KFaceK_{Face} obtained from ViT-WholeBody.

3D whole-body mesh recovery annotation: Different from previous optimization-based annotation that may output implausible poses, we use our proposed OSX to estimate the SMPL-X parameters from human images as a proper 3D initialization to provide pseudo-3D constraints. Benefiting from current 2D keypoint localization that tends to be more accurate, we additionally supervise the projected 2D whole-body keypoints by the above annotated 2D whole-body keypoints as a way to train OSX. More importantly, to avoid performance degradation from not accurate enough initial labeling and consistently push up the 3D annotation quality, we propose an iterative training-labeling-revision loop for every 30 epochs to train 120 epochs in total.

E Inter-Scene Benchmark on UBody dataset

Due to the page limit, we further provide another data protocol comparison to show the usage of the proposed UBody. Table S-4 presents the performance comparisons of existing 3D whole-body methods. Inter-scene test shows large errors than the intra-scene test due to the different motion and gesture distributions. The model finetuned on AGORA still has a significant gap than trained on the COCO dataset. Furthermore, we also train Hand4Whole and UBody on our training set, we can find a consistent improvement compared to the original pretrained model, indicating that UBody can serve to bridge the gap among these downstream real-life scenes. Moreover, different from single-frame AGORA and EHF, UBody provides videos, which can drive progress in spatial-temporal modeling on such edit media sources.

F Qualitative with SOTA method

Qualitative comparisons on AGORA: We compare the mesh quality on the AGORA dataset in Figure S-2. Agora is a synthetic dataset with many challenging factors like heavy occlusion, dark environment, and unnatural multi-person interaction. It only has limited actions, e.g., taking phones, walking, sitting, etc. We can see OSX outperforms ExPose and Hand4Whole consistently in terms of global body orientations, whole-body poses, and hand pose.

Qualitative comparisons on EHF: The visual comparisons of whole-body mesh recovery quality on the EHF dataset can be found in Figure S-3. As can be seen, OSX estimates the most accurate whole-body poses, in which the body parts like hands, feet, and hands are better aligned with the person in the image.

Qualitative comparisons on UBody: The qualitative comparison on our UBody is in Figure S-4. UBody focuses more on the expressive upper body part. Hand4Whole and our OSX produces better body mesh recoveries than ExPose . Close inspection of the hand part shows that our hand recovery is more accurate than Hand4Whole.

Visualization of our annotation on UBody: The visualizations of our SMPL-X annotation in our UBody can be found in Figure S-5, S-6, and S-7. Our annotation produces high-quality ground truth. In many challenging cases of expressive hand poses, our estimated mesh can capture fine-level details.