PyMAF-X: Towards Well-aligned Full-body Model Regression from Monocular Images
Hongwen Zhang, Yating Tian, Yuxiang Zhang, Mengcheng Li, Liang An, Zhenan Sun, Yebin Liu
Introduction
Recent years have witnessed the rise of the regression-based paradigm in recovering body , hand , face , and even full-body models from monocular images. These methods learn to predict model parameters directly from images in a data-driven manner. Despite the high efficiency and promising results, regression-based methods typically suffer from coarse alignment between the predicted meshes and image observations.
When recovering the parametric body or full-body models , minor rotation errors accumulated along the kinematic chain may lead to noticeable drifts in joint positions (see the top-left example in Fig. 1), since joint poses are represented as relative rotations w.r.t. their parent joints. In order to generate well-aligned results, optimization-based methods include data terms in the objective function so that the alignment between the projection of meshes and 2D evidence can be optimized explicitly. Similar strategies are also exploited in regression-based methods to impose 2D supervisions upon the projection of estimated meshes in the training procedure. However, during testing, these deep regressors either are open-loop or simply include an Iterative Error Feedback (IEF) loop in their architectures. As shown in Fig. 2(a), IEF reuses the same global feature in its feedback loop, making the regressor hardly perceive the mesh-image misalignment in the inference phase.
As suggested in previous works , neural networks tend to retain high-level information and discard detailed local features when reducing the spatial size of feature maps. In order to leverage spatial information in the regression networks, it is essential to extract pixel-wise contexts for fine-grained perception. Several attempts have been made to leverage pixel-wise representation such as part segmentation or dense correspondences in their regression networks. Though pixel-level evidence is considered, it is still challenging for those methods to learn structural priors and get hold of spatial details simultaneously based merely on high-resolution contexts.
Motivated by the above observation, we design a Pyramidal Mesh Alignment Feedback (PyMAF) loop in our regression network to exploit multi-scale and position-sensitive contexts for better mesh-image alignment. The central idea of our approach is to correct parametric deviation explicitly and progressively based on the alignment status. In PyMAF, mesh-aligned evidence will be extracted from the spatial features according to the 2D projection of the estimated mesh and then fed back to the regressors for parameter updates. As illustrated in Fig. 2, the mesh alignment feedback loop takes advantage of more informative features for parameter correction compared with the commonly used iterative error feedback loop . In order to leverage multi-scale contexts, mesh-aligned evidence is extracted from a feature pyramid so that the coarse-aligned meshes can be corrected with large step sizes based on the lower-resolution features. To enhance these mesh-aligned features, an auxiliary task is imposed on the highest-resolution feature to infer pixel-wise dense correspondence, guiding the image encoder to preserve the most related information in the spatial feature maps. Meanwhile, a spatial alignment attention mechanism is introduced to fuse the grid and mesh-aligned features so that the regressor could be aware of the whole image contexts.
Since the SMPL family includes the hand and face models, PyMAF can be easily modified to reconstruct the hand and face meshes. We leverage three part-specific PyMAF networks as part experts to predict body, hand, and face parameters, and propose PyMAF-X for expressive full-body mesh recovery. Benefiting from the well-aligned results of each PyMAF-based expert, PyMAF-X can produce plausible full-body mesh results in common scenarios even using the most naive integration strategy . However, as shown in Fig. 1, the naive “Copy-Paste” integration may lead to unnatural wrist poses under challenging cases. To address this issue, we propose an adaptive integration strategy to adjust the twist rotation of the elbow poses so that the elbow and wrist poses could be more compatible. In this way, the updated twist rotation of the elbow joint serves as compensation for the wrist joint and helps to produce natural wrist poses in the full-body model. Moreover, since the twist component of the elbow poses is the rotation around the elbow-to-wrist bone, it barely changes the position of the body and hand joints, which is the key to maintaining the well-aligned performances of body and hand experts. Different from existing full-body solutions , our method do not rely on additional networks to infer the wrist poses, and hence bypass the learning issue raised by insufficient full-body mesh annotations.
The contributions of this work can be summarized as follows:
A mesh alignment feedback loop is proposed for regression-based human mesh recovery, where mesh-aligned evidence is exploited to correct parametric errors explicitly so that the estimated meshes can be better aligned with the input images.
A feature pyramid is incorporated with the mesh alignment feedback loop so that the regression network can leverage multi-scale contexts. This yields the Pyramidal Mesh Alignment Feedback (PyMAF) loop, a novel architecture for human mesh recovery.
An auxiliary pixel-wise supervision and spacial alignment attention are introduced in PyMAF to enhance the mesh-aligned features such that they can be more informative, relevant, and aware of the whole image contexts.
PyMAF is further extended as PyMAF-X for full-body mesh recovery, where an adaptive integration strategy with the elbow-twist compensation is proposed to avoid unnatural wrist poses while maintaining the alignment of the body and hand estimations.
An early version of this work has been published as a conference paper . We have made significant extensions to our previous work from three aspects. First, PyMAF is improved to be more accurate with the newly introduced spatial alignment attention, which effectively enhances the feature learning and further improves the mesh-image alignment. Second, PyMAF goes beyond body mesh recovery and is extended to reconstruct hand and full-body models from monocular images. The well-aligned performance of the body- and hand-specific PyMAF makes it more promising to produce well-aligned full-body meshes. Third, an adaptive integration strategy is proposed to assemble predictions from body and hand experts. Such a strategy effectively addresses the unnatural wrist issues while maintaining the part-specific alignment. Based on these updates, our final method PyMAF-X achieves new state-of-the-art results both qualitatively and quantitatively, contributing novel solutions towards the well-aligned and natural recovery of full-body models from monocular images.
Related Work
Monocular recovery of human meshes has been actively studied in recent years. Aiming at the same goal of producing well-aligned and natural results, two different paradigms for human mesh recovery have been investigated in the research community. In this subsection, we give a brief review of these two paradigms and refer readers to for a more comprehensive survey.
Optimization-based Approaches. Pioneering work in this field mainly focus on the optimization process of fitting parametric models (e.g., SCAPE and SMPL ) to 2D observations such as keypoints and silhouettes . In their objective functions, prior terms are designed to penalize the unnatural shape and pose, while data terms measure the fitting errors between the re-projection of meshes and 2D evidence. Based on this paradigm, different updates have been investigated to incorporate information such as 2D/3D body joints , silhouettes , part segmentation , dense correspondences in the fitting procedure. Despite the well-aligned results obtained by these optimization-based methods, their fitting process tends to be slow and sensitive to initialization. Recently, Song et al. exploit the learned gradient descent in the fitting process. Though this solution leverages rich 2D pose datasets and alleviates many issues in traditional optimization-based methods, it still relies on the accuracy of 2D poses and breaks the end-to-end learning. Alternatively, our solution supports end-to-end learning and is also able to leverage rich 2D datasets thanks to the progress (e.g., SPIN , EFT , and NeuralAnnot ) in the generation of more precise pseudo 3D ground-truth for 2D datasets .
Regression-based Approaches. Alternatively, taking advantage of the powerful nonlinear mapping capability of neural networks, recent regression-based approaches have made significant advances in predicting human models directly from monocular images. These deep regressors take 2D evidence as input and learn model priors implicitly in a data-driven manner under different types of supervision signals during the learning procedure. To mitigate the learning difficulty of the regressor, different network architectures have also been designed to leverage proxy representations such as silhouette , 2D/3D joints , segmentation and dense correspondences . Such strategies can benefit from synthetic data and the progress in the estimation of proxy representations . In these regressors, though supervision signals are imposed on the re-projected models to penalize the mismatched predictions during training, their architectures can hardly perceive the misalignment during the inference phase. In comparison, the proposed PyMAF is a close-loop for both training and inference, which enables a feedback loop in our regressor to leverage spatial evidence for better mesh-image alignment of the estimated human models.
Directly regressing model parameters from images is very challenging, even for neural networks. Existing methods have also offered non-parametric solutions to reconstruct human body models. Among them, volumetric representation , mesh vertices , and position maps have been adopted as regression targets. Using non-parametric representations as the regression targets is more readily to leverage high-resolution features but needs further processing to retrieve parametric models from the outputs. Besides, the mesh surfaces of non-parametric outputs tend to be rough and more sensitive to occlusions without additional structure priors. In our solution, the deep regressor uses spatial features at multiple scales for both high-level and fine-grained perception. It produces parametric models directly with no further processing required.
Recently, there are also numerous efforts devoted to achieving or handling multi-person recovery , video inputs , occlusions , more accurate shape , ambiguities , camera estimation , imbalanced data , pseudo ground-truth generation , and clothed human reconstruction . Our work is complementary to them and focuses on the design of regressor architectures for single-image well-aligned body and full-body mesh recovery.
2 Full-body Mesh Recovery
Compared with the large number of solutions for the body-only , hand-only , and face-only mesh recovery, the full-body mesh recovery receives less attention due to its challenging nature and the lack of annotated datasets. Similar to the developments of body-only mesh recovery algorithms, the research in the field of full-body mesh recovery begins with the proposal of full-body models, including Frank , Adam , SMPL-X , and GHUM , etc., and their corresponding optimization-based methods . Recently, several regression-based methods have been proposed to overcome the slow and unnatural issues of optimization-based methods.
Following the pioneering work ExPose , regression-based methods typically consist of three part-specific modules, namely part experts, to predict parameters of body, hand, and face from the corresponding part images cropped from original inputs. They differ mainly in the architecture of the part experts and the strategy to integrate part estimations. As the part experts are basically chosen from the body- or hand-only mesh recovery solutions, the integration strategy to sew up independent estimations becomes an essential aspect of a regression-based full-body method. The most straightforward strategy to integrate the body and hand estimations would be the “Copy-Paste” . To obtain more natural integration results, learning-based strategies are proposed in recent state-of-the-art methods . For instance, FrankMocap learns to correct the arm poses based on the distance between the wrist positions predicted by body and hand experts. Zhou et al. incorporate body features in the learning of the hand expert so that the predicted hand poses could be more compatible with the arm. PIXIE introduces a learnable moderator to merge body and hand features for the regression of wrist and finger poses. All the above solutions rely on additional networks to predict or correct the wrist poses with the condition of body information, which is typically inferior to the original hand poses predicted by the hand expert, resulting in degraded alignment on the hand parts. Recently, Hand4Whole proposes to learn wrist poses based on the positions of selected hand joints but does not consider the compatibility of arm poses. In contrast to existing solutions, PyMAF-X resorts to the adjustment of the twist components of wrist and elbow poses, which produces natural wrist rotations while maintaining the well-aligned performances of each part expert during the integration. Besides, our motivation and method also differ from the previous work that decomposes the twist components in the inverse kinematics problem.
3 Iterative Fitting in Regression Tasks
Strategies for incorporating fitting processes along with regression tasks have also been investigated in the literature. For human mesh recovery, Kolotouros et al. combine an iterative fitting procedure with the training procedure to generate more accurate ground truths for better supervision. Several attempts have been made to deform human meshes so that they can be aligned with the intermediate estimates such as depth maps , part segmentation , and dense correspondences . These approaches adopt intermediate estimations as fitting objectives and hence rely on their quality. In contrast, our approach uses the currently estimated meshes to extract deep features for refinement, enabling the fully end-to-end learning of the deep regressor.
In a broader view, remarkable efforts have been made to involve iterative fitting strategies in other computer vision tasks, including facial landmark localization , human/hand pose estimation , etc. For generic objects, Pixel2Mesh progressively deforms an initial ellipsoid by leveraging perceptual features. Following the spirit of these works, we exploit new strategies to extract fine-grained evidence and contribute novel solutions in the context of human mesh recovery.
Method
In this section, we will elaborate technical details of our approach. We first present PyMAF, a powerful model for regression-based human mesh recovery, then extend it to PyMAF-X for full-body mesh recovery.
As illustrated in Fig. 3, PyMAF consists of a feature pyramid for mesh recovery in a coarse-to-fine fashion. Coarse-aligned predictions will be improved by utilizing the mesh-aligned evidence extracted from spatial feature maps. In order to enhance the mesh-aligned evidence, an auxiliary dense prediction task is imposed on the image encoder while a spatial alignment attention is applied to fuse the grid and mesh-aligned features.
Our image encoder aims to generate a pyramid of spatial features from coarse to fine granularities, which provide descriptions of the posed person at different scale levels. The feature pyramid will be used in subsequent predictions of the SMPL model with the pose, shape, and camera parameters .
where denotes the feature sampling and processing operations, denotes the concatenation, and is the MLP. After that, a parameter regressor takes features and the current estimation of parameters as inputs and outputs the parameter residual. Parameters are then updated as by adding the residual to . For the level , adopts the mean parameters calculated from training data.
where is the squared L2 norm, , , and denote the ground truth 2D keypoints, 3D joints, and model parameters, respectively.
1.2 Mesh Alignment Feedback Loop
As mentioned in HMR , directly regressing mesh parameters in one go is challenging. To tackle this issue, HMR uses an Iterative Error Feedback (IEF) loop to iteratively update by taking the global features and the current estimation of as input. Though the IEF strategy reduces parameter errors progressively, it uses the same global features each time for parameter update, which lacks fine-grained information and is not adaptive to new predictions. By contrast, we propose a Mesh Alignment Feedback (MAF) loop so that mesh-aligned evidence can be leveraged in our regressor to rectify current parameters and improve the mesh-image alignment of the estimated model.
Though the mesh-aligned features are position-sensitive, these features are confined to the re-projection regions of the current mesh result. To enable the perception of the relative positions in the whole image context, we further design spatial alignment attention to fuse the information from both grid and mesh-aligned features. Considering that both these two features are extracted from the same spatial feature map, we adopt a self-attention module to process them. Specifically, the point-wise features extracted based on the grid-pattern points and the mesh-aligned points are first concatenated together as :
where is the total number of the grid-pattern and mesh-aligned points. Then, spatial alignment attention is applied to learn attentive relations among so that the mesh-aligned features can be more effectively enhanced with the spatial information in the grid features. In our solution, a self-attention module is employed to process the features :
where , , and are the learnable matrices used to generate different subspace representations of the query, key, and value features , respectively, denotes the scaled dot-product attention function with softmax. In this way, the messages of the grid and mesh-aligned features can be fully fused together since the self-attention mechanism captures the relationships between all elements of the features . After that, the enhanced mesh-aligned features are obtained by reducing the dimension of and concatenating them together. Finally, the enhanced mesh-aligned features are fed into the regressor for parameter update:
1.3 Auxiliary Dense Supervision
As depicted in the second row of Fig. 4, spatial features tend to be affected by noisy inputs, since raw images may contain a large amount of unrelated information such as occlusions, appearance, and illumination variations. To improve the reliability of the mesh-aligned cues extracted from spatial features, we impose an auxiliary pixel-wise prediction task on the spatial features at the last level. Specifically, during training, the spatial feature maps will go through a convolutional layer to generate dense correspondence maps with pixel-wise supervision. Dense correspondences encode the mapping relationship between foreground pixels on the 2D image plane and mesh vertices in 3D space. In this way, the auxiliary supervision provides mesh-image correspondence guidance for the image encoder to preserve the most related information in the spatial feature maps.
In our implementation, we adopt the IUV maps defined in DensePose as the dense correspondence representation, which consists of the part index and UV values of the mesh vertices. Note that we do not use DensePose annotations in the dataset but render IUV maps based on the ground-truth SMPL models . During training, classification and regression losses are applied on the part index and channels of dense correspondence maps, respectively. Specifically, for the part index channels, a cross-entropy loss is applied to classify a pixel belonging to either background or one among body parts. For the channels, a smooth L1 loss is applied to regress the corresponding values of the foreground pixels. Only the foreground regions are taken into account in the regression loss, i.e., the estimated channels are firstly masked by the ground-truth part index channels before applying the regression loss. Overall, the loss function for the auxiliary pixel-wise supervision is written as
where denotes the mask operation. Note that the auxiliary prediction is required in the training phase only.
Fig. 4 visualizes the spatial features of the encoder trained with and without auxiliary supervision, where the feature maps are simply added along the channel dimension as grayscale images and visualized with colormap. We can see that the spatial features are more neat and robust to input variations after applying auxiliary supervision. Note that the dense correspondence is not limited to the IUV representation, the Projected Normalized Coordinate Code (PNCC) can be also adopted as dense correspondences when IUV is not defined in the mesh model. More discussions about the choice of dense correspondences can be found in the Supplementary Material.
2 PyMAF-X for Full-body Mesh Recovery
The body-specific PyMAF can be easily modified to reconstruct hand and face meshes by simply changing the SMPL model in the above formulation to the MANO and FLAME models. Based on the regression power of PyMAF, we extend it to PyMAF-X for full-body mesh recovery.
Following previous works , PyMAF-X consists of three experts, i.e., three part-specific PyMAFs, to predict the parameters of body, hand, and face, as illustrated in Fig. 5. To ensure high-resolution observations of part regions, part experts perform individual predictions on the body, hand, and face images cropped from the original inputs. At each iteration of the mesh alignment feedback loop, the predictions of the body-, hand-, and face-specific PyMAF are collected and integrated as the parameters of the full-body model SMPL-X , where , , and denotes the pose, shape, and facial expression parameters, respectively. The pose parameters consist of the rotational poses of 55 joints in total, including 22 joints for the body, 30 finger joints for the hands, and 3 jaw joints for the face. The camera parameters are taken from the predictions of the body-specific PyMAF and used to project body, hand, and face vertices on the image plane. Moreover, considering that the positions of hand and face are susceptible to inaccurate body pose estimations, we align the center of their re-projected points to the image center of hand and face to ensure their mesh-aligned features are meaningful.
After individual regression of each part, we need to figure out the rotation of wrist joints to integrate the body and hand meshes. The most straightforward strategy would be the naive “Copy-Paste” integration . Specifically, the poses of the wrist joints are calculated based on the body poses predicted by the body expert and the global orientation of hands predicted by the hand expert. Let be the global orientation of the left or right hand, which is also the global rotation of the wrist joint. The wrist pose of the full-body model can be solved by first computing the global rotation of the elbow joint and then the relative rotation of the wrist joint, i.e.,
where denotes the relative rotation of the -th body joint, the ordered set of joint ancestors of the elbow joint and itself in the kinematic tree, and the inverse global rotation of the hand. Benefiting from the well-aligned results of each part, PyMAF-X can produce plausible results in common scenarios using such a simple integration strategy.
As pointed out in previous work , the body expert hardly perceives the hand poses due to the small proportion of hand region in the body images. It may lead to incompatible configurations of the arm and hand poses predicted individually by the body and hand experts, resulting in unnatural wrist poses of the full-body model, as illustrated in Fig. 6. Previous work alleviates this issue by learning wrist poses from the body and hand features but typically degrades the accuracy of the wrist poses and alignment. In our work, we propose an adaptive integration strategy to correct the elbow poses directly based on the solved wrist poses such that the elbow and wrist poses could be more compatible. To maintain the mesh-image alignment, we only correct the twist rotation of the elbow joints as it is the rotation along the elbow-wrist bone and barely affects the position of the body and hand joints. To this end, we first compute the twist angle of the wrist poses w.r.t. the elbow-to-wrist bone, then update the elbow and wrist poses by adding and subtracting the compensated twist rotation, respectively.
Step 1: Computing the original twist angle. The twist component around the elbow-to-wrist vector can be decomposed from the wrist poses. Without loss of generality, let the quaternion representation of the left or right wrist pose solved in Eq. (8) be . By using Huyghe’s method , the quaternion of the twist rotation around the normalized elbow-to-wrist vector can be calculated as:
where in is the projection vector of the normalized onto . Let be the first element of the twist quaternion , then the twist rotation angle can be computed as .
As shown in our experiments, with the twist compensation from the elbow joint, the wrist pose becomes more natural while maintaining the mesh-image alignment of the body and hands. In practice, the adaptive integration is not applied for those invisible hands since the global orientation predicted by the hand expert is not reliable when the hand is invisible. In our implementation, the hand expert of PyMAF-X also predicts the confidence of the visibility status of hands. When the hand is invisible, the full-body model simply adopts the wrist poses predicted by the body expert and the mean poses of hands.
Experiments
The part-specific PyMAF primarily adopts ResNet-50 as the backbone of the image encoder. We also follow ExPose and PIXIE to adopt HRNet-W48 as the encoder backbone for the body model regression. For each part-specific PyMAF, the image encoder takes a image as input and produces spatial feature maps with resolutions of . When generating mesh-aligned features, the vertex number of body, hand, and face meshes is down-sampled to , , and , respectively. The mesh-aligned features extracted from feature maps of each point will be processed by MLPs so that their channel dimensions will be reduced to . Hence, the mesh-aligned feature vector for the body model has a length of , which is similar to the length of the global features used in HMR . The maximum number is set to , which is equal to the iteration number used in HMR. For the grid features used at , they are uniformly sampled from with a grid pattern, i.e., the point number is which is approximate to the vertices number after mesh downsampling. The regressors have the same architecture as the one in HMR, except that they have slightly different input dimensions. The twist angle constraint is empirically set to in our implementation. Following previous work , we adopt the continuous 6D representation for pose parameters in the regressor. Following PARE , the body encoder is initialized with the model pretrained on 2D pose datasets . During training, we use the Adam optimizer and set the learning rate to without decay. The part-specific PyMAFs are first pre-trained individually and then assembled together for finetuning on full-body datasets. Similar to PARE , we also observed a slight performance gain when removing the auxiliary supervision at the final stage of training, but we do not apply such a strategy in our experiments for more consistent ablation studies of our newly introduced components. More details of the implementation can be found in our code and the Supplementary Material.
Camera Setting. We follow previous work to use a weak perspective camera with a pre-defined focal length of 5,000 by default for training and evaluation. When running experiments on AGORA , we use a perspective camera with the focal lengths estimated by SPEC as there are stronger perspective distortions in this dataset. Incorporating the camera setting of SPEC with our method for more accurate mesh recovery is left for future work.
Runtime. The PyTorch implementation of the body-only PyMAF takes about 22 ms to process one sample on the machine with an NVIDIA RTX 3090 GPU. For full-body mesh recovery, PyMAF-X takes about 80 ms to process one sample, which is on par with existing regression-based approaches . In our current implementation, the part-specific backbone networks run in sequence to process the body, hand, and face images. Running them in parallel would further reduce the runtime.
2 Datasets
Following the practices of previous work , the body expert is trained on a mixture of data from several datasets with 3D and 2D annotations, including Human3.6M , MPI-INF-3DHP , MPII , LSP , LSP-Extended , and COCO . For the hand expert, we use images from FreiHAND , InterHand2.6M and COCO-Wholebody for training. For the face expert, we use the images from VGGFace2 for training. Detailed descriptions of the datasets can be found in the Supplementary Material.
Following previous work , the SMPL/SMPL-X models fitted in EFT and ExPose are used as pseudo ground-truth annotations for the training of body and full-body model regressors. For the training of the face expert, we use DECA and a face alignment algorithm FAN to generate pseudo ground-truth FLAME models and facial landmarks on the training set of VGGFace2 .
Note that we do not use the DensePose annotations in COCO for auxiliary supervision but render dense correspondence maps based on the pseudo ground-truth meshes using the method described in .
3 Evaluation Metrics
We report the results of our approach in various evaluation metrics for quantitative comparisons with existing state-of-the-art methods, where all metrics are computed in the same way as previous work in literature.
To quantitatively evaluate the performance of the 3D pose estimation, PVE, MPJPE, PA-PVE, and PA-MPJPE are adopted as the primary evaluation metrics. They are all reported in millimeters (mm) by default. Among these metrics, PVE denotes the mean Per-vertex Error, defined as the average point-to-point Euclidean distance between the predicted and ground truth mesh vertices, while MPJPE denotes the Mean Per Joint Position Error. PA-PVE and PA-MPJPE denote the PVE and MPJPE after rigid alignment of the prediction with the ground truth using Procrustes Analysis. Note that the metrics PA-PVE and PA-MPJPE are not aware of the global rotation and scale errors since they are calculated after rigid alignment.
4 Comparison with the State of the Art
We first evaluate our approach on the 3D human pose and shape estimation task and make comparisons with previous state-of-the-art regression-based methods. We present evaluation results for quantitative comparison on 3DPW and Human3.6M datasets in Table I. Our PyMAF achieves competitive or superior results among previous approaches, including frame-based and temporal approaches. Note that the approaches reported in Table I are not strictly comparable since they may use different training data, pseudo ground-truths, learning rate schedules, training epochs, etc. For a fair comparison, we report our baseline results in Table I, which is trained under the same setting as PyMAF. The baseline approach has the same network architecture with HMR and also adopts the 6D rotation representation for pose parameters. Under the setting of using ResNet-50 backbone and without training on 3DPW, PyMAF reduces the MPJPE over the baseline by 4.7 mm and 5.5 mm on 3DPW and Human3.6M datasets, respectively.
From Table I, we can see that PyMAF has more notable improvements on the metrics MPJPE and PVE. We would argue that the metric PA-MPJPE can not reveal the mesh-image alignment performance since it is calculated as the MPJPE after rigid alignment. As depicted in the Supplementary Material, a reconstruction result with a smaller PA-MPJPE value can have a larger MPJPE and worse alignment between the reprojected mesh and the input image.
We evaluate 2D human pose estimation performance on the COCO validation set to measure the mesh-image alignment quantitatively in real-world scenarios. During the evaluation, we project keypoints from the estimated mesh on the image plane and compute the Average Precision (AP) based on the keypoint similarity with the ground truth 2D keypoints. The results of keypoint localization APs are reported in Table II. OpenPose , a widely-used 2D human pose estimation algorithm, is also included for reference. We can see that the COCO dataset is very challenging for approaches to human mesh recovery as they typically have much worse performances in terms of the 2D keypoint localization accuracy. In Table II, we also include the results of the optimization-based SMPLify by fitting the SMPL model to the ground-truth 2D keypoints with 1,500 optimization iterations. As pointed out in previous work , SMPLify may produce well-aligned but unnatural results. Moreover, SMPLify is much more time-consuming than regression-based solutions. Among approaches to recovering 3D human mesh, PyMAF outperforms previous regression-based methods by remarkable margins, making it the most competitive mesh recovery method in comparison with OpenPose . Under the backbone of ResNet-50, PyMAF brings significant improvements over our baseline by 8.5% and 6.2% on AP and , respectively. Qualitative comparisons can be found in the supplementary material.
4.2 Evaluation on Hand-only Reconstruction
We compare the hand-only PyMAF with state-of-the-art approaches on the FreiHAND dataset. As shown in Table III, PyMAF outperforms the baseline and previous full-body methods and is comparable with recent hand-only methods . It is also worth noting that full-body methods typically adopt the parametric representation of the hand mesh, which tends to be numerically inferior to the non-parametric representation used in recent hand-only methods , as pointed out in previous works .
4.3 Evaluation on Face-only Reconstruction
Following previous work , we compare the face-only PyMAF with state-of-the-art face reconstruction approaches on the test set of Stirling3D and NoW datasets. Table IV reports the performances of different methods in Point-to-Surface after Procrustes Alignment (PA-P2S). It shows that the PyMAF outperforms the face expert of previous full-body methods ExPose and PIXIE , while achieving similar results compared with the strong face-only method DECA . Qualitative comparisons of face reconstruction results are visualized in Fig. 7. For more consistent comparisons, we only show the intermediate parametric FLAME model predicted by DECA without using detail displacements. We can see that PyMAF is able to capture expressive face shapes and has competitive results against DECA .
4.4 Evaluation on Full-body Mesh Recovery
Following previous work on full-body mesh recovery, we evaluate the performance of PyMAF-X on two benchmark datasets, i.e., EHF and AGORA .
Table V reports the results of different methods for full-body mesh recovery, including the optimization-based MTC and SMPLify-X , and the regression-based ExPose , FrankMocap , Zhou et al. , PIXIE , and Hand4Whole . We can see that PyMAF-X achieves the best results among existing solutions on most metrics, especially on the evaluation of the body and full-body reconstruction.
Table VI compares the results of PyMAF-X and other full-body methods on the test set of AGORA , where all the evaluation results are taken from the official evaluation platform. Recent state-of-the-art approaches to body-only mesh recovery are also included in Table VI for comprehensive comparisons. Note that the evaluation on AGORA is also affected by the detection results as the predictions are first matched with the ground truth and then used to calculate the reconstruction error. We use an off-the-shelf tool OpenPifPaf to detect persons and the corresponding hands and face regions, of which the person detection result is slightly worse than the recent solutions Hand4Whole and BEV . For matched predictions, PyMAF-X outperforms other methods, especially in the metrics for hand and full-body reconstruction on this challenging dataset.
Qualitative comparisons of different full-body methods on real-world images are shown in Fig. 8, where we can see that PyMAF-X produces more accurate body, hand, and wrist poses than recent state-of-the-art approaches, including FrankMocap , PIXIE , and Hand4Whole . The video results of PyMAF-X and other full-body methods can be found on our project page and supplementary materials.
5 Ablation Study
In this part, we will perform ablation studies under various settings to validate the key components proposed in PyMAF and PyMAF-X. The efficacy of the mesh-aligned features, pyramidal design, auxiliary dense supervision, and spatial alignment attention proposed in PyMAF will be validated on Human3.6M . As the Human3.6M dataset includes large-scale amounts of images and the corresponding ground-truth 3D labels, ablation approaches of PyMAF are trained and evaluated on Human3.6M. As for the proposed adaptive integration in PyMAF-X, ablation approaches are evaluated on EHF , where different approaches are trained under the same setting.
In PyMAF, mesh-aligned features provide the current mesh-image alignment information in the feedback loop, which is essential for better mesh recovery. To verify this, we alternatively replace the mesh-aligned features with the global features or the grid features uniformly sampled from spatial features as the input for parameter regressors. Table VII reports the performance of approaches using different types of features in the feedback loop. The results under the non-pyramidal setting are also included in Table VII, where the grid and mesh-aligned features are extracted from the feature maps with the highest resolution (i.e., ), and the mesh-aligned features are extracted on the reprojected points of the mesh under the mean pose at . Note that all approaches in Table VII do not use auxiliary supervision.
Unsurprisingly, using mesh-aligned features yields the best performance under both non-pyramidal and pyramidal designs. The approach using the grid features sampled from spatial feature maps has better results than global features but is worse than the mesh-aligned counterpart. The mesh-aligned solution achieves even more performance gain when using pyramidal feature maps since multi-scale mesh-alignment evidence is leveraged in the feedback loop. Though the grid features contain primary spatial cues on uniformly distributed pixel positions, they can not reflect the alignment status of the current estimation. This implies that mesh-aligned features are the most informative ones for the regressor to rectify the current mesh parameters.
The auxiliary pixel-wise supervision helps to enhance the reliability of the mesh-aligned evidence extracted from spatial features. Using alternative pixel-wise supervision such as part segmentation rather than dense correspondences is also possible in our framework. In our approach, these auxiliary predictions are solely needed for supervision during training since the point-wise features are extracted from feature maps. For more in-depth analyses, we have also tried extracting point-wise features from the auxiliary predictions, i.e., the input type of regressors are intermediate representations such as part segmentation or dense correspondences. Table VIII compares different auxiliary supervision settings and input types for regressors during training. Using part segmentation is slightly worse than our dense correspondence solution. Compared with the part segmentation, the dense correspondences preserve clean and rich information in foreground regions. Moreover, using feature maps for point-wise feature extraction consistently performs better than auxiliary predictions. This can be explained by the fact that using intermediate representations as input for regressors hampers the end-to-end learning of the whole network. Under the auxiliary supervision strategy, the spatial feature maps are learned with the signal backpropagated from both auxiliary prediction and parameter correction tasks. In this way, the background features can also contain information for mesh parameter correction since the deep features have larger receptive fields and are trained in an end-to-end manner. As shown in Table VIII, when the mesh-aligned features are masked with the foreground region of part segmentation predictions, the performance degrades from 75.5 mm to 77.6 mm on MPJPE.
In our approach, Spatial Alignment Attention (SAA) is designed to enable the awareness of the whole image context in the regressor. To validate its efficacy, we replace the spatial alignment attention with fully-connected layers to fuse the grid and mesh-aligned features. As reported in Table IX, simply fusing the grid features (the second row) only brings marginal improvements in comparison with the approach using the spatial alignment attention (the third row). The performances of PyMAF with or without spatial alignment attention across each refinement iteration are reported in Table X, where the PyMAF with spatial alignment attention improves the reconstruction results more quickly.
In PyMAF-X, an elbow-twist compensation is used to adaptively correct the elbow poses in the integration of body and hand estimations. Such an adaptive integration strategy could produce physically-plausible wrist poses while preserving the mesh-image alignment. We investigate different integration strategies and compare our solution with two alternatives: i) a learned integration strategy similar to PIXIE , which predicts the wrist poses based on the fused features of body and hand features; ii) the naive copy-paste integration strategy , which calculates the wrist poses based on the estimated body and hand poses. Table XI reports the performances of the three different integration strategies on the EHF dataset. Here, we use the MPJPE of body and hand joints to measure the mesh-image alignment and the PA-PVE of wrist vertices to measure the physical plausibility of the wrist joint. As shown in the first row, the learned integration strategy can also produce natural wrist poses but degrade the alignment of hand parts in the full-body model. Compared with the learned and copy-paste strategies, the proposed adaptive integration produces both well-aligned and natural poses of the body, hand, and wrist parts. Fig. 9 provides a visual comparison of different integration strategies under challenging cases in real-world scenarios. We can see that our adaptive integration maintains the alignment and effectively improves the plausibility of the wrist poses by leveraging the twist compensation from elbow joints.
Conclusion
In this paper, we first present Pyramidal Mesh Alignment Feedback (PyMAF) for regression-based human mesh recovery and further extend it as PyMAF-X for full-body mesh recovery. PyMAF is primarily motivated by the observation of the reprojection misalignment between the parametric mesh results and the input images. At the core of PyMAF, the parameter regressor leverages spatial information from a feature pyramid to correct the parameter deviation explicitly in a feedback loop based on the alignment status of the currently estimated meshes. To achieve this, given a coarse-aligned mesh estimation, the mesh-aligned features are first extracted from the spatial feature maps and then fed back into the regressor for parameter rectification. Moreover, an auxiliary dense supervision is employed to enhance the learning of mesh-aligned features while spatial alignment attention is introduced to enable the awareness of the global contexts in our deep regressor. When extending PyMAF for full-body model recovery, an adaptive integration with the elbow-twist compensation strategy is proposed in PyMAF-X to produce natural wrist poses while maintaining the alignment performances of part-specific PyMAF. The efficacy of PyMAF and PyMAF-X is validated on indoor and in-the-wild datasets, where our approaches effectively improve the mesh-image alignment over the baseline and previous regression-based solutions.
In our experiments, we found that PyMAF-X fails to reconstruct interacting hands due to the separated regression of two hands. Meanwhile, when handling images with strong perspective distortions or with only upper-torso observations, common issues such as bent legs and erroneous limb poses remain unsolved in this work. We will leave these issues for future work to incorporate the merits of recent datasets and solutions such as SPEC , PIXIE , Hand4Whole , and IntagHand into our framework.
Moreover, similar to existing methods , the full-body alignment performance of PyMAF-X heavily relies on the pose and shape estimation of the body expert. Due to the lack of full-body mesh annotations, the estimated body shapes are typically inaccurate in challenging cases, resulting in erroneous bone lengths of arms and coarse alignment of hands. Combining PyMAF-X with SPIN , EFT , or NeuralAnnot for the generation of more precise pseudo 3D ground-truth full-body mesh annotations on in-the-wild data would be interesting future work. Besides, the elbow-twist rotations are adjusted empirically in PyMAF-X based on the twist components of wrist poses. Learning the compensation angle via networks is also possible when large-scale full-body mesh annotations are available.
Acknowledgments
This work was supported by the National Key R&D Program of China (2021ZD0113501), the National Natural Science Foundation of China (No.62125107, U1836217, and 62276263), and the China Postdoctoral Science Foundation (No.2022M721844). We would like to thank Xinchi Zhou, Wanli Ouyang, and Limin Wang for their help, feedback, and discussions in the early work of this paper.
References
Appendix A More Implementation Details
We follow Mesh Graphormer to implement the attention module for the grid and mesh-aligned feature fusion. We simply use a single attention block for each iteration because we found that a single block brought enough performance grains while using more attention blocks merely increased memory consumption.
For body and hand meshes, we simply use the down-sampling matrix provided in GraphCMR and Mesh Graphormer to reduce the vertex number for mesh-aligned feature extraction. For face meshes, we manually select the vertices in the front face region considering that the expression information concentrates on the front face. For body-only PyMAF, the dimension of the mesh-aligned features is reduced to match the feature dimension in HMR for more fair comparisons. For hand- and face-only PyMAF, we do not strictly abide by this rule but simply replace SMPL with MANO/FLAME during the implementation. Fig. 10 visualizes the selected vertices on the body, hand, and face meshes.
In the auxiliary dense prediction task, the dense correspondence and part segmentation (for ablation experiments) used for supervision are rendered based on the pseudo ground truth meshes. Examples of the rendered dense correspondence are visualized in Fig. 11. We choose to use the rendered dense correspondence mainly based on the following reasons:
i) The dense correspondence can be regarded as a more fine-grained part segmentation. In our method, we use the IUV representation as dense correspondence for the body expert, while using the Projected Normalized Coordinate Code (PNCC) for hand and face experts. Dense correspondences are more general since PNCC does not need to split the mesh into manually defined parts.
ii) The body part definition may vary from different datasets while the rendered one is consistent. Moreover, the rendering process can be done efficiently in batch using the tools like PyTorch3D . The costs of rendering part segmentation and dense correspondences are almost the same for the PyTorch3D renderer.
iii) There is only a limited number of datasets providing the annotated part segmentation and dense correspondence, while the pseudo ground truth meshes are typically available for common datasets thanks to previous work such as SPIN and EFT .
We simply use three FC layers to predict hand visibility based on the hand-only mesh-aligned features. The pseudo visibility confidence used for supervision is calculated based on the proportion of the visible hand keypoints annotated in the whole-body COCO dataset. For example, assume that there are 15 visible keypoints and the total number of hand keypoints is 21, then the pseudo visibility confidence of this hand is about 0.714 (15/21).
Appendix B About Metrics
Though the PA-PVE and PA-MPJPE are widely adopted in the 3D pose estimation task, these two metrics can not fully reveal the mesh-image alignment performance since they are calculated after rigid alignment. As depicted in Fig. 12, a reconstruction result with a lower PA-MPJPE value can have a higher MPJPE value and worse alignment between the reprojected mesh and image.
Appendix C About Datasets
Following the practices of previous work , we train our body model regression network on several datasets with 3D or 2D annotations, including Human3.6M , MPI-INF-3DHP , LSP , MPII , COCO . For hand-only and full-body model regression, FreiHAND , InterHand2.6M , FFHQ , and COCO-WholeBody are also used for training. Here, we provide more descriptions of the datasets to supplement the main manuscript.
Human3.6M is commonly used as the benchmark dataset for 3D human pose estimation, consisting of 3.6 million video frames captured in the controlled environment. The ground truth SMPL parameters in Human3.6M are generated by applying MoSH to the sparse 3D MoCap marker data, as done in Kanazawa et al. . The original videos are down-sampled from 50fps to 10fps, resulting in 312,188 frames for training. Following the common protocols , our experiments use five subjects (S1, S5, S6, S7, S8) for training and two subjects (S9, S11) for evaluation. The original videos are also down-sampled from 50 fps to 10 fps to remove redundant frames, resulting in 312,188 frames for training and 26,859 frames for evaluation.
3DPW is captured in challenging outdoor scenes with IMU-equipped actors under various activities. This dataset provides accurate shape and pose ground truth annotations. Following the protocol of previous work , we do not use its data for training by default unless specified in the table.
MPI-INF-3DHP is a 3D human pose dataset covering more actor subjects and poses than Human3.6M. The images of this dataset were collected under both indoor and outdoor scenes, and the 3D annotations were captured by a multi-camera marker-less MoCap system. Hence, there is some noise in the 3D ground truth annotations. The training set includes 8 subjects and there are 96,507 frames down-sampled from videos used for training.
LSP and LSP-Extended are 2D human pose benchmark datasets, containing person images with challenging poses. There are 14 visible 2D keypoint locations annotated for each image and 10,428 samples used for training.
MPII is a standard benchmark for 2D human pose estimation. There are 25,000 images collected from YouTube videos covering a wide range of activities. We discard those images without complete keypoint annotations, producing 14,667 samples for training.
COCO and COCO-WholeBody contain a large scale of person images labeled with 17 body keypoints, 42 hand keypoints, and 68 face keypoints. We use COCO to train body-only PyMAF and leverage the hand keypoints in COCO-WholeBody during the training of hand- and face-only PyMAF and PyMAF-X. Since this dataset does not contain ground-truth meshes, we conduct a quantitative evaluation on the 2D keypoint localization task using its validation set, which consists of 50,197 samples.
EHF contains 100 testing images of one subject captured in lab environments. For each image, the corresponding 3D scans and ground-truth SMPL-X meshes are provided. EHF is used for testing only and is commonly adopted as a full-body evaluation benchmark dataset in literature .
AGORA is a synthetic dataset with accurate SMPL-X models fitted to 3D scans. Since the ground-truth labels of its test set are not publicly available, the evaluation is performed on the official platformhttps://agora-evaluation.is.tuebingen.mpg.de. For evaluation on AGORA, we use the training set of AGORA to finetune our model.
FreiHAND contains 130,240 samples for training and 3960 samples images for evaluation. For each sample in the training set, the MANO parameters recovered from multi-view images are provided. We use this dataset for the training and evaluation of the hand expert.
InterHand2.6M is a large-scale real-captured hand dataset, providing accurate MANO parameters of interacting hands. We crop single-hand images from this dataset for the training of the hand expert.
VGGFace2 is a large-scale face dataset. The images of this dataset are downloaded from Google and have large variations in pose, age, and ethnicity. It contains about 3 million images from training. We run the method of FAN and DECA on its training set to generate the pseudo ground truth facial landmarks and FLAME models for the training of the face expert.
Stirling3D provides facial images with the ground-truth 3D scans. The test set contains 2,000 facial images in neutral expressions, including 1,344 low-quality (LQ) images and 656 high-quality (HQ) images. We follow previous work to use it for evaluation only.
NoW contains the facial images captured with an iPhone X, and a separate 3D scan for each subject. Its test set contains 1,702 images for evaluation. Since the ground-truth scans of the test set are not publicly available, the evaluation is performed by following the instructions on the official websitehttps://now.is.tue.mpg.de/index.html. We follow previous work to use this dataset for evaluation only.
Appendix D More Qualitative Results
We provide more qualitative results of our method in this section. In Fig 13, we visualize the estimated meshes after each iteration, where it can be seen that PyMAF can correct the drift of body parts progressively and result in better-aligned human models. In Fig. 14, the body mesh recovery results of different methods on COCO are depicted for qualitative comparisons, where PyMAF convincingly performs better than competitors and our baseline by producing better-aligned and natural results. In Fig. 15, we provide more full-body model reconstruction results on the COCO validation set, where PyMAF-X can produce well-aligned full-body model under challenging cases. In Fig.16, we further visualize the reconstructed full-body models from different viewpoints.
Cases under Occlusions. As pointed out in the main paper, the adaptive integration is not applicable when the hand part is invisible. To handle this, the visibility status of hands is also predicted by the hand expert in PyMAF-X. In cases of invisible hands, the full-body model adopts the default hand poses and the wrist poses estimated by the body expert. Fig. 17 shows the example results of PyMAF-X when the body or hands are occluded. We can see that PyMAF-X produces reasonable full-body meshes under these cases.
Failure Cases. Due to the rotational pose representation of the kinematic model, the full-body alignment of PyMAF-X heavily relies on the accuracy of body pose estimation. Moreover, the misalignment may also occur when the body shape is inaccurate since it affects the body bone length. Besides, it is still challenging for PyMAF-X to handle challenging hand poses or interacting hands. Fig. 18 visualizes some erroneous results of our approach, where PyMAF-X produces misaligned results due to the issues mentioned above.