Collaborative Regression of Expressive Bodies using Moderation

Yao Feng, Vasileios Choutas, Timo Bolkart, Dimitrios Tzionas, Michael J. Black

Introduction

To model human behavior, we need to capture how people look, how they feel, and how they interact with each other. To facilitate this, our goal is to reconstruct whole-body 3D shape and pose, facial expressions, and hand gestures from an RGB image. This is challenging, as humans vary in shape and appearance, they are highly articulated, they wear complex clothing, they are often occluded, and their face and hands are small, yet highly deformable. For these reasons, the community studies the body , hands and face mostly separately.

Recent whole-body statistical models enable approaches to address the problem holistically, by jointly capturing the body, face and hands. ExPose reconstructs SMPL-X meshes from an RGB image, using “expert” sub-networks for the body, face and hands. However, ExPose’s part experts operate completely independently, as they only “see” their respective part image. Thus, they do not exploit the correlations between parts to overcome challenges like occlusion or motion blur.

Face-only methods are well studied and recover accurate facial shape, albedo and geometric details, which are important to capture emotions. However, they need a tight crop around the face and struggle with extreme viewing angles and faces that are small, low-resolution or occluded. While whole-body methods handle these challenges well, they estimate average-looking face shapes, without face albedo and fine geometric details.

To get the best of all worlds, we introduce PIXIE (“Pixels to Individuals: eXpressive Image-based Estimation”). PIXIE estimates expressive whole-bodied 3D humans from an RGB image more realistically than existing work. To do so, it pushes the state of the art in three ways.

First, PIXIE learns not only experts for the body, face and hands, but also a novel moderator that estimates their confidence in each sub-image, and fuses their features weighted by this. The learned fusion helps improve whole-body shape, using SMPL-X’s shared shape space across all body parts. Moreover, it helps to robustly estimate head and hand pose when these are ambiguous (e.g. occlusions or blur) by using full-body context; see Fig. 2 for examples.

Second, PIXIE significantly improves “gendered” body shape realism. While human shape is highly correlated with gender, existing work ignores this and estimates inaccurate body shapes – often with the wrong gender or with a gender-neutral shape. An exception is SMPLify-X, but it uses an offline gender classifier and fits a gender-specific SMPL-X model. Instead, using a single unisex SMPL-X model enables end-to-end training of neural nets. PIXIE adopts this approach, and learns to implicitly reason about shape. For this, we define male, female, and non-binary body-shape priors within the SMPL-X shape space. At training time, given automatically created gender labels for input images, we train PIXIE to output plausible shape parameters for the specified gender. At inference time, PIXIE needs no gender labels, is applicable to any in-the-wild image, and supports non-binary genders. Note that this approach is general and is relevant for the broader community (face, body, whole-body). Body shape is also correlated with face shape . Thus, we do the same “gendered” training for our face expert; this allows PIXIE to use face information to inform body shape. This training and network architecture significantly improves body shape both qualitatively and quantitatively.

Third, PIXIE’s face expert additionally infers facial albedo and dense 3D facial-surface displacements. For this, we draw inspiration from Feng et al. , and go beyond them in three ways: (1) We use a whole-body shape space, rather than a face-only space, to capture correlations between the body and face shape. (2) We use photometric and identity losses on faces to inform whole-body shape. (3) We use the inferred geometric details only when the face expert is confident, as judged by the moderator. As shown in Fig. 1, this results in whole-body 3D humans with detailed faces that can be fully animated.

To summarize, here we make three key contributions: (1) We train a novel moderator, that infers the confidence of body-part experts and fuses their features weighted by this. This improves shape and pose inference under ambiguities. (2) We train the network to implicitly reason about gender, i.e. without gender labels at test time, with a novel “gendered” 3D shape loss that encourages likely body shapes. (3) We extend our face expert with branches that estimate facial albedo and 3D facial-surface displacements, enabling whole-body animation with a realistic face. PIXIE is a step towards automatic, accurate and realistic 3D avatar creation from a single RGB image. Models and code are available for research purposes mat pixie.is.tue.mpg.de.

Related work

Body reconstruction: For years, the community focused on the prediction of 2D or 3D landmarks for the body , face and hands , with a recent shift towards estimating 3D model parameters or 3D surfaces . One line of work simplifies the problem by using proxy representations like 2D joints , silhouettes , part labels or dense correspondences . These are then “lifted” to 3D, either as part of an energy term or using a regressor . To overcome ambiguities, they use priors such as known limb lengths , joint angle limits , or a statistical body model like SMPL or SMPL-X . While these approaches benefit from 2D annotations, they cannot overcome errors in the proxy features and do not fully exploit image context. The alternative is to directly regress 3D skeletons , statistical model parameters , 3D meshes , depth maps , 3D voxels or distance fields from the image pixels.

Face reconstruction: Most modern monocular 3D face reconstruction methods estimate the parameters of a pre-computed statistical face model . Similar to the body literature, this problem is tackled with both optimization and regression methods . Many learning-based approaches follow an analysis-by-synthesis strategy , which jointly estimates geometry, albedo, and lighting, to render a synthetic image that is compared with the input. Recent work further employs face-recognition terms during training to reconstruct more accurate facial geometry. Even geometric details, such as wrinkles, can be learned from large collections of in-the-wild images . We refer to Egger et al. for a comprehensive overview. The major downsides of face-specific approaches are their need for tightly cropped face images and their inability to handle non-frontal images. The latter is mainly due to the lack of supervision; 2D landmarks may be missing or the face might not even be detected at all, in which case the photometric term is not applicable. By integrating face and body regression, PIXIE regresses head pose and shape robustly in situations where face-only methods fail and lets the face contribute to whole-body shape estimation.

Hand reconstruction: While hand pose estimation is most often performed from RGB-D data, there has been a recent shift towards the use of monocular RGB images . Similar to the body, we split these into methods that predict 3D joints , parameters of a statistical hand model , such as MANO , or a 3D surface .

Whole-body reconstruction: Recent methods approach the problem of human reconstruction holistically. Some of these estimate 3D landmarks for the body, face and hands , but not their 3D surface. This is addressed by whole-body statistical models , that jointly capture the 3D surface for the body, face and hands.

SMPLify-X fits SMPL-X to 2D body, hand and face keypoints estimated in an image. Xiang et al. estimate both 2D keypoints and a part orientation field and fit Adam to these. Xu et al. fit GHUM to detected body-part image regions. While these methods work, they are based on optimization, consequently they are slow and do not scale up to large datasets.

Deep-learning methods tackle these limitations, and quickly regress SMPL-X parameters from an image. ExPose uses “expert” sub-networks for the body, face and hands; the body expert estimates the body and rough part (hand/face) pose from the full-body image, while part experts refine the rough part poses using only local image information (hand/face crop). ExPose merges the output of its experts by always trusting them. Instead, we evaluate the confidence of each expert for each sub-image and fuse body/face and body/hand features weighted by this. To account for different body-part sizes, we use ExPose’s body-driven attention, and multiple data sources for both part-only and whole-body supervision.

FrankMocap is similar to ExPose and adds an (optional) optimization step to better align the estimated SMPL-X mesh with the image. Zhou et al. train a network to regress a body-and-hands (SMPL+H) model and the detailed MoFA face model from an RGB image, following a body-part attention mechanism and multi-source training like ExPose. Note that SMPL+H and MoFA are disparate models, which are (offline) manually cut-and-stitched together. Instead, we use the whole-body SMPL-X model that captures the shape of all body parts together, thus no stitching is required. Zhou et al. fuse only hand-body features in a “binary” fashion, while their face model is “disconnected” from the body. Instead, we fuse both face-body and hand-body features in a “fully analog” fusion, and thus our face expert can inform the whole-body shape. Zhou et al. have no face camera, and need PnP-RANSAC and Procrustes to align their face to the image. Instead, we infer a face-specific camera and need no extra steps. Zhou et al. use a complicated architecture, with several modules that are trained separately, and is applicable only to whole bodies. Instead, we use no intermediate tasks to avoid possible sources of error and train our model end to end. Our full model is applicable to whole bodies, but the part experts are also (separately) applicable to part-only data.

Method

Here we introduce PIXIE, a novel model for reconstructing SMPL-X humans with a realistic face from a single RGB image. It uses a set of expert sub-networks for body, face/head, and hand regression, and combines them in a bigger network architecture with three main novelties: (1) We use a novel moderator that assesses the confidence of part experts and fuses their features weighted by this, for robust inference under ambiguities, like strong occlusions. (2) We use a novel “gendered” shape loss, to improve body shape realism by learning to implicitly reason about gender. (3) In addition to the albedo predicted by our face expert, we employ the surface details branch of Feng et al. .

2 PIXIE Architecture

PIXIE uses the architecture of Fig. 3, and is trained end to end. All model components are described below.

Input images: Given an image II with full resolution, we assume a bounding box around the body. We use this to crop and downsample the body to I_{{\color[rgb]{0,0,0}b}} to feed our network. However, this makes hands and faces too low resolution. We thus use an attention mechanism to extract from II high-resolution crops for the face/head, I_{{\color[rgb]{0,0,0}f}}, and hand, I_{{\color[rgb]{0,0,0}h}}.

Feature encoding: We feed \{I_{{\color[rgb]{0,0,0}b}},I_{{\color[rgb]{0,0,0}f}},I_{{\color[rgb]{0,0,0}h}}\} to separate expert encoders \{E_{{\color[rgb]{0,0,0}b}},E_{{\color[rgb]{0,0,0}f}},E_{{\color[rgb]{0,0,0}h}}\} to extract features \{F_{{\color[rgb]{0,0,0}b}},F_{{\color[rgb]{0,0,0}f}},F_{{\color[rgb]{0,0,0}h}}\}. We use ResNet-50 for the face/head and hand experts to generate F_{{\color[rgb]{0,0,0}f}},F_{{\color[rgb]{0,0,0}h}}\in\mathcal{R}^{2048}. The body expert E_{{\color[rgb]{0,0,0}b}} uses HRNet , followed by convolutional layers that aggregate the multi-scale feature maps, to generate F_{{\color[rgb]{0,0,0}b}}\in\mathcal{R}^{2048}.

Feature fusion (moderator): We identify the expert pairs of {body, head} and {body, hand} as complementary, and learn the novel moderators \{\mathcal{M}_{{\color[rgb]{0,0,0}f}},\mathcal{M}_{{\color[rgb]{0,0,0}h}}\} that build “fused” features \{F_{{\color[rgb]{0,0,0}f}}^{{\color[rgb]{0,0,0}\text{fused}}},F_{{\color[rgb]{0,0,0}h}}^{{\color[rgb]{0,0,0}\text{fused}}}\} and feed them to face/head and hand regressors \{\mathcal{R}_{{\color[rgb]{0,0,0}f}}^{{\color[rgb]{0,0,0}\text{fused}}},\mathcal{R}_{{\color[rgb]{0,0,0}h}}^{{\color[rgb]{0,0,0}\text{fused}}}\} (described below) for more informed inference. A moderator is implemented as a multi-layer perceptron (MLP) and gets the body, F_{{\color[rgb]{0,0,0}b}}, and part, F_{{\color[rgb]{0,0,0}p}} (F_{{\color[rgb]{0,0,0}f}} or F_{{\color[rgb]{0,0,0}h}}), features and fuses them with a weighted sum:

where \mathcal{M}_{{\color[rgb]{0,0,0}p}} (\mathcal{M}_{{\color[rgb]{0,0,0}f}} or \mathcal{M}_{{\color[rgb]{0,0,0}h}}) is the part moderator, w_{{\color[rgb]{0,0,0}p}} (w_{{\color[rgb]{0,0,0}f}} or w_{{\color[rgb]{0,0,0}h}}) is the expert’s confidence, and F_{{\color[rgb]{0,0,0}b}}^{{\color[rgb]{0,0,0}p}} (F_{{\color[rgb]{0,0,0}b}}^{{\color[rgb]{0,0,0}f}} or F_{{\color[rgb]{0,0,0}b}}^{{\color[rgb]{0,0,0}h}}) is the body feature F_{{\color[rgb]{0,0,0}b}} transformed by the respective “extractor”, i.e. the linear layer \mathcal{L}^{{\color[rgb]{0,0,0}p}} (\mathcal{L}^{{\color[rgb]{0,0,0}h}} or \mathcal{L}^{{\color[rgb]{0,0,0}f}}) between the body encoder E_{{\color[rgb]{0,0,0}b}} and part moderator \mathcal{M}_{{\color[rgb]{0,0,0}p}}. Finally, tt is a learned temperature weight, jointly trained with all network weights with the losses of Sec. 3.3, with no tt-specific supervision.

Parameter regression: We use two main regressor types: (1) We use the body, face/head, and hand \{\mathcal{R}_{{\color[rgb]{0,0,0}b}},\mathcal{R}_{{\color[rgb]{0,0,0}f}},\mathcal{R}_{{\color[rgb]{0,0,0}h}}\} regressors, that get features only from the respective expert encoder {F_{{\color[rgb]{0,0,0}b}}, F_{{\color[rgb]{0,0,0}f}}, F_{{\color[rgb]{0,0,0}h}}}. \mathcal{R}_{{\color[rgb]{0,0,0}b}} infers the camera C_{{\color[rgb]{0,0,0}b}}=(s_{{\color[rgb]{0,0,0}b}},\bm{t}_{{\color[rgb]{0,0,0}b}}), and body rotation and pose \bm{\theta}_{{\color[rgb]{0,0,0}b}} up to (excluding) the head and wrist. \mathcal{R}_{{\color[rgb]{0,0,0}f}} infers the camera C_{{\color[rgb]{0,0,0}f}}=(s_{{\color[rgb]{0,0,0}f}},\bm{t}_{{\color[rgb]{0,0,0}f}}), face albedo \bm{\alpha}_{{\color[rgb]{0,0,0}f}}, and lighting \textbf{l}_{{\color[rgb]{0,0,0}f}}. \mathcal{R}_{{\color[rgb]{0,0,0}h}} infers the camera C_{{\color[rgb]{0,0,0}h}}=(s_{{\color[rgb]{0,0,0}h}},\bm{t}_{{\color[rgb]{0,0,0}h}}). (2) We use the face/head, \mathcal{R}_{{\color[rgb]{0,0,0}f}}^{{\color[rgb]{0,0,0}\text{fused}}}, and hand, \mathcal{R}_{{\color[rgb]{0,0,0}h}}^{{\color[rgb]{0,0,0}\text{fused}}}, regressors that get from moderators the “fused” features, F_{{\color[rgb]{0,0,0}f}}^{{\color[rgb]{0,0,0}\text{fused}}} and F_{{\color[rgb]{0,0,0}h}}^{{\color[rgb]{0,0,0}\text{fused}}}. \mathcal{R}_{{\color[rgb]{0,0,0}h}}^{{\color[rgb]{0,0,0}\text{fused}}} infers the wrist θwrist\bm{\theta}_{\text{wrist}} and finger pose θfingers\bm{\theta}_{\text{fingers}}. \mathcal{R}_{{\color[rgb]{0,0,0}f}}^{{\color[rgb]{0,0,0}\text{fused}}} infers expressions ψ\bm{\psi}, head rotation θhead\bm{\theta}_{\text{head}}, and jaw pose θjaw\bm{\theta}_{\text{jaw}}. Importantly, \mathcal{R}_{{\color[rgb]{0,0,0}f}}^{{\color[rgb]{0,0,0}\text{fused}}} also infers body shape β\bm{\beta}, letting our face expert contribute to whole-body shape.

Detail capture: We use the fine geometric details branch R_{{\color[rgb]{0,0,0}d}} of Feng et al. that, given a face image I_{{\color[rgb]{0,0,0}f}}, estimates dense 3D displacements on top of FLAME’s surface. We convert the displacements from FLAME’s to SMPL-X’s UV map, and apply them on PIXIE’s inferred head shape. However, inferring geometric details from full-body images is not trivial; faces tend to be much noisier in these compared to face-only images. We account for this with our moderator, and use the inferred displacements only when the face/head expert is confident.

3 Training Losses

To train PIXIE we use body, hand and face losses:

defined as follows; the hat (e.g. x^\hat{\bm{x}}) denotes ground truth.

Body losses: Following , we use a combination of a 2D re-projection, a 3D joint, and a SMPL-X parameter loss:

Hand losses: We employ a similar set of losses to train the 3D hand pose and shape estimation network:

defined similarly to L2D/3D-JointsbodyL_{\text{2D/3D-Joints}}^{\text{body}} and LparamsbodyL_{\text{params}}^{\text{body}} of the body, but using the hand joints and pose parameters θwrist\bm{\theta}_{\text{wrist}} and θfingers\bm{\theta}_{\text{fingers}}.

Face losses: We adopt standard losses used by the 3D face estimation community :

The landmark loss penalizes the difference between detected target 2D landmarks mj^\hat{\bm{m}_{j}} and respective model landmarks (lying on M_{{\color[rgb]{0,0,0}f}}) projected on the image plane, mj\bm{m}_{j}:

Following , we also compute a loss for the set EE of landmarks on the upper, lower eyelid and upper, lower lip:

The face parameter loss LparamsfaceL_{\text{params}}^{\text{face}} follows LparamsbodyL_{\text{params}}^{\text{body}}, but for face pose θface\bm{\theta}_{\text{\text{face}}} only. This loss is only used for face crops from body data, when the target face pose is available.

Given the predicted 3D face mesh M_{{\color[rgb]{0,0,0}f}} as a subset of MM, face albedo \bm{\alpha}_{{\color[rgb]{0,0,0}f}} and lighting \textbf{l}_{{\color[rgb]{0,0,0}f}}, we render a synthetic image I_{{\color[rgb]{0,0,0}r}} for the input subject using the differentiable renderer from Pytorch3D . We then minimize the difference between the input face image I_{{\color[rgb]{0,0,0}f}} and the rendered image I_{{\color[rgb]{0,0,0}r}}:

where S is a binary face mask with value 11 in the face skin region, and elsewhere, and ⊙\odot denotes the Hadamard product. The segmentation mask prevents errors from non-face regions influencing the optimization, and we use the segmentation network of Nirkin et al. to extract S. The image formation process is the same as in Feng et al. .

Following , we use a pre-trained face recognition network , fidf_{\text{id}}, to compute embeddings for the rendered image I_{{\color[rgb]{0,0,0}r}} and the input I_{{\color[rgb]{0,0,0}f}}. We then maximize the cosine similarity between the two identity embeddings

Priors: Due to the difficulty of the problem, we use additional priors to constrain PIXIE to generate plausible solutions. For expression parameters, we use a Gaussian prior:

We also add soft regularization on jaw and face pose:

All these priors are “standard” regularizers, empirically found to discourage implausible configurations (extreme values, unrealistic shape/pose, inter-penetrations, etc).

Gender: As gender strongly affects body shape, we use a gender-specific shape prior during training, when gender labels are available. For this, we register SMPL-X to CAESAR scans, and compute the mean μ\bm{\mu} and covariance Σ\Sigma of shape parameters for each gender. We then use:

When gender is unknown, we use a Gaussian prior computed over all scans/registrations, irrespective of gender. Please note that we do not need gender labels for inference.

Feature update loss: We encourage the transformed body features F_{{\color[rgb]{0,0,0}b}}^{{\color[rgb]{0,0,0}p}} (F_{{\color[rgb]{0,0,0}b}}^{{\color[rgb]{0,0,0}f}} or F_{{\color[rgb]{0,0,0}b}}^{{\color[rgb]{0,0,0}h}}) to match F_{{\color[rgb]{0,0,0}p}}^{\text{fused}} with a loss that was empirically found to stabilize network training:

4 Implementation Details

Training data: For whole-body data we use the curated SMPL-X fits of , and SMPL-X fits to whole-body COCO data . For hand-only data we use FreiHAND and Total Motion . For face/head data we use VGGFace2 and detect Nlmk=68N_{\text{lmk}}=68 2D landmarks with the method of Bulat et al. . We get gender annotations by running the method of Rothe et al. on many photos per identity and using majority voting to improve robustness. For data augmentation, see Sup. Mat.

Network training: We do multi-step training that empirically aids stability. We pre-train on part-only data, and train on whole-body data end to end; for details see Sup. Mat.

Experiments

EHF : We evaluate whole-body accuracy on this. It has 100100 RGB images of 11 minimally-clothed subject in a lab setting with ground-truth SMPL-X meshes and 3D scans.

AGORA : We evaluate whole-body and body-only accuracy on this, using its body-face-hands (BFH) subset. It has rendered photo-realistic images of 3D human scans in scenes . It has SMPL-X ground truth recovered from scans, images and semantic labels .

3DPW : We evaluate main-body accuracy on this. It captures 55 subjects in indoor/outdoor videos with SMPL pseudo ground truth, recovered from images and IMUs.

NoW : We use it to evaluate face/head-only accuracy. It contains 3D head scans for 100100 subjects, and 20542054 images with various viewing angles and facial expressions.

FreiHAND : We evaluate hand-only accuracy on this. It has 37k37k hand/hand-object images of 3232 subjects, with MANO ground truth, recovered from multi-view images.

2 Evaluation Metrics

Mesh alignment: Prior to computing a metric, we align estimated meshes to ground-truth ones. The prefix “ PA” denotes Procrustes Alignment (solving for scale, rotation and translation), while “ TR” denotes translation alignment. “ TR” is stricter, as it does not factor out scale and rotation. When reporting hand-/face-only metrics for the full body, we align each part separately.

Mean Per-Joint Position Error ( MPJPE): We report the mean Euclidean distance between the estimated and ground-truth joints. For the body-only metric, we compute the 1414 LSP-common joints as a common skeleton across different body models, using a linear joint regressor on the estimated and ground-truth vertices. This is a standard metric, but is too sparse; it cannot capture errors in full 3D shape (i.e. surface), or all limb rotation errors.

Vertex-to-Vertex ( V2V): For methods that infer meshes with the same topology as the ground-truth ones, e.g. SMPL(-X) estimations and SMPL(-X) ground truth, we compute the mean per-vertex error by taking into account all vertices. This is not possible for methods with different topology, e.g. SMPL estimations for SMPL-X ground truth, and vice versa. For such cases, we compute a main-body variant of V2V, i.e. without the hands and head, as SMPL and SMPL-X share the same topology for the main body. FB- V2V is the weighted sum of body (B), hand (LH, RH) and face (F) errors: FB=B+LH+RH+F3\text{FB}=\text{B}+\frac{\text{LH}+\text{RH}+\text{F}}{3}. V2V is stricter than MPJPE; it also captures 3D shape errors and unnatural limb rotations (for the same joint positions).

Point-to-Surface ( P2S): To compare PIXIE with methods that use a different mesh topology to SMPL(-X), e.g. MTC , we measure the mean distance from ground-truth vertices to the surface of the estimated mesh. P2S is stricter than MPJPE; it captures errors in 3D shape, but not unnatural limb rotations (for the same joint positions).

3 Quantitative Evaluation

Whole-body. In Tab. 1 - 2 we report whole-body metrics (“All”), by taking into account the body, face and hands jointly. We add body-only (“Body”), hand-only (“L/R hand”), and face-only (“Face”) variants for completeness.

EHF : Table 1 compares PIXIE to three baseline sets: (1) the optimization-based SMPLify-X and MTC that infer SMPL-X and Adam, (2) the regression-based SPIN that infers SMPL, and (3) the regression-based ExPose and FrankMocap that infer SMPL-X. Note that MTC does not estimate the face. PIXIE outperforms optimization methods on most metrics, while being significantly faster. Moreover, it is on par with regression methods, both in terms of error metrics and runtime, which drops to 0.080.08 sec for known body-part crops.

AGORA : Figure 4 compares PIXIE to whole-body and body-only regressors, for a varying occlusion degree. PIXIE outperforms all methods, and is competitive on body-only metrics even to the occlusion-aware PARE . Note that AGORA is much more complex and natural than EHF, making the results more representative of real-world scenarios.

Ablation for moderators: Table 2 compares PIXIE to naive whole-body regression (no body-part experts) and the “copy-paste” fusion strategy. The latter copies pose parameters from the part experts (see ), as well as shape parameters from the face expert, to the whole body.

The naive version does not benefit from the expertise of the part experts. “Copy-paste” fusion can lead to erroneous hand/face orientation inference, since the respective experts lack global context. Moreover, estimating whole-body shape from a face image is not always reliable, e.g. when a person faces away from the camera (Fig. 2). PIXIE fuses “global” body and “local” part features with its moderators. In this way, it estimates more accurate 3D bodies and is more robust to challenging ambiguities (blur, occlusion) than existing whole-body regressors, especially on stricter metrics without Procrustes alignment. Ablation for “gendered” shape loss on 3DPW : By removing our “gendered” shape loss, the PA- V2V error increases from 50.950.9 to 51.751.7 mm. A qualitative ablation is shown in Fig. 5; learned implicit reasoning about gender gives more realistic body shapes. SMPL-X’s shared shape space for the whole body lets parts contribute to the whole.

Parts-only: For completeness, we use standard benchmarks for body-only, face-only, and hand-only evaluation.

Body-only on 3DPW : Table 3 shows that PIXIE performs on par FrankMocap and ExPose and is worse than SPIN , for the PA- MPJPE metric, but outperforms them all in the stricter TR- MPJPE (joints) and V2V (surface) metrics.

Face-only on NoW : Table 4 shows that PIXIE outperforms not only the expressive whole-body method ExPose , but also strong and dedicated face-only methods, except for the recent work of Feng et al. . Hand-only on FreiHAND : Table 5 shows that our hand expert performs on par with the whole-body ExPose , is a bit worse than the hand-specific “MANO CNN” , but outperforms the hand expert of Zhou et al. .

4 Qualitative Evaluation

Figure 6 compares PIXIE with FrankMocap and ExPose , which also regresses SMPL-X. Both baselines fail when the hand expert faces ambiguities (row 22); PIXIE gains robustness by using the full-body context. Both baselines give body shapes that look average (rows 11, 44) or have the wrong gender (rows 22, 33); PIXIE gives the most realistic shapes due to its “gendered” shape loss. FrankMocap fails for strong occlusions (rows 11, 33). Lastly, ExPose struggles with accurate facial expressions, and FrankMocap with head rotations (rows 11, 33); PIXIE outperforms both with its strong face/head expert and predicts a more realistic face.

Figure 7 compares PIXIE with Zhou et al. , recent work that also estimates a textured face. PIXIE gives more accurate poses (see how hands and faces align to the image), as it fuses both face-body and hand-body expert features, weighted by their confidence. PIXIE also gives more realistic body shapes, both due to its gendered shape loss and due to part experts contributing to whole-body shape, using SMPL-X’s shared body, hand and face shape space.

Future work: Mesh-to-image misalignment is a common limitation of regressors that pool “global” features from the image, losing local information. This could be tackled with “pixel-aligned” features . Moreover, SMPL-X models bodies without clothing; adding clothing models is a challenging but promising avenue. Furthermore, due to the formulation of the photometric term the model prefers to explain image evidence using lighting, rather than albedo, which leads to wrong skin tone predictions. Future work could further improve cases with self-contact , or other extreme ambiguities.

Conclusion

We present PIXIE, a novel expressive whole-body reconstruction method that recovers an animatable 3D avatar with a detailed face from a single RGB image. PIXIE uses body-driven attention to leverage dedicated body, head and face experts. It learns a novel moderator that reasons about the confidence of each expert, to fuse their features according to confidence, and exploit their complementary strengths. It uses the best practices from the face community for accurate faces with realistic albedo and geometric details. The face expert can contribute to more realistic whole-body shapes, by using a shared face-body shape space. To further improve shape, PIXIE uses implicit reasoning about gender, to encourage likely “gendered” body shapes. Qualitative results show natural and expressive humans, with improved body shape, well articulated hands, and realistic faces, comparable to the best face-only methods. We believe that PIXIE will be useful for many applications that need expressive human understanding from images. Acknowledgments:We thank Victoria Fernández Abrevaya, Yinghao Huang, Yuliang Xiu, Radek Danecek for discussions and Priyanka Patel for AGORA experiments. This work was partially supported by the Max Planck ETH Center for Learning Systems. Disclosure: https://files.is.tue.mpg.de/black/CoI/3DV2021.txt

References

Appendix A Implementation Details

Data augmentation: For training data, we use image crops around the body, face and hands. We augment our training image crops, following mainly , as described below. First, we use standard techniques, namely random horizontal flipping, random image rotations, color noise addition and random translation of the crop’s center. However this is not enough, as there is a significant domain gap between face-only and hand-only datasets, and the respective image crops extracted from full-body images; the former have significantly higher resolution. To account for this, we also randomly down-sample and up-sample the head and hand image crops, to simulate various lower resolutions. Finally, inspired by , we add synthetic motion blur to face and hand crops, to simulate the motion blur that is common in full-body images.

Training details: We use PyTorch to implement our pipeline. We follow a three-step training procedure: (1) We pre-train the model with body-only, face-only and hand-only datasets; for each dataset we train only the respective parameters. Since these datasets are captured independently, there is no body image that corresponds to a face-only or hand-only image. Consequently, for this step we cannot apply feature fusion, and body-part features go directly to the respective regressor(s) (bypassing the moderators), to estimate the respective body-part parameters. Similar to existing work, we train only a right hand regressor; for images of a left hand, we flip the image horizontally to use the right hand regressor, and mirror the predictions to get a left hand. (2) Then, using the same data, we freeze the feature encoders and proceed with training the regressors and extractors (see Fig. 33 of the paper the linear layers \mathcal{L}^{{\color[rgb]{0,0,0}h}} and \mathcal{L}^{{\color[rgb]{0,0,0}f}} between the body encoder E_{{\color[rgb]{0,0,0}b}} and moderators \mathcal{M}_{{\color[rgb]{0,0,0}h}} and \mathcal{M}_{{\color[rgb]{0,0,0}f}} respectively). This step encourages features F_{{\color[rgb]{0,0,0}b}}^{{\color[rgb]{0,0,0}h}} and F_{{\color[rgb]{0,0,0}b}}^{{\color[rgb]{0,0,0}f}} from body images to be in the same space as features F_{{\color[rgb]{0,0,0}h}} and F_{{\color[rgb]{0,0,0}f}} from part-only images, so that regressors \mathcal{R}_{{\color[rgb]{0,0,0}f}}^{{\color[rgb]{0,0,0}\text{fused}}} and \mathcal{R}_{{\color[rgb]{0,0,0}h}}^{{\color[rgb]{0,0,0}\text{fused}}} work for both feature types. (3) Finally, we train the full network, including the moderators \mathcal{M}_{{\color[rgb]{0,0,0}h}} and \mathcal{M}_{{\color[rgb]{0,0,0}f}}, but this time using training images with full SMPL-X ground truth, to extract part crops from full-body images as well. However, there are two problems. First, for these images there is no skin mask available, consequently we remove the loss for body shape β\bm{\beta} and do not apply a photometric and identity loss on head crops. Second, localizing the hands with body-driven attention is much harder compared to the head, due to the longer kinematic chain, consequently we freeze the hand regressor \mathcal{R}_{{\color[rgb]{0,0,0}h}} to avoid fine-tuning it with invalid inputs.

All parameters are optimized using Adam with a learning rate of 0.00010.0001. For training the body, hand and face sub-networks, we use a batch size of 1616 , 1616, and 88, respectively. The moderator is a fully connected network with the following structure: FC (2048, 1024), ReLU, FC (1024, 1). All input images are resized to 224×224224\times 224 pixels before feeding them to our network. During inference, we extract the hand/face crops using the hand and face locations from RbR_{b}’s output. Hand and face cameras are ignored when estimating full body pose.

Global to relative pose: The regressors \mathcal{R}_{{\color[rgb]{0,0,0}f}}^{{\color[rgb]{0,0,0}\text{fused}}} and \mathcal{R}_{{\color[rgb]{0,0,0}h}}^{{\color[rgb]{0,0,0}\text{fused}}} estimate the absolute head and wrist orientation θg\bm{\theta}_{\text{g}}, i.e. irrespective of the (parent) main body’s pose. However, to “apply” these θg\bm{\theta}_{\text{g}} estimates on a SMPL-X body that is already posed by \mathcal{R}_{{\color[rgb]{0,0,0}b}} with \bm{\theta}_{{\color[rgb]{0,0,0}b}} (up to the wrist and neck, excluding them), we need to express them relative to their parent in the kinematic skeleton:

where Γ\bm{\Gamma} is the chain transformation function according to SMPL-X’s kinematic skeleton hierarchy.

Appendix B Evaluation

PIXIE gives more realistic body shapes, not only due to its gendered shape loss, but also thanks to the shared body, hand and face shape space of SMPL-X. This allows PIXIE’s face expert to – uniquely – contribute to whole-body shape. To verify this, we apply our face expert on face-only images and get the whole-body shapes of Fig. A.1. These are not only correctly “gendered”, but also have a plausible BMI. For the sumo wrestler in Fig. A.1, PIXIE predicts a body with higher BMI (26.9) than the mean shape (26.1). PIXIE is the only 3D whole-body estimation method that explores such face-body shape correlations explicitly. We believe that this is a useful insight and points the community towards a new direction.

B.2 Qualitative Evaluation

Comparison with MTC: In Fig. A.2 we compare PIXIE with MTC . PIXIE is two orders of magnitude faster and predicts more accurate 3D body shapes. However, when 2D joint estimations are accurate, optimization-based methods, such as MTC and SMPLify-X , tend to estimate bodies that are better aligned with the image.

Expressive body reconstruction: We compare our method, PIXIE, with other state-of-the-art expressive body reconstruction methods in Fig. A.3.

PIXIE is more robust to challenging ambiguities (blur, occlusion) than existing whole-body regressors , since its moderators fuses “global” body and “local” part.

Qualitative results: Finally, in Fig. A.4, A.5 and A.6 we provide more standalone PIXIE results. Overall, PIXIE produces visually plausible body shapes with detailed facial expressions.

Failure cases: Although the gender prior loss and the shared whole-body shape space result in better 3D shape predictions, they are not sufficient for perfectly estimating full-body 3D shape. Furthermore, the employed photometric term often causes the model to prefer to explain image evidence using lighting, rather than albedo, which leads to incorrect skin tone predictions. These points highlight important directions for improving PIXIE. Representative failure cases can be seen in Fig. A.7.