3D Multi-bodies: Fitting Sets of Plausible 3D Human Models to Ambiguous Image Data

Benjamin Biggs, Sébastien Ehrhadt, Hanbyul Joo, Benjamin Graham, Andrea Vedaldi, David Novotny

Introduction

We are interested in reconstructing 3D human pose from the observation of single 2D images. As humans, we have no problem in predicting, at least approximately, the 3D structure of most scenes, including the pose and shape of other people, even from a single view. However, 2D images notoriously do not contain sufficient geometric information to allow recovery of the third dimension. Hence, single-view reconstruction is only possible in a probabilistic sense and the goal is to make the posterior distribution as sharp as possible, by learning a strong prior on the space of possible solutions.

Recent progress in single-view 3D pose reconstruction has been impressive. Methods such as HMR , GraphCMR and SPIN formulate this task as learning a deep neural network that maps 2D images to the parameters of a 3D model of the human body, usually SMPL . These methods work well in general, but not always (fig. 2). Their main weakness is processing heavily occluded images of the object. When a large part of the object is missing, say the lower body of a sitting human, they output reconstructions that are often implausible. Since they can produce only one hypothesis as output, they very likely learn to approximate the mean of the posterior distribution, which may not correspond to any plausible pose. Unfortunately, this failure modality is rather common in applications due to scene clutter and crowds.

In this paper, we propose a solution to this issue. Specifically, we consider the challenge of recovering 3D mesh reconstructions of complex articulated objects such as humans from highly ambiguous image data, often containing significant occlusions of the object. Clearly, it is generally impossible to reconstruct the object uniquely if too much evidence is missing; however, we can still predict a set containing all possible reconstructions (see fig. 1), making this set as small as possible. While ambiguous pose reconstruction has been previously investigated, as far as we know, this is the first paper that looks specifically at a deep learning approach for ambiguous reconstructions of the full human mesh.

Our primary contribution is to introduce a principled multi-hypothesis framework to model the ambiguities in monocular pose recovery. In the literature, such multiple-hypotheses networks are often trained with a so-called best-of-MM loss — namely, during training, the loss is incurred only by the best of the MM hypothesis, back-propagating gradients from that alone . In this work we opt for the best-of-MM approach since it has been show to outperform alternatives (such as variational auto-encoders or mixture density networks) in tasks that are similar to our 3D human pose recovery, and which have constrained output spaces .

A major drawback of the best-of-MM approach is that it only guarantees that one of the hypotheses lies close to the correct solution; however, it says nothing about the plausibility, or lack thereof, of the other M−1M-1 hypotheses, which can be arbitrarily ‘bad’. Theoretically, best-of-MM can minimize its loss by quantizing optimally (in the sense of minimum expected distortion) the posterior distribution, which would be desirable for coverage. However, this is not the only solution that optimizes the best-of-MM training loss, as in the end it is sufficient that one hypothesis per training sample is close to the ground truth. In fact, this is exactly what happens; for instance, during training hypotheses in best-of-MM are known to easily become degenerate and ‘die off’, a clear symptom of this problem. Not only does this mean that most of the hypotheses may be uninformative, but in an application we are also unable to tell which hypothesis should be used, and we might very well pick a ‘bad’ one. This has also a detrimental effect during learning because it makes gradients sparse as prediction errors are back-propagated only through one of the MM hypotheses for each training image.

In order to address these issues, our first contribution is a hypothesis reprojection loss that forces each member of the multi-hypothesis set to correctly reproject to 2D image keypoint annotations. The main benefit is to constrain the whole predicted set of meshes to be consistent with the observed image, not just the best hypothesis, also addressing gradient sparsity.

Next, we observe that another drawback of the best-of-MM pipelines is to be tied to a particular value of MM, whereas in applications we are often interested in tuning the number of hypothesis considered. Furthermore, minimizing the reprojection loss makes hypotheses geometrically consistent with the observation, but not necessarily likely. Our second contribution is thus to improve the flexibility of best-of-MM models by allowing them to output any smaller number n<Mn<M of hypotheses while at the same time making these hypotheses more representative of likely poses. The new method, which we call nn-quantized-best-of-MM, does so by quantizing the best-of-MM model to output weighed by a explicit pose prior, learned by means of normalizing flows.

To summarise, our key contributions are as follows. First, we deal with the challenge of 3D mesh reconstruction for articulated objects such as humans in ambiguous scenarios. Second, we introduce a nn-quantized-best-of-MM mechanism to allow best-of-MM models to generate an arbitrary number of n<Mn<M predictions. Third, we introduce a mode-wise re-projection loss for multi-hypothesis prediction, to ensure that predicted hypotheses are all consistent with the input.

Empirically, we achieve state-of-the-art monocular mesh recovery accuracy on Human36M, its more challenging version augmented with heavy occlusions, and the 3DPW datasets. Our ablation study validates each of our modelling choices, demonstrating their positive effect.

Related work

There is ample literature on recovering the pose of 3D models from images. We break this into five categories: methods that reconstruct 3D points directly, methods that reconstruct the parameters of a 3D model of the object via optimization, methods that do the latter via learning-based regression, hybrid methods and methods which deal with uncertainty in 3D human reconstruction.

Several papers have focused on the problem of estimating 3D body points from 2D observations . Of these, Martinez et al. introduced a particularly simple pipeline based on a shallow neural network. In this work, we aim at recovering the full 3D surface of a human body, rather than only lifting sparse keypoints.

Fitting 3D models via direct optimization.

Several methods fit the parameters of a 3D model such as SMPL or SCAPE to 2D observations using an optimization algorithm to iteratively improve the fitting quality. While early approaches such as required some manual intervention, the SMPLify method of Bogo et al. was perhaps the first to fit SMPL to 2D keypoints fully automatically. SMPL was then extended to use silhouette, multiple views, and multiple people in . Recent optimization methods such as have significantly increased the scale of the models and data that can be handled.

Fitting 3D models via learning-based regression.

More recently, methods have focused on regressing the parameters of the 3D models directly, in a feed-forward manner, generally by learning a deep neural network . Due to the scarcity of 3D ground truth data for humans in the wild, most of these methods train a deep regressor using a mix of datasets with 3D and 2D annotations in form of 3D MoCap markers, 2D keypoints and silhouettes. Among those, HMR of Kanazawa et al. and GraphCMR of Kolotouros et al. stand out as particularly effective.

Hybrid methods.

Other authors have also combined optimization and learning-based regression methods. In most cases, the integration is done by using a deep regressor to initialize the optimization algorithm . However, recently Kolotouros et al. has shown strong results by integrating the optimization loop in learning the deep neural network that performs the regression, thereby exploiting the weak cues available in 2D keypoints.

Modelling ambiguities in 3D human reconstruction.

Several previous papers have looked at the problem of modelling ambiguous 3D human pose reconstructions. Early work includes Sminchisescu and Triggs , Sidenbladh et al. and Sminchisescu et al. .

More recently, Akhter and Black learn a prior over human skeleton joint angles (but not directly a prior on the SMPL parameters) from a MoCap dataset. Li and Lee use the Mixture Density Networks model of to capture ambiguous 3D reconstructions of sparse human body keypoints directly in physical space. Sharma et al. learn a conditional variational auto-encoder to model ambiguous reconstructions as a posterior distribution; they also propose two scoring methods to extract a single 3D reconstruction from the distribution.

Cheng et al. tackle the problem of video 3D reconstruction in the presence of occlusions, and show that temporal cues can be used to disambiguate the solution. While our method is similar in the goal of correctly handling the prediction uncertainty, we differ by applying our method to predicting full mesh of the human body. This is arguably a more challenging scenario due to the increased complexity of the desired 3D shape.

Finally, some recent concurrent works also consider building priors over 3D human pose using normalizing flows. Xu et al. release a prior for their new GHUM/GHUML model, and Zanfir et al. build a prior on SMPL joint angles to constrain their weakly-supervised network. Our method differs as we learn our prior on 3D SMPL joints.

Preliminaries

Before discussing our method, we describe the necessary background, starting from SMPL.

Predicting the SMPL parameters from a single image.

Normalizing flows.

The idea of normalizing flows (NF) is to represent a complex distribution p(X)p(X) on a random variable XX as a much simpler distribution p(z)p(z) on a transformed version z=f(X)z=f(X) of XX. The transformation ff is learned so that p(z)p(z) has a fixed shape, usually a Normal p(z)∼N(0,1)p(z)\sim\mathcal{N}(0,1). Furthermore, ff itself must be invertible and smooth. In this paper, we utilize a particular version of NF dubbed RealNVP . A more detailed explanation of NF and RealNVP has been deferred to the supplementary.

Method

We start from a neural network architecture that implements the function G(I)=(θ,β,γ,t)G(I)=(\theta,\beta,\gamma,t) described above. As shown in SPIN , the HMR architecture attains state-of-the-art results for this task, so we use it here. However, the resulting regressor G(I)G(I), given an input image II, can only produce a single unique solution. In general, and in particular for cases with a high degree of reconstruction ambiguity, we are interested in predicting set of plausible 3D poses rather than a single one. We thus extend our model to explicitly produce a set of MM different hypotheses Gm(I)=(θm,βm,γm,tm)G_{m}(I)=(\theta_{m},\beta_{m},\gamma_{m},t_{m}), m=1,…,Mm=1,\dots,M. This is easily achieved by modifying the HMR’s final output layer to produce a tensor MM times larger, effectively stacking the hypotheses. In what follows, we describe the learning scheme that drives the monocular predictor GG to achieve an optimal coverage of the plausible poses consistent with the input image. Our method is summarized in fig. 3.

For learning the model, we assume to have a training set of NN images {Ii}i=1,…,N\{I_{i}\}_{i=1,\dots,N}, each cropped around a person. Furthermore, for each training image IiI_{i} we assume to know (1) the 2D location YiY_{i} of the body joints (2) their 3D location XiX_{i}, and (3) the ground truth SMPL fit (θi,βi,γi)(\theta_{i},\beta_{i},\gamma_{i}). Depending on the set up, some of these quantities can be inferred from the others (e.g. we can use the function JJ to convert the SMPL parameters to the 3D joints XiX_{i} and then the camera projection to obtain YiY_{i}).

Given a single input image, our network predicts a set of poses, where at least one should be similar to the ground truth annotation XiX_{i}. This is captured by the best-of-MM loss :

where X^m(Ii)=J(Gm(V(Ii)))\hat{X}^{m}(I_{i})=J(G_{m}(V(I_{i}))) are the 3D joints estimated by the mm-th SMPL predictor Gm(Ii)G_{m}(I_{i}) applied to image IiI_{i}. In this way, only the best hypothesis is steered to match the ground truth, leaving the other hypotheses free to sample the space of ambiguous solutions. During the computation of this loss, we also extract the best index mi∗m^{*}_{i} for each training example.

Limitations of best-of-M𝑀M.

As noted in section 1, best-of-MM only guarantees that one of the MM hypotheses is a good solution, but says nothing about the other ones. Furthermore, in applications we are often interested in modulating the number of hypotheses generated, but the best-of-MM regressor G(I)G(I) only produces a fixed number of output hypothesis MM, and changing MM would require retraining from scratch, which is intractable.

We first address these issues by introducing a method that allows us to train a best-of-MM model for a large MM once and leverage it later to generate an arbitrary number of n<Mn<M hypotheses without the need of retraining, while ensuring that these are good representatives of likely body poses.

n𝑛n-quantized-best-of-M𝑀M

Formally, given a set of MM predictions X^M(I)={X^1(I),...,X^M(I)}\mathcal{\hat{X}}^{M}(I)=\{\hat{X}^{1}(I),...,\hat{X}^{M}(I)\} we seek to generate a smaller nn-sized set Xˉn(I)={Xˉ1(I),...,Xˉn(I)}\mathcal{\bar{X}}^{n}(I)=\{\bar{X}^{1}(I),...,\bar{X}^{n}(I)\} which preserves the information contained in X^M\mathcal{\hat{X}}^{M}. In other words, Xˉn\mathcal{\bar{X}}^{n} optimally quantizes X^M\mathcal{\hat{X}}^{M}. To this end, we interpret the output of the best-of-MM model as a set of choices X^M(I)\mathcal{\hat{X}}^{M}(I) for the possible pose. These poses are of course not all equally likely, but it is difficult to infer their probability from (1). We thus work with the following approximation. We consider the prior p(X)p(X) on possible poses (defined in the next section), and set:

This amounts to using the best-of-MM output as a conditioning set (i.e. an unweighted selection of plausible poses) and then use the prior p(x)p(x) to weight the samples in this set. With the weighted samples, we can then run KK-means to further quantize the best-of-MM output while minimizing the quantization energy EE:

This can be done efficiently on GPU — for our problem, K-Means consumes less than 20% of the execution time of the entire forward pass of our method.

Learning the pose prior with normalizing flows.

In order to obtain p(X)p(X), we propose to learn a normalizing flow model in form of the RealNVP network ff described in section 3 and the supplementary. RealNVP optimizes the log likelihood Lnf(f)\mathcal{L}_{\text{nf}}(f) of training ground truth 3D skeletons {X1,...XN}\{X_{1},...X_{N}\} annotated in their corresponding images {I1,...,IN}\{I_{1},...,I_{N}\} :

D re-projection loss.

Since the best-of-MM loss optimizes a single prediction at a time, often some members of the ensemble X^(I)\mathcal{\hat{X}}(I) drift away from the manifold of plausible human body shapes, ultimately becoming ‘dead’ predictions that are never selected as the best hypothesis m∗m^{*}. In order to prevent this, we further utilize a re-projection loss that acts across all hypotheses for a given image. More specifically, we constrain the set of 3D reconstructions to lie on projection rays passing through the 2D input keypoints with the following hypothesis re-projection loss:

Note that many of our training images exhibit significant occlusion, so YY may contain invisible or missing points. We handle this by masking Lri\mathcal{L}_{\text{ri}} to prevent these points contributing to the loss.

SMPL loss.

The final loss terms, introduced by prior work , penalize deviations between the predicted and ground truth SMPL parameters. For our method, these are only applied to the best hypothesis mi∗m_{i}^{*} found above:

Note here we use Lrb\mathcal{L}_{\text{rb}} to refer to a 2D re-projection error between the best hypothesis and ground truth 2D points YiY_{i}. This differs from the earlier loss Lri\mathcal{L}_{\text{ri}}, which is applied across all modes to enforce consistency to the visible input points. Note that we could have used eqs. 6 and 7 to select the best hypothesis mi∗m_{i}^{*}, but it would entail an unmanageable memory footprint due to the requirement of SMPL-meshing for every hypothesis before the best-of-MM selection.

Overall loss.

where m∗m^{*} is given in eq. 1 and λri,λbest,λθ,λβ,λV,λrb\lambda_{\text{ri}},\lambda_{\text{best}},\lambda_{\theta},\lambda_{\beta},\lambda_{\text{V}},\lambda_{\text{rb}} are weighing factors. We use a consistent set of SMPL loss weights across all experiments λbest=25.0,λθ=1.0,λβ=0.001,λV=1.0,\lambda_{\text{best}}=25.0,\lambda_{\theta}=1.0,\lambda_{\beta}=0.001,\lambda_{\text{V}}=1.0, and set λri=1.0\lambda_{\text{ri}}=1.0. Since the training of the normalizing flow ff is independent of the rest of the model, we train ff separately by optimizing Lnf\mathcal{L}_{\text{nf}} with the weight of λnf=1.0\lambda_{\text{nf}}=1.0. Samples from our trained normalizing flow are shown in fig. 4

Experiments

In this section we compare our method to several strong baselines. We start by describing the datasets and the baselines, followed by a quantitative and a qualitative evaluation.

Our evaluation focuses on the Human3.6m (H36M) and 3DPW datasets . H36M is one of the largest datasets of humans annotated with 3D pose using MoCap sensors.

As common practice, we train on subjects S1, S5, S6, S7 and S8, and test on S9 and S11. 3DPW is only used for evaluation and, following , we evaluate on its test set.

Our evaluation is consistent with - we report two metrics that compare the lifted dense 3D SMPL shape to the ground truth mesh: Mean Per Joint Position Error (MPJPE), Reconstruction Error (RE). For H36M, all errors are computed using an evaluation scheme known as “Protocol #2”. Please refer to supplementary for a detailed explanation of MPJPE and RE.

Multipose metrics.

MPJPE and RE are traditional metrics that assume a single correct ground truth prediction for a given 2D observation. As mentioned above, such an assumption is rarely correct due to the inherent ambiguity of the monocular 3D shape estimation task. We thus also report MPJPE-nn/RE-nn an extension of MPJPE RE used in , that enables an evaluation of nn different shape hypotheses. In more detail, to evaluate an algorithm, we allow it to output nn possible predictions and, out of this set, we select the one that minimizes the MPJPE/RE metric. We report results for n∈{1,5,10,25}n\in\{1,5,10,25\}.

Ambiguous H36M/3DPW (AH36M/A3DPW).

Since H36M is captured in a controlled environment, it rarely depicts challenging real-world scenarios such as body occlusions that are the main source of ambiguity in the single-view 3D shape estimation problem.

Hence, we construct an adapted version of H36M with synthetically-generated occlusions (fig. 5) by randomly hiding a subset of the 2D keypoints and re-computing an image crop around the remaining visible joints. Please refer to the supplementary for details of the occlusion generation process.

While 3DPW does contain real scenes, for completeness, we also evaluate on a noisy, and thus more challenging version (A3DPW) generated according to the aforementioned strategy.

Baselines

Our method is compared to two multi-pose prediction baselines. For fairness, both baselines extend the same (state-of-the-art) trunk architecture as we use, and all methods have access to the same training data.

SMPL-MDN follows and outputs parameters of a mixture density model over the set of SMPL log-rotation pose parameters. Since a naïve implementation of the MDN model leads to poor performance (≈\approx 200mm MPJPE-n=5n=5 on H36M), we introduced several improvements that allow optimization of the total loss eq. 8. SMPL-CVAE, the second baseline, is a conditional variational autoencoder combined with our trunk network. SMPL-CVAE consists of an encoding network that maps a ground truth SMPL mesh VV to a gaussian vector zz which is fed together with an encoding of the image to generate a mesh V′V^{\prime} such that V′≈VV^{\prime}\approx V. At test time, we sample nn plausible human meshes by drawing z∼N(0,1)z\sim\mathcal{N}(0,1) to evaluate with MPJPE-nn/RE-nn. More details of both SMPL-CVAE and SMPL-MDN have been deferred to the supplementary material.

For completeness, we also compare to three more baselines that tackle the standard single-mesh prediction problem: HMR , GraphCMR , and SPIN , where the latter currently attain state-of-the-art performance on H36M/3DPW. All methods were trained on H36M , MPI-INF-3DHP , LSP , MPII and COCO .

1 Results

Table 1 contains a comprehensive summary of the results on all 3 benchmarks. Our method outperforms the SMPL-CVAE and SMPL-MDN in all metrics on all datasets. For SMPL-CVAE, we found that the encoding network often “cheats” during training by transporting all information about the ground truth, instead of only encoding the modes of ambiguity. The reason for a lower performance of SMPL-MDN is probably the representation of the probability in the space of log-rotations, rather in the space of vertices. Modelling the MDN in the space of model vertices would be more convenient due to being more relevant to the final evaluation metric that aggregates per-vertex errors, however, fitting such high-dimensional (dim=6890×36890\times 3) Gaussian mixture is prohibitively costly.

Furthermore, it is very encouraging to observe that our method is also able to outperform the single-mode baselines on the single mode MPJPE on both H36M and 3DPW. This comes as a surprise since our method has not been optimized for this mode of operation. The difference is more significant for 3DPW which probably happens because 3DPW is not used for training and, hence, the normalizing flow prior acts as an effective filter of predicted outlier poses. Qualitiative results are shown in fig. 6.

We further conduct an ablative study on 3DPW that removes components of our method and measures the incurred change in performance. More specifically, we: 1) ablate the hypothesis reprojection loss; 2) set p(X∣I)=Uniformp(X|I)=\text{Uniform} in eq. 3, effectively removing the normalizing flow component and executing unweighted K-Means in nn-quantized-best-of-MM. Table 2 demonstrates that removing both contributions decreases performance, validating our design choices.

Conclusions

In this work, we have explored a seldom visited problem of representing the set of plausible 3D meshes corresponding to a single ambiguous input image of a human. To this end, we have proposed a novel method that trains a single multi-hypothesis best-of-MM model and, using a novel nn-quantized-best-of-MM strategy, allows to sample an arbitrary number n<Mn<M of hypotheses.

Importantly, this proposed quantization technique leverages a normalizing flow model, that effectively filters out the predicted hypotheses that are unnatural. Empirical evaluation reveals performance superior to several strong probabilistic baselines on Human36M, its challenging ambiguous version, and on 3DPW. Our method encounters occasional failure cases, such as when tested on individuals with unusual shape (e.g. obese people), since we have very few of these examples in the training set. Tackling such cases would make for interesting and worthwhile future work.

The authors would like to thank Richard Turner for useful technical discussions relating to normalizing flows, and Philippa Liggins, Thomas Roddick and Nicholas Biggs for proof reading. This work was entirely funded by Facebook AI Research.

Broader impact

Our method improves the ability of machines to understand human body poses in images and videos. Understanding people automatically may arguably be misused by bad actors. However, importantly, our method is not a form of biometric as it does not allow the identification of people. Rather, only their overall body shape and pose is reconstructed, but these details are insufficient for unique identification. In particular, individual facial features are not reconstructed at all.

Furthermore, our method is an improvement of existing capabilities, but does not introduce a radical new capability in machine learning. Thus our contribution is unlikely to facilitate misuse of technology which is already available to anyone.

Finally, any potential negative use of a technology should be balanced against positive uses. Understanding body poses has many legitimate applications in VR and AR, medical, assistance to the elderly, assistance to the visual impaired, autonomous driving, human-machine interactions, image and video categorization, platform integrity, etc.

References

Appendix A Evaluation metrics

More details related to the evaluation metrics, briefly outlined in section 5, are provided in this section.

We report two metrics that compare the lifted dense 3D SMPL shape to the ground truth mesh: Mean Per Joint Position Error (MPJPE), Reconstruction Error (RE). All errors are computed using an evaluation scheme known as “Protocol #1”, as explained below.

For each H36M test skeleton, MPJPE calculates the mean distance between 14 ground truth skeleton 3D joints and the predicted joints obtained by using a fixed linear regressor that maps the array of 6890 3D coordinates of the predicted dense mesh to the skeleton 3D joint coordinates. We report an average of all MPJPE errors measured for each test skeleton. The reconstruction error (RE) is a modification of MPJPE which consists of finding an additional rigid Procrustes alignment between the pair of assessed poses before evaluating the inter-joint distances.

Appendix B Ambiguous H36m/3DPW

In this section we give a detailed explanation of the generation of the Ambiguous H36m/3DPW datasets (briefly explained in section 5).

We begin with the full size image with a set of 2D joints and apply synthetic occlusions to the subject’s body parts by randomly hiding a subset of the 2D keypoints and re-computing a slightly padded image crop around the joints that remained visible. For each image, we randomly choose one of 4 possible strategies for hiding the keypoints: 1) Hiding arm and head keypoints; 2) legs; 3) head; 4) no keypoints hiddenThe selection probabilities are p(1)=p(2)=p(3)=0.3p(1)=p(2)=p(3)=0.3, p(4)=0.1p(4)=0.1.

Appendix C Multi-hypothesis baselines

Here, we describe SMPL-MDN and SMPL-CVAE (section 5) in more detail.

SMPL-MDN predicts parameters of a Gaussian mixture defined over the log-rotation parameters θ\theta of the SMPL kinematic tree. Here, mm-th Gaussian in the mixture is parametrized with a mean μm\mu_{m}, covariance matrix σm\sigma_{m} and mixture weight ωm\omega_{m}. As noted in section 5, for SMPL-MDN, it was crucial to enable optimization of the total loss (8) in addition to optimizing the log-likelihood of the predicted Gaussian mixture. While the mixture log-likelihood optimizes directly the mixture parameters, the total loss requires a single prediction of θ\theta. In order to obtain a single estimate of θ\theta that can enter the total loss, similar to the Best-of-MM loss, we utilize the Gaussian mixture parameters and the ground truth angles θ\theta to generate a virtual prediction θ^\hat{\theta} that lies close to the ground truth in the sense of the posterior probability of θ\theta. More specifically, the virtual θ^\hat{\theta} is defined as a weighted combination of mixture means μm\mu_{m}, where the weights are the posterior probabilities of θ\theta being assigned to mm-th mixture component:

where αm,σm,μm\alpha_{m},\sigma_{m},\mu_{m} are the weight, variance and mean of the mm-th mixture component respectively; and N(θ∣μm,σm)N(\theta|\mu_{m},\sigma_{m}) is an evaluation of the multivariate normal distribution with mean μm\mu_{m} and variance σm\sigma_{m} at θ\theta. Note that the best quantitative results were obtained with fixing ∀m:αm=1M,σm=0.001\forall m:\alpha_{m}=\frac{1}{M},\sigma_{m}=0.001 and only allowing μm\mu_{m} to learn.

This way, the SMPL-MDN regressor GMDNG_{MDN} is altered to generate a single prediction GMDN(I)=(θ^,β,γ,t)G_{MDN}(I)=(\hat{\theta},\beta,\gamma,t) that enters the total loss (8).

At test-time, following , the predicted hypotheses are randomly sampled per-mode predictions {(μm,β,γ,t)}m=1M\{(\mu_{m},\beta,\gamma,t)\}_{m=1}^{M} rather than random samples from the mixture. We observed that randomly sampling from the mixture density gave worse quantitative results.

SMPL-CVAE

SMPL-CVAE consists of a pair of encoder and decoder networks. The encoder network takes as input the ground truth SMPL mesh VV and outputs a Gaussian vector zz whose goal is to encode the mode of ambiguity. The decoder GCVAE(I,z)=(θ,β,γ,t)G_{CVAE}(I,z)=(\theta,\beta,\gamma,t) then takes as input the zz, together with the input image II, in order to generate the standard tuple of SMPL parameters. At train-time, the network minimizes the total loss (8) and a KL divergence between the predicted distribution of zz and a standard multivariate normal distribution N(0,1)N(0,1).

Appendix D Training details

Our network is trained in two stages. First we train the original HMR model until convergence according to the training protocol from . We then convert the model to our nn-quantized-best-of-MM architecture and continue training until convergence with an Adam optimizer with an initial learning rate of 10−510^{-5}.

Appendix E Performance analysis

A single inference pass takes on average 0.14s per image on NVIDIA V100 GPU. The overall training time (including the HMR pre-training step) is 5 days on a single gpu.

Appendix F Normalizing Flows

The idea of normalizing flows is to represent a complex distribution p(X)p(X) on a random variable XX as a much simpler distribution p(z)p(z) on a transformed version z=f(X)z=f(X) of XX. The transformation ff is learned so that p(z)p(z) has a fixed shape, usually a Normal p(z)∼N(0,1)p(z)\sim\mathcal{N}(0,1). Furthermore, ff itself must be invertible and smooth. In this case, the relation between p(θ)p(\theta) and p(z)p(z) is given by a change of variable

The challenge is to learn ff from data in a way that maintains its invertibility and smoothness. This is done by decomposing z=fL∘⋯∘f1(X)z=f_{L}\circ\dots\circ f_{1}(X) in nn layers, where Xl=fl(Xl−1)X_{l}=f_{l}(X_{l-1}), x=Xnx=X_{n} and X=X0X=X_{0}, and each layer is in turn smooth and invertible. Then one can write