Exemplar Fine-Tuning for 3D Human Model Fitting Towards In-the-Wild 3D Human Pose Estimation
Hanbyul Joo, Natalia Neverova, Andrea Vedaldi
Abstract
Differently from 2D image datasets such as COCO, large-scale human datasets with 3D ground-truth annotations are very difficult to obtain in the wild. In this paper, we address this problem by augmenting existing 2D datasets with high-quality 3D pose fits. Remarkably, the resulting annotations are sufficient to train from scratch 3D pose regressor networks that outperform the current state-of-the-art on in-the-wild benchmarks such as 3DPW. Additionally, training on our augmented data is straightforward as it does not require to mix multiple and incompatible 2D and 3D datasets or to use complicated network architectures and training procedures. This simplified pipeline affords additional improvements, including injecting extreme crop augmentations to better reconstruct highly truncated people, and incorporating auxiliary inputs to improve 3D pose estimation accuracy. It also reduces the dependency on 3D datasets such as H36M that have restrictive licenses. We also use our method to introduce new benchmarks for the study of real-world challenges such as occlusions, truncations, and rare body poses. In order to obtain such high quality 3D pseudo-annotations, inspired by progress in internal learning, we introduce Exemplar Fine-Tuning (EFT). EFT combines the re-projection accuracy of fitting methods like SMPLify with a 3D pose prior implicitly captured by a pre-trained 3D pose regressor network. We show that EFT produces 3D annotations that result in better downstream performance and are qualitatively preferable in an extensive human-based assessment.
Introduction
While 3D human pose estimation has progressed significantly in the past few years, the availability of suitable training datasets is still problematic. Obtaining 3D annotations for realistic benchmarks collected under everyday scenarios is difficult. Thus, most of the available 3D annotations are for datasets such as H36M and MPI-INF-3DHP , captured mostly indoors or in laboratory conditions. In order to improve generalization, previous approaches propose relatively complicated training procedures such as adversarial learning or SMPLify-in-the-loop that combine these 3D datasets with more realistic images with 2D keypoint annotations. For the same reason, using current 3D datasets for evaluation is also not necessarily representative of performance in the real world. Finally, many recent state-of-the-art approaches rely on SMPL fits on the H36M dataset which, due to licensing restrictions, make reproducibility problematic.
In this paper, we address these limitations by augmenting existing large-scale 2D datasets such as COCO , MPII , PoseTrack , LSPet with pseudo-ground truth (GT) 3D annotations, as shown in fig. 1. While still inferred from single images, the quality of thus produces annotations is surprisingly high. To show this, we demonstrate that the new data (e.g., COCO with our pseudo annotations) is sufficient to train state-of-the-art 3D pose regressor networks by itself, outperforming previous methods trained on the combination of datasets with 3D and 2D ground-truth on challenging benchmarks such as 3DPW. Notably, since the new dataset directly contains 3D annotations, which are unified across different benchmarks, training becomes straightforward without requiring complicated training techniques. Our datasets and codes are publicly available in our webpage.
Naturally, the success of our strategy depends on the quality of the pseudo-annotations that we can generate. Prior attempts at generating pseudo annotations for human pose such as UP-3D adopted established techniques such as SMPLify that fit the parameters of a 3D human model to the location of 2D keypoints, manually or automatically annotated in images . However, they regularize the solution by means of a “view-agnostic” 3D pose prior, which incurs limitations: fitting ignores the RGB images themselves, and balancing between the 3D prior, learned separately in laboratory conditions, and the data term (e.g., 2D keypoint error) is difficult. In practice, we show that the quality of the resulting fits is insufficient for our purposes.
To address these difficulties, we introduce Exemplar Fine-Tuning (EFT), an alternative pose fitting method inspired by recent progress in internal learning approaches . The idea is to start from an image-based 3D pose regressor such as . The pretrained regressor implicitly embodies an informal pose prior as a function , sending the image to the parameters of the 3D body. This function can be interpreted as a conditional re-parameterization of pose, where the conditioning factor is the image , and the new parameters are the weights of the underlying neural network regressor. We can then fit 2D annotations similar to SMPLify, but using this as a new pose parameterization by fine-tuning the weights rather than directly optimizing . While similar techniques have been used in to improve test-time performance, they have not been rigorously assessed against classical optimization approaches such as SMPLify, nor they have been used to generate pseudo-labels. Our extensive quantitative and qualitative analysis demonstrate that EFT results in high quality fits simply by minimizing the 2D reprojection loss without the use of an explicit 3D pose prior.
Using our 3D pseudo-annotations simplifies training 3D pose regressors and makes it easier to incorporate additional improvements, of which we explore two. First, we introduce extreme crop augmentation to train regressors that work better with truncated human bodies (e.g., upper body only) . Second, we demonstrate that auxiliary inputs, such as color-coded segmentation maps or DensePose IUV encodings , can further improve the 3D human pose estimation accuracy, outperforming previous state-of-the-art approaches simply by training on the COCO data. Compared to the previous approaches , these techniques are much easier to implement on top of our pseudo-GT dataset.
We also show that our pseudo-annotations are useful to benchmark models, not just to train them. For this, we introduce a new 3D benchmark representative of real-world scenarios not captured by existing indoor or outdoor datasets such as 3DPW . We include: (1) COCO validation images representative of various in-the-wild scenes, (2) LSPet images for gymnastic poses, and (3) OCHuman for severe occlusions. We use EFT to generate pseudo-GT annotations and assess them via a large human study conducted on Amazon Mechanical Turk (AMT). Furthermore, we propose a new protocol for the 3DPW dataset to assess performance when only part of the human body is visible (truncation). Overall, the scenarios in the new 3D benchmark complement 3DPW; furthermore, they contain a far large number of different subjects, and thus much greater diversity. Assessed against this difficult data, our networks still outperform competitors.
Summarizing, our contributions are: (1) providing large scale and high quality pseudo-GT 3D pose annotations that are sufficient to train state-of-the-art regressors without indoor 3D datasets; (2) running an extensive analysis of the quality of the pseudo-annotations generated via EFT against alternative approaches; (3) demonstrating the benefits of integrating the pseudo-annotations with extreme crop augmentation and auxiliary input representations; (4) introducing new 3D human pose benchmarks to assess regressors in less studied real-world scenarios.
Related Work
Deep learning has significantly advanced 2D pose recognition , facilitating the more challenging task of 3D reconstruction , which is our focus.
Single-view 3D Human Pose Estimation. Single-view 3D pose reconstruction methods differ in how they incorporate a 3D pose prior and in how they perform the prediction. Fitting-based methods assume a 3D body model such as SCAPE , SMPL , SMPL-X, and STAR , and use an optimization algorithm to fit it to the 2D observations. While early approaches required manual input, starting with SMPLify the process has been fully automatized, then improved in to use silhouette annotations, and eventually extended to multiple people . A recent method replaces the gradient descent update rule by a learned deep network . Regression-based methods, on the other hand, predict 3D pose directly. The work of uses sparse linear regression that incorporates a tractable but somewhat weak pose prior. Later approaches use instead deep neural networks, and differ mainly in the nature of their inputs and outputs . Some works start from a pre-detected 2D skeleton while others start from raw images or other proxy representation (e.g., DensePose outputs) . Using a 2D skeleton relies on the quality of the underlying 2D keypoint detector and discards appearance details that could help fitting the 3D model to the image. Using raw images can potentially make use of this information, but training such models from current 3D indoor datasets might fail to generalize to unconstrained images. Hence several papers combine 3D indoor datasets with 2D in-the-wild ones . Methods also differ in their output, with some predicting 3D keypoints directly , some predicting the parameters of a 3D human body model , and others volumetric heatmaps for the body joints or meshes . Finally, hybrid methods such as SPIN or MTC combine fitting and regression approaches.
3D Reconstruction Without Paired 3D Ground-truth. Fitting methods such as SMPLify only require 2D keypoint annotations and a human body model acquired independently on mocap data, but the output quality depends on the initialization, which is problematic for atypical poses. These methods can also use empirical pose priors by collecting a large number of samples in laboratory conditions , but those may lack realism. While most regression methods require paired 3D ground-truth, exceptions include that instead use 2D datasets and motion capture sequences to induce an empirical 3D prior enforced by means of adversarial learning or by using a geometric constraint on bone length ratios .
Human Pose Datasets. There are several in-the-wild datasets with sparse 2D pose annotations, including COCO , MPII , Leeds Sports Pose Dataset (LSP) , PennAction , and Posetrack . DensePose-COCO offers dense surface point annotations instead. Annotating 3D human poses is much more challenging as both specialized hardware and manual work are usually required. Examples include the Human3.6M dataset , Human Eva , Panoptic Studio , and MPI-INF-3DHP , which are hardly realistic. Two exceptions are 3DPW and PedX that capture 3D ground-truth outdoors, but these are collected using specialized hardware (IMUs and LiDAR ) which limits data diversity. Previous attempts at creating 3D pseudo-annotations include UP-3D , using SMPLIfy, ExPose , using SMPLIfy-X, and the work of Arnab et al. , using a multi-frame optimization. Their use of traditional fitting methods limits the quality and scope of the pseudo annotations, due to the challenge in modeling generic human pose priors.
Exemplar Fine Tuning
where the data term is the re-projection error between reconstructed and observed 2D keypoints. Since the viewpoint of the image is unknown, the camera parameters must be optimized together with the body parameters . The pose prior term prioritizes plausible solutions, compensating for the fact that 2D keypoints do not contain sufficient information to infer 3D pose uniquely, and is a scalar to adjust the weight of the prior term. SMPLify uses multiple prior terms including the pose prior expressed via a mixture of Gaussians or naïve thresholding, using a separate 3D motion capture dataset to learn the prior parameters beforehand, as well as an angle prior to penalize unnatural bending. is a shape prior term and is its weight. However, the success in optimizing (1) depends strongly on the quality of the heuristic used for initialization (e.g., a multi-stage approach by first aligning the torso and then optimizing the limbs together) and balancing the weights between the data term and multiple prior terms is crucial.
In contrast, regression-based approaches predict the model parameters directly from image-based cues such as raw RGB values and sparse and dense keypoints. The mapping is implemented by a neural network . Learning uses a number of example images with ground truth annotations for 2D keypoints , 3D keypoints and SMPL model parameters , which may in whole or in part be available, depending on the dataset (e.g. for 3D and for 2D annotations). Learning then optimizes the following training objective:
where and are loss-balancing coefficients which can be set to zero for samples that do not have 3D annotations. The camera projection is often predicted as additional output of the neural network .
During training, the regressor learns implicitly a prior on possible human poses, conditioned on the 2D input cues . This is arguably stronger than the prior used by fitting methods, which are learned separately, on different 3D data, and, being unconditional, tend to regress to a mean solution. On the other hand, fitting methods explicitly minimize the 2D re-projection error at test time, thus resulting in better 2D fits than regression methods that only minimize this term during training.
Exemplar Fine-Tuning (EFT) combines the advantages of fitting and regression methods. The idea is to interpret the network as a re-parameterization of the model as a function of the network parameters . With this, we can rewrite eq. 1 as where
The second term requires the network parameters to be close to the pre-trained regressor parameters . The shape regularizer favors the default SMPL shape parameters (as in eq. 1), but in practice its effect is very small (, see table 5).
Compared to eq. 1, we dropped the explicit pose prior , hypothesizing that a prior is implicitly captured by the pre-trained network, with the added advantage of accounting for the input image . Concretely, consider using a large value of for the prior in eqs. 1 and 3. Traditional fitting methods would then fall back to predicting the mean pose implied by the prior , ignoring the input. In contrast, in EFT a large value of falls back to the ‘best guess’ of the pre-trained regressor network for the pose of the observed input. Therefore, adjusting is much easier in EFT, and, in supp. material, we demonstrate that EFT is not sensitive to . In particular, fitting the pose prior in methods such as SMPLify may nullify the advantage of using a better regressor for initializing the pose. Instead, using better pose regressors in EFT generally result in better performance (see table 5 and supp. material). In our experiments, we employ EFT by setting to produce our pseudo annotations, using instead early stopping for regularization, starting from (sec. 6.6). In practice, we use less than 100 EFT iterations for all fits.
Note that, while eqs. 2 and 3 are optimized in a similar manner, the goals are very different: while eq. 2 is averaged over the training set and minimized to learn the parameters of the model, eq. 3 is optimized on a single example and only to find a better fit of the model to it. After an individual update is obtained, is discarded.
Implementation details. For the network , we use HMR pre-trained using SPIN . For robustness, we change the perspective projection model used in SPIN back to the weak-perspective projection of HMR and fine-tune the SPIN network to work correctly with this model (this step does not noticeably affect the performance of the model). The input RGB image is a crop, 224 224, around the 2D keypoint annotations. eq. 3 is optimized using Adam with the default PyTorch parameters and a learning rateThis learning rate is chosen as a smaller value than the learning rate used for training 3D pose regressor. of , which is sufficient to cover both good and bad initializations. We switch off batch normalization and dropout. We iterate until the average 2D re-projection loss is less than 3 pixels, or up to a maximum of 50 iterations (100 for OCHuman as the initial regressed pose tend to contain larger errors). Although the input 2D keypoints are manually annotated, they still contain non-negligible errors. In particular, the heights of hips and the ankles are not consistently annotated, causing foreshortening (examples are shown in supp. material). To compensate, the hip locations are ignored while optimizing eq. 3 and we add a loss term to match the orientation of the lower leg, encouraging the reconstruction of the orientation of the vector connecting the knee to the ankle. We use in eq. 3 unless otherwise mentioned. Finally, we discard samples as unreliable by checking the maximum value of the SMPL shape parameters and of the 2D keypoint loss, rejecting if these are larger than 5 and 0.01 respectively.
The EFT Training and Validation Datasets
The main application of EFT is augmenting existing 2D human pose datasets with 3D pseudo-ground truth annotations for downstream model training and validation (benchmarking). We experiment with augmenting the COCO , MPII , LSPet , PoseTrack , and OCHuman datasets. Most of these datasets come with a Train and Val split, which we use.We make a random split for LSPet, and for OCHuman we use Val set as training and Test set as testing data. We discard samples that fail the EFT sanity check discussed in sec. 3. For COCO-Train and COCO-Val, we only retain samples with at least 6 keypoints annotations. We also consider a subset of COCO-Train, named “COCO-Part”, used by , which more conservatively only uses instances with 12 keypoint annotations for which all limbs are present. For COCO-Val, LSPet and OCHuman we further carry out a manual filtering step, described below. We use the notation to denote an EFT-augmented version of a dataset and thus obtain the , , , , , and datasets, summarized in table 1. All datasets are publicly available in our webpage.
Human Study and Validation. We conduct a human study and validation on Amazon Mechanical Turk (AMT) to assess the quality of the EFT fits. First, we use A/B testing to compare EFT and SMPLify fits. To this end, we show 500 randomly-chosen images from the MPII, COCO and LSPet datasets to human annotators in AMT and ask them whether they prefer the EFT or the SMPLify reconstruction. Each sample is evaluated by three different annotators, showing the input image and two views of the 3D reconstructions, from the same viewpoint as the image and from the side. Examples are shown in supp. material. Our method was preferred 61.8% of the times with a majority of at least 2 votes out of 3, and obtained 59.6% favorable votes by considering the 1500 votes independently. We found that in many cases the perceptual difference between SMPLify and EFT is subtle, but SMPLify suffers more from bad initialization due to occlusions and challenging body poses.
We further conduct a full human assessment of COCO-Val, LSPet and OCHuman to validate the annotations before using them for benchmarking purposes. To this end, we show each sample to three annotators at random, asking whether the estimated 3D annotation accurately describes the pose of the target subject in the scene or not. Conservatively, we accept a sample only if all three annotators agree to accept it. The final number of samples and rejection rates are shown in table 1. The rejection rates () are relatively high due to the strict selection criteria (the rate would be only if were to accept samples with only two ‘accept’ votes).
Truncated 3DPW Dataset. In order to assess the robustness of algorithms to people only partially visible due to view truncation (a very common case in applications as similarly addressed in ), we also propose a new protocol for the 3DPW dataset using a pre-defined set of aggressive image crops (see fig. 2). As shown in fig. 2 (d), crops are generated by computing bounding boxes in between the full bounding box (level 7) of a person and a much smaller bounding box containing only his/her face (level 0) whose 2D locations is computed by projecting the SMPL fits on images. For each human instance, we generate eight intermediate bounding boxes in this manner (levels 0 to 7).
Training Pose Regressors with EFT datasets
Our EFT datasets allow us to train 3D pose regressors from scratch using a simple and clean pipeline. Specifically, we train the HMR model (without the discriminator) using the same hyper parameters as used in SPIN . A key advantage in using the pseudo-GT annotations is that the training pipeline becomes straightforward, providing opportunities to apply additional techniques to improve performance. We explore two such techniques: applying extreme crop augmentation and inputting auxiliary inputs.
Augmentation by extreme cropping. A shortcoming of previous 3D pose estimation methods is that it is assumed that most of the human body is visible in the input image . However, humans captured in real-world videos are often truncated, so that only the upper body or even just the face is visible, dramatically increasing the ambiguity of pose reconstruction. In order to train models that are more robust to truncation, we propose to augment the training data with extreme cropping. While the same challenge has been recently addressed by , our approach is more straightforward because we already have full 3D annotations — we only need to randomly crop the training images. In practice, we generate random crops in the same way as described in sec. 4 for the Truncated 3DPW dataset. During the training, we trigger crop augmentation with 30% chance, randomly choosing a truncated bounding box among the pre-computed ones shown in fig. 2 (d).
Auxiliary Input for 3D Pose Regressor. Recent methods demonstrate that other types of input encoding such as DensePose or body part segmentation can improve 3D pose regressors . Training with such auxiliary inputs is also straightforward in our pipeline. To show this, we train a pose regressor by concatenating the standard RGB input with an additional input encoding. We test color-coded segmentation maps and DensePose map, by pre-computing these representations for all images in the datasets using Detectron2 , as shown in fig. 2 (b,c). To accept 6 channel inputs (RGB concatenated with a color-coded auxiliary input), we modify the first layer of HMR (i.e. the first layer of ResNet50) by duplicating the initial weights.
Results
Our key result is showing that the quality of the EFT-generated pseudo GT is sufficient to train from scratch state-of-the-art 3D pose regressors, outperforming previous methods (sec. 6.2 and 6.3) in the in-the-wild benchmark. We also use EFT to generate benchmark representative of difficult scenarios of practical importance and run a large comparison of state-of-the-art regressors using it (sec. 6.4 and 6.4). Finally, we use EFT as a post-processing step with off-the-shelf regressors (sec. 6.5), and provide additional analysis in (sec. 6.6).
In addition to the EFT datasets of sec. 4, we also test the datasets with ground-truth 3D pose annotations, including H36M , MPI-INF-3DHP , collected in laboratory conditions. For these datasets, we use the existing SMPL fittings on H36M and MPI-INF-3DHP . We use the 3DPW test set as the major target benchmark that is captured outdoor and comes with 3D ground truth obtained by using IMUs and cameras.
2 EFT Datasets for Learning Models
In sec. 4 we have assessed the EFT-augmented datasets via a human study. We now assess them based on how well they work when used to train pose regressors from scratchResNet50 of HMR is initialized with ImageNet pretrained weights.. To this end, we follow the procedure described in sec. 5 and train the HMR model using the EFT datasets in isolation as well as in combination with other 3D datasets such as H36M. In all cases, we use as single training loss the prediction error against 3D annotations (actual or pseudo), with a major simplification compared to approaches that mix, and thus need to balance, 2D and 3D supervision.
The results are summarized in table 2. The table (bottom) evaluates the regressor trained from scratch using straight 3D supervision on standard 3D datasets (H36M, MPI-INF-3DHP, 3DPW), the EFT-lifted datasets , , and , as well as various combinations. Following , performance is measured in terms of reconstruction errors (PA-MPJPE) in after rigid alignment on two public benchmarks: 3DPW and H3.6MSee the supp. material for the result on the MPI-INF-3DHP benchmark.. The models trained with indoor 3D pose datasets (H36M and MPI-INF-3DHP) perform poorly on the 3DPW dataset, which is collected outdoors. By comparison, training exclusively on the EFT datasets performs much better. Notably, the model trained only on is already comparable to the previous state-of-the art method, SPIN , on this benchmark. Combining multiple training datasets improves performance further; the model trained with outperforms SPIN achieving (57.5 ) and the model trained with + is comparable to VIBE that is the previous best method using a video input (using temporal cues). Combining both EFT datasets and 3D datasets shows even better results with the result (54.7 ) achieved by the + H36M + MPI-INF-3DHP combination, and the best one (51.6 ) by including the 3DPW training data.
We also compare EFT to SMPLify for generating the pseudo-annotations. In we use the SMPLify output to post-process the output of the fully trained SPIN model, and in we use the fitting outputs that SPIN generates during learning , which are publicly available. We include the result of ExPose in table 2 that similarly built a pseudo-gt dataset using SMPLify-X with a manual filtering process by human annotators. As shown, the models trained on the EFT datasets perform better than the ones trained on the SMPLify-based ones, suggesting that the quality of the pseudo GT generated by EFT is better.
This result shows the significance of the indoor-outdoor domain gap and highlights the importance of developing in-the-wild 3D datasets, as these are usually closer to applications, thus motivating our approach. As it might be expected, networks trained exclusively on indoor dataset with GT 3D annotations (H36M, MPI-INF-3DHP, and 3DPW-train) show poor performance on 3DPW. Similarly, the networks exclusively trained on the the EFT-based datasets, which are ‘in the wild’, are not as good when tested on H36M. However, the error decreases markedly once the networks are also trained using the H36M data (H36M, +H36M, and + H36M +MPI-INF-3DHP).
3 Learning Models with Auxiliary Inputs
Next, we test the performance of models trained with RGB and auxiliary inputs. In table 2 bottom, w/ DensePose and w/ DensePose show results obtained by using the concatenation of RGB and DensePose IUV map as input. These additional inputs improve the accuracy on in-the-wild scenes, previously shown by . In particular, the models trained with only outperform the state-of-the-art video-based regressor VIBE (56.1 vs 56.5 ). In the sup. mat. we show that, instead, the segmentation encodings do not noticeably improve performance.
4 New 3D Human Pose Benchmarks
We evaluate the performance of various models on our new benchmark datasets with pseudo GT, OCHuman and LSPet . These datasets have challenging body poses, camera viewpoints, and occlusions. We found that most models trained on other datasets are struggling in these benchmarks, showing more than 100 errors, as shown in table 3. The models trained with pseudo annotations on similar data (training sets of LSPet and OCHuman) show better performance (less than 100 ).
Testing with Truncated Input on 3DPW.
We use the protocol defined in sec. 4 on 3DPW to assess performance on truncated body inputs. Among 8 different bounding box levels, we consider levels 1,2,4 that are shown in yellow in fig. 2 (Right). We use PA-MPJPE as in the original benchmark test. The result is shown in table 4, where are the datasets where we apply crop augmentations. For comparison, we include a recent work of Rockwell et al. that is trained to handle similar partial view scenes.
While HMR, SPIN, VIBE, and our network trained with without crop augmentation work poorly, we found that our models trained with show better performance even without crop augmentation, even outperforming the work of Rockwell et al. . This is because already includes many such samples with severe occlusions. Note that includes all samples with 6 or more valid 2D keypoint annotations. Our models trained with crop augmentation show much better performance even for scenes with extreme truncation (crop level 1 and 2). Note that crop augmentation does not degrade the performance of our models in the original 3DPW benchmark, as shown in table 2.
5 EFT as A Post-processing Method
EFT can also be be used as a drop-in post-processing step on top of any given pose regressor to improve its test-time performance. We compare EFT post-processing against performing the same refinement using a traditional fitting method such as SMPLify, starting from the same regressor for initialization. In this experiment, we use the ‘gold-standard setting’ for each algorithm: for EFT, this is the same setup as in the other experiments, and for SMPLify, we use the best hyper-parameter setting determined in SPIN. See supp. mat. for details. In the sup. mat. we also report results when the settings are exactly the same, except for the fitting method, and arrive to analogous conclusions.
In table 5 we carry out a quantitative evaluation on the 3DPW dataset. For initialization, we use various regressors pre-trained on a number of different datasets, providing different initialization states. We then use EFT and SMPLify post-processing to fit the same set of 2D keypoint annotations, obtained automatically by means of the OpenPose detector . For SMPLify, we performed an ablation study by turning off multi-stage trick (M) and prior terms (P), compared to the standard setting (M+D+P) used in SPIN and the original SMPLify literature . Similarly, we tested EFT performance without shape regularizer (i.e., ). The key observation is that the performance of SMPLify post-processing depends on the initialization quality, the use of their multi-stage trick, and balancing between data and prior terms. Specifically, if the initialization is already good, as in the last row of table 5, SMPLify degrades the performance, particularly when the pose prior is used. This illustrates the difficulty in balancing data and prior terms, which may be difficult to do in practice. In contrast, EFT does not suffer from such issues, improving the accuracy in all cases, regardless of the quality of the initialization. This is also true when the shape regularizer is removed entirely.
6 Overfitting Analysis of EFT
As we do not use any explicit pose prior terms in our EFT optimization ( in eq. 3), by overfitting to a single sample EFT could ‘break’ the implicit pose prior originally captured in the regressor, potentially causing the results drifting away from the true poses in the later EFT iterations. To test this, we can check the generalization capabilities of the overfitted pose regressor from EFT procedure on each single sample, by evaluating the accuracy of the overfitted regressor on the entirety of the 3DPW test set. We repeat this test for 500 different samples by using 20 and 100 EFT iterations respectively and report the results in fig. 3. The performance of SPIN (blue line) and HMR (cyan line) are also shown for comparison. As can be noted, the overfitted regressors still perform very well overall, suggesting that the network retains its good properties despite EFT. In particular, the performance is at least as good as the HMR baseline , and occasionally the overall performance improves after overfitting a single sample. The effect is different for different samples — via inspection, we found that samples that are more likely to ‘disrupt’ the regressor contain significant occlusions or annotation errors. Note that the overfitted network is discarded after applying EFT on a single example — this analysis is only meant to illustrate the effect of fine-tuning.
7 Qualitative Evaluation on Internet Videos
We demonstrate the performance of our 3D pose regressor models trained using our EFT datasets on various challenging real-wold Internet videos, containing cropping, blur, fast motion, multiple people, and other challenging effects. Example results are shown in the supp. videos.
Discussion
We provide, via EFT, large-scale and high-quality pseudo-GT 3D pose annotations that are sufficient to train state-of-the art regressors. We expect out ‘EFT datasets’ to be of particular interest to the research community, stripping the training of 3D pose regressors from complicated preprocessing or balancing techniques. Our 3D annotations on the popular 2D datasets such as COCO can also provide more opportunities to relate 3D human pose estimation with other computer vision tasks such as object detection and human-object interaction.
References
Supplementary Material
In this supplementary material, we provide additional experiments and other details. We describe implementations of the corresponding “gold-standard” settings of SMPLify and EFT in sec. A, and then further compare these two algorithms using an identical set of hyper parameters in sec. B. More extensive quantitative evaluation on public benchmarks is provided in sec. C. We also include other details that were not provided in main manuscript due to space constraints.
A Implementation Details of Gold-standard Setting (main paper sec. 6.5)
Below we provide more details on implementations of “gold-standard settings” for each algorithm benchmarked in Sec. 6.5 of the main manuscript.
For SMPLify, we use the setting from SPIN See the public SMPLify code from SPIN for further details : https://github.com/nkolot/SPIN/blob/master/smplify/smplify.py with minor modifications. For the data term, the 2D reprojection error is computed in pixel space of 224224 input images, with Geman-McClure robust function (=100). Prior terms include pose prior with GMM, angle prior to penalize unnatural bending, and shape prior, with respective weights of 4.78, 15.2, and 5.0. We use Adam optimizer with learning rate of . The SMPL parameters are first initialized with predictions of a pre-trained 3D pose regressor, same as in EFT. Then a multi-stage approach is employed by optimizing camera translation and body orientation first, and then all parts later.
We made a few modifications to the original code: (1) we use weak perspective projection; (2) we do not use openpose input (see sec. K for justification); and (3) we stop once the average 2d reprojection error is less than 3 pixels (max 50 iterations), which is the same setting as in EFT.
EFT.
EFT uses the standard L2 loss for 2d keypoints data term (no robust loss function). Note that, as EFT follows the standard neural network training procedure, the reprojection loss is computed in a normalized image space (-1 to 1). As mentioned in the main manuscript, the hip locations are ignored when computing reprojection errors, and we add a loss term to match the orientation of lower legs (with weight of ), encouraging the reconstruction of the orientation of the vector connecting the knee to the ankle to be similar to the one from 2D annotations. The key motivation underpinning the use of leg orientation loss is based on the observation that ankle annotations are particularly noisy, causing erroneous foreshortening during fitting process (see sec. H).
We use Adam optimizer with learning rate (instead of used for training the 3D pose regressor (also used in SPIN)), see more discussion below. We stop iterating once the average 2d error 3 pixels (max 50 iterations), which corresponds the same setting as SMPLify.
Note that, EFT computes the 2D reprojection error in a normalized space instead of the image pixel space as in the SMPLify. However, this does not make difference because Adam is re-scaling invariant. For example, we can compute the 2D reprojection error in either space after re-scaling 2D distance and adjusting the weights for the other terms accordingly, for which Adam provides the same updates.
Justification of Different Learning Rates.
The learning rate of EFT should not be directly compared to the one in SMPLify . In SMPLify, the optimizer directly updates the 85-dimensional vector of SMPL parameters, while in EFT it optimizes the weights of the neural network (27 millions) that produces SMPL parameters via a series of non-linear computations (including Batch Normalization). In fact, no same learning rate can be used for both optimizations — we confirmed that with learning rate of EFT changes SMPL parameters too much in each iteration, while with learning of SMPLify almost does not change SMPL parameters. For balancing, we instead observed reprojection error changes in the first few iterations, and found that the current learning rate settings produce similar step size in updating SMPL parameters, which is also qualitatively confirmed by 3D visualizions.
Discussion.
The performance of SMPLify is largely affected by hyper parameters such as weights between data term and prior terms, and also by the initialization states that motivates the use of a multi-stage technique. Given the excellent achievement of SPIN in the use of SMPLify on the same application with ours, we treat this setting as a “gold-starndard”. The results shown in sec. 6.5 of our main manuscript should be understood as a comparison with their own best settings between EFT and SMPLify. In the following section, we perform an additional comparison in the same setting.
B Comparison SMPLify and EFT in the Same Data Term Setting (main paper sec. 6.5)
In this section, we perform further comparison between EFT and SMPLify with the same set of hyper parameters. Concretely, the exact same data term is used for both methods, and an ablation study is performed by adding prior terms for each method with varying weights.
To employ the same data term for EFT and SMPLify, we make the following modifications. We remove Geman-McClure robust function in SMPLify and use standard L2 loss, computing the reprojection error in the normalized space (-1 to 1) as in EFT. For EFT, we remove the leg orientation data term. The exact same body keypoints, ignoring hips, are used in computing 2D reprojection loss for both methods. We keep the same learning rates as in previous experiments, for SMPLify and EFT .
After performing a comparison with data-term only ( in table 6), we further investigate by adding prior terms of each method (). For SMPLify, we only consider the pose prior term, omitting the angle prior term and the shape prior term. The original weight for the pose prior term (4.78) of SMPLify is also adjusted to compensate the scaling change of the data term (i.e., in table 6). With varying , we analyze the performance changes of SMPLify.
For EFT, we add the neural network weight regularizer term (the second term in eq. (3) of our main manuscript), by varying its weight .
Results.
As in sec. 6.5 and table 5 of our main manuscript, we carry out a quantitative evaluation by using the 3DPW dataset. As in the previous experiment, we use various regressors to provide different initialization states, then use EFT and SMPLify for post-processing. The results are shown in table 6, showing analogous conclusions to the result of our main paper (sec. 6.5 and table 5).
When no prior term is used (i.e., ), SMPLify produces mixed results, depending on the quality of initial states. For example, given a relatively accurate initialization (the last row), the post-processing makes the result slightly worse, and with intermediate models (the second and the third) it improves the final accuracy. With very poor initializations (the first low), the result becomes worse. Adding pose prior on SMPLify also shows mixed performances. It is advantageous to have the pose prior when initializations are poor. However, the use of prior degrades the accuracy if the initialization from regressor is accurate enough. Note that, the post-processing performance changes quite significantly with small amount of weight changes (e.g., 51.72 to 57.64 in the last row, and more changes in other models). This particular result indicates that the weight needs to be adjusted carefully to obtain the best performance in each scenario. To explore further, we investigate the performance with an extremely strong with fixed 100 iterations without early stopping (shown in ), and, as expected, the results converge to a certain accuracy (around 86), presumably the mean pose defined by the pose prior term.
Now, we see the results from EFT. When we use the data-term only , the same setting we use throughout our paper, it improves accuracy in all cases. Furthermore, by performing experiments with varying , we confirmed that the results of EFT are not sensitive to the change of . Notably, adding prior does not degenerate the initial accuracy from regressor in all cases. This is still true even though an extremely large is used with fixed 100 iterations (similar setting with SMPLify as shown in ), where the regularizer term almost overshadows the data term. In this case, as expected, the results fall back to the ‘best guess’ of the pre-trained regressor network. The use of prior term sometimes helps to reduce the accuracy further, but the advantage is marginal. It shows less than 1 difference in most cases, providing a justification for our default setting with no pose prior term.
Discussion
As demonstrated above, traditional fitting methods such as SMPLify suffer from the balancing issue between the data term and the “view-agnostic” 3D pose prior term. In contrast, EFT does not suffer from such issues, and we demonstrate that it can be applicable even without any prior terms. EFT effectively leverages better pose regressors, showing better performance in all cases.
Note that this experiment is intended to explore the major difference between EFT and SMPLify, rather than finding the best settings for each method. The best SMPLify setting on 3DPW is not necessarily the best in other cases, and, often, it is difficult to determine the optimal set of hyper parameters. For example, we found that the use of prior terms or multi-stage optimization sometimes degrade the accuracy, but they can be still advantageous on other cases with poor initialization, such as OCHuman . In general, the gold standard setting proposed in SMPLify and SPIN is recommended for practitioners. Importantly, as demonstrated, EFT can provide a better option, showing better performance without complicated parameter tuning.
C Further Evaluation on Standard Benchmarks (main paper sec. 6.2)
The full evaluation results on the standard benchmarks including MPI-INF-3DHP are shown in table 10. All results are measured in terms of reconstruction errors (PA-MPJPE) in after rigid alignment, following . We also report Per-Vertex Error in table 7. Here, we highlight several noticeable results.
The model trained with our EFT dataset is also competitive on the MPI-INF-3DHP dataset. The models trained by only, outperform HMR . Combining 3D datasets with our EFT dataset improves the performance further.
H36M Protocol-1.
We also report quantitative results on H36m protocol 1. In this evaluation, we use all available camera views in H36M testing set. The overall errors are slightly higher than protocol-2 (frontal camera only), but the general tendency is similar to the results of H36M protocol-2.
A Baseline Model Trained with 2D Loss Only.
As a naïve baseline method, we train a 3D pose regressor by just using 2D annotations (that is, with only 2D keypoint loss) without including other 3D losses from real or pseudo-ground truth 3D datasets. The results are shown as and in table 10. As expected, the model cannot estimate accurate 3D pose, showing very poor performance. However, interestingly, we found the 2D projection of the estimated 3D keypoints to still be quite accurately aligned to the target individual’s 2D joints.
PVE Error:
We report Per-Vertex Error in Table 7 between the ground truth mesh vertices and the vertices from predictions. While PVE error would be affected by shape parameters which is not the major focus of our work, its tendency is aligned to the performance quantified by MPJPE errors, showing lower PVE errors with the model with lower MPJPE errors.
D Training Models with Auxiliary Inputs (main paper sec. 6.3)
In addition to the results shown in section 6.3, we explore the performance of the models trained with other auxiliary inputs. We analyze the results of the models trained with two different segmentation maps: Panoptic Segmentation that has both things and stuff classes and PointRend that has only thing classes but with more accurate contours. The results are shown in table 9, along with the original results by RGB only and the results with DensePose . As can be seen, the color-coded segmentation encodings do not noticeably improve performance, compared to the results with DensePose.
E Overfitting Analysis of EFT: Examples (main paper sec. 6.6)
As shown in fig. 3 of our main manuscript, overfitting the network to individual samples with EFT usually has a small effect on the overall regression performance, suggesting that the network retains its good properties despite fine-tuning on exemplars. The effect is different for different samples, and we show the examples with a strong effect in fig. 7 that contain significant occlusions or annotation errors.
F More Qualitative Comparisons Between EFT vs. SMPLIfy (main paper sec. 4)
As addressed in our main paper, our EFT method was preferred 61.8% of the times with a majority of at least 2 votes out of 3 in Amazon Mechanical Turk (AMT) study. Here, we visualize the examples with 0 vote and 3 votes, as shown in fig. 9 and 10. Note that in our AMT study, we show the meshes over white backgrounds to reduce the bias cased by 2D localization quality.
There are 47 / 500 samples where SMPLify outputs are favored by all three annotators, and examples are shown in fig. 9. Surprisingly the difference is quite minimal with no obvious failure patterns.
When annotators prefer EFT?
There were 132 / 500 samples where our EFT outputs were favored by all three annotators. Examples are shown in fig. 10. In most cases, EFT produces more convincing 3D poses leveraging the learned pose prior conditioned on the target raw image, while SMPLify tends to produce outputs similar to its prior poses (e.g., knees tend to be bent).
G Computation Time
We compare computation time between our EFT and SMPLify processes. The computation times are computed for their gold-standard settings by using a single GeForce RTX 2080 GPU.
A single EFT iteration takes about 0.04 sec, and a whole EFT process for a sample data (including reloading the pre-trained network model) takes about 0.82 sec with 20 iterations.
SMPLify.
Camera optimization takes about 0.01 sec per iteration, and, thus, it takes 0.5 sec per sample (with 50 iterations). Optimizing body pose takes about 0.02 sec/iteration, and it takes 1 sec with 50 iterations. Thus whole SMPLify process takes about 1.5 sec per sample with 50 iterations.
H Noisy 2D keypoint issue in GT data (Implementation details in sec. 3)
The keypoints of hips and ankles are more difficult to correctly localize by annotators than other keypoints due to loose clothing and occlusions. This can be also observed by checking the standard deviation of annotated 2D keypoints in COCO dataset (called OKS sigma values) https://github.com/cocodataset/cocoapi/blob/8c9bcc3cf640524c4c20a9c40e89cb6a2f2fa0e9/PythonAPI/pycocotools/cocoeval.py, where hips (1.07) and ankles (0.89) have higher values than shoulders (0.79) and wrists (0.62). We show examples from COCO dataset in fig. 4, where in the annotations the lower leg lengths are shorter than it should be. We empirically found that such noisy 2D localization causes artifacts in 3D, motivating us to use 2d orientations for the lower leg parts rather than GT locations.
I Statistics of our EFT datasets
Since we now have high quality pseudo 3D annotations, we can check and compare the 3D pose variations among datasets. In fig. 5, we compare the variations of 3D pose parameters of the pseudo GT data produced from three in-the-wild datasets, COCO (red), MPII (green), and LSPet (blue). The top-left figure shows the standard deviation of pose parameters (we choose the X-axis in angle-axis representation) for each body joint, including necks, hips, shoulders, and elbows. We can see that the LSPet dataset has the most diverse pose variations compared to MPII and COCO. In the later figures, we visualize the histograms of the pose parameters of a particular joint (Left Hip), to see the pose variations and distributions of all samples. As shown, the 3D poses of LSPet have higher portion of samples that have large angle axis values (less than -1.0), showing that it has more gymnastic poses. In contrast, the majority of COCO dataset has the values around zeros, we is closer to the rest standing pose.
J Sampling Ratios across Datasets During Training
Following the approach of SPIN , we use fixed sampling ratio across datasets while training the 3D pose regressor. The ratios used in our experiments are shown in table 8. denotes our EFT dataset. If multiple EFT datasets are used (e.g., COCO and MPII), we sample from them uniformly.
K Analysis on the Failures of SMPLify Fittings in SPIN
We found erroneous cases in the publicly available SMPL fitting outputs from SPIN that are related to its SMPLify implementation. In SPIN, both 2D keypoint annotations and OpenPose estimations are used for SMPLify, assuming that the OpenPose estimations have better localization accuracy if OpenPose is applied to the training data. However, we found many cases where the OpenPose outputs from different individuals are mistakenly associated to the target person, as shown in fig. 6. Due to this reason, we exclude the OpenPose estimations and only use the GT annotations for our SMPLify process.
L Potential Applications for EFT Datasets
Our EFT datasets provide high-quality pseudo-GT 3D pose data on in-the-wild images. This can potentially open up several new research opportunities.
We can automatically render DensePose annotations from our pseudo GT data. Examples are shown in fig. 8. As a potential research direction, these automatically generated densepose annotations can be used for training.
Nearest Neighbor Search.
Our pseudo-GT datasets allow us to further analyse diverse 3D human poses in the context of outdoor environments. For example, by using a simple nearest neighbor we can find similar 3D human poses in the COCO dataset. The examples are shown in fig. 11.