PaMIR: Parametric Model-Conditioned Implicit Representation for Image-based Human Reconstruction
Zerong Zheng, Tao Yu, Yebin Liu, Qionghai Dai
Introduction
Image-based parsing of human bodies is a popular topic in computer vision and computer graphics. Among all the tasks of image-based human parsing, recovering 3D humans from a single RGB image attracts more and more interests given its wide applications in VR/AR content creation, image and video editing, telepresence and virtual dressing. However, as an ill-posed problem, recovering 3D humans from a single RGB image is very challenging due to the lack of depth information and the variations of body shapes, poses, clothing types and lighting conditions.
Benefiting from the huge progress of deep learning techniques, recent studies have tried to address these challenges using learning-based methods. According to their 3D representations, these methods can be roughly classified into two categories: parametric methods and non-parametric methods. Parametric methods like HMD and Tex2Shape utilize a statistical body template model (e.g., SMPL) as the geometrical prior, and learn to deform the template according to the silhouette and the shading information. However, the low dimensional parametric model limits the performance when handling different clothing types such as long dresses and skirts. Non-parametric methods use various free-form representations, including voxel grids, silhouettes, implicit fields or depth maps, to represent 3D human models. Although being able to describe arbitrary clothes, current non-parametric methods still suffer from challenging poses and self-occlusions due to the lack of semantic information and depth ambiguities.
Recently, DeepHuman pioneered in conditioning the non-parametric volumetric representation on the parametric SMPL model, and demonstrates robust reconstruction for challenging poses and diverse garments. However, it struggles to recover high-quality geometric details due to the resolution limitation of the regular occupancy volume. More importantly, conditioning the non-parametric representation on SMPL models raises a crucial problem: how can we obtain accurate SMPL model estimation given a single image at testing time? In DeepHuman, the SMPL models at testing time are estimated using learning-based methods and/or optimization-based methods. Although the recent years have witness a renaissance in the context of single-image body shape estimation, we observe that there is still an accuracy gap between the training and testing stages: during training, we can obtain well-aligned SMPL models for the training data by fitting the SMPL template to the high quality scans, while for testing, the SMPL estimations provided by the state-of-the-art single-image methods still cannot align with the keypoints or the silhouettes as perfectly as the SMPL models in the training dataset.
As a result, the neural network trained with the ground-truth image-SMPL pairs cannot generalize well to the testing images which have no ground-truth SMPL annotations. A simple solution is to replace the ground-truth SMPL models with the predicted ones while still using the ground-truth surface scans for training supervision, preventing the network from heavy dependence on accurate SMPL annotations. However, due to depth ambiguities, it is impossible to guarantee good alignments between the predicted SMPL models and the ground-truth scans along the depth axis. Consequently, by doing so we are forcing the networks to make “correct” geometry inference given a “wrong” SMPL reference along the depth axis, which will generate even worse results.
To achieve high-quality geometric detail reconstruction while maintaining robustness to challenging pose and clothing styles, in this paper, we propose Parametric Model-Conditioned Implicit Representation, dubbed PaMIR, to incorporate the parametric SMPL model and the free-form implicit surface function into a unified learning and optimization framework (Fig.2). The implicit surface function overcomes the resolution limit of volumetric representation, and enables detailed surface reconstruction capability. To fill the accuracy gap between the ground truth training data and the inaccurate testing input, we further propose a new depth-ambiguity-aware training loss and a body reference optimization module in the PaMIR-based framework, and achieve surface detail reconstruction even under imperfect body reference initialization. Specifically, the three technical contributions are summarized as follow:
PaMIR. Our PaMIR representation has the ability to condition the implicit field on the SMPL prediction, which is realized by a novel network architecture converting an image feature map and the corresponding SMPL feature volume into an implicit surface representation. The SMPL feature volume is directly encoded from the SMPL prediction provided by GraphCMR, and serves as a soft constraint for handling challenging poses and/or self-occlusions.
Training Losses. We use the predicted SMPL models at both the training and the testing stages. In order to make the network to be more robust against the inaccurate body reference along the depth direction, instead of using a traditional reconstruction loss, we propose a depth-ambiguity-aware reconstruction loss that adaptively adjusts the reconstruction supervision based on the current SMPL estimation while respecting 2D image observations.
Body Reference Optimization. We also propose a body reference optimization method for testing process to further refine the body reference. Our optimization method directly utilizes the network itself as well as its outputs to construct an efficient energy function for body fitting. The optimization further alleviates the error caused by the inaccurate initial SMPL estimation, and consequently closing the accuracy gap of SMPL annotation between training and testing.
Overall, in our framework, the underlying body template and the outer surface are mutually beneficial: on one hand, benefiting from the proposed depth-ambiguity-aware reconstruction loss, even the imperfect body template can be used to provide strong semantic information for implicit surface reconstruction; on the other hand, the deep implicit function of the outer surface is also used to optimize the underlying body template by minimizing the body fitting error directly.
Furthermore, thanks to the usage of the common underlying SMPL model as body reference, our PaMIR representation implicitly builds a correspondence relationship across different models. As a result, our method can be easily extended to multi-image setups like multi-view inputs or video inputs without the requirement for calibration and synchronization. Experiments show that our method is able to reconstruct high-quality human models with various body shapes, poses and clothes, and outperforms state-of-the-art methods in terms of accuracy, robustness and generalization capability.
Related Work
Human Reconstruction from multi-view images. Previous studies focused on using multi-view images for human model reconstruction . Shape cues like silhouette, stereo and shading have been integrated to improve the reconstruction performance. State-of-the-art multi-view real-time systems include . Extremely high-quality reconstruction results have also been demonstrated with tens or even hundreds of cameras . To capture detailed motions of multiple interacting characters, more than six hundred cameras have been used to overcome the self-occlusion challenges.
In order to reduce the difficulty of system setup, human model reconstruction from sparse camera views has recently been investigated by using CNNs for learning silhouette cues and stereo cues . These systems require about 4 camera views for a coarse-level surface detail capture. Note also that although temporal deformation systems using lightweight camera setups have been developed for dynamic human model reconstruction using skeleton tracking () or human mesh template deformation , these systems require a pre-scanned subject-specific 3D template for deformation optimization.
Image-based Parametric Body Estimation. Dense 3D parsing from a single image has attracted substantial interest recently because of the emergence of human statistical models like SCAPE and SMPL. For example, by fitting the SMPL model to the 2D keypoint detections and other dense shape cues, the shape and pose parameters can be automatically obtained from a single image. Instead of optimizing mesh and skeleton parameters, recent approaches proposed to train deep neural networks that directly regress the 3D shape and pose parameters from a single image . The estimation accuracy of these methods are further improved by performing fitting optimization after network inference, introducing model optimization into the training loop, incorporating adversarial prior in temporal domain, or combining global and local features to produce fine-grained human body poses. However, the parametric body models can only capture the shape and pose of a minimally clothed body, thus lack the ability to represent a 3D human model with more general clothing layers.
Non-parametric Human Reconstruction from a Single Image. Regarding single-image human model reconstruction using non-parametric models, recent studies have adopted techniques based on silhouette estimation, template-based deformation, depth estimation and volumetric reconstruction . Although they have achieved promising results, typical limitations still exist when using a single 3D representation: silhouette-based methods like is suffer from lack of details and view inconsistency, template-based deformation methods are unable to handle loose clothes, depth-based methods cannot handle self-occlusions naturally, and volumetric methods cannot recover high-frequency details due to their cubically growing memory consumption.
Implicit Representation. How to represent 3D surface is a core problem in 3D learning. Explicit representations like point clouds, voxel grids, triangular meshes have recently been explored to replicate the success of deep learning techniques on 2D images. However, the loss of structure information in point cloud, the memory requirement of regular voxel grid, and the fixed topology of meshes make these explicit representations unsuitable to represent arbitrary high-quality 3D surfaces in neural networks. Implicit surface representations, on the other hand, overcome these challenges and demonstrate the best flexibility and expressiveness for topology representation. It defines a surface as the level set of a function of occupancy probability or surface distance. Recent works have shown promising results on generative models for 3D object shape inference based on deep implicit representations.
Deep Implicit Representation for Human Reconstruction. The success of deep implicit representations in general object modeling has inspired research in 3D human reconstruction. For example, PIFu proposed to regress an deep implicit function using pixel-aligned image features and is able to reconstruct high-resolution results, which is then accelerated to real-time framerate in . PIFuHD extended PIFu to capture more local details by adding a fine-level feature extraction network. However, both PIFu and PIFuHD are prone to reconstruction artifacts in cases of challenging poses and self-occlusions. In contrast, our method achieve more robust performance under these challenging scenarios by introducing SMPL model as an additional semantic condition for the implicit representation. ARCH and IP-Net are two concurrent works that also use SMPL model for semantic reconstruction. In particular, ARCH proposed to regress animation-ready 3D avatars in a canonical pose (A-pose), but fails to generate accurate results, especially for loose clothes and human-object interaction (e.g., holding a camera with two hands as in Fig.8) because it is ambiguous to determine the position of the objects and accessories in A-pose. IP-Net jointly predict the outer 3D surface of a human model, the inner body surface and the body part labels, but is restricted to point cloud inputs. In contrast, our method can naturally handle human-object interaction and can be flexibly adopted in both single-image and multi-image setups.
Surface Representation
where is the condition variable that encodes the overall shape of a specific surface and can be custom designed in accordance of applications. Intuitively, predicts the continuous inside/outside probability field of a 3D model, in which iso-surface can be easily extracted
In PIFu, the authors combined the condition variable with the point coordinate and formulate a pixel-aligned implicit function as:
where represents the image feature map from the deep image encoder , the 2D projection of on the feature map , is the sampling function used to sample the value of at pixel using bilinear interpolation, and is the depth value of in the camera coordinate space. A weak perspective camera is assumed in both PIFu and this paper. Intuitively, given a pixel on the image, PIFu casts a ray through that pixel along the direction of z-axis and estimates the values of occupancy probability on that ray based on the local image feature of the given pixel. Thanks to the usage of pixel-level features as condition variable, PIFu can reconstruct fine-scale detailed surfaces that are well aligned to the input images.
PIFu demonstrates high-quality human digitization for fashion poses, like standing or walking. However, we argue that purely relying on 2D image feature is insufficient to handle severe occlusions and large pose variations for single-image 3D human recovery, especially when ideal high quality dataset (covering sufficient poses, shapes and cloth types) is inaccessible. Specifically, the reasons are two-folds. On one hand, under the scenarios of self-occlusions, there may be multiple peaks of occupancy probability along the z-rays, but it is hard to infer these types of changes consistently and accurately only based on the local pixel-level 2D features. On the other hand, without the awareness of the underlying body shape and pose, the reconstruction results are prone to seriously incomplete or asymmetric bodies, like breaking arms and unequal legs, in the cases of challenging poses.
2 Parametric body Model
Inspired by DeepHuman, we integrate the parametric body model, SMPL , to regularize the human reconstruction. SMPL is a function that maps pose and shape to a mesh of vertices:
where linear blend-skinning with skinning weights poses the T-pose template based on its skeleton joints . As shown in DeepHuman, a SMPL model fitted to the input image can serve as a strong geometry prior, and guarantee pose and shape aware human model reconstructions. However, estimating SMPL model given a single image is a fundamentally ill-posed problem. Recent studies have explored to learn from large-scale datasets to address this challenge. In this work we use GCMR, one of the state-of-the-art networks as the backbone for body shape inference. Note that GCMR can be replaced with other equivalent networks such as SPIN for single images and VIBE for videos, as we do not make any assumptions on how the initial body is estimated and our method is able to deal with inaccurate body initialization thanks to our body reference optimization (Sec.4.3).
3 PaMIR: Parametric Model-Conditioned Implicit Representation
To combine the strengths of parametric body models and non-parametric implicit field, we introduce Parametric Model-Conditioned Implicit Representation (PaMIR). Specifically, in PaMIR, we define the definition of in Eqn.(3) as:
where has the same definition as Eqn.(3), while is the feature volume and represents the voxel-aligned volumetric feature at sampled from . The feature volume is obtained by firstly converting the SMPL mesh into an occupancy volume through mesh voxelization and then encoding it through a 3D encoder , i.e., . Given the per-pixel feature vector of as well as its per-voxel feature vector , we learn an implicit function in Eqn.(2) that can classify whether is inside or outside the surface. Note that although the introduction of 3D dense feature volume may increase memory consumption, in practice, we found that a relatively small feature volume is enough for soft semantic regularization. Moreover, we omit in since already contains 3D coordinate information. The illustration of the proposed PaMIR representation and the corresponding network is shown in Fig.2.
In PaMIR, the free-form implict representation is regularized by the semantic features of the parametric model. As we can see in the following sections, introducing SMPL model as a body reference has several advantages:
Pose Generalization. For surface geometry reconstruction, the SMPL input can be regarded as an initial guess for the network, which helps to factor out pose changes of the subject and eliminate depth ambiguities, thus making the network mainly focusing on surface detail reconstruction. As a result, our method has more robust performance compared with PIFu especially under challenging poses.
Multi-modal Prediction. Unlike PIFu or Occupancy Network that only condition on the input images, our PaMIR representation also conditions on the underlying bodies. Thus, in contrast to reconstruct one model for each input image, our method can reconstruct various plausible models given different but plausible body poses (Fig.12).
Easy Extension to Multi-image Setups. With the common underlying SMPL model as a body reference, our PaMIR representation also implicitly builds a correspondence relationship across different models and images. As a result, our method can be easily extended to multi-image settings like multi-view input and video input without explicit calibration and synchronization.
Method
Our PaMIR-based 3D human reconstruction is implemented as a neural network. An overview of the network is illustrated in Fig.2. Given a single image as input, our method first feeds it into the GCMR network to estimate an initial SMPL model, which is then converted into an occupancy volume through voxelization. In the feature extraction step, the input image is encoded into a feature map by a 2D convolution network, while the occupancy volume are encoded into a feature volume by a 3D convolution network. For each point in the 3D space, its pixel-aligned image feature and voxel-aligned volume feature are sampled in the feature map and the feature volume, respectively. The two feature vectors are then concatenated and translated to an occupancy probability value by a feature-to-occupancy decoder as formulated in Eqn.(2) and (5).
To alleviate the reliance on accurate SMPL annotations, we use the training images, the SMPL models estimated by GCMR and the corresponding ground-truth meshes to train the other parts of the network. To deal with the depth inconsistency between the predicted SMPL models and the ground-truth scans, we carefully design a depth-ambiguity-aware reconstruction loss as elaborated in Sec.4.2. Moreover, for inference, we propose body reference optimization, and optimize the human body template as well as the implicit function in an iterative manner in Sec.4.3. To obtain the final reconstruction results, we densely sample the occupancy probability field over the 3D space and extract the iso-surface of the probability field at threshold 0.5 using the Marching Cube algorithm. Texture inference can be performed in a similar way to geometry inference (Sec.4.4). Our framework can be easily extended to multi-image setups, which is described in Sec.5.
2 Depth-ambiguity-aware Reconstruction Loss
To train our network, we sample 3D points in 3D space around the human model, infer their occupancy probabilities and construct a per-point reconstruction loss. The traditional reconstruction loss is defined as the mean square error between the predicted occupancy probability of the point samples and the ground-truth ones. However, we argue that with a single-image setup, depth ambiguity is inevitable for many body poses and hence it is unpractical to force the network output to be perfectly identical with the ground-truth. Furthermore, if we utilize the predicted SMPL models as reference of 3D shape and pose, using traditional reconstruction loss will lead to negative impact on the reconstruction accuracy because the predicted models may not be well aligned with the ground-truth along the z-axis. To deal with this issue, we propose a novel depth-ambiguity-aware reconstruction loss defined as:
where is the nearest SMPL vertex set of and , the corresponding blending weight, the weight normalizer, and are the -th vertex of the predicted SMPL model and the corresponding ground-truth SMPL model, respectively. The blending weight is defined according to the distance between the point sample and its neighboring SMPL vertex :
Intuitively, with the depth-ambiguity-aware loss, we are guiding the network to output plausible surface corresponding to the predicted SMPL model but not the exact ground-truth occupancy volume. In this way the network is able to learn to infer an implicit field registered with the predicted 3D pose. The illustration of depth-ambiguity-aware reconstruction loss is shown in Fig.3.
3 Body Reference Optimization
Although our training scheme already prevents our network from being heavily dependent on the accuracy of SMPL estimation, the inconsistency between the image observation and the SMPL estimation may still lead to reconstruction artifacts. Fortunately, the predictions of our method can be further refined at inference time. Specifically, we can improve the accuracy of SMPL estimation by minimizing:
where is the body fitting loss used to encourage the alignment of the predicted implicit function and SMPL model and is a regularization term penalizing the difference between and the initial prediction. The body fitting loss is defined as following:
where is the number of SMPL vertices, is the -th predicted SMPL vertex and is the penalty loss defined as:
where . In other words, applies more penalization on the vertices that fall outside of the predicted surface while allowing the body to shrink into the surface in case of loose clothes like skirts and dresses. The regularization term is defined as
where is the initial SMPL parameters estimated by GCMR. Note that other constraints like 2D keypoint detection results can also be added to further improve the fitting performance.
The insight behind the formulation of body fitting loss is that the predicted SMPL may not be perfectly aligned with the image observation and consequently, the output implicit function is a compromise between these two information. Thus, by minimizing the body fitting loss, we can eliminate the inconsistency between the SMPL prediction and the image observation. As a result, more consistent SMPL prediction as well as more accurate surface inference can be obtained. The idea of our body reference optimization is illustrated in Fig.4.
Although body model optimization has been proposed in SMPLify and HoloPose, our optimization scheme are substantially different from that in previous works. Existing methods fit body models to image observation such as 2D keypoints detection, 3D keypoints estimation and/or dense correspondences, while ours directly utilizes the human reconstruction network to fit SMPL model into the implicit surface via minimizing the loss in Eqn.(10). Our scheme guarantees that the image observation, the estimated body model and the reconstructed outer surface are aligned with each other, and eliminates the requirements of sparse/dense keypoint detection (although they can be used as additional constraints).
We visualize the optimization process in Fig.5. In this figure, the initial pose of the right arm is incorrect, which leads to reconstruction artifacts. However, as the optimization is carried out, the right arm gradually moves towards the right position. In the meantime, the reconstructed mesh becomes more and more plausible. To clarify, the reconstruction results in this experiment do not contradict the results in Fig.14: in this experiment, the right arm is almost invisible and consequently it is difficult for the network to infer the correct geometry with an inaccurate arm pose.
4 Texture Inference
Following the practice of Texture Field and PIFu, we also regard the surface texture as a vector function defined in the space near the surface, which can support texturing of shapes with arbitrary topology and self-occlusion. To perform texture inference, we make some simple modification to our network. Specifically, we define the output of the decoder in Fig.2 as an RGB vector field instead of a scalar field. The RGB value is the network prediction of the color of a specific point on the mesh surface, while the alpha channel is used to blend the predicted value with the observed one:
where is the color prediction provided by the network, i.e., the first three channels of the decoder output, is the last channel of the decoder output and is the final color prediction. Thus, the reconstruction loss in Eqn.6 is modified accordingly to:
where is the ground-truth vertex texture and L1 loss is used to avoid color over smoothing. Intuitively, the network learns to infer the color of the whole surface and also determine which part of the surface is visible so that we can directly sample color observation on the image. We show the effect of the alpha channel in Fig.6.
Note that we do not decompose shading component from the original color observation to obtain the “real” surface texture. To obtain shading-free texture, one can use state-of-the-art albedo estimation method like to firstly remove shading from the input images and then feed the processed images to our texture inference network.
Extension to Multi-image Setup
With the common underlying SMPL model as body reference, our PaMIR representation also implicitly builds a correspondence relationship across different models. As a result, our method can be easily extended to multi-image settings like multi-view input or video input. Specifically, given images of the same person , we first use the optimization method described in Sec.4.3 to obtain the SMPL models aligned to each input images, . Note that the optimization step is performed on each individual image independently; for future work more constraints like shape consistency and pose consistency constraints can be incorporated to the energy function. With these SMPL template parameters we can establish correspondences across two images. For example, for a point in the model space of the reference image , its corresponding point in the model space of can be calculated as
where and are the linear blending skinning (LBS) matrices of the -th SMPL vertex in the reference frame and the -th frame, respectively, , and share the same definition as Eqn.(7).
To perform geometry inference using multiple image frames, we follow the practice of PIFu and decompose the feature-to-occupancy network in Eqn.2 into two components:
where is a feature embedding network while is an occupancy reasoning network. For a 3D point in the reference frame , we first sample its feature vector as described in Sec.3.3. Then we calculate the corresponding point in the -th frame and also sample its corresponding feature vector , where . The feature embedding network takes these feature vectors and encodes them into latent feature embedding vectors . The latent embedding vectors across different frames are aggregated into one embedding using mean pooling, which is then mapped to the target implicit field by the occupancy reasoning network . The multi-image network is fine-tuned from the single-image network using multi-view rendering of 3D human models.
Note that PIFu also demonstrates results given mutli-view inputs. However, the multi-view input should be well calibrated and synchronized. In contrast, neither calibration or synchronization is necessary in our method because we can utilize the SMPL estimation to build the correspondence across different views. Moreover, our method can handle the cases where body poses are not identical across images, e.g., video input. This feature enables video-based personalized avatar creation. A similar application has been demonstrated in , but our method can support more challenging cloth topologies and recover more surface details. Overall, in terms of multi-image setup, our method is more general and practical than state-of-the-art methods. A comparison between the single-image result and the multi-image result is presented in Fig.7. As shown in the figure, by adding four more video frames as input, our method is able to recover the surface details on the back that is invisible in the first frame.
Since we consider only pose changes across views and neglect surface detail deformations (e.g., wrinkle movements) across frames, challenging pose deviation and inconsistent geometry between frames will cause inaccurate feature fusion and reconstruction artifacts. Besides, the multi-image network is trained using only multi-view rendering of human models without pose deviations. For future work, we can introduce 4D data to resolve these issues.
Experiments
In this section we evaluate our approach. Details about the implementation and the evaluation dataset are given in Sec.6.1 and Sec.6.2 respectively. In Sec.6.3 we demonstrate that our method is able to reconstruct human models with challenging poses. We then compare our method against state-of-the-art methods in Sec.6.4. We also evaluate our contributions both qualitatively and quantitatively in Sec.6.5. The quantitative evaluation results are given in Tab.III. We mainly use geometry reconstruction performance for evaluation, which is our main focus.
Network Architecture. For image feature extraction, we adapt the same 2D image encoders in PIFu (i.e., Hourglass Stack for geometry and CycleGAN for texture), which take as input an image of 512512 resolution and outputs a 256-channel feature map with a resolution of 128128. For volumetric feature extraction, we use a 3D convolution network which consists of two convolution layers and three residual blocks. Its input resolution is 128128128 and its output is a 32-channel feature volume with a resolution of 323232. We replace batch normalization with group normalization to improve the training stability. The feature decoder is implemented as a multi-layer perceptron, where the number of neurons is , where for the geometry network while for the texture network.
Training Data. To achieve state-of-the-art reconstruction quality, we collect 1000 high-quality textured human scans with various clothing, shapes, and poses from Twindomhttps://web.twindom.com/. We randomly split the Twindom dataset into a training set of 900 scans and a testing of 100 scans. To augement pose variety, we also randomly sample 600 models from DeepHuman dataset. We render the training models from multiple view points using Lambertian diffuse shading and spherical harmonics lighting. We render the images with a weak perspective camera and image resolution of 512 512. To obtain the ground-truth SMPL annotations for the training data, we apply MuVS to the multi-view images for pose computation and then solve for the shape parameters to further register the SMPL to the scans. For point sampling during training, we use the same scheme proposed in PIFu, in which the authors combine uniform sampling and adaptive sampling based on the surface geometry and use embree algorithm for occupancy querying.
Network training. We use Adam optimizer for network training with the learning rate of , the batch size of 3, the number of epochs of 10, and the number of sampled points of 5000 per subject. The learning rate is decayed by the factor of 0.1 at every 10000-th iteration. We also combine predicted and ground-truth SMPL models during network training. Specifically, in each batch, we randomly select one image to replace the predicted SMPL model with the ground-truth when constructing the depth-ambiguity-aware reconstruction loss. In this way we can guarantee the best performance is obtained once the underlining SMPL becomes more accurate. The multi-image network is fine-tuned from the models trained for single-view network with three random views of the same subject using a learning rate of and a batch size of 1. Before training the PaMIR network, we first fine-tune the pre-trained GCMR network on our training set.
Network testing. When testing, our network only requires an RGB image as input and outputs both the parametric model and the reconstructed surface with texture. To maximize the performance, We run the body reference optimization step for all results unless otherwise stated. Fifty iterations are needed for the optimization and take about 40 seconds.
Network complexity. We report the module complexity in Tab.I. To reconstruct the 3D human model given an RGB image, we use the hierarchical SDF querying method proposed in OccNet to reduce network querying times. Overall, taking the body reference optimization step into account, it takes about 50s to reconstruct the geometry of the 3D human model and 1s to recover its surface color.
2 Evaluation Dataset
For qualitative evaluation on single-image reconstruction, we utilize real-world full-body images collected from the DeepFashion dataset and from the Internet. We remove the background of these real-world image using neural semantic segmentation followed by Grabcut refinement. For qualitative evaluation on multi-image reconstruction, we use the open-source dataset from VideoAvatar which contains 24 RGB sequences of different subjects turning 360 degree with a rough A-pose in front of a camera. The testing split of the Twindom dataset is used for quantitative comparison. In addition, for quantitative comparison and evaluation we also use the BUFF dataset, which provides 26 4D human sequences with various clothes and over 9000 scans in total. To reduce computation burden, We select 300 scans of different poses from the buff dataset. We render both the Twindom testing data and the BUFF data from 12 views spanning every 30 degrees in yaw axis using the same method in Sec.6.1.
3 Results
We demonstrate our approach for single-image 3D human reconstruction in Fig.1 and Fig.8. The input images in Fig.1 and Fig.8 covers various body poses (dancing, Kungfu, sitting and running), and also covers different clothes (loose pants, skirts, sports suits and casual clothes). The results demonstrate the ability of our method to reconstruct high-quality 3D human models and its robust performance to tackle various human poses and clothing styles.
We also test our performance for multi-image human model reconstruction using the VideoAvatar dataset in Fig.9. We uniformly sample 5 frames for each sequence for surface reconstruction. Note that for each subject, the poses in different frames are different due to the body movements and no camera extrinsic parameter is provided, so PIFu is not suitable to reconstruct human models in this case. In contrast, our method can reconstruct full-body models with high-resolution details, proving the generalization capability of our PaMIR representation. We also present two example results of applying multi-image feature fusion on images with moderate pose deviation in Fig.10. As the figure shows, our method is still able to reconstruct the overall shapes of the subjects in these cases.
4 Comparison
We qualitatively compare our method with several state-of-the-art methods including HMD, Tex2Shape, Moulding Humans, DeepHuman and PIFu. Among them, HMD and Tex2Shape are parametric methods based on SMPL model deformation, PIFu uses a deep implicit function as geometry representation, Moulding Humans uses the combination of a front depth map and a back depth map for representation, and DeepHuman combines volumetric representation with the SMPL model. Note that Tex2Shape only computes shapes, so we use the pose parameters obtained by our method to skin the its results. For simplicity, we omit comparisons with the works that have already been compared like BodyNet and SiCloPe. In our experiments, DeepHuman, PIFu and Moulding Humans are all retrained on the our dataset, while parametric methods like HMD and Tex2shape are not because our dataset contains loose garments like dresses which will deteriorate the performance of those methods.
We conduct qualitative comparison in Fig.11. As shown in the figure, HMD and Tex2Shape have difficulties dealing with loose clothes and cannot reconstruct the surface geometry accurately; Moulding Human fails to handle challenging poses (due to the lack of semantic constraints), and produces broken body parts when self-occlusions occur, which is the essential limitation of its double depth representation; DeepHuman cannot recover high-frequency geometrical details although it succeeds to reconstruct the rough shapes from the images; PIFu struggles to reconstruct the models in challenging poses and also suffer from self-occlusions. In contrast, our method is able to reconstruct plausible 3D human models under challenging body poses and various clothing styles. In terms of surface quality and pose generalization capacity, our method is superior to other state-of-the-art methods.
We quantitatively compare our method with the state-of-the-art methods using both Twindom testing dataset and BUFF rendering dataset to evaluate the geometry reconsturction accuracy. Similar to the experiments in PIFu, we use point-to-surface error as well as the Chamfer distance as error metric. The numerical results are presented in Tab.II. The quantitative comparison shows that our method outperforms the state-of-the-art methods in terms of surface reconstruction accuracy. We also provide the errors when ground-truth SMPL annotations are available to present the upper limit of our reconstruction accuracy if the SMPL estimation is perfect. Overall, our method is more general, more robust and more accurate than HMD, Moulding Humans, DeepHuman and PIFu.
For multi-image setups, we also conduct qualitative comparison against state-of-the-art methods and . The method proposed in is an optimization-based method which deforms the SMPL template according to the silhouettes. The optimization is performed on the whole video sequence. In contrast, the method in is a learning-based method that deforms the SMPL template based on only 8 views of images. The comparison results are shown in Fig.9. From the results we can see that although all methods are able to reconstruct the overall shapes correctly, our method is able to recover more surface details than the other two methods. This is because our non-parametric representation allows more flexible surface reconstruction. We do not perform quantitative comparison because there is no such benchmark available.
5 Evaluation
The superiority of our PaMIR representation is already proved through the comparison experiments in Sec.6.3 and 6.4. In this evaluation, we demonstrate the advantage of our PaMIR representation from another aspect: it can support multi-modal outputs. To be more specific, our method can output multiple possible human models corresponding to different body pose hypotheses. Two examples are shown in Fig.12. In both experiments we manually adjust one part of the body in order to generate multiple pose hypotheses. As a result, the reconstruction results change in accordance with the input SMPL models. In contrast, non-parametric method like PIFu or Moulding Humans can only generate a specific mesh for each image, showing their limited and overfit capability.
5.2 Depth-ambiguity-aware Loss
To conduct ablation study on our depth-ambiguity-aware loss, we implement a baseline network that is trained using the traditional reconstruction loss and compare it against our network. After the same number of training epochs, we test both network on real-world input images. Two example results are shown in Fig.13. As we can see in the figure, with our depth-ambiguity-aware reconstruction loss, the network is able to learn to reconstruct 3D human models that have consistent poses with the estimated SMPL models. On the contrary, the models reconstructed by the baseline network without our depth-ambiguity-aware loss are less consistent with the SMPL models and have more artifacts. The quantitative study is presented in Tab.III. From the numeric comparison between the 3rd and 4th rows of Tab.III, we can conclude that our depth-ambiguity-aware loss avoids the negative impact of the depth inconsistency between predicted SMPL models and ground-truth scans, and improves the accuracy of the reconstruction results.
5.3 Training Scheme
We also implement another baseline network that is trained using the ground-truth SMPL annotations to evaluate our training scheme. We compare this baseline network with our network on different input images. In Fig.14, we present some cases in which the predicted SMPL models are not well aligned with image keypoints and silhouettes. As we can see in Fig.14, under the scenarios of inaccurate SMPL estimation, the baseline network fails to reconstruct complete full-body models, while our network is able to provide plausible results in accordance with image observations. Note that this feature is of vital importance. For instance, in the last example of Fig.14, the right forearm of the baseline model is completely missing. In this case, the reference body optimization step could not work because no matter how the right forearm moves, the occupancy probability value of its vertices are always closed to zero. The quantitative study is presented in Tab.III. From the numerical comparison between the 2nd and 4th rows of Tab.III, we can conclude that our training scheme improves the accuracy of our reconstruction results at inference time when no accurate SMPL annotation is available.
5.4 Reference Body Optimization
To evaluate the effectiveness of our reference body optimization step, we compare the body fitting results before and after optimization using the evaluation images in Sec.6.5.3. The results are presented in Fig.15. As shown in the figure, the optimization step can further register the SMPL model to the image observation, resulting into more accurate body pose estimation. This is also proven in the quantitiative evaluation in Tab.IV. From the numerical results in the last two rows of Tab.III, we can also see that the mesh reconstruction is also improved after reference body optimization.
5.5 Multi-image Setup
To evaluate the detail changes using more or less images, we conduct a quantitative evaluation in Tab.V. Specifically, we perform feature fusion with 1, 2, 3 and 4 views of the same model and extract the meshes. Then we measure the normal reprojection error from 8 uniform viewpoints to evaluate the detail improvements. The numerical results presented in Tab.V suggest that more geometric details are recovered after fusing information from multiple input images.
Discussion
Conclusion. Modeling 3D humans accurately and robustly from a single RGB image is an extremely ill-posed problem due to the varieties of body poses, clothing types, view points and other environment factors. Our key idea to overcome these challenges is factoring out pose estimation from surface reconstruction. To this end, we have contributed a deep learning-based framework to combine the parametric SMPL model and the non-parametric deep implicit function for 3D human model reconstruction from a single RGB image. Benefiting from the proposed PaMIR representation, the depth-ambiguity-aware reconstruction loss and the reference body optimization algorithm, our method outperforms state-of-the-art methods in terms of both robustness and surface details. As shown in the supplementary video, we can obtain temporally consistent reconstruction results by applying our method to video frames individually. We believe that our method will enable many applications and further stimulate the research in this area.
From a more general perspective of view, we have made a step forward towards integrating semantic information into the free-form implicit fields which have attracted more and more attention from the research community for its flexibility, representation power and compact nature. For example, similar topics include the reconstruction of hand and face using such PaMIR representation. We also believe our semantic implicit fields would become one important future topic and inspire many other studies in 3D vision. For example, the multi-modality of our method (Fig.12) can inspire research on image-based geometry editing; the practice of integrating geometry template into implicit functions can be used to improve the robustness of other 3D recovery tasks. Moreover, the semantic implicit functions can also be adopted in research on 3D perception, parsing and understanding.
Limitation and Future Work. Our method needs high-quality human scans for training. However, it is highly costly and time consuming to obtain a large-scale dataset of high-quality human scans. Moreover, the currently available scanners for human bodies require the subject to keep static poses in a sophisticated capturing environment, which makes them incapable to capture real-world human motions in the wild and consequently our training data is biased towards simple static poses like standing. Therefore, although the proposed method already makes a step forward in terms of generalization capability, it still fails in the cases of extremely challenging poses, see Fig.16. One important future direction is to alleviate the reliance on ground-truth by exploring large-scale image and video dataset for unsupervised training.