Multi-Garment Net: Learning to Dress 3D People from Images
Bharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt, Gerard Pons-Moll
Introduction
The 3D reconstruction and modelling of humans from images is a central problem in computer vision and graphics. Although a few recent methods attempt reconstruction of people with clothing, they lack realism and control. This limitation is in great part due to the fact that they use a single surface (mesh or voxels) to represent both clothing and body. Hence they can not capture the clothing separately from the subject in the image, let alone map it to a novel body shape.
In this paper, we introduce Multi-Garment Network (MGN), the first model capable of inferring human body and layered garments on top as separate meshes from images directly. As illustrated in Fig. 1 this new representation allows full control over body shape, texture and geometry of clothing and opens the door to a range of applications in VR/AR, entertainment, cinematography and virtual try-on.
Compared to previous work, MGN produces reconstructions of higher visual quality, and allows for more control: 1) we can infer the 3D clothing from one subject, and dress a second subject with it, (see Fig. 1, 8) and 2) we can trivially map the garment texture captured from images to any garment geometry of the same category (see Fig.7).
To achieve such level of control, we address two major challenges: learning per-garment models from 3D scans of people in clothing, and learning to reconstruct them from images.
We define a discrete set of garment templates (according to the categories long/short shirt, long/short pants and coat) and register, for every category, a single template to each of the scan instances, which we automatically segmented into clothing parts and skin. Since garment geometry varies significantly within one category (e.g. different shapes, sleeve lengths), we first minimize the distance between template and the scan boundaries, while trying to preserve the Laplacian of the template surface. This initialization step only requires solving a linear system, and nicely stretches and compresses the template globally, which we found crucial to make subsequent non-rigid registration work. Using this, we compile a digital wardrobe of real 3D garments worn by people, (see Fig. 3). From such registrations, we learn a vertex based PCA model per garment. Since garments are naturally associated with the underlying SMPL body model, we can transfer them to different body shapes, and re-pose them using SMPL.
From the digital wardrobe, MGN is trained to predict, given one or more images of the person, the body pose and shape parameters, the PCA coefficients of each of the garments, and a displacement field on top of PCA that encodes clothing detail. At test time, we refine this bottom-up estimates with a new top-down objective that forces projected garments and skin to explain the input semantic segmentation. This allows more fine-grained image matching as compared to standard silhouette matching. Our contributions can be summarized as:
A novel data driven method to infer, for the first time, separate body shape and clothing from just images (few RGB images of a person rotating in front of the camera).
A robust pipeline for 3D scan segmentation and registration of garments. To the best of our knowledge, there are no existing works capable of automatically registering a single garment template set to multiple scans of real people with clothing.
A novel top-down objective function that forces the predicted garments and body to fit the input semantic segmentation images.
We demonstrate several applications that were not previously possible such as dressing avatars with predicted garments from images, and transfer of garment texture and geometry.
We will make publicly available the MGN to predict clothing from images, the digital wardrobe, as well as code to “dress” SMPL with it.
Related Work
In this section we discuss the two branches of work most related to our method, namely capture of clothing and body shape and data-driven clothing models.
Performance Capture. The classical approach to bring dynamic sequences into correspondence is to deform meshes non-rigidly or volumetric shape representations to fit multiple image silhouettes. Without a pre-scanned template, fusion trackers incrementally fuse geometry and appearance to build the template on the fly. Although flexible, these require multi-view , one or more depth cameras , or require the subject to stand still while turning the cameras around them . From RGB video, Habermann et al. introduced a real time tracking system to capture non-rigid clothing dynamics. Very recently, SimulCap allows multi-part tracking of human performances from a depth camera.
Body and cloth capture from images and depth. Since current statistical models can not represent clothing, most works are restricted to inferring body shape alone. Model fits have been used to virtually dress and manipulate people’s shape and clothing in images . None of these approaches recover 3D clothing. Estimating body shape and clothing from an image has been attempted in , but it does not separate clothing from body and requires manual intervention . Given a depth camera, Chen et al. retrieve similar looking synthetic clothing templates from a database. Daněřek et al. use physics based simulation to train a CNN but do not estimate garment and body jointly, require pre-specified garment type, and the results can only be as good as the synthetic data.
Closer to ours is the work of Alldieck et al. which reconstructs, from a single image or a video, clothing and hair as displacements on top of SMPL, but can not separate garments from body, and can not transfer clothing to new subjects. In stark contrast to , we register the scan garments (matching boundaries) and body separately, which allows us to learn the mapping from images to a multi-layer representation of people.
Data-driven clothing. A common strategy to learn efficient data-driven models is to use off-line simulations for generating data. These approaches often lack realism when compared to models trained using real data. Very few approaches have shown models learned from real data. Given a dynamic scan sequence, Neophytou et al. learn a two layer model (body and clothing) and use it to dress novel shapes. A similar model has been recently proposed , where the clothing layer is associated to the body in a fuzzy fashion. Other methods focus explicitly on estimating the body shape under clothing. Like these methods, we treat the underlying body shape as a layer, but unlike them, we segment out the different garments allowing sharp boundaries and more control. For garment registration, we build on the ideas of ClothCap , which can register a subject specific multi-part model to a 4D scan sequence. By contrast, we register a single template set to multiple scan instances– varying in garment geometry, subject identity and pose, which requires a new solution. Most importantly, unlike all previous work , we learn per-garment models and train a CNN to predict body shape and garment geometry directly from images.
Method
In order to learn a model to predict body shape and garment geometry directly from images, we process a dataset of scans of people in varied clothing, poses and shapes. Our data pre-processing (Sec. 3.1) consists of the following steps: SMPL registration to the scans, body aware scan segmentation and template registration. We obtain, for every scan, the underlying body shape, and the garments of the person registered to one of the garment template categories: shirt, t-shirt, coat, short-pants, long-pants. The obtained digital wardrobe is illustrated in Fig. 3. The garment templates are defined as regions on the SMPL surface; the original shape follows a human body, but it deforms to fit each of the scan instances after registration. Since garment registrations are naturally associated to the body represented with SMPL, they can be easily reposed to arbitrary poses. With this data, we train our Multi-Garment Network to estimate the body shape and garments from one or more images of a person, see Sec. 3.2.
Unlike ClothCap which registers a template to a scan sequence of a single subject, our task is to register single template across instances of varying styles, geometries, body shapes and poses. Since our registration follows the ideas of , we describe the main differences here.
We first automatically segment the scans into three regions: skin, upper-clothes and pants (we annotate the garments present for every scan). Since even SOTA image semantic segmentation is inaccurate, naive lifting to 3D is not sufficient. Hence, we incorporate body specific garment priors and segment scans by solving an MRF on the UV-map of the SMPL surface after non-rigid alignment.
After solving the MRF on the SMPL UV map, we can segment the scans into 3 parts by transferring the labels from the SMPL registration to the scan.
Garment Template We build our garment template on top of SMPL+D, , which represents the human body as a parametric function of pose(), shape(), global translation() and optional per-vertex displacements ():
The basic principle of SMPL is to apply a series of linear displacements to a base mesh with vertices in a T-pose, and then apply standard skinning . Specifically, models pose-dependent deformations of a skeleton , and models the shape dependent deformations. represents the blend weights.
Consequently, we can obtain the garment shape (unposed), for a new shape and pose as
To pose the vertices of a garment, each vertex uses the skinning function in Eq. 1 of the associated SMPL body vertex.
where is a constant ( in our experiments), is the vertex of and is the point closest to on .
Our garment formulation allows us to freely repose the garment vertices. We can use this to our advantage for applications such as animating clothed virtual avatars, garment re-targeting etc. However, posing is highly non-linear and can lead to undesired artefacts, specially when re-targeting garments across subjects with very different poses. Since we re-target the garments in unposed space, we reduce distortion by forcing distances from garment vertices to the body to be preserved after unposing:
where is the distance between point and surface . and denote garment vertex and body surface in unposed space, using Eq. 5 and 1 respectively.
Dressing SMPL The SMPL model has proven very useful for modelling unclothed shapes. Our idea is to build a wardrobe of digital clothing compatible with SMPL to model clothed subjects. To this end we propose a simple extension that allows to dress SMPL. Given a garment , we use Eq. 3, 4, 5 to pose and skin the garment vertices. The dressed body including body shape (encoded as ) will be given by stacking the individual garment vertices . We define the function which returns the posed and shaped vertices for the skin, and each of the garments combined. See Fig. 5 and supplementary for results on re-targeting garments using MGN across different SMPL bodies.
2 From Images to Garments
From registrations, we learn a shape space of garments, and generate a synthetic training dataset with pairs of images and body+3D garment pairs. From this data we train MGN:Multi-Garment Net, which maps images to 3D garments and body shape.
MGN: Multi-Garment Net The input to the model is a set of semantically segmented images, , and corresponding joint estimates, , where is the number of images used to make the prediction. Following , we abstract away the appearance information in RGB images and extract semantic garment segmentation to reduce the risk of over-fitting, albeit at the cost of disregarding useful shading signal. For simplicity, let now denote both the joint angles and translation .
The base network, , maps the 2D poses , and image segmentations , to per frame latent code corresponding to 3D poses
and to a common latent code corresponding to body shape and garments by averaging the per frame codes
From the shape and pose latent codes , we predict body shape parameters and pose respectively, using a fully connected layer. Using the predicted body shape and geometry we compute displacements as in Eq. 3:
Consequently, the final predicted vertices posed for the frame are obtained with , from which we render 2D segmentation masks
where is a differentiable renderer , the rendered semantic segmentation image for frame , and denotes the camera parameters that are assumed fixed while the person moves. The rendering layer in Eq. (14) allows us to compare predictions against the input images. Since MGN predicts body and garments separately, we can predict a semantic segmentation image, leading to a more fine-grained loss, which is not possible using a single mesh surface representation . Note that Eq. 14 allows to train with self-supervision.
3 Loss functions
The proposed approach can be trained with supervision on vertex coordinates, and with self supervision in the form of segmented images. We use upper-hat for variables that are known and used for supervision during training. We use the following losses to train the network in an end to end fashion:
vertex loss in the canonical T-pose ():
where, represents zero-vector corresponding to zero pose.
segmentation loss: Unlike we do not optimize silhouette overlap, instead we jointly optimize the projected per-garment segmentation against the input segmentation mask. This ensures that each garment explains its corresponding mask in the image:
Intermediate losses: We further impose losses on intermediate pose, shape and garment parameter predictions: where are the number of images and garments respectively. are the ground truth PCA garment parameters. While such losses are a bit redundant, they stabilize learning.
4 Implementation details
We use a CNN to map the input set to the body shape, pose and garment latent spaces. It consists of five, convolutions followed by max-pooling layers. Translation invariance, unfortunately, renders CNNs unable to capture the location information of the features. In order to reproduce garment details in 3D, it is important to leverage 2D features as well as their location in the 2D image. To this end, we adopt a strategy similar to , where we append the pixel coordinates to the output of every CNN layer. We split the last convolutional feature maps into three parts to individuate the body shape, pose and garment information. The three branches are flattened out and we append 2D joint estimates to the pose branch. Three fully connected layers and average pooling on garment and shape latent codes, generate and respectively. See supplementary for more details.
Garment Network (): We train separate garment networks for each of the garment classes. The garment network consists of two branches. The first predicts the overall mesh shape, and second one adds high frequency details. From the garment latent code (), the first branch, consisting of two fully connected layers (sizes=1024, 128), regresses the PCA coefficients. Dot product of these coefficients with the PCA basis generates the base garment mesh. We use the second fully connected branch (size = ) to regress displacements on top of the mesh predicted in the first branch. We restrict these displacements to to ensure that overall shape is explained by the PCA mesh and not these displacements.
Dataset and Experiments
We use 356 3D scans of people with various body shapes, poses and in diverse clothing. We held out 70 scans for testing and use the rest for training. Similar to , we also restrict our setting to the scenario where the person is turning around in front of the camera. We register the scans using multi-mesh registration, SMPL+G. This enables further data augmentation since the registered scans can now be re-posed and re-shaped. We adopt the data pre-processing steps from including the rendering and segmentation. We also acknowledge the scale ambiguity primarily present between the object size and the distance to the camera. Hence we assume that the subjects in 3D have a fixed height and regress their distance from the camera. Same as , we also ignore the effect of camera intrinsics.
1 Experiments
In this section we discuss the merits of our approach both qualitatively and quantitatively. We also show real world applications in the form of texture transfer (Fig. 7), where we maintain the original geometry of the source garment but map novel texture. We also show garment re-targeting from images using MGN in Fig. 8.
Qualitative comparisons: We compare our method against on our scan dataset. For fair comparison we re-train the models proposed by Alldieck et al. on our dataset and compare against our approach (Dataset used by is not publicly available). Figure 6 indicates the advantage of incorporating the garment model in structured prediction over simply modelling free form displacements. Explicit garment modelling allows us to predict sharper garment boundaries and minimize distortions (see Fig. 6). More examples are shown in the supplementary material.
Quantitative Comparison: In this experiment we do a quantitative analysis of our approach against the state of the art prediction method, . We compute a symmetric error between the predicted and GT garment surfaces similar to . We report per-garment error, (supplementary), and overall error, i.e. mean of over all the garments
where is the number of meshes with garment . and denote the set of vertices and the surface of the predicted mesh respectively, belonging to garment . Operator denotes GT values. computes the distance between the vertex and surface .
This criterion is slightly different than because we do not evaluate error on the skin parts. We reconstruct the 3D garments with mean vertex-to-surface error of 5.78 mm with 8 frames as input. We re-train octopus on our dataset and the resulting error is 5.72mm.
We acknowledge the slightly better performance of and attribute it to the fact that the single mesh based approaches do not bind vertices to semantic roles, i.e these approaches can pull vertices from any part of the mesh to explain deformations where as our approach ensures that only semantically correct vertices explain the shape.
It is also worth noting that MGN predicts garments as linear function (PCA coefficients) of latent code, whereas deploys GraphCNN. PCA based formulation though easily tractable is inherently biased towards smooth results. Our work paves the way for further exploration into building garment models for modelling the variations in garment geometry over a fixed topology.
We report the results for using varying number of frames in the supplementary.
GT vs Predicted pose: The vertex predictions are a function of pose and shape. In this experiment we do an ablation study to isolate the effect of errors in pose estimation on vertex predictions. This experiment is important to better understand the strengths and weaknesses of the proposed approach in shape estimation by marginalizing over the errors due to pose fitting. We study two scenarios, first where we predict the 3D pose and second, where we have access to GT pose. We report mean vertex-to-surface error of 5.78mm with GT poses and 11.90mm with our predicted poses.
2 Re-targeting
Our multi-mesh representation essentially decouples the underlying body and the garments. This opens up an interesting possibility to take garments from source subject and virtually dress a novel subject. Since the source and the target subjects could be in different poses, we first unpose the source body and garments along with the target body. We drop the notation for the unposed space in the following section for clarity. Below we propose and compare two garment re-targeting approaches. After re-targeting the target body and re-targeted garments are re-posed to their original poses.
Naive re-targeting: The simplest approach to re-target clothes from source to target is to extract the garment offsets, from the source subject using Eq. 13 and dress a target subject using Eq. 5.
Body aware re-targeting: The naive approach is problematic because it relies on non-local pre-set vertex association between the garment and the body (). This results in inaccurate association between the body blend shapes, and the garment vertices. This eventually leads to incorrect estimation of source offsets, and in turn leads to higher inter-penetrations between the re-targeted garment and the body (see supplementary). In order to mitigate this issue, we compute the new target garment vertex location, as follows
where is the source garment vertex, is the vertex (indexed by ) among the source body vertices, , closest to and is the corresponding vertex among the target body vertices.
MGN allows us to predict separable body shape and garments in 3D, allowing us to do garment re-targeting (as described above) using just images. To the best of our knowledge this is the first method to do so. See Fig. 8 for results on garment re-targeting by MGN. See supplementary for more results.
Conclusion and Future Works
We introduce MGN, the first model capable of jointly reconstructing from few images, body shape and garment geometry as layered meshes. Experiments demonstrate that this representation has several benefits: it is closer to how clothing layers on top of the body in the real world, which allows control such as re-dressing novel shapes with the reconstructed clothing. Additionally, we introduce for the first time, a dataset of registered real garments from real scans obtained with a robust registration pipeline. When compared to more classical single mesh representations, it allows more control and qualitatively the results are very similar. In summary, we think that MGN provides a first step in a promising research direction. We will release the MGN model and the digital wardrobe to stimulate research in this direction. Further discussion on limitations and future works in supplementary.
Acknowledgements This work is partly funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) - 409792180 (Emmy Noether Programme, project: Real Virtual Humans) and Google Faculty Research Award. We thank twindom (https://web.twindom.com/) for providing scan data, Thiemo Alldieck for providing code for texture/segmentation stitching, and Verica Lazova for discussions.