MonoClothCap: Towards Temporally Coherent Clothing Capture from Monocular RGB Video

Donglai Xiang, Fabian Prada, Chenglei Wu, Jessica Hodgins

Introduction

Dynamic capture of detailed human geometry and motion from monocular images and videos is attracting increasing attention in the computer vision and computer graphics community. High-quality human capture would enable applications in virtual and augmented reality, games, and movies. In recent years, great progress has been made on the estimation of general body shape from a single image or a monocular video . However, capturing the detailed deformation of clothing as it moves on the human body is still far from a solved problem.

Capturing a temporally coherent shape for clothing from monocular RGB imagery is an extremely challenging task, due to the fundamental ambiguity of single-view 3D reconstruction and the large deformation space of clothing. Previous work utilizes a 3D personalized actor model as a shape prior to track the dynamic clothing deformation. This model is acquired by multi-view reconstruction on an additional video of the same actor wearing the same clothing and rotating in a T-pose. However, such a model is generally unavailable for in-the-wild videos. The need for a pre-scanned template model limits the applicability of these approaches.

With the development of deep neural networks, other efforts have been made to regress a clothed human shape directly from a single input image with supervised learning . These methods produce plausible results for individual input images of common human poses. However, it is difficult to extend them to capture temporally coherent dynamic clothing deformation from monocular videos for the following reasons. First, these methods are not robust to the variety of human motion due to the limited diversity of training data. They can easily produce incomplete geometry that is difficult to fix via post-processing. Second, it is non-trivial to estimate the temporal correspondence from the output of individual frames due to the data representation used (voxel , depth map or implicit function ). This limits the application of these methods in scenarios that require correspondence, such as clothing retargeting or image editing.

In this work, we present a novel method to capture dynamic clothing deformation from a monocular RGB video in a temporally coherent manner, as illustrated by Fig. LABEL:fig:teaser. To the best of our knowledge, it is the first attempt to solve this challenging problem without the prerequisite of a pre-scanned personalized template .

Our method is based on the following observations. First, a deformation model of the clothing that provides a statistical shape prior is key to solving the problem. It not only reduces the ambiguity of single-view 3D reconstruction, but also helps to estimate temporal correspondence across frames. While clothing models have been investigated in the existing literature for the purpose of clothing shape generation, our work is the first study that fully demonstrates the value of a clothing model for RGB-based clothing captureDue to the limitation in types of available clothing data to train our model, in this paper, we assume that the subject to be captured wears a T-shirt on the upper body and shorts or pants on the lower body.. Second, to solve the clothing capture problem, we make use of human appearance information including silhouette, segmentation, texture and surface normal. We present a novel method to integrate all those image measurements using a differentiable renderer . Our method captures the realistic dynamic of clothing in a temporally coherent manner including fine-grained wrinkle details from various videos.

Our Contributions. (1) We present the first approach for temporally coherent clothing capture from a monocular RGB video without using a pre-scanned template of the subject. (2) We propose a novel method to capture clothing deformation by fitting statistical clothing models to image measurements including silhouette, segmentation, texture and surface normal with a differentiable renderer.

Related Work

Single-Image Human Pose and Shape Estimation. Most previous work in human pose estimation focuses on the position of body keypoints in 2D and 3D . Because estimating 3D pose from single images is highly ill-posed, deformable human models including SMPL , SMPL-X and Adam are used to help with the problem by fitting the models to images . These models not only provide a strong prior for body pose, but also enable estimation of 3D body shape from single images. Deformable human models can be further integrated in deep neural network architectures . These networks are usually trained in a weakly-supervised manner without full 3D supervision.

Because deformable human models are not able to express clothing shape, all the work above only estimates body shapes with minimal clothing. Detailed clothing shape has been largely ignored in the previous literature, except a few papers . These methods use deep neural networks to infer dense clothed human shapes in various data representation including voxels , depth maps , point clouds and implicit functions , all with supervised learning. However, because the amount of available training data is very limited, these methods are not robust to human motion. In addition, it is non-trivial to estimate correspondence across frames required for clothing capture due to their data representation. Our method achieves temporally coherent body and clothing capture in terms of both geometry and correspondence with the help of a statistical clothing model.

Garment Modeling and Reconstruction. Human clothing, especially physically based simulation of garments , has been extensively studied due to its important role in animation. Recently, there is growing interest in modeling garments in a data-driven manner. Pons-Moll et al. proposes a method to automatically segment 4D clothed human scans into different garment pieces, and track the deformation of each piece over time. The captured clothing data can be further used to train a deformable model, either a linear model or a deep neural network . In those methods, the clothing models are primarily used for shape generation, while we use the model to track clothing deformation from a monocular video.

Another line of work reconstructs clothing shape from images by allowing per-vertex deformation on top of the SMPL body model. Alldieck et al. builds clothed human avatar from videos of a person slowly rotating in A-pose. This is further improved to use only images of several different views or even a single image . However, these methods reconstruct clothing as static objects without considering the temporal dynamics. By contrast, in this paper, we address the challenging problem of capturing clothing dynamics from a monocular video.

Monocular Human Performance Capture. Motion capture and performance capture refer to the capture of space-time coherent human motion sequences in the form of sparse 3D joints and surface geometry respectively. Many approaches have been developed to enable motion and performance capture from multi-view inputs . Here we focus on monocular-based capture methods. Mehta et al. proposes systems to capture body skeleton motion from a single RGB video in real time. Some performance capture methods are proposed to capture dense human body and clothing geometry from a monocular RGB-D video using a double-layer representation. Most relevant to our work are performance capture methods from monocular RGB videos . These methods, however, require a pre-scanned mesh template of the subject, which restricts the applications where they can be used. Habermann et al. further proposes to train a deep neural network to deform a pre-scanned mesh template to match the surface deformation in the video. This method, requires the mesh template and multi-view images of the subject for network training. Our method relaxes the constraint to scenarios such as in-the-wild videos where neither pre-scanned templates nor multi-view images are available.

Method Overview

In this section, we present an overview of our approach. Our goal is to capture the dynamic deformation of three types of garments, T-shirt, shorts and pants, along with the underlying body shape from a monocular video. Our method takes as input a sequence of images, denoted as {Ii}i=1F\{\mathbf{I}_{i}\}_{i=1}^{F}, where FF is the number of frames in the sequence. The subject is assumed to be wearing a T-shirt for the upper body. The clothing for the lower body is manually identified as either short pants or long pants. Our method outputs a sequence of mesh pairs {Mib,Mic}i=1F\{\mathbf{M}^{b}_{i},\mathbf{M}^{c}_{i}\}_{i=1}^{F}, where Mib\mathbf{M}^{b}_{i} denotes the body mesh and Mic\mathbf{M}^{c}_{i} denotes the clothed mesh. {Mib}i=1F\{\mathbf{M}^{b}_{i}\}_{i=1}^{F} and {Mic}i=1F\{\mathbf{M}^{c}_{i}\}_{i=1}^{F} are both temporally coherent with fixed topology across time. Mib\mathbf{M}^{b}_{i} and Mic\mathbf{M}^{c}_{i} share the same vertex positions except for the clothing region.

Our method makes use of linear clothing deformation models defined in the canonical pose. We briefly describe our model formulation and model building procedure in Section 4. Our pipeline to capture clothing from a monocular video consists of four stages, explained in Section 5. First, we estimate the underlying body pose and shape of subject (Section 5.1). Then, we run sequential tracking of the clothing using our linear clothing models. This step is followed by a batch optimization stage including all the frames to produce temporally coherent dynamic clothing deformation (Section 5.2). In the final stage, we add fine-grained wrinkle detail to our results (Section 5.3). A visualization of this pipeline is shown in Fig. 1.

Statistical Clothing Deformation Model

Statistical models of clothing have been investigated for clothing shape generation in the previous literature , but have yet to be exploited for capturing clothing from a monocular video. In this section we give the mathematical formulation of our clothing deformation models and briefly describe the procedure to learn these models from data.

where WW is the Linear Blend Skinning (LBS) function; T(β,θ)T(\boldsymbol{\beta},\boldsymbol{\theta}) is the rest pose body shape; J(β)J(\boldsymbol{\beta}) is the locations of 2424 kinematic joints; W\mathcal{W} is the blend weights. In particular, the unposed shape T(β,θ)T(\boldsymbol{\beta},\boldsymbol{\theta}) is defined as the sum of template shape T‾\overline{T}, shape dependent deformation BS(β)B^{S}(\boldsymbol{\beta}) and pose dependent deformation BP(θ)B^{P}(\boldsymbol{\theta}),

On top of the SMPL model, an extra additive offset field DD is introduced to account for clothing deformation in rest pose, i.e.,

where z={zu,zl}\mathbf{z}=\{\mathbf{z}^{u},\mathbf{z}^{l}\} is the collection of clothing parameters. A visual illustration of our clothing model formulation is shown in Figure 2.

2 Model Building

where ⊙\odot denotes the element-wise multiplication. We use a standard PCA training algorithm based on Singular Value Decomposition (SVD), leaving nz=50n_{z}=50 bases in our model. We refer readers to the original papers for details on scan registration and underlying body shape estimation.

Monocular Clothing Capture

Given the pre-trained clothing models, we now present our approach for temporally coherent clothing capture from only a monocular video.

In particular, E2dbE^{b}_{\text{2d}} is the squared L2L_{2} error between projected SMPL joints and 2D keypoint detection from OpenPose . EdpbE^{b}_{\text{dp}}, also used in Guler et al. , is an energy term for dense correspondence estimation from DensePose . Specifically, for any pixel p\mathbf{p} in the image with DensePose prediction, we identify the corresponding SMPL vertex index j(p)j(\mathbf{p}) and optimize an energy term defined as

where Π\Pi denotes the projection function determined by the camera intrinsics K\mathbf{K}. EsilbE^{b}_{\text{sil}} is the silhouette matching term. We extract silhouettes Si\mathbf{S}_{i} of our SMPL body mesh with a differentiable renderer , and obtain the target silhouette S^i\hat{\mathbf{S}}_{i} from a clothing segmentation algorithm . We use an Intersection-over-Union error

EpofbE^{b}_{\text{pof}} is an error term based on 3D orientation between adjacent joints in the body skeleton hierarchy. We match the spatial orientation of SMPL body joints to the prediction of Part Orientation Field (POF) similar to . We refer readers to the original papers for details. We also apply regularization on our estimation, denoted by EregbE^{b}_{\text{reg}}, which consists of a Mixture of Gaussian prior for body pose {θi}i=1F\{\boldsymbol{\theta}_{i}\}_{i=1}^{F} , L2L_{2} regularization on the shape parameters β\boldsymbol{\beta}, and temporal smoothness terms to reduce motion jitters.

After solving the energy optimization, we obtain a temporally consistent body mesh for every frame by Mib=M(β,θi)\mathbf{M}^{b}_{i}=M(\boldsymbol{\beta},\boldsymbol{\theta}_{i}). We fix the SMPL parameters β,{θi,ti}i=1F\boldsymbol{\beta},\{\boldsymbol{\theta}_{i},\mathbf{t}_{i}\}_{i=1}^{F} and camera parameters K\mathbf{K} during later stages of our pipeline. The estimated body meshes provide a strong guidance for the subsequent estimation of clothing deformation.

2 Clothing Deformation Capture

We now illustrate our proposed method to capture clothing deformation. Compared to previous work where a pre-scanned template of the subject is assumed, this problem is significantly more challenging due to the lack of strong shape prior to resolve the single-view 3D ambiguity, and the lack of a pre-defined personalized texture that provides correspondence for surface tracking. To solve this problem, we (1) exploit the deformation space learned in our clothing models and (2) progressively extract a personalized texture from the input image sequence to enable surface tracking across time and reduce drifting.

We perform clothing capture in a sequential manner. For each frame ii, we estimate per-frame clothing parameters for clothing on the upper and lower body zi={ziu,zil}\mathbf{z}_{i}=\{\mathbf{z}^{u}_{i},\mathbf{z}^{l}_{i}\} given the input image Ii\mathbf{I}_{i}, initializing from the result of the previous frame zi−1\mathbf{z}_{i-1}. We formulate the task as solving an energy optimization problem, formally,

Now we explain each cost term individually. An illustration of the different cost terms is shown in Fig. 3.

Silhouette matching term EsilcE^{c}_{\text{sil}}: Similar to EsilbE^{b}_{\text{sil}} in the first stage, we use a differentiable renderer to match the silhouette of our rendering output with a target silhouette extracted from the original image. Differently, here what we compare with target silhouette is the silhoutte of human shape with clothing Mc(β,θi,zi)M^{c}(\boldsymbol{\beta},\boldsymbol{\theta}_{i},\mathbf{z}_{i}), instead of bare body shape M(β,θi)M(\boldsymbol{\beta},\boldsymbol{\theta}_{i}) in the previous stage.

Clothing segmentation term EsegcE^{c}_{\text{seg}}: Clothing segmentation provides not only the overall silhouette of the person, but also the boundary between different garment regions and exposed skin in the image. We utilize this information by penalizing the clothing offset on vertices whose projection falls outside the segmentation region. Concretely, the differentiable renderer is used to render the per-vertex offset fields Du(ziu),Dl(zil)D^{u}(\mathbf{z}^{u}_{i}),D^{l}(\mathbf{z}^{l}_{i}) on the clothed mesh Mc(zi)M^{c}(\mathbf{z}_{i})SMPL body parameters β,θi\boldsymbol{\beta},\boldsymbol{\theta}_{i} and ti\mathbf{t}_{i} are fixed in this stage thus omitted here.; we denote the output by R(Dk(zik))\mathcal{R}\left(D^{k}(\mathbf{z}^{k}_{i})\right), where kk represents either uu for upper clothing or ll for lower clothing. In addition, we denote the segmentation masks for the clothing region k∈{u,l}k\in\{u,l\} by S^ik\hat{\mathbf{S}}^{k}_{i}. Then we have

where p\mathbf{p} iterates over all pixels in the image. Effectively, for each clothing type we penalize the clothing offset outside the corresponding clothing region, where S^ik\hat{\mathbf{S}}^{k}_{i} is . The gradient in the image domain is propagated to the mesh by the differentiable renderer R\mathcal{R}.

Photometric tracking term EphotocE^{c}_{\text{photo}}: This term is introduced to estimate temporal correspondence more accurately, especially when the garments we capture have high-contrast texture. We progressively build a personalized RGB texture image Ti\mathbf{T}_{i} in a pre-defined UV space of SMPL model, along with a binary mask Ti′\mathbf{T}^{\prime}_{i} that indicates texels where RGB values in Ti\mathbf{T}_{i} have been identified. The photometric tracking term is defined to compare the rendered output of our clothing models using Ti\mathbf{T}_{i} with the input image. For this purpose we use a differential renderer that works with UV texture , and denotes the rendered output as R(Ti)\mathcal{R}\left(\mathbf{T}_{i}\right). We also render the mesh with Ti′\mathbf{T}^{\prime}_{i} to indicate pixels where texture from Ti\mathbf{T}_{i} is available. Formally, we have

where the summation is taken over the pixels in the image. After the optimization in Eq. 12 is solved for frame ii, we update Ti,Ti′\mathbf{T}_{i},\mathbf{T}^{\prime}_{i} to obtain Ti+1,Ti+1′\mathbf{T}_{i+1},\mathbf{T}^{\prime}_{i+1}, which will be used for solving optimization for frame i+1i+1. To achieve this, we project Ii\mathbf{I}_{i} to the mesh surface and fill in new RGB values to UV texels in Ti\mathbf{T}_{i} where no previous values have been identified, indicated by s in Ti′\mathbf{T}^{\prime}_{i}. Corresponding texels in Ti′\mathbf{T}^{\prime}_{i} are also set to 11 to obtain Ti+1′\mathbf{T}^{\prime}_{i+1}. This process is initialized by setting T1\mathbf{T}_{1} and T1′\mathbf{T}^{\prime}_{1} to ; in other words, the photometric tracking term takes no effect for the first frame in the sequence, since no texture has been extracted.

Regularization term EregcE^{c}_{\text{reg}}: Our clothing deformation models are PCA-based linear models. They may produce unreasonable shapes when the parameters z\mathbf{z} are large. Therefore we apply regularization on the cloth parameters using an adaptive cost function ρ\rho that penalizes large input values:

The sequential tracking stage is then followed by a batch optimization stage that optimize for all FF frames in the sequence together. The energy function we use is the same as the previous stage (Eq. 12) with an additional term that penalizes too drastic temporal change of clothing parameters, which helps to produce temporally stable results. This term is defined as

The output of batch optimization stage is a sequence of capture results with clothing {Mic=Mc(β,θi,zi)}i=1F\{\mathbf{M}^{c}_{i}=M^{c}(\boldsymbol{\beta},\boldsymbol{\theta}_{i},\mathbf{z}_{i})\}_{i=1}^{F}.

3 Wrinkle Detail Extraction

Up to the batch optimization stage, we can capture large clothing deformation. However, the results are limited by low mesh resolution and unable to capture the fine-grained wrinkles on the clothing. Therefore, the last stage of our approach is to extract wrinkle details from the original images and apply them to our coarsely tracked meshes.

Traditionally, such wrinkle details are captured with Shape from Shading (SfS) . For in-the-wild monocular clothing capture, we empirically find it difficult to extract wrinkles reliably by SfS due to complex garment albedo, large variation of lighting conditions and self-shadowing. Recently, we observed the success of learning-based approaches in estimating accurate surface normal for human appearance using neural networks . The estimated surface normal provides strong and direct clues on how the wrinkles should be added to our clothing capture results in order to match the original images.

Formally, let us denote the output of a surface normal estimation network for frame ii to be Iin\mathbf{I}^{n}_{i}, a 3-channel normal map for each pixel in the original image. We first subdivide the mesh Mic\mathbf{M}^{c}_{i} with Loop subdivision to increase the spatial resolution, with the subdivided mesh denoted by Mis\mathbf{M}^{s}_{i}. Then, we solve for a deformed mesh Oi\mathbf{O}_{i} whose rendered normal map matches the estimated normal map Iin\mathbf{I}^{n}_{i} in the garment region. We denote the rendered normal output by Rn(Oi)\mathcal{R}^{n}(\mathbf{O}_{i}), where Rn\mathcal{R}^{n} is the differential renderer. We solve the following optimization problem for each frame ii individually:

where ∇\nabla denotes the image gradient operator, and L\mathbf{L} denotes the mesh Laplacian operator, S^ic=S^iu⋃S^il\hat{\mathbf{S}}^{c}_{i}=\hat{\mathbf{S}}^{u}_{i}\bigcup\hat{\mathbf{S}}^{l}_{i} is the union of all pixels in the clothing segmentation masks. Here we penalize the difference between normal maps in the image gradient domain to be more tolerant to error in absolute normal direction from the neural network. We also restrict the deformation of Oi\mathbf{O}_{i} from Mis\mathbf{M}^{s}_{i} be in the direction of the camera rays. Here, we use the implementation of differentiable renderer in and surface normal network in . The final results are the deformed meshes {Oi}i=1F\{\mathbf{O}_{i}\}_{i=1}^{F}.

Quantitative Evaluation

In this section, we present the results of quantitative evaluation. We use a benchmark sequence from the MonoPerfCap dataset and video sequences rendered from BUFF dataset to test the performance of our method.

Experiment Setting. We follow previous work to use the Pablo sequence in their dataset to perform quantitative comparison. Surface meshes and 3D joints obtained by a multi-view performance capture method are provided as the ground truth in the dataset. We compare our method with a state-of-the-art template-based performance capture method Monocular capture results are provided by the author. and single-image human reconstruction methods . Body pose is not estimated in , so we apply our estimated body pose to their T-pose results.

Evaluation of Clothing Surface Reconstruction. We first evaluate our method using a surface reconstruction metric. Due to the intrinsic depth-scale ambiguity of single-view reconstruction, we compute a global scaling factor from our result to the ground truth, which is applied to our result before comparison. Following we align our results to the ground truth with a translation to eliminate the global depth offsets. We compute the average point-to-surface distance from all the ground truth vertices in the clothing region to the output mesh as the evaluation metric. The clothing region (the T-shirt and shorts) is obtained by manual segmentation on the ground truth surface mesh. The same procedure is applied to all the methods under evaluation. A visualization of our results is shown in Fig. 4.

We report the mean surface error averaged across all frames in the middle column of Table 1. For the visualization of per-frame error curves please refer to our supplementary material. Our method achieves a significantly lower surface error compared to all previous single-image surface reconstruction methods. Our performance even comes close to the template-based tracking method which requires a pre-scanned personalized template that provides strong prior information about the body and clothing shape of the subject. By contrast, our method does not require a pre-processed template, and therefore can be applied to a wider range of videos.

Evaluation of 3D Pose Estimation. Although body pose estimation is not a focus of this paper, we follow to validate our method on the metric of 3D joint error on the Pablo sequence. Average per-joint 3D position error after alignment with translation is reported in Table 1 (right). Our method achieves an error of 77.377.3 mm, significantly lower than 118.7118.7 mm in . This verifies the effectiveness of our body pose initialization that utilizes various image measurements including 2D joints, dense correspondences, silhouette, etc.

2 Evaluation on BUFF Dataset

Experiment Setting. BUFF is a dataset of high-resolution 4D textured scan sequences of five people. In this experiment, we sample a test sequence from the BUFF dataset (00096-shortlong_hips, first 200 frames) and train a pair of upper and lower clothing models with the data of four other people. We render the sequence from three views: front, left and front-left, as visualized in Fig. 5. We evaluate our method in four stages: body initialization, sequential tracking, batch optimization and wrinkle extraction The evaluation protocol is the same as Section 6.1: we rigidly align the estimated and ground truth meshes with a global scaling and translation, and compute average distance from ground truth clothing vertices to our results.

Results. The quantitative results are shown in Table 2. First, from all three viewpoints, results with clothing consistently achieve lower reconstruction error than body only. This verifies that our method captures clothing shape that cannot be explained by the SMPL body shape space. Second, we can see that temporal smoothing and wrinkle extraction, which improve the visual quality as shown in qualitative results, have little influence on the reconstruction error. Third, our results show similar range of error in the clothing region across different views, implying that our method is not very sensitive to the viewpoint variation.

Qualitative Evaluation

We qualitatively evaluate our method on various videos including public benchmark and in-the-wild videos where no pre-scanned template is available. Example results are shown in Figure 6. Please see our supplementary video for full results and qualitative comparison with other work.

As shown in the supplementary video, our result not only demonstrates better temporal robustness than the single-image 3D human reconstruction methods in terms of reconstructed surfaces, but also provides 3D temporal correspondences effectively shown by the re-rendering of our output mesh with a consistent texture map. This is hard to obtain by methods that regress 3D shape in voxels , depth maps or implicit functions . Template-based monocular performance capture methods rely heavily on non-rigid surface regularization such as As-Rigid-As-Possible (ARAP), which often prevents those methods from capturing natural dynamics of the clothing deformation. In comparison, our method is able to capture more realistic dynamics of the garment with regularization provided by the clothing models.

In addition, we perform extensive ablation studies on various loss terms used in our pipeline. Please refer to our supplementary document for the results.

Conclusion and Future Work

In this paper, we have presented a method to capture temporally coherent dynamic deformation of clothing from a monocular video. To the best of our knowledge, we have shown the first result of temporally coherent clothing capture from a monocular RGB video without using a pre-scanned template. Our results on various in-the-wild videos endorse the effectiveness and robustness of our method.

Our method is limited by the types of garments in the available training data. We have demonstrated results on several types of tight clothing. Treatment of free-flowing garments like skirts requires collection of more data and additional design of the clothing model. Our method is constrained in the ability to capture drastically changing deformations due to the limited expressiveness of our models, which may be addressed by using higher-capacity models like a deep neural network. We also would like to further incorporate physics into the clothing models to enable more physically realistic clothing capture.

Acknowledgements. We would like to thank Eric Yu for his help with the rendering of our results using Blender.

References

Appendix A Further Ablation Studies

In this section, we conduct more ablation studies on various loss terms we use in the energy optimization for clothing capture and body shape estimation.

We first study the loss terms used for clothing capture in Section 5.2 (Eq. 12). In the experiments below, we compare the results of the batch optimization stage with different loss terms, initialized from the same body capture and sequential tracking results.

Clothing segmentation term (Eq. 13). In order to study the effect of the clothing segmentation term, we run an ablative experiment where the weight for the segmentation term is set to , while all other terms remain the same. To better visualize the effect, we render the output meshes in three colors: grey for skin, yellow for upper clothing and green for lower clothing. We consider a vertex jj as a skin vertex if the length of the clothing offset for this vertex is below a certain threshold ε\varepsilon, or

where DjD_{j} is defined in Eq. 4 in the main paper. We consider a vertex as belonging to the upper clothing if

or, similarly, as belong to the lower clothing if

The result of this experiment is shown in Fig. 7. In each frame, we observe that the boundary between the upper and lower clothing is more consistent with the original image in the result with segmentation term than the result without segmentation term. Our method adopts a combination of upper clothing and lower clothing models, which might both have non-zero offsets around the body waist. It is important for our method to produce both offsets with correct relative length to realistically reconstruct the spatial arrangement of the T-shirt and trousers in the original images. This result proves the effectiveness and necessity of the clothing segmentation term.

Photometric tracking term (Eq. 14). Similarly, we run an ablative experiment where the weight for the photometric tracking term is set to and other terms remain the same. To visualize its effect, we render the output tracked mesh with the final texture extracted in the sequential tracking stage (see Section 5.2 of the main paper for detail), and compare the results with and without the photometric tracking term with the original images.

The result of this experiment is shown in Fig. 8. Notice that the same final texture image is used to render all the results. In order to assist visual comparison of the rendered pattern, we draw several auxiliary horizontal dashed lines in red. We can observe that the results with photometric tracking term is more consistent with the original image than the result without photometric tracking term, in terms of the location of the white strip on the T-shirt and the boundary between the T-shirt and trousers. This demonstrates that our photometric tracking loss can help to obtain better temporal correspondence across different frames in the video.

Silhouette matching term. We now compare the results with and without the silhouette matching term. We render both results and align them with the original images to visualize how well the silhouette matches.

The result of this experiment is shown in Fig. 9. We observe that the result with silhouette matching term achieves a better alignment of silhouette with the original image. This suggests that the silhouette matching term can help to reconstruct the accurate shape of the clothing in the video.

A.2 Losses Terms for Body Shape Estimation

Although body pose and shape estimation is not a focus of this paper, we conduct ablative studies on the loss terms used in body shape estimation in Section 5.1. (Eq. 9). In each of the experiment in this section, the weight for the loss term under study is set to , and all other terms stay the same as the full results. We render the estimated body shapes and compare them with the full results.

Silhouette term. The result of this experiment is shown in Fig. 10. We can observe in the result that silhouette provides critical information for the estimation of body shape and pose in the following two ways. First, the projection of human body should always lie in the interior of the overall silhouette in the image, which includes the region of body and clothes. Second, in the top-right and bottom-left examples, an arm of the subject is occluded by the torso. There is no available information to reason about the location of the arm from the 2D keypoints or DensePose results. In this situation, only the silhouette can constrain the position of the arm to be behind the torso in the camera view. This proves the importance of the silhouette term for accurate estimation of human body and shape.

DensePose term. The result of this experiment is shown in Fig. 11. The use of DensePose together with SMPL model for accurate body estimation was first proposed in . In our work, we find that the DensePose term helps to estimate the hand orientation more accurately, as fingers are usually not included in the hierarchy of 2D body pose output.

POF term. The result of this experiment is shown in Fig. 12. The use of POF together with deformable human body model was first proposed in . We find that the POF term can help to eliminate the ambiguity of 3D body pose given only 2D keypoints in the front view, and therefore help to estimate more accurate body pose in 3D.

Appendix B Quantitative Comparison with Monocular 3D Pose Estimation Methods

In the first stage of our pipeline, we use a standard model-fitting method to estimate 3D body pose from the video. Although we do not claim any contribution or novelty in this aspect, we still provide a quantitative comparison with recent state-of-the-art approaches that estimate 3D body pose with SMPL model from a monocular view. In particular, we evaluate all methods on the Pablo sequence using the same protocol as Section 6.1 in the main paper. The evaluation results are shown in Table 3. As a part of our pipeline, our estimation of 3D body pose is highly accurate even when compared with recent state-of-the-art approaches that focus on 3D body pose only. This lays a solid foundation for the following clothing capture stages.

Appendix C Runtime Analysis

In this section, we present the runtime information of our approach. Our method runs on a Linux server with 40 CPU cores and 4 GTX TITAN X GPUs. Our approach requires the memory of 4 GPUs in order to run the batch optimization on a video of around 250 frames together. For optimization, we use the L-BFGS solver implemented in PyTorch . We measure the average time consumed for each frame in every stage, and the results are shown in Table 4.

Appendix D Complete Quantitative Evaluation Results

In this section, we present the figures for complete per-frame results of the quantitative experiments conducted in Section 6 of the main paper.

Evaluation of Clothing Surface Reconstruction. The complete per-frame results corresponding to the surface error in Table 1 in the main paper are shown in Fig. 13.

Evaluation of 3D Pose Estimation. The complete per-frame results corresponding to the joint error in Table 1 in the main paper are shown in Fig. 14.

D.2 Evaluation on BUFF Dataset

The complete per-frame results corresponding to Table 2 in the main paper are shown in Fig. 15.