Learning-Based Animation of Clothing for Virtual Try-On

Igor Santesteban, Miguel A. Otaduy, Dan Casas

Acknowledgments.

We would like to thank Rosa M. Sánchez-Banderas and Héctor Barreiro for their help in editing the supplementary video, and Gerard Pons-Moll and Sergi Pujades for providing us the ClothCap [PMPHB17] meshes. Igor Santesteban was supported by the Predoctoral Training Programme of the Department of Education of the Basque Government (PRE_2018_1_0307), and Dan Casas was supported by a Marie Curie Individual Fellowship, grant agreement 707326. The work was also funded in part by the European Research Council (ERC Consolidator Grant no. 772738 TouchDesign).

Introduction

Clothing plays a fundamental role in our everyday lives. When we choose clothing to buy or wear, we guide our decisions based on a combination of fit and style. For this reason, the majority of clothing is purchased at brick-and-mortar retail stores, after physical try-on to test the fit and style of several garments on our own bodies. Computer graphics technology promises an opportunity to support online shopping through virtual try-on animation, but to date virtual try-on solutions lack the responsiveness of a physical try-on experience. Beyond online shopping, responsive animation of clothing has an impact on fashion design, video games, and interactive graphics applications as a whole.

One approach to produce animations of clothing is to simulate the physics of garments in contact with the body. While this approach has proven capable of generating highly detailed results [KJM08, SSIF09, NSO12, CLMMO14], it comes at the expense of significant runtime computational cost. On the other hand, it bears no or little preprocessing cost, hence it can be quickly deployed on almost arbitrary combinations of garments and body shapes and motions. To fight the high computational cost, interactive solutions sacrifice accuracy in the form of coarse cloth discretizations, simplified cloth mechanics, or approximate integration methods. Continued progress on the performance of solvers is bringing the approach closer to the performance needs of virtual try-on [TWL∗18].

An alternative approach for cloth animation is to train a data-driven model that computes cloth deformation as a function of body motion [WHRO10, dASTH10]. This approach succeeds to produce plausible cloth folds and wrinkles when there is a strong correlation between body pose and cloth deformation. However, it struggles to represent the nonlinear behavior of cloth deformation and contact in general. Most data-driven methods rely to a certain extent on linear techniques, hence the resulting wrinkles deform in a seemingly linear manner (e.g., with blending artifacts) and therefore lack realism.

Most previous data-driven cloth animation methods work for a given garment-avatar pair, and are limited to representing the influence of body pose on cloth deformation. In virtual try-on, however, a garment may be worn by a diverse set of people, with corresponding avatar models covering a range of body shapes. In this paper, we propose a learning-based method for cloth animation that meets the needs of virtual try-on, as it models the deformation of a given garment as a function of body motion and shape. Other methods that account for changes in body shape do not deform the garment in a realistic way, and either resize the garment while preserving its style [GRH∗12, BSBC12], or retarget cloth wrinkles to bodies of different shapes [PMPHB17, LCT18].

We propose a two-level strategy to learn the complex nonlinear deformations of clothing. On one hand, we learn a model of garment fit as a function of body shape. And on the other hand, we learn a model of local garment wrinkles as a function of body shape and motion. Our two-level strategy allows us to disentangle the different sources of cloth deformation.

We compute both the garment fit and the garment wrinkles using nonlinear regression models, i.e., artificial neural networks, and hence we avoid the problems of linear data-driven models. Furthermore, we propose the use of recurrent neural networks to capture the dynamics of wrinkles. Thanks to this strategy, we avoid adding an external feedback loop to the network, which typically requires a dimensionality reduction step for efficiency reasons [CO18].

Our learning-based cloth animation method is formulated as a pose-space deformation, which can be easily integrated into skeletal animation pipelines with little computational overhead. We demonstrate example animations such as the ones in Figure Learning-Based Animation of Clothing for Virtual Try-On, with a runtime cost of just 44ms per frame (more than 10001000x speed-up over a full simulation) for cloth meshes with thousands of triangles, including collision postprocessing.

To train our learning-based model, we leverage state-of-the-art physics-based cloth simulation techniques [NSO12], together with a parametric human model [LMR∗15] and publicly available motion capture data [CMU, VRM∗17]. In addition to the cloth animation model, we have created a new large dataset of dressed human animations of varying shapes and motions.

Related Work

Physics-based simulation of clothing entails three major processes: computation of internal cloth forces, collision detection, and collision response; and the total simulation cost results from the combined influence of the three processes. One attempt to limit the cost of simulation has been to approximate dynamics, such as in the case of position-based dynamics [BMO∗14]. While approximate methods produce plausible and expressive results for video game applications, they cannot transmit the realistic cloth behavior needed for virtual try-on.

Another line of work, which tries to retain simulation accuracy, is to handle efficiently both internal forces and collision constraints during time integration. One example is a fast GPU-based Gauss-Seidel solver of constrained dynamics [FTP16]. Another example is the efficient handling of nonlinearities and dynamically changing constraints as a superset of projective dynamics [OBLN17]. Very recently, Tang et al. [TWL∗18] have designed a GPU-based solver of cloth dynamics with impact zones, efficiently integrated with GPU-based continuous collision detection.

A different approach to speed up cloth simulation is to apply adaptive remeshing, focusing simulation complexity where needed [NSO12]. Similar in spirit, Eulerian-on-Lagrangian cloth simulation applies remeshing with Eulerian coordinates to efficiently resolve the geometry of sharp sliding contacts [WPLS18].

Data-Driven Models.

model surface deformations as a function of pose. some existing data-driven methods for clothing animation also use the underlying kinematic skeletal model to drive the garment deformation [KV08, WHRO10, GRH∗12, XUC∗14, HTC∗14]. Kim and Vendrovsky [KV08] first introduced a pose-space deformation approach that uses a skeletal pose as subspace domain. Hahn et al. [HTC∗14] went one step further and performed cloth simulation in pose-dependent dynamic low-dimensional subspaces constructed from precomputed data. Wang et al. [WHRO10] used a precomputed database to locally enhance a low-resolution clothing simulation based on joint proximity.

Other methods produce detailed cloth animations by augmenting coarse simulations with example-based wrinkle data. Rohmer et al. [RPC∗10] used the stretch tensor of a coarse animation output as a guide for wrinkle placement. Kavan et al. \shortciteKavan2011 used example data to learn an upsampling operator that adds fine details to a coarse cloth mesh. Zurdo et al. [ZBO13] proposed a mapping between low and high-resolution simulations, employing tracking constraints [BMWG07] to establish a correspondence between both resolutions. More recently, Oh et al.[OLL18] have shown how to train a deep neural network to upsample low-resolution cloth simulations.

A different approach for cloth animation is to approximate full-space simulation models with coarse data-driven models. James and Fatahalian [JF03] used efficient precomputed low-rank approximations of physically-based simulations to achieve interactive deformable scenes. De Aguiar et al. [dASTH10] learned a low-dimensional linear model to characterize the dynamic behavior of clothing, including an approximation to resolve body-cloth collisions. Kim et al. [KKN∗13] performed a near-exhaustive precomputation of the state of a cloth throughout the motion of a character. At run-time a secondary motion graph explored to find the closest cloth state for the current pose. Despite its efficient implementation, the method cannot generalize to new motions. Xu et al. [XUC∗14] used a precomputed dataset to mix and match parts of different samples to synthesize a garment mesh that matches the current pose.

As discussed in the introduction, virtual try-on requires cloth models that respond to changes in body pose and shape. However, this is a scarce feature in data-driven cloth animation methods. Guan et al.\shortciteGuan2012 dressed a parametric character and independently modeled cloth deformations due to shape and pose. However, they relied on a linear model that struggles to generate realistic wrinkles, specially under fast motions. Moreover, they accounted for body shape by resizing the cloth model. Other works also apply a scaling factor to the garment to fit a given shape, without realistic deformation [YFHWW18, PMPHB17, LCT18].

Performance Capture Re-Animation.

Taking advantage of the recent improvements on performance capture methods [BPS∗08, ZPBPM17, PMPHB17], virtual animation of real cloth that has been previously captured (and not simulated) has become an alternative. Initial attempts fit a parametric human model to the captured 3D scan to enable the re-animation of the captured data, without any explicit cloth layer [JTST10, FCS15]. More elaborate methods extract a cloth layer from the captured 3D scan and fit a parametric model to the actor [YFHWW18, NH13, PMPHB17, LCT18]. This allows editing the actor’s shape and pose parameters while keeping the same captured garment or even changing it. However, re-animated motions lack realism since they cannot predict the nonrigid behavior of clothing under unseen poses or shapes, and are usually limited to copying wrinkles across bodies of different shapes [PMPHB17, LCT18].

Image-Based Methods.

Cloth animation and virtual try-on methods have also been explored from an image-based point of view [SM06, ZSZ∗12, HSR13, HFE13, HWW∗18]. These methods aim to generate compelling 2D images of dressed characters, without dealing with any 3D model or simulation of any form. Hilsmann et al. [HFE13] proposed a pose-dependent image-based method that interpolates between images of clothes. More recently, Han et al. [HWW∗18] have shown impressive photorealistic results using convolutional neural networks. However, these image-based methods are limited to 2D static images and fixed camera positions, and cannot fully convey the 3D fit and style of a garment.

Clothing Animation

In this section, we describe our learning-based data-driven method to animate the clothing of a virtual character. Figure 1 shows an overview of the method, separating the preprocessing and runtime stages.

In Section 3.1 we overview the components of our shape-and-pose-dependent cloth deformation model. The two key novel ingredients of our model are: (i) a Garment Fit Regressor (Section 3.2), which allows us to apply global body-shape-dependent deformations to the garment, and (ii) a Garment Wrinkle Regressor (Section 3.3), which predicts dynamic wrinkle deformations as a function of body shape and pose.

We denote as MbM_{\text{b}} a deformed human body mesh, determined by shape parameters β\beta (e.g., the principal components of a database of body scans) and pose parameters θ\theta (e.g., joint angles). We also denote as McM_{\text{c}} a deformed garment mesh worn by the human body mesh. A physics-based simulation would produce a cloth mesh Sc(β,θ)S_{\text{c}}(\beta,\theta) as the result of simulating the deformation and contact mechanics of the garment on a body mesh with shape β\beta and pose θ\theta. Instead, we approximate ScS_{\text{c}} using a data-driven model.

Based on the observation that most garments closely follow the deformations of the body, we design our clothing model inspired by the Pose Space Deformation (PSD) literature [LCF00] and subsequent human body models [ASK∗05, FCS15, LMR∗15]. We assume that the body mesh is deformed according to a rigged parametric human body model,

where R\scaletoG4pt()R_{\scaleto{\text{G}}{4pt}}() and R\scaletoL4pt()R_{\scaleto{\text{L}}{4pt}}() represent two nonlinear regressors, which take as input body shape parameters and shape and pose parameters, respectively.

The final cloth skinning step can be formally expressed as

we define the skinning weight matrix Wc\mathcal{W}_{\text{c}} by projecting each vertex of the template cloth mesh onto the closest triangle of the template body mesh, and interpolating the body skinning weights Wb\mathcal{W}_{\text{b}}.

The pipeline Figure 1 shows the template body mesh Tˉb\mathbf{\bar{T}}_{\text{b}} wearing the template cloth mesh Tˉc\mathbf{\bar{T}}_{\text{c}} (Figure 1-a), and then the template cloth mesh in isolation (Figure 1-b), with the addition of garment fit (Figure 1-c), with the addition of garment wrinkles (Figure 1-d), and the final deformation after the skinning step (Figure 1-e).

By training regressors with collision-free data, our data-driven model learns naturally to approximate contact interactions, but it does not guarantee collision-free cloth outputs. In particular, when the garments are tight, interpenetrations with the body can become apparent. After the skinning step, we apply a postprocessing step to cloth vertices that collide with the body, by pushing them outside their closest body primitive. An example of collision postprocessing is shown in Figure 2.

2 Garment Fit Regressor

Our learning-based cloth deformation model represents corrective displacements on the unposed cloth state, as discussed above. We observe that such displacements are produced by two distinct sources. On one hand, the shape of the body produces an overall deformation in the form of stretch or relaxation, caused by tight or oversized garments, respectively. As we show in this section, we capture this deformation as a static global fit, determined by body shape alone. On the other hand, body dynamics produce additional global deformation and small-scale wrinkles. We capture this deformation as time-dependent displacements, determined by both body shape and motion, as discussed later in Section 3.3. We reach higher accuracy by training garment fit and garment wrinkles separately, in particular due to their static vs. dynamic nature.

where Sc(β,0))S_{\text{c}}(\beta,\mathbf{0})) represents a simulation of the garment on a body with shape β\beta and pose θ=0\theta=\mathbf{0}, and ρ\rho represents a smoothing operator.

See Figure 3(a) for a visualization of the garment fit regression. Notice how the original template mesh is globally deformed but lacks pose-dependent wrinkles.

3 Garment Wrinkle Regressor

See Figure 3(b) for a visualization of the garment wrinkle regression. Notice how the garment obtained in the first step of our pipeline is further deformed and enriched with pose-dependent dynamic wrinkles.

Training Data and Regressor Settings

In this section, we give details on the generation of synthetic training sequences and the extraction of ground-truth data to train the regressor networks. In addition, we discuss the network settings and the hyperparameters used in our results.

To produce ground-truth data for the training of the Garment Fit Regressor and the Garment Wrinkle Regressor, we have created a novel dataset of dressed character animations with diverse motions and body shapes. Our prototype dataset has been created using only one garment, but it can be applied to other garments or their combinations.

As explained in Section 3.1, our approach relies on the use of a parametric human model. In our implementation, we have used SMPL [LMR∗15]. We have selected 1717 training body shapes, as follows. For each of the 44 principal components of the shape parameters β\beta, we generate 44 samples, leaving the rest of the parameters in β\beta as . To these 1616 body shapes, we add the nominal shape with β=0\beta=0.

As animations, we have selected character motions from the CMU dataset [CMU], applied to the SMPL body model [VRM∗17]. Specifically, we have used 5656 sequences containing 7,1177,117 frames in total (at 3030 fps, downsampled from the original CMU dataset of 120120 fps). We have simulated each of the 5656 sequences for each of the 1717 body shapes, wearing the same garment mesh (i.e., the T-shirt shown throughout the paper, which consists of 8,7108,710 triangles).

All simulations have been produced using the ARCSim physics-based cloth simulation engine [NSO12, NPO13], with remeshing turned off to preserve the topology of the garment mesh. ARCSim requires setting several material parameters. In our case, since we are simulating a T-shirt, we have chosen an interlock knit with 60%60\% cotton and 40%40\% polyester, from a set of measured materials [WRO11]. We have executed all simulations using a fixed time step of 3.33ms, with the character animations running at 3030 fps and interpolated to each time step. We have stored in the output database the simulation results from 1 out of every 10 time steps, to match the frame rate of the character animations. This produces a total of 120,989120,989 output frames of cloth deformation.

ARCSim requires a valid collision-free initial state. To this end, we manually pre-position the garment mesh once on the template body mesh Tˉb\mathbf{\bar{T}}_{\text{b}}. We run the simulation to let the cloth relax, and thus define the initial state for all subsequent simulations. In addition, we apply a smoothing operator ρ(⋅)\rho(\cdot) to this initial state to obtain the template cloth mesh Tˉc\mathbf{\bar{T}}_{\text{c}}.

The generation of ground-truth garment fit data requires the simulation of the garment worn by unposed bodies of various shapes. We do this by incrementally interpolating the shape parameters from the template body mesh to the target shape, while simulating the garment from its collision-free initial state. Once the body reaches its target shape, we let the cloth rest, and we compute the ground-truth garment fit displacements Δ\scaletoG4pt\scaletoGT4pt\Delta^{\scaleto{\text{GT}}{4pt}}_{\scaleto{\text{G}}{4pt}} according to Equation 4.

Similarly, to simulate the garment on animations with arbitrary pose and shape, we incrementally interpolate both shape and pose parameters from the template body mesh to the shape and initial pose of the animation. Then, we let the cloth rest before starting the actual animation. The simulations produce cloth meshes Sc(β,θ)S_{\text{c}}(\beta,\theta), and from these we compute the ground-truth garment wrinkle displacements Δ\scaletoL4pt\scaletoGT4pt\Delta^{\scaleto{\text{GT}}{4pt}}_{\scaleto{\text{L}}{4pt}} according to Equation 5.

2 Network Implementation and Training

We have implemented the neural networks presented in Sections 3.2 and 3.3 using Tensorflow [AAB∗15]. The MLP network for garment fit regression contains a single hidden layer with 20 hidden neurons, which we found enough to predict the global fit of the garment. The GRU network for garment wrinkle regression also contains a single hidden layer, but in this case we obtained the best fit of the test data using 1500 hidden neurons. In both networks, we have applied dropout regularization to avoid overfitting the training data. Specifically, we randomly disable 20%20\% of the hidden neurons on each optimization step. Moreover, we shuffle the training data at the beginning of each training epoch.

During training, we use the Adam optimization algorithm [KB14] for 2000 epochs with an initial learning rate of 0.0001. For the garment fit MLP network, we use for training the ground-truth data from all 1717 body shapes. For the garment wrinkle GRU network, we use for training the ground-truth data from 5252 animation sequences, leaving 44 sequences for testing purposes. When training the GRU network, we use a batch size of 128. Furthermore, to speed-up the training process of the GRU network, we compute the error gradient using Truncated Backpropagation Through Time (TBPTT), with a limit of 9090 time steps.

Evaluation

In this section, we discuss quantitative and qualitative evaluation of the results obtained with our method. We compare our results with other state-of-the-art methods, and we demonstrate the benefits of our method for virtual try-on, in terms of both visual fidelity and runtime performance.

We have implemented our method on an Intel Core i7-6700 CPU, with a Nvidia Titan X GPU with 32GB of RAM. Table 1 shows average per-frame execution times of our implementation including garment fit regression, garment wrinkle regression, and skinning, with and without collision postprocessing. For reference, we also include simulation timings of a CPU-based implementation of full physics-based simulation using ARCSim[NSO12].

The low computational cost of our method makes it suitable for interactive applications. Its memory footprint is as follows: 1.1MB for the Garment Fit Regressor MLP, and 108.1MB for the Garment Wrinkle Regressor GRU, both without any compression.

2 Quantitative Evaluation

In Figure 4, we compare the fitting quality of our nonlinear regression method vs. linear regression (implemented using a single-layer MLP neural network ), on a training sequence. While our method retains the rich and history-dependent wrinkles, linear regression suffers smoothing and blending artifacts.

Generalization to new body shapes.

In Figure 5, we quantitatively evaluate the generalization of our method to new shapes (i.e., not in the training set). We depict the per-vertex mean error on a static pose (top) and a dynamic sequence (bottom), as we change the body shape over time. To provide a quantitative comparison to existing methods, we additionally show the error suffered by cloth retargeting [LCT18, PMPHB17]. As discussed in Section 2, such retargeting methods scale the garment in a way analogous to the body to retain the garment’s style. As we show in the accompanying video, even if retargeting produces appealing results, it does not suit the purpose of virtual try-on, and produces larger error w.r.t. a physics-based simulation of the garment. This is clearly visible in Figure 5, where the error with retargeting increases as the shape deviates from the nominal shape, while it remains stable with our method.

Generalization to new body poses.

In Figure 6, we depict the per-vertex mean error of our method in 2 test motion sequences with constant body shape but varying pose. In particular, we validate our cloth animation results on the CMU sequences 01_01 and 55_27 [CMU], which were excluded from the training set, and exhibit complex motions including jumping, dancing and highly dynamic arm motions. Additionally, we show the error suffered by two baseline methods for cloth animation. On one hand, Linear Blend Skinning (LBS), which consists of applying the kinematic transformations of the underlying skeleton directly to the garment template mesh. On the other hand, a Linear Regressor (LR) that predicts cloth deformation directly as a function of pose, implemented using a single-layer MLP neural network without nonlinear activation function. The results demonstrate that our two-step approach, with separate nonlinear regression of garment fit and garment wrinkles, outperforms the linear approach. This is particularly evident in the accompanying video, where the linear regressor exhibits blending artifacts.

3 Qualitative Evaluation

In Figure 7, we show the clothing deformations produced by our approach on a static pose while changing the body shape over time. We compare results with a physics-based simulation and with retargeting techniques [LCT18, PMPHB17]. Notice how our method successfully reproduces ground-truth deformations, including the overall drape (i.e., how the T-shirt slides up the belly due to the stretch caused by the increasingly fat character) and mid-scale wrinkles.

We also compare our method to state-of-the-art data-driven methods that account for changes in both body shape and pose. Figure 8 shows the result of DRAPE [GRH∗12] when the same garment is worn by two avatars with significantly different body shapes. DRAPE approximates the deformation of the garment by scaling it such that it fits the target shape, which produces plausible but unrealistic results. In contrast, our method deforms the garment in a realistic manner.

In Figure 9, we compare our model to ClothCap [PMPHB17]. However, their retargeting lacks realism because cloth deformations are simply copied across different shapes. In contrast, our method produces realistic pose- and shape-dependent deformations.

Generalization to new poses.

We visually evaluate the quality of our model in Figure 10, where we compare ground-truth physics-based simulation and our data-driven cloth deformations on a test sequence. The overall fit and mid-scale wrinkles are successfully predicted using our data-driven model, with a performance gain of three orders of magnitude. Similarly, in Figure 11, we show more frames of a test sequence. Notice the realistic wrinkles in the belly area that appear when the avatar crouches. Please see the accompanying video for animated results and further comparisons.

Conclusions and Future Work

We have presented a novel data-driven method for animation of clothing that enables efficient virtual try-on applications at over 250 fps. Given a garment template worn by a human model, our two-level regression scheme independently models two distinct sources of deformation: garment fit, due to body shape; and garment wrinkles, due to shape and pose. We have shown that this strategy, in combination with the ability of the regressors to represent nonlinearities and dynamics, allows our method to overcome the limitations of previous data-driven approaches.

We believe our approach makes an important step towards bridging the gap between the accuracy and flexibility of physics-based simulation methods and the computational efficiency of data-driven methods. Nevertheless, there are a number of limitations that remain open for future work.

First, our method requires independent training per . In particular, Moreover, garment animations would not capture correctly the interactions between garments. Mix-and-match virtual try-on requires training each possible combination of test garments.

Second, collisions are not fully handled by our method. Our regressors are trained with collision-free data, and therefore our model implicitly learns to approximate contact, but it is not guaranteed to be collision-free. Future work could address this limitation by imposing low-level collision constraints as an explicit objective for the regressor.

Our results show that our method succeeds to predict the overall drape and mid-scale wrinkles of garments, but it smooths excessively high-frequency wrinkles, both spatially and temporally. We wish to investigate alternative methods of recursion to handle accurately both history-dependent draping and highly dynamic wrinkles.

Finally, our model is rooted the assumption that most garments follow closely the body. This assumption may not be valid for loose clothing, and the decomposition of the deformation into a static fit and dynamic wrinkles would not lead to accurate results. It remains to test our method under such conditions.

References