SelfRecon: Self Reconstruction Your Digital Avatar from Monocular Video
Boyi Jiang, Yang Hong, Hujun Bao, Juyong Zhang
Introduction
Clothed body reconstruction has been an important and challenging research topic in the community for years. In the film and gaming industry, high-fidelity human reconstruction usually requires pre-captured templates, multi-camera systems, controlled studios, and long-term works of talented artists. However, these requirements exceed the application scenarios of general customers, such as personalized avatars for telepresence, AR/VR, anthropometry, and virtual try-on, etc. Therefore, directly reconstruction high-fidelity digital avatar from monocular video will have significant practical application value.
The state-of-the-art marker-less monocular human performance capture approaches are mainly designed based on explicit mesh representation. They require actor-specific rigged templates and utilize detected 2D/3D joints and silhouettes to estimate per frame’s posture and non-rigid deformation. DeepCap additionally uses multi-view information during training to resolve deep ambiguity and improve tracking accuracy for monocular inference. The explicit representation has some advantages, including space-time coherence and compatibility with existing graphics control pipelines, like texture editing and reposing. Moreover, skinning deformation is suitable under this paradigm to model the body’s large-scale articulated deformations. However, actor-specific templates limit the extension of these methods to unseen human sequences. For videos of self-rotation humans under rough A-pose, VideoAvatar can estimate general clothed humans with the SMPL+D parametric representation , while it can not recover folds and loose clothing, like skirts.
Recently, some neural implicit representation based monocular human reconstruction approaches have demonstrated compelling results . These methods can handle various topologies, and thus can represent various clothing and hairstyles. However, they require high-quality 3D data for supervision, and they only reconstruct for a specific frame and can not keep the space-time coherence of surface vertices for the whole sequence. A simple solution to guarantee the coherence and correct body structure is to maintain an implicit template surface in the canonical space, and then utilize backward deformation fields to map current points to canonical space to assist their implicit function queries. The backward deformation strategy has been widely applied recently and works well for small-scale deformations . However, it is not very suitable for articulated skinning deformation due to its irreversibility in some parts of current space . To this end, technologies such as pose-related skinning weights prediction and specific inverse articulated deformation design are proposed at the cost of high complexity and poor generalization.
In this work, we propose SelfRecon, which combines the explicit and implicit representations together to reconstruct high-fidelity digital avatar from a monocular video. Specifically, SelfRecon utilizes a learnable signed distance field (SDF) rather than a template with fixed topology to represent the canonical shape. To improve the generalization of the deformation and reduce the optimization difficulty, we adopt the forward deformation to map canonical points to the current frame space . During optimization, we periodically extract the explicit canonical mesh and warp it to each frame with the deformation fields. For these meshes, we utilize mask loss and smooth constraints to recover the overall shape. For the implicit part, a differential formulation is designed to intersect the deformed surface and follow IDR’s neural rendering to refine the geometry. A consistency loss is designed to match both geometric representations as close as possible.
SelfRecon alleviates the dependence on actor-specific templates and extracts a space-time coherent mesh sequence from a monocular video. Extensive evaluations on self-rotating human videos demonstrate that it outperforms existing methods. We believe that SelfRecon will inspire more studies on combining implicit and explicit representations for 3D reconstruction for articulated object.
Related Work
Implicit Human Reconstruction. PIFu adopts a deep network to extract image features and concatenates pixel’s feature and its corresponding 3D point depth information as the input of a Multi-Layer Perceptron (MLP) to obtain high-fidelity 3D clothed human occupancy field. However, it may generate incorrect body structures for humans under challenging poses. StereoPIFu aims at binocular images, utilizes volume alignment feature and predicted high-precision depth to guide implicit function prediction, can effectively alleviate the depth ambiguity and restore absolute scale information. PIFuHD utilizes higher resolution features and predicted normal information to refine the geometric details of PIFu. PaMIR utilizes parameterized human body to decrease the influence of deep ambiguity in implicit function training, reduces the occurrence of abnormal human body structure, and improves the reconstruction accuracy. These methods train an MLP to represent the human’s implicit geometry from single or several images and achieve impressive results. However, they require the corresponding high-quality 3D data of color images to train the model, which is hard to obtain and thus limits their generalization to in-the-wild images.
Besides, overfitting the implicit neural representation of a person’s movement sequence to acquire actor-specific reconstructions becomes popular. NASA coarsely models the naked body as the union of articulated parts, and each part is an implicit occupancy field. SCANimate proposes an end-to-end trainable framework that turns raw 3D scans of a clothed human into an animatable avatar. SNARF learns a forward deformation field to improve its generalization for unseen human poses. All these methods need 4D scan data to train their clothed body representation, and thus are difficult to be widely used for general image data.
Recently, some implicit representation methods, which can extract geometry and synthesize novel views based on multi-view images, attract researchers’ attention. NeuralBody reconstructs per frame’s NeRF field conditioned at body structured latent codes and utilizes the NeRF field to synthesize new images. However, the extracted geometry from NeRF suffers from noise. H-NeRF utilizes an implicit parametric model to reconstruct the temporal motion of humans. Neural Actor integrates texture map features to refine volume rendering. IDR combines implicit signed distance field and differential neural rendering to generate high-quality rigid reconstruction from multi-view images. Concurrent IMAvatar expand IDR to learn implicit head avatars from monocular videos.
Explicit Human Reconstruction. With the help of human statistical model , some works utilize image cues to automatically obtain model parameters . To represent human clothing, some methods add displacements on SMPL vertices to model tight clothing . However, this SMPL+D representation can only support tight clothing types and recover coarse level geometry shape. To improve the representation ability, some works adopt separate clothing representation and combine with SMPL body to do reconstruction , but they need clothing type and high-quality 3D supervision.
Besides, to capture the performance of a specific person, many prior works use an actor-specific template to assist tracking. Monoperfcap optimizes the deformation of template mesh to match 2D cues. LiveCap refines the optimization pipeline and achieves real-time tracking for a specific person with monocular RGB input. DeepCap adopts a network to predict per frame’s template deformation for a specific person. However, the requirement for pre-defined templates limits their broader applications.
Method
SelfRecon aims to reconstruct a high-fidelity and space-time coherent clothed body shape from a monocular video depicting a self-rotating person, and the whole algorithm pipeline is given in Fig. 1. Both explicit and implicit geometric representations are utilized to achieve the above target. Specifically, we utilize the forward deformation field to generate space-time coherent explicit meshes. The deformation fields are decomposed into two parts, where the first one represents per frame’s non-rigid deformation with a learnable MLP, and the second is the skinning deformation field. Differentiable masks, regular and smooth losses are adopted to control the shape of explicit meshes. To update the shape of implicit neural representation, we use non-rigid ray-casting (sec 3.3) to find the differentiable intersection points of rays and the deformed implicit surface. Then, the implicit rendering network (sec 3.4) will utilize the rays’ color information to improve the geometry. Unless otherwise indicated, we also utilize the predicted normal map to refine the details. Finally, a consistency loss is designed to match both representations.
For a self-rotating video with frames, we adopt the method described in VideoAvatar to generate the initial shape parameter and per-frame’s pose parameters of SMPL model. We pre-defined a template pose and generate initial canonical SMPL body mesh with and this pose parameter. Our implicit and explicit representations are both initialized with . In the following, we present the algorithm details of each component.
In the similar work of VideoAvatar , they adopt the SMPL+D representation for clothed human body. However, SMPL+D has limited resolution and representation ability, and thus it can not represent high-fidelity geometry shape and various clothing types. In this work, we represent the canonical template shape as the zero isosurface of a SDF, which is expressed by an MLP with learnable weights :
To avoid unexpected solution, we use IGR to initialize as the initial canonical body .
2 Deformation Fields
Following prior works , we utilize skeleton skinning transformation to control human body’s large-scale movements due to the articulated structure. However, garments’ non-rigid deformation cannot be fully represented by skinning transformation. Therefore, we extend to model non-rigid deformation with another MLP.
Non-rigid Deformation Field. We use an MLP with learnable weights to represent the non-rigid deformation field. For -th frame, takes its optimizable conditional variable as input and deform points in the canonical space with -th frame’s specific non-rigid deformation.
Skinning Transformation Field. Given -th frame’s pose parameter , we have to define a canonical-to-current space skinning transformation field . As the initial template body has well-defined skinning weights relate to its SMPL skeleton, an intuitive idea is to expand the skinning weights of ’s vertices to the whole canonical space to define the skinning transformation field. Specifically, we first pre-defined a sparse grid containing in the canonical space. For each grid point, we find its nearest vertices on and average their skinning weights with IDW (inverse distance weight) as its initial weight. Then, we smooth all grid points’ weights with Laplace smoothing. Finally, given a point in the canonical space, we compute its skinning weights by trilinear interpolation in the grid. During our optimization, the grid is pre-computed and fixed. This forward deformation design avoids the trouble of inverse skinning transformation and provides a regular constrain for human articulated movement.
Finally, by compositing and , we get the final deformation field . It takes -th frame’s conditional variable and SMPL pose parameter as input, and transform canonical points to the -th frame space. For brevity of description, we use to denote -th frame’s deformation field, for -th frame’s zero isosurface and for ’s optimizable parameters .
3 Differentiable Non-rigid Ray-casting
For rigid scenes, the sphere tracing algorithm is widely used to find the intersection point of a ray and the SDF. However, it is not feasible here due to the deformation fields. Inspired by the method in , which proposes a strategy to render a deformed SDF, we utilize the explicit mesh to help find the intersection point of a ray and .
As shown in Fig 2, we extract an explicit template mesh from the canonical surface . With deformation , we can get -th frame’s mesh . Theoretically, is a piecewise linear approximation of . Therefore, consider a ray emitted from the camera position along the direction , its first intersection with is a good approximation of its intersection with . Moreover, with the intersected triangle on , we can find ’s corresponding point on the template by consistent barycentric weights. Obviously, is close to and is a good approximation of . With as good initialization, we can find a point on , whose deformed point is exactly the intersection point of the ray and . Specifically, we solve by:
where the first item constrains to be close with and the second item restricts on the ray. In our implementation, we set and execute gradient descent iterations to solve . To guarantee accuracy, we reject those samples with large losses.
Differentiable Formula. The above-mentioned solving process of is an iterative optimization process, which is not differentiable. For the ray in -th frame, camera position , view direction , ’s parameters and ’s parameter uniquely determine the . Therefore, can be seen as a function of these parameters, and we need to compute partial derivatives of to all these parameters. For brevity, we only clarify our calculation for here, and other partial derivatives are computed similarly.
Through the above analysis, satisfies the surface and ray constraints: and . We differentiate these two equations w.r.t to get:
where and is ’s cross product matrix. We concatenate these two equations to get a linear system, then is computed by solving its normal equation.
4 Implicit Rendering Network
IDR proposes an MLP to approximate the rendering equation and demonstrates certain disentangle ability of lighting and material. In their rigid configuration, takes the zero isosurface’s point, its normal, the view direction and its global geometry feature vector as input to estimate the point’s color along the view direction. We subtly transfer their design to non-rigid scenarios by converting related current frame’s attributes to canonical space. As shown in Fig. 2, considering a ray emitted from camera center along a sampled pixel, whose direction is determined by camera’s intrinsic parameters , we compute its intersection point on with the algorithm described in sec 3.3. In the meantime, we compute its normal by gradient calculation. Then, the view direction in canonical space can be computed by transferring with the Jacobian matrix of the deformed point w.r.t . As for the global geometry feature, we similarly use a larger MLP to additionally compute it, which implies the geometry information around and can be used to help the prediction of global shadow . Finally, we use an MLP with learnable weights to compute ’s color , formulated as:
where related symbols have been described above. It can be seen that the color along direction of the deformed point in -th frame is determined by the MLP weights and , camera parameters and deformation field parameters .
5 Loss Function
According to the description above, for a -frame self-rotating video, the set of all optimizable parameters is:
which includes camera parameters, the learnable weights of MLPs shared by the whole sequence and per frame’s specific pose parameters and non-rigid deformation field’s conditional variable. Our target is to design a loss function and optimize to match the mask and RGB images of the input video. Besides, a predicted normal map is added to the optimization. As SelfRecon maintains both explicit and implicit geometry, the loss terms can be divided into two parts.
During the computation of explicit losses, we temporarily regard the canonical mesh vertices as an optimizable variable and compute its gradient together with . Then in the consistency loss, we associate its variations with our implicit representation. At present, explicit losses mainly include mask loss, deformation regularization loss, and the smoothness loss of the skeleton.
Mask Loss. We utilize a differentiable renderer based on point cloud to render the mask of -th frame’s mesh with camera parameters, and compute the IoU loss with target mask :
where and are the operators that perform element-wise product and sum respectively.
Deformation Regularization Loss. As stated in Sec. 3.2, the -th frame’s deformation field contains variable and fixed . represents the deformation that can not be represented by skinning transformation , and this deformation should be relatively small. To associate the skeleton pose, we design the following regularization loss:
where is a vertex coordinate of , is the vertices number of , and is the Geman-McClure robust loss .
Finally, the loss for the explicit representation is:
and adjust the weights of related losses. After each iteration, we reserve ’s gradients and wait for the implicit loss iteration to update together. For canonical mesh vertices, we use SGD to update to , which will be used in consistency loss to match two representations.
5.2 Implicit Loss
We sample pixels within the ground truth mask and utilize Sec. 3.3 to get the ray’s intersection on and its corresponding ground truth color and predicted normal if available. Then, based on this sampled points set P, we construct two losses.
Color Loss. By referring to Eq. (4), we formulate the color loss as:
Here, we use to substitute related parameters in Eq. (4). Intuitively, this loss requires that the rendered images should match the input images.
Normal Loss. We utilized the predicted normal map by PIFuHD to further refine the geometry shape. Referring to Eq. (4), we can easily compute ’s normal . Besides, we need to transform the corresponding predicted normal from the space of current frame to canonical space, which can be computed with , where is the Jacobian matrix of the forward deformation field at . Thus, the normal loss is:
Here, means to normalize the vector. is the weight defined by the cosine of angle between and corresponding view direction. Since the predicted normals are noisy and inconsistent between frames, we use this weights to alleviate the impact of normal that deviates from its view direction and avoid geometry artifacts.
We also design regular losses for the implicit representation, and these losses are defined on the set of sampled points S near the implicit surface .
Rigidity Loss. We require the first deformation field to be as rigid as possible to avoid distortion. Following Park et al. , we design our loss as:
where is the singular value diagonal matrix of the Jacobian of on and is the robust function .
Eikonal Loss. We adopt the regular loss of IGR to make to be sign distance function:
where is obtained by differentiating at .
Finally, the implicit loss can be represented as:
where , and are balancing weights.
5.3 Explicit/Implicit Consistency
After explicit iteration, the canonical mesh has been updated to , to make the implicit SDF consistent with the updated explicit mesh during implicit iteration, we design a consistency loss:
where is a vertex coordinate of . Intuitively, the loss requires to match the implicit surface .
In each optimization step, we first perform the explicit iteration to obtain and reserve ’s gradients. Then, we compute implicit and consistency losses to accumulate new ’s gradients. Finally, Adam is utilized to update with computed gradients.
Experiments
We conduct quantitative and qualitative experiments to demonstrate the effectiveness of SelfRecon. For quantitative evaluation, we synthesize several sequences with commercial software due to the lack of high-quality geometry data of human body with general clothing. For qualitative evaluation, we mainly utilize the PeopleSnapshot dataset and several real sequences collected by ourselves. We also present ablation study for the loss term design and present an avatar generation application.
We synthesize data to quantitatively evaluate our reconstruction algorithm. Specifically, we use Blender to design the self-rotating motions for male and female avatars. Then, we utilize CLO3D to design several clothes and animate the clothed body with motions. Finally, we synthesized two sets of male and three sets of female dressing sequences. We reconstruct these sequences with VideoAvatar and our method, and report the registration error for the canonical posture results in Tab. 1. Compared with VideoAvatar, our method significantly reduces the values of various error metrics. In Fig. 3, we also present four group results and their error maps. Intuitively, our results capture the overall shape and have some reasonable details. As VideoAvatar is based on the SMPL+D representation, it has plausible results for tighter clothing, like the male examples, but lacks detailed reconstruction ability. Moreover, it can not correctly reconstruct loose clothing, especially for the females dressed in skirts.
2 Qualitative Evaluation
We also qualitatively compare SelfRecon with multi-frame prediction algorithm PaMIR , optimization method VideoAvatar and NeRF based neural rendering method NeuralBody on several sequences of PeopleSnapshot dataset. In Fig. 4, we present the first frame of input video, our rendered image, and reconstruction results of all methods from two perspectives. And we compare with PaMIR in the first two rows and the others with NeuralBody. As we can see, VideoAvatar based on SMPL+D can only approximately capture the overall shape, but details such as hairstyle and clothing folds are lost. PaMIR uses multi-frame input to improve its results, but still suffers from the deep ambiguity. As in the second example, its reconstructed human is not upright from the side view. Besides, its results have some details but miss facial features, while our method has better details and can recover certain facial features. Similar with SelfRecon, NeuralBody also inputs a video to do self-supervised optimization. It mainly focuses on novel view synthesis, but still can extract geometry from underlying NeRF representation. We can see that its reconstructions alleviate deep ambiguity and conform to the human’s overall structure while suffering from large noise on the surface, which may be caused by excessive freedom of volume rendering. Different from their method, implicit surface representation based SelfRecon can recover high-fidelity geometry shape without noises.
Fig. 5 shows reconstruction results on our collected videos with smartphones. For each group, we present the first frame of the video, our rendered image and reconstruction. Our results have high-fidelity geometry shapes for kinds of clothing and body, and our neural rendered images are also quite close to the input images.
3 Ablation Study
Our complete algorithm requires color images, masks and normal maps as inputs. Fig. 6 shows ablation experiments of two examples on three inputs. As the results show, if only using the mask loss, the recovered geometry shape is inside the convex hull formed by the silhouettes but lacks details and has noticeable concavities. After adding the color loss, it significantly improves the details and reduces the unnatural concavities. For the second example, the result has been very close to the result obtained by adding the normal loss. However, for the first example, it can not completely eliminate the depressed geometry without normal loss. This may be caused by lacking rich texture and multi-view observations in these areas. With the normal loss, our results are further improved, and the unnatural pits are eliminated while the details are preserved.
Since the normal prediction network is trained with synthetic images, its prediction may be not accurate for real tests and may be not consistent for different frames. As shown in Fig. 7, without the adaptive weights in Eq. (10), the normal loss might result unexpted results.
4 Avatar Generation
Thanks to our forward deformation field design, we can extract a mesh sequence with consistent topology. Based on the tracking results, we can extract a texture template mesh from the images with bound skinning weights from the skinning transformation field. Then, an animatable avatar is generated and can be driven with SMPL pose parameters. For texture extraction, we follow the method of VideoAvatar . Fig. 8 shows two examples of texture generation and driving from the PeopleSnapshot dataset. Our method recovers better geometric details like facial, shoes and clothing folds thanks to more accurate tracking results. Besides, our driving results look plausible and may be of sufficient quality for some applications.
Conclusion and Discussion
We proposed SelfRecon, a self-supervised reconstruction method based on neural implicit representation and neural rendering. With forward deformation, our method can be easily applied to body movement and recover space-time coherent surfaces, which is convenient for downstream applications. Moreover, combining the explicit representation, we proposed a non-rigid ray casting algorithm, which makes it possible for differentiable intersecting with the deformed implicit surface. SelfRecon can reconstruct high-fidelity clothed body shape from a self-rotating video without pre-computed templates. We also show high-fidelity avatar generation with our tracking results, demonstrating potential applications of SelfRecon.
SelfRecon still has several limitations. First, it requires relatively long time to optimize, which limits its convenient applications. However, this problem can be alleviated with the help of body priors and the fast growing field of neural rendering. Second, current method relies on the predicted normal maps to improve the geometric details. How to recover the geometric details directly from the self-supervised rendering loss is worthy of future study. Third, the proposed method mainly works well for self-rotating motions, and it is worthy of study for more general motion sequences.
This research was supported by National Natural Science Foundation of China (No. 62122071), the Youth Innovation Promotion Association CAS (No. 2018495), “the Fundamental Research Funds for the Central Universities”(No. WK3470000021).