SelfRecon: Self Reconstruction Your Digital Avatar from Monocular Video

Boyi Jiang, Yang Hong, Hujun Bao, Juyong Zhang

Introduction

Clothed body reconstruction has been an important and challenging research topic in the community for years. In the film and gaming industry, high-fidelity human reconstruction usually requires pre-captured templates, multi-camera systems, controlled studios, and long-term works of talented artists. However, these requirements exceed the application scenarios of general customers, such as personalized avatars for telepresence, AR/VR, anthropometry, and virtual try-on, etc. Therefore, directly reconstruction high-fidelity digital avatar from monocular video will have significant practical application value.

The state-of-the-art marker-less monocular human performance capture approaches are mainly designed based on explicit mesh representation. They require actor-specific rigged templates and utilize detected 2D/3D joints and silhouettes to estimate per frame’s posture and non-rigid deformation. DeepCap additionally uses multi-view information during training to resolve deep ambiguity and improve tracking accuracy for monocular inference. The explicit representation has some advantages, including space-time coherence and compatibility with existing graphics control pipelines, like texture editing and reposing. Moreover, skinning deformation is suitable under this paradigm to model the body’s large-scale articulated deformations. However, actor-specific templates limit the extension of these methods to unseen human sequences. For videos of self-rotation humans under rough A-pose, VideoAvatar can estimate general clothed humans with the SMPL+D parametric representation , while it can not recover folds and loose clothing, like skirts.

Recently, some neural implicit representation based monocular human reconstruction approaches have demonstrated compelling results . These methods can handle various topologies, and thus can represent various clothing and hairstyles. However, they require high-quality 3D data for supervision, and they only reconstruct for a specific frame and can not keep the space-time coherence of surface vertices for the whole sequence. A simple solution to guarantee the coherence and correct body structure is to maintain an implicit template surface in the canonical space, and then utilize backward deformation fields to map current points to canonical space to assist their implicit function queries. The backward deformation strategy has been widely applied recently and works well for small-scale deformations . However, it is not very suitable for articulated skinning deformation due to its irreversibility in some parts of current space . To this end, technologies such as pose-related skinning weights prediction and specific inverse articulated deformation design are proposed at the cost of high complexity and poor generalization.

In this work, we propose SelfRecon, which combines the explicit and implicit representations together to reconstruct high-fidelity digital avatar from a monocular video. Specifically, SelfRecon utilizes a learnable signed distance field (SDF) rather than a template with fixed topology to represent the canonical shape. To improve the generalization of the deformation and reduce the optimization difficulty, we adopt the forward deformation to map canonical points to the current frame space . During optimization, we periodically extract the explicit canonical mesh and warp it to each frame with the deformation fields. For these meshes, we utilize mask loss and smooth constraints to recover the overall shape. For the implicit part, a differential formulation is designed to intersect the deformed surface and follow IDR’s neural rendering to refine the geometry. A consistency loss is designed to match both geometric representations as close as possible.

SelfRecon alleviates the dependence on actor-specific templates and extracts a space-time coherent mesh sequence from a monocular video. Extensive evaluations on self-rotating human videos demonstrate that it outperforms existing methods. We believe that SelfRecon will inspire more studies on combining implicit and explicit representations for 3D reconstruction for articulated object.

Related Work

Implicit Human Reconstruction. PIFu adopts a deep network to extract image features and concatenates pixel’s feature and its corresponding 3D point depth information as the input of a Multi-Layer Perceptron (MLP) to obtain high-fidelity 3D clothed human occupancy field. However, it may generate incorrect body structures for humans under challenging poses. StereoPIFu aims at binocular images, utilizes volume alignment feature and predicted high-precision depth to guide implicit function prediction, can effectively alleviate the depth ambiguity and restore absolute scale information. PIFuHD utilizes higher resolution features and predicted normal information to refine the geometric details of PIFu. PaMIR utilizes parameterized human body to decrease the influence of deep ambiguity in implicit function training, reduces the occurrence of abnormal human body structure, and improves the reconstruction accuracy. These methods train an MLP to represent the human’s implicit geometry from single or several images and achieve impressive results. However, they require the corresponding high-quality 3D data of color images to train the model, which is hard to obtain and thus limits their generalization to in-the-wild images.

Besides, overfitting the implicit neural representation of a person’s movement sequence to acquire actor-specific reconstructions becomes popular. NASA coarsely models the naked body as the union of articulated parts, and each part is an implicit occupancy field. SCANimate proposes an end-to-end trainable framework that turns raw 3D scans of a clothed human into an animatable avatar. SNARF learns a forward deformation field to improve its generalization for unseen human poses. All these methods need 4D scan data to train their clothed body representation, and thus are difficult to be widely used for general image data.

Recently, some implicit representation methods, which can extract geometry and synthesize novel views based on multi-view images, attract researchers’ attention. NeuralBody reconstructs per frame’s NeRF field conditioned at body structured latent codes and utilizes the NeRF field to synthesize new images. However, the extracted geometry from NeRF suffers from noise. H-NeRF utilizes an implicit parametric model to reconstruct the temporal motion of humans. Neural Actor integrates texture map features to refine volume rendering. IDR combines implicit signed distance field and differential neural rendering to generate high-quality rigid reconstruction from multi-view images. Concurrent IMAvatar expand IDR to learn implicit head avatars from monocular videos.

Explicit Human Reconstruction. With the help of human statistical model , some works utilize image cues to automatically obtain model parameters . To represent human clothing, some methods add displacements on SMPL vertices to model tight clothing . However, this SMPL+D representation can only support tight clothing types and recover coarse level geometry shape. To improve the representation ability, some works adopt separate clothing representation and combine with SMPL body to do reconstruction , but they need clothing type and high-quality 3D supervision.

Besides, to capture the performance of a specific person, many prior works use an actor-specific template to assist tracking. Monoperfcap optimizes the deformation of template mesh to match 2D cues. LiveCap refines the optimization pipeline and achieves real-time tracking for a specific person with monocular RGB input. DeepCap adopts a network to predict per frame’s template deformation for a specific person. However, the requirement for pre-defined templates limits their broader applications.

Method

SelfRecon aims to reconstruct a high-fidelity and space-time coherent clothed body shape from a monocular video depicting a self-rotating person, and the whole algorithm pipeline is given in Fig. 1. Both explicit and implicit geometric representations are utilized to achieve the above target. Specifically, we utilize the forward deformation field to generate space-time coherent explicit meshes. The deformation fields are decomposed into two parts, where the first one represents per frame’s non-rigid deformation with a learnable MLP, and the second is the skinning deformation field. Differentiable masks, regular and smooth losses are adopted to control the shape of explicit meshes. To update the shape of implicit neural representation, we use non-rigid ray-casting (sec 3.3) to find the differentiable intersection points of rays and the deformed implicit surface. Then, the implicit rendering network (sec 3.4) will utilize the rays’ color information to improve the geometry. Unless otherwise indicated, we also utilize the predicted normal map to refine the details. Finally, a consistency loss is designed to match both representations.

For a self-rotating video with NN frames, we adopt the method described in VideoAvatar to generate the initial shape parameter β\boldsymbol{\beta} and per-frame’s pose parameters {θi∣i∈{1,...,N}}\{\boldsymbol{\theta}_{i}|i\in\{1,...,N\}\} of SMPL model. We pre-defined a template pose and generate initial canonical SMPL body mesh B\mathcal{B} with β\boldsymbol{\beta} and this pose parameter. Our implicit and explicit representations are both initialized with B\mathcal{B}. In the following, we present the algorithm details of each component.

In the similar work of VideoAvatar , they adopt the SMPL+D representation for clothed human body. However, SMPL+D has limited resolution and representation ability, and thus it can not represent high-fidelity geometry shape and various clothing types. In this work, we represent the canonical template shape Sη\mathcal{S}_{\eta} as the zero isosurface of a SDF, which is expressed by an MLP ff with learnable weights η\eta:

To avoid unexpected solution, we use IGR to initialize Sη\mathcal{S}_{\eta} as the initial canonical body B\mathcal{B}.

2 Deformation Fields

Following prior works , we utilize skeleton skinning transformation to control human body’s large-scale movements due to the articulated structure. However, garments’ non-rigid deformation cannot be fully represented by skinning transformation. Therefore, we extend to model non-rigid deformation with another MLP.

Non-rigid Deformation Field. We use an MLP dd with learnable weights ϕ\phi to represent the non-rigid deformation field. For ii-th frame, dd takes its optimizable conditional variable hi\mathbf{h}_{i} as input and deform points in the canonical space with ii-th frame’s specific non-rigid deformation.

Skinning Transformation Field. Given ii-th frame’s pose parameter θi\boldsymbol{\theta}_{i}, we have to define a canonical-to-current space skinning transformation field W\mathcal{W}. As the initial template body B\mathcal{B} has well-defined skinning weights relate to its SMPL skeleton, an intuitive idea is to expand the skinning weights of B\mathcal{B}’s vertices to the whole canonical space to define the skinning transformation field. Specifically, we first pre-defined a sparse grid containing B\mathcal{B} in the canonical space. For each grid point, we find its nearest 3030 vertices on B\mathcal{B} and average their skinning weights with IDW (inverse distance weight) as its initial weight. Then, we smooth all grid points’ weights with Laplace smoothing. Finally, given a point in the canonical space, we compute its skinning weights by trilinear interpolation in the grid. During our optimization, the grid is pre-computed and fixed. This forward deformation design avoids the trouble of inverse skinning transformation and provides a regular constrain for human articulated movement.

Finally, by compositing dd and W\mathcal{W}, we get the final deformation field D=W(d(⋅))\mathcal{D}=\mathcal{W}(d(\cdot)). It takes ii-th frame’s conditional variable hi\mathbf{h}_{i} and SMPL pose parameter θi\boldsymbol{\theta}_{i} as input, and transform canonical points to the ii-th frame space. For brevity of description, we use Di\mathcal{D}_{i} to denote ii-th frame’s deformation field, Si\mathcal{S}_{i} for ii-th frame’s zero isosurface Di(Sη)\mathcal{D}_{i}(\mathcal{S}_{\eta}) and ψi\psi_{i} for Di\mathcal{D}_{i}’s optimizable parameters {ϕ,hi,θi}\{\phi,\mathbf{h}_{i},\boldsymbol{\theta}_{i}\}.

3 Differentiable Non-rigid Ray-casting

For rigid scenes, the sphere tracing algorithm is widely used to find the intersection point of a ray and the SDF. However, it is not feasible here due to the deformation fields. Inspired by the method in , which proposes a strategy to render a deformed SDF, we utilize the explicit mesh to help find the intersection point of a ray and Si\mathcal{S}_{i}.

As shown in Fig 2, we extract an explicit template mesh T\mathbf{T} from the canonical surface Sη\mathcal{S}_{\eta}. With deformation Di\mathcal{D}_{i}, we can get ii-th frame’s mesh Ti\mathbf{T}_{i}. Theoretically, Ti\mathbf{T}_{i} is a piecewise linear approximation of Si\mathcal{S}_{i}. Therefore, consider a ray emitted from the camera position c\mathbf{c} along the direction v\mathbf{v}, its first intersection x^\hat{\mathbf{x}} with Ti\mathbf{T}_{i} is a good approximation of its intersection with Si\mathcal{S}_{i}. Moreover, with the intersected triangle on Ti\mathbf{T}_{i}, we can find x^\hat{\mathbf{x}}’s corresponding point p^\hat{\mathbf{p}} on the template T\mathbf{T} by consistent barycentric weights. Obviously, p^\hat{\mathbf{p}} is close to Sη\mathcal{S}_{\eta} and is a good approximation of Di−1(x^)\mathcal{D}_{i}^{-1}(\hat{\mathbf{x}}). With p^\hat{\mathbf{p}} as good initialization, we can find a point p\mathbf{p} on Sη\mathcal{S}_{\eta}, whose deformed point x=Di(p)\mathbf{x}=\mathcal{D}_{i}(\mathbf{p}) is exactly the intersection point of the ray r\mathbf{r} and Si\mathcal{S}_{i}. Specifically, we solve p\mathbf{p} by:

where the first item constrains p^\hat{\mathbf{p}} to be close with Sη\mathcal{S}_{\eta} and the second item restricts Di(p^)\mathcal{D}_{i}(\hat{\mathbf{p}}) on the ray. In our implementation, we set ω=3.05\omega=3.05 and execute 1010 gradient descent iterations to solve p\mathbf{p}. To guarantee accuracy, we reject those samples with large losses.

Differentiable Formula. The above-mentioned solving process of p\mathbf{p} is an iterative optimization process, which is not differentiable. For the ray in ii-th frame, camera position c\mathbf{c}, view direction v\mathbf{v}, Di\mathcal{D}_{i}’s parameters ψi\psi_{i} and ff’s parameter η\eta uniquely determine the p\mathbf{p}. Therefore, p\mathbf{p} can be seen as a function of these parameters, and we need to compute partial derivatives of p\mathbf{p} to all these parameters. For brevity, we only clarify our calculation for η\eta here, and other partial derivatives are computed similarly.

Through the above analysis, p\mathbf{p} satisfies the surface and ray constraints: f(p)≡0f(\mathbf{p})\equiv 0 and (Di(p)−c)×v≡0(\mathcal{D}_{i}(\mathbf{p})-\mathbf{c})\times\mathbf{v}\equiv 0. We differentiate these two equations w.r.t η\eta to get:

where x=Di(p)\mathbf{x}=\mathcal{D}_{i}(\mathbf{p}) and [v]×[\mathbf{v}]_{\times} is v\mathbf{v}’s cross product matrix. We concatenate these two equations to get a 4×34\times 3 linear system, then ∂p∂η\frac{\partial\mathbf{p}}{\partial\eta} is computed by solving its normal equation.

4 Implicit Rendering Network

IDR proposes an MLP MM to approximate the rendering equation and demonstrates certain disentangle ability of lighting and material. In their rigid configuration, MM takes the zero isosurface’s point, its normal, the view direction and its global geometry feature vector as input to estimate the point’s color along the view direction. We subtly transfer their design to non-rigid scenarios by converting related current frame’s attributes to canonical space. As shown in Fig. 2, considering a ray emitted from camera center c\mathbf{c} along a sampled pixel, whose direction v\mathbf{v} is determined by camera’s intrinsic parameters τ\tau, we compute its intersection point p\mathbf{p} on Sη\mathcal{S}_{\eta} with the algorithm described in sec 3.3. In the meantime, we compute its normal np=∇f(p;η)\mathbf{n}_{\mathbf{p}}=\nabla f(\mathbf{p};\eta) by gradient calculation. Then, the view direction vp\mathbf{v}_{\mathbf{p}} in canonical space can be computed by transferring v\mathbf{v} with the Jacobian matrix Jx(p)J_{\mathbf{x}}(\mathbf{p}) of the deformed point x=Di(p)\mathbf{x}=\mathcal{D}_{i}(\mathbf{p}) w.r.t p\mathbf{p}. As for the global geometry feature, we similarly use a larger MLP F(p;η)=(f(p;η),z(p;η))F(\mathbf{p};\eta)=(f(\mathbf{p};\eta),\mathbf{z}(\mathbf{p};\eta)) to additionally compute it, which implies the geometry information around p\mathbf{p} and can be used to help the prediction of global shadow . Finally, we use an MLP MM with learnable weights γ\gamma to compute p\mathbf{p}’s color Lp(η,ψi,γ,τ)L_{\mathbf{p}}(\eta,\psi_{i},\gamma,\tau), formulated as:

where related symbols have been described above. It can be seen that the color along direction v\mathbf{v} of the deformed point x\mathbf{x} in ii-th frame is determined by the MLP weights η\eta and γ\gamma, camera parameters τ\tau and deformation field parameters ψi\psi_{i}.

5 Loss Function

According to the description above, for a NN-frame self-rotating video, the set X\mathcal{X} of all optimizable parameters is:

which includes camera parameters, the learnable weights of MLPs shared by the whole sequence and per frame’s specific pose parameters and non-rigid deformation field’s conditional variable. Our target is to design a loss function and optimize X\mathcal{X} to match the mask and RGB images {Oi,Ii∣i∈1,...,N}\{O_{i},I_{i}|i\in{1,...,N}\} of the input video. Besides, a predicted normal map {Ni∣i∈1,...,N}\{N_{i}|i\in{1,...,N}\} is added to the optimization. As SelfRecon maintains both explicit and implicit geometry, the loss terms can be divided into two parts.

During the computation of explicit losses, we temporarily regard the canonical mesh vertices T\mathbf{T} as an optimizable variable and compute its gradient together with X\mathcal{X}. Then in the consistency loss, we associate its variations with our implicit representation. At present, explicit losses mainly include mask loss, deformation regularization loss, and the smoothness loss of the skeleton.

Mask Loss. We utilize a differentiable renderer based on point cloud to render the mask O(Ti)O(\mathbf{T}_{i}) of ii-th frame’s mesh Ti=Di(T)\mathbf{T}_{i}=\mathcal{D}_{i}(\mathbf{T}) with camera parameters, and compute the IoU loss with target mask OiO_{i}:

where ⊗\otimes and ⊕\oplus are the operators that perform element-wise product and sum respectively.

Deformation Regularization Loss. As stated in Sec. 3.2, the ii-th frame’s deformation field Di\mathcal{D}_{i} contains variable dd and fixed W\mathcal{W}. dd represents the deformation that can not be represented by skinning transformation W\mathcal{W}, and this deformation should be relatively small. To associate the skeleton pose, we design the following regularization loss:

where t\mathbf{t} is a vertex coordinate of T\mathbf{T}, ∣T∣|\mathbf{T}| is the vertices number of T\mathbf{T}, and ρ\rho is the Geman-McClure robust loss .

Finally, the loss for the explicit representation is:

λe1\lambda_{e1} and λe2\lambda_{e2} adjust the weights of related losses. After each iteration, we reserve X\mathcal{X}’s gradients and wait for the implicit loss iteration to update together. For canonical mesh vertices, we use SGD to update T\mathbf{T} to T^\hat{\mathbf{T}}, which will be used in consistency loss to match two representations.

5.2 Implicit Loss

We sample pixels within the ground truth mask and utilize Sec. 3.3 to get the ray’s intersection p\mathbf{p} on Sη\mathcal{S}_{\eta} and its corresponding ground truth color IpI_{\mathbf{p}} and predicted normal NpN_{\mathbf{p}} if available. Then, based on this sampled points set P, we construct two losses.

Color Loss. By referring to Eq. (4), we formulate the color loss as:

Here, we use X\mathcal{X} to substitute related parameters in Eq. (4). Intuitively, this loss requires that the rendered images should match the input images.

Normal Loss. We utilized the predicted normal map by PIFuHD to further refine the geometry shape. Referring to Eq. (4), we can easily compute p\mathbf{p}’s normal np\mathbf{n}_{\mathbf{p}}. Besides, we need to transform the corresponding predicted normal NpN_{\mathbf{p}} from the space of current frame to canonical space, which can be computed with Jx(p)TNpJ_{\mathbf{x}}(\mathbf{p})^{T}N_{\mathbf{p}}, where Jx(p)J_{\mathbf{x}}(\mathbf{p}) is the Jacobian matrix of the forward deformation field at p\mathbf{p} . Thus, the normal loss is:

Here, unit(⋅)unit(\cdot) means to normalize the vector. ωp\omega_{\mathbf{p}} is the weight defined by the cosine of angle between np\mathbf{n}_{\mathbf{p}} and corresponding view direction. Since the predicted normals are noisy and inconsistent between frames, we use this weights to alleviate the impact of normal that deviates from its view direction and avoid geometry artifacts.

We also design regular losses for the implicit representation, and these losses are defined on the set of sampled points S near the implicit surface .

Rigidity Loss. We require the first deformation field dd to be as rigid as possible to avoid distortion. Following Park et al. , we design our loss as:

where Σp\Sigma_{\mathbf{p}} is the singular value diagonal matrix of the Jacobian of dd on p\mathbf{p} and ρ\rho is the robust function .

Eikonal Loss. We adopt the regular loss of IGR to make ff to be sign distance function:

where np\mathbf{n}_{\mathbf{p}} is obtained by differentiating ff at p\mathbf{p}.

Finally, the implicit loss can be represented as:

where λi1\lambda_{i1}, λi2\lambda_{i2} and λi3\lambda_{i3} are balancing weights.

5.3 Explicit/Implicit Consistency

After explicit iteration, the canonical mesh has been updated to T^\hat{\mathbf{T}}, to make the implicit SDF consistent with the updated explicit mesh during implicit iteration, we design a consistency loss:

where t^\hat{\mathbf{t}} is a vertex coordinate of T^\hat{\mathbf{T}}. Intuitively, the loss requires T^\hat{\mathbf{T}} to match the implicit surface Sη\mathcal{S}_{\eta}.

In each optimization step, we first perform the explicit iteration to obtain T^\hat{\mathbf{T}} and reserve X\mathcal{X}’s gradients. Then, we compute implicit and consistency losses to accumulate new X\mathcal{X}’s gradients. Finally, Adam is utilized to update X\mathcal{X} with computed gradients.

Experiments

We conduct quantitative and qualitative experiments to demonstrate the effectiveness of SelfRecon. For quantitative evaluation, we synthesize several sequences with commercial software due to the lack of high-quality geometry data of human body with general clothing. For qualitative evaluation, we mainly utilize the PeopleSnapshot dataset and several real sequences collected by ourselves. We also present ablation study for the loss term design and present an avatar generation application.

We synthesize data to quantitatively evaluate our reconstruction algorithm. Specifically, we use Blender to design the self-rotating motions for male and female avatars. Then, we utilize CLO3D to design several clothes and animate the clothed body with motions. Finally, we synthesized two sets of male and three sets of female dressing sequences. We reconstruct these sequences with VideoAvatar and our method, and report the registration error for the canonical posture results in Tab. 1. Compared with VideoAvatar, our method significantly reduces the values of various error metrics. In Fig. 3, we also present four group results and their error maps. Intuitively, our results capture the overall shape and have some reasonable details. As VideoAvatar is based on the SMPL+D representation, it has plausible results for tighter clothing, like the male examples, but lacks detailed reconstruction ability. Moreover, it can not correctly reconstruct loose clothing, especially for the females dressed in skirts.

2 Qualitative Evaluation

We also qualitatively compare SelfRecon with multi-frame prediction algorithm PaMIR , optimization method VideoAvatar and NeRF based neural rendering method NeuralBody on several sequences of PeopleSnapshot dataset. In Fig. 4, we present the first frame of input video, our rendered image, and reconstruction results of all methods from two perspectives. And we compare with PaMIR in the first two rows and the others with NeuralBody. As we can see, VideoAvatar based on SMPL+D can only approximately capture the overall shape, but details such as hairstyle and clothing folds are lost. PaMIR uses multi-frame input to improve its results, but still suffers from the deep ambiguity. As in the second example, its reconstructed human is not upright from the side view. Besides, its results have some details but miss facial features, while our method has better details and can recover certain facial features. Similar with SelfRecon, NeuralBody also inputs a video to do self-supervised optimization. It mainly focuses on novel view synthesis, but still can extract geometry from underlying NeRF representation. We can see that its reconstructions alleviate deep ambiguity and conform to the human’s overall structure while suffering from large noise on the surface, which may be caused by excessive freedom of volume rendering. Different from their method, implicit surface representation based SelfRecon can recover high-fidelity geometry shape without noises.

Fig. 5 shows reconstruction results on our collected videos with smartphones. For each group, we present the first frame of the video, our rendered image and reconstruction. Our results have high-fidelity geometry shapes for kinds of clothing and body, and our neural rendered images are also quite close to the input images.

3 Ablation Study

Our complete algorithm requires color images, masks and normal maps as inputs. Fig. 6 shows ablation experiments of two examples on three inputs. As the results show, if only using the mask loss, the recovered geometry shape is inside the convex hull formed by the silhouettes but lacks details and has noticeable concavities. After adding the color loss, it significantly improves the details and reduces the unnatural concavities. For the second example, the result has been very close to the result obtained by adding the normal loss. However, for the first example, it can not completely eliminate the depressed geometry without normal loss. This may be caused by lacking rich texture and multi-view observations in these areas. With the normal loss, our results are further improved, and the unnatural pits are eliminated while the details are preserved.

Since the normal prediction network is trained with synthetic images, its prediction may be not accurate for real tests and may be not consistent for different frames. As shown in Fig. 7, without the adaptive weights in Eq. (10), the normal loss might result unexpted results.

4 Avatar Generation

Thanks to our forward deformation field design, we can extract a mesh sequence with consistent topology. Based on the tracking results, we can extract a texture template mesh from the images with bound skinning weights from the skinning transformation field. Then, an animatable avatar is generated and can be driven with SMPL pose parameters. For texture extraction, we follow the method of VideoAvatar . Fig. 8 shows two examples of texture generation and driving from the PeopleSnapshot dataset. Our method recovers better geometric details like facial, shoes and clothing folds thanks to more accurate tracking results. Besides, our driving results look plausible and may be of sufficient quality for some applications.

Conclusion and Discussion

We proposed SelfRecon, a self-supervised reconstruction method based on neural implicit representation and neural rendering. With forward deformation, our method can be easily applied to body movement and recover space-time coherent surfaces, which is convenient for downstream applications. Moreover, combining the explicit representation, we proposed a non-rigid ray casting algorithm, which makes it possible for differentiable intersecting with the deformed implicit surface. SelfRecon can reconstruct high-fidelity clothed body shape from a self-rotating video without pre-computed templates. We also show high-fidelity avatar generation with our tracking results, demonstrating potential applications of SelfRecon.

SelfRecon still has several limitations. First, it requires relatively long time to optimize, which limits its convenient applications. However, this problem can be alleviated with the help of body priors and the fast growing field of neural rendering. Second, current method relies on the predicted normal maps to improve the geometric details. How to recover the geometric details directly from the self-supervised rendering loss is worthy of future study. Third, the proposed method mainly works well for self-rotating motions, and it is worthy of study for more general motion sequences.

This research was supported by National Natural Science Foundation of China (No. 62122071), the Youth Innovation Promotion Association CAS (No. 2018495), “the Fundamental Research Funds for the Central Universities”(No. WK3470000021).

References