TAVA: Template-free Animatable Volumetric Actors

Ruilong Li, Julian Tanke, Minh Vo, Michael Zollhofer, Jurgen Gall, Angjoo Kanazawa, Christoph Lassner

Introduction

Ever since the first 3D vector graphics games in the 1980s, we are striving to build better representations of 3D objects and humans. With increasing processing power, we can afford to capture, reconstruct and encode increasingly realistic representations. This makes exploring neural representations for graphical objects particularly appealing—it is a representation that has proven powerful , even though still being in its infancy. Recent methods for neural 3D representations go beyond capturing plain texture by modeling radiance fields , achieving more photo-realistic results than rasterization-based approaches . However, it is unclear how their representational power can be used to not only capture static, but also dynamic scenes that can be animated in a meaningful way, making the representations useful for capturing actors that can be “driven” post-capture. However, due to the high-dimensional nature of pose configurations, it is generally neither possible nor practical to capture all pose variations in one capture. This poses a new problem absent in static settings: generalization to out-of-distribution (OOD) poses. Furthermore, the neural 3D representation is desired to be editable as the classical representation like mesh and textures. Some works explored this aspects in the static settings , but it is not clear on how to edit neural representations on dynamic actors.

In this paper, we propose TAVA, a novel approach for Template-free Animatable Volumetric Avatars (illustrated in Fig. 1). We propose to use coordinate-based radiance fields to capture appearance, leading to high quality, faithful renderings. We extend the radiance capture with a carefully designed deformation model: while it requires solely 3D skeleton information at training time, it captures non-linear pose-dependent deformations and exhibits stable generalization behavior to unseen poses thanks to being anchored in an LBS formulation. The radiance field and deformation model are optimized jointly and end-to-end, leading to a simple-to-use and powerful representation: creating it requires only a tracked skeleton and multi-view photometric data, no template mesh or artist-designed rigging; the appearance and the deformation model can complement each other for highest quality results. These designed properties make TAVA suitable for content creation and editing as well as correspondence-based matching.

In our experiments, we demonstrate that the proposed approach outperforms state-of-the art approaches for animating and rendering human actors on the ZJU motion capture dataset . Thanks to being template-free, our approach is not limited to capturing humans: we present a detailed evaluation and ablation study on two synthetically rendered animals. This demonstrates the flexibility of the proposed approach and allows us to show additional applications in content-creation and editing.

Related Work

Deformable Neural Scene Representations: Coordinate-based neural scene representations produce impressive results in encoding shape and appearance . These methods train a coordinate-based neural network to model various properties of a scene, e.g., occupancy , distance to the closest surface , or density and color . However, making implicit scene representations deformable and animatable remains a challenging research problem. Nerfies and Neural 3D Video Synthesis handle changes in the scene by optimizing a deformation field and a latent code for each frame. HyperNeRF extends this by additionally creating a hyper-space which allows topology changes of the scene. Non-Rigid Neural Radiance Fields optimize a rigidity model in addition to a deformation field. While these methods produce impressive results on dynamic scenes, they are designed to only memorize the scene and cannot control the scene beyond interpolations.

Animatable Neural Radiance Fields: Recently, many approaches for controllable animatable NeRFs have been proposed. Neural Actor uses a pose-dependent radiance field by warping rays into the canonical space of a template body model while using 2D texture maps to model fine detail. NeuralBody anchors latent codes on the vertices of a deformable mesh controlled by LBS. The follow-up work Animatable-NeRF establishes a transformation between view and canonical space through optimizing the inverse deformation field. Other works like NARF, A-NeRF predict the radiance field at a given 3D location based on its relative coordinates to the bones. Most recently, a concurrent work HumanNeRF produces a free-viewpoint rendering of a human by modeling the inverse deformation as a mixture of affine fields . Yet many of these methods do not have a 3D canonical space that preserves correspondences across different poses, which is required for content-creation or editing. Some are built on top of the SMPL body template, which prohibits them to be applied to creatures beyond humans. Moreover, most of the aforementioned methods either introduce latent codes to better memorize the seen poses , or represent the deformation in the inverse direction from view space to the canonical space . Thus they do not generalize well to the unseen poses because the existence of pose-conditioned MLPs. In contrast, our approach is template-free, enables editing, and is designed to be robust to unseen poses. We provide an overview of the comparison between our method and those previous works in Tab. 1.

Animatable Shapes: Non-rigid shape reconstruction often utilizes a canonical space that is fixed across frames, with a deformation model to create a mapping between the canonical and the deformed space. Traditionally, this has been achieved by extracting a low dimensional articulated mesh , such as SMPL , or by extracting a rigged mesh via post-processing. Several methods have been proposed to optimize blend weights and rigs from data. ARCH deforms an estimated implicit representation to fit to a clothed human using a single image. Recent approaches model inverse deformation fields , which map points from pose-dependent global space to pose-independent canonical space where the surface is represented. For example, SCANimate regularizes the inverse skinning by using a cycle consistency loss. The main drawback of these inverse deformation approaches is that the inverse transformation is pose dependent and may not generalize well to previously unseen poses. SNARF addresses this by learning a forward deformation field instead, mapping points from canonical to pose-dependent deformed space. However, unlike our appraoch, these methods require 3D geometry supervision and most do not optimize for appearance.

Method

Our goal is to create an animatable neural actor from multi-view images with known 3D skeleton information without requiring a body template. Similar to a traditional personalized body rig, we want to build a representation that not only represents the shape and appearance of the actor but also allows to animate it while maintaining correspondence among different poses and views. TAVA is designed to achieve the above goals with three components: (1) a canonical representation of the actor in neutral pose, (2) deformation modeling based on forward skinning, and (3) volumetric neural rendering with pose-dependent shading. Fig. 2 illustrates an overview of our method. To employ volumetric neural rendering in the view space, our method first deforms the samples along a ray back to the canonical space through inverting the forward skinning via root-finding, then queries their colors and densities in the canonical space, as well as the pose-dependent effects. Below, we first establish preliminaries, then discuss each of the components.

NeRF is a groundbreaking technique for novel view synthesis of a static scene. It models the geometry and view-dependent appearance of the scene by using a multi-layer perceptron (MLP). Given a 3D coordinate x=(x,y,z)\mathbf{x}=(x,y,z) and the corresponding viewing direction (θ,ϕ)(\theta,\phi), NeRF queries the emitted color c=(r,g,b)\mathbf{c}=(r,g,b) and material density σ\sigma at that location using the MLP. A pixel color C(r)C(\mathbf{r}) can then be computed by accumulating the view-dependent colors along the ray r\mathbf{r}, weighted by their densities:

Please refer to the original papers for more details.

2 Canonical Neural Actor Representation

We represent an articulated subject as a volumetric neural actor in its canonical space. The representation includes a Lambertian neural radiance field FΘrF_{\Theta_{r}} to represent the geometry and appearance of this actor, and a neural blend skinning function FΘsF_{\Theta_{s}}, which describes how to animate the actor:

Discussion. This formulation not only models the canonical geometry and appearance of an avatar, but also describes its dynamic attributes through the skinning weights w\mathbf{w}. Unlike previous works, such as SNARF , which models a pose-dependent geometry in the canonical space, and NARF and A-NeRF , which entirely skips canonical space modeling, our method is based on a canonical representation that fully eliminates any effects of pose on the geometry and appearance. Moreover, the skinning weights learnt in the canonical space remain valid for a large range of poses, meaning that the actor is ready to be animated in novel poses outside of the training distribution (see Sec. 4.2 for validation of its robustness to out-of-distribution novel poses). Last but not least, our representation eases the correspondence finding problem across different poses and views, because the matching can be done in the pose-independent canonical space (see Sec. 4.2 for results).

3 Skinning-based Deformation

3.2 Inverse Skinning.

To render this model, we need to query color and density in the view space. So it is required to find the the correspondence xc\mathbf{x}_{c} in the canonical space for each xv\mathbf{x}_{v} in the view space. As our forward skinning in Eq. 5 is defined through neural networks, there is no analytical form for the inverse skinning. So, inspired by SNARF , we pose this as a root finding problem:

and solve it numerically using Newton’s method:

The gradients of the network parameters Θs\Theta_{s} and ΘΔ\Theta_{\Delta} can then be analytically computed for the inverse skinning :

Please refer to the supplemental material for their derivations.

4 Deformation-based Neural Rendering

However, for a dynamic object, the shading on the surface may change depending on pose due to self-occlusion. This can lead to colors in the view space being darker than the colors in the canonical space, providing inconsistent supervision signals. However, it is non-trivial to accurately model this self-occlusion without ray tracing (including secondary rays) and known global illumination. A simple but effective estimator, widely used in modern rendering engines like Unreal and Blender is ambient occlusion, in which the shading caused by occlusion is modeled by a scaling factor multiplied with the color values, where the value is calculated by the percentage of view directions being occluded around each point on the surface. Since it is an attribute defined at each coordinate that depends on the global geometry of the actor, we model this shading effect use a coordinate based MLP FΘaF_{\Theta_{a}} conditioned on the pose P\mathbf{P} of the actor:

where h\mathbf{h} is an intermediate activation from FΘrF_{\Theta_{r}}, and ai∗a_{i}^{*} is the ambient occlusion at this location under pose P\mathbf{P}. Note that only the ambient occlusion ai∗a_{i}^{*} is pose-conditioned, which makes sure the actor (geometry and appearance) is represented in a canonical space that is pose-independent, as described in Sec. 3.2.

then we use (c=av∗cv,σ=σv)(\mathbf{c}=a_{v}*\mathbf{c}_{v},\sigma=\sigma_{v}) as the final emitted color and density in the view space, for the volumetric rendering in Eq. 1.

Note that in general there is no way to guarantee that the inverse root finding converges. In practice, root finding fails for 1% to 8% of the points in the view space, making it impossible to query their attributes. For these points, an option is to just simply set their densities to zero, which would only be problematic if the points are close to the surface. A slightly better way is to estimate the color and density for those points by interpolating the attributes from their nearest valid neighbors along the ray. We conduct experiments on both strategies in Sec. 4.2, which results in slightly better performance. We choose the second strategy in our full model.

5 Establishing Correspondences

As our method is endowed with a 3D canonical space, we have the ability to trace surface correspondences across different views and poses. When rendering an image using Eq. 1, besides accumulating colors {cm}\{\mathbf{c}_{m}\} of the samples {xv,m}\{\mathbf{x}_{v,m}\} along the ray r\mathbf{r}, we also accumulate the corresponding canonical coordinates {xc,m}\{\mathbf{x}_{c,m}\}:

6 Training loss

Besides the image loss Lim\mathcal{L}_{im} defined in Eq. 2, we also employ two auxiliary losses that help the training. Due to the fact that all the points along a bone should have the same transformation, we encourage the skinning weights w\mathbf{w} of samples xˉc\mathbf{\bar{x}}_{c} on the bones to be one-hot vectors w^\mathbf{\hat{w}} (noted as Lw\mathcal{L}_{w}). We also encourage the non-linear deformations Δv\mathbf{\Delta}_{v} of those samples to be zero given any pose P\mathbf{P} (noted as LΔ\mathcal{L}_{\Delta}). We use MSEMSE to calculate both, Lw=∣∣w(xˉc)−w^)∣∣22\mathcal{L}_{w}=||\mathbf{w}(\mathbf{\bar{x}}_{c})-\mathbf{\hat{w}})||^{2}_{2} and LΔ=∣∣Δw(xˉc,P)−0)∣∣22\mathcal{L}_{\Delta}=||\Delta_{w}(\mathbf{\bar{x}}_{c},\mathbf{P})-\mathbf{0})||^{2}_{2}. Our final loss is: L=Lim+λLw+βLΔ,\mathcal{L}=\mathcal{L}_{im}+\lambda\mathcal{L}_{w}+\beta\mathcal{L}_{\Delta}, where λ\lambda is set to 1.01.0 and β\beta is set to 0.10.1 in all our experiments.

Experiments

We conduct experiments on 1) four human subjects (313, 315, 377, 386) in the ZJU-Mocap dataset , a public multi-view video dataset for human motion, and 2) two synthetic animal subjects (Hare, Wolf) introduced in this paper, rendered from multiple views using Blender.

Data Splits. Prior works create the train and val sets on the ZJU-Mocap dataset by simply splitting each video with 500∼2200500\sim 2200 frames into two splits, where the training set has 60∼30060\sim 300 frames and the validation set has 300∼1000300\sim 1000 frames. This is not an ideal split to evaluate pose synthesis performance because 1) a training set with 6060 consecutive frames in a 3030fps video does not sufficiently cover pose variation to learn from, and 2) due to the repetitive motion of the actors, quite often similar poses are in both the training and validation sets, which should be avoided for evaluating a method on pose generalization. Therefore we establish a new protocol to split the dataset by clustering the frames based on pose similarity. Specifically, for each subject, we first randomly withhold a chunk of consecutive frames to be the test set, for the purpose of the final evaluation. Then, we use the K-Medoids algorithm on the remaining frames to cluster them into K=10K=10 clusters, based on pose similarity measured by the V2V Euclidean distance using ground-truth mesh. The most different cluster is selected as the valposeood\text{val}_{\text{pose}}^{\text{ood}} set, in which the frames are all considered to contain the out-of-distribution poses. For the remaining 99 clusters, we randomly split each cluster 2:12:1 to form train and valposeind\text{val}_{\text{pose}}^{\text{ind}} sets, where the frames in valposeind\text{val}_{\text{pose}}^{\text{ind}} still contains new poses which are considered to be in the distribution of the training set. For the view splits, we follow the protocol from for ZJU-Mocap, where 44 views are used for training and 1717 views for testing. The animal subjects have 1010 random views for training and 1010 for testing. We denote valview\text{val}_{\text{view}} as our novel-view synthesis evaluation set, which contains all the training poses but rendered from different viewpoints. Please see the supplemental material for more details.

2 Evaluation and Comparison

Baselines. We compare our work with two types of previous methods: 1) Template-free methods, including NARF and A-NeRF , as well as 2) SMPL-based methods, including Animatable-NeRF and NeuralBody . As our baseline, we use Pose-NeRF: we slightly modify Mip-NeRF to learn the density and color conditioned on pose. We conduct experiments for all the moethods above on ZJU-Mocap, but exclude Animatable-NeRF and NeuralBody for the animal subjects (they require a template 3D model). Although code is available for each method, we noticed that each is using a different set of hyper-parameters for neural rendering (e.g., number of MLP layers, number of samples, near and far planes) and different training schedules, all of which are not related to method design but can greatly affect the performance. To make as-fair-as-possible comparisons, we integrated the template-free methods, NARF and A-NeRF, into our code base, which shares the same set of hyper-parametersFor NARF, our re-implementation achieves better performance than it’s official implementation. Please refer to the supplmental material for further details.. For Animatable-NeRF and NeuralBody, we use the original implementations since their designs are based on the SMPL body template.

Novel-view Synthesis. In this task, we conduct experiments on both, the ZJU-Mocap dataset and the two animal subjects Hare and Wolf, using valview\text{val}_{\text{view}} set. As shown in Tab. 3 and Tab. 2, our method outperforms other template-free methods measured by PSNR and SSIM. On the ZJU-Mocap dataset, our method achieves comparable performance with two template-based methods, Animatable-NeRF and NeuralBody, which greatly benefit from the SMPL body template, but do not work on other creatures like animals. See Fig. 4 for a qualitative comparison.

Novel-pose Synthesis. Due to the high interdependency of appearance changes caused by pose and motion, novel-pose synthesis is a more challenging task than novel-view synthesis, especially for poses that are out of the training distribution. To carefully study this problem, we conduct experiments on both, in-distribution (InD) novel poses, using the valposeind\text{val}_{\text{pose}}^{\text{ind}} set, and out-of-distribution(OOD) novel poses, using the valposeood\text{val}_{\text{pose}}^{\text{ood}} set. Our experiments reveal that for InD poses, the performance of nearly all the methods are consistent with their performance on the novel-view task, as shown in Tabs. 2, 3. However, there is a huge drop in performance from InD poses to OOD pose (0.670.67db∼2.66\sim 2.66db on ZJU-Mocap; 1.681.68db∼5.73\sim 5.73db on animals). This is not surprising if the method contains neural networks that directly infer appearance information from pose input: generalization to vastly different pose inputs can not be expected. One of the main goals in this paper is to reduce this reliance of the neural networks to the pose input, for improving the robustness of the method to the OOD poses. Our method benefits from explicitly incorporating the forward LBS. We observe only an 1.681.68db performance drop comparing InD to OOD poses on the animal subjects, whereas other methods suffer from ∼5\sim 5db performance drops, as shown in Tab. 3. Since these two synthetic subjects are not rendered with pose-dependent shading, and do not have “clothing” deformations, we disabled the ambient occlusion aa (set to 1.1.) and non-linear deformation Δv\Delta_{v} (set to 0.0.) terms in our method during both, training and inference. These synthetic subjects allows us to study the underlying formulation of the articulation deformation, and we show here the forward LBS-based deformation is more reliable than the inverse deformation used in the baselines which takes pose as input to the MLP. The results on the ZJU-Mocap dataset in Tab. 2 show that our method outperforms both, the template-free and template-based methods, on the OOD poses. All the methods are prone to nearly the same drop in performance on the ZJU-Mocap dataset comparing InD to OOD poses. This is, because currently all of these methods, including ours, are still implicitly modeling pose-dependent shading effects (e.g., self-occlusion) as either a neural network or a latent code during training, which does not generalize well to OOD poses. Our method, though, provides a possibility to factor out the shading effects during inference and reveal the albedo color of the actor, which yields better generalization but is not suitable for evaluation comparing to the ground-truth, as shown in Fig. 8. See Figs. 3, 4 for qualitative comparisons.

Pixel-to-Pixel Correspondences. We quantitatively evaluate correspondence on the animal subjects against Pose-NeRF, A-NeRF and NARF . We show qualitative results on ZJU-Mocap since no ground-truth correspondences are available (see Fig. 5). Even though neither of the baseline methods demonstrate that they can establish correspondences, we still tried our best to create a valid comparisonPose-NeRF, A-NeRF and NARF all query the color and density of (xv,P){(\mathbf{x}_{v},\mathbf{P})} in a higher dimensional (>3>3) space, where we do the nearest neighbor matching for them using our approach as described in Sec. 3.5. Please refer to the supp.mat. for further details.. For quantitative evaluation, we randomly sample 20002000 image pairs (A,BA,B) in valposeood\text{val}_{\text{pose}}^{\text{ood}} set, and use the ground truth mesh to establish ground-truth pixel-to-pixel correspondences (χA→χB)(\chi_{A}\rightarrow\chi_{B}) for every pair of images (A→BA\rightarrow B), where χA\chi_{A} and χB\chi_{B} are the corresponding image coordinates. Then, we use each method to render this pair of images, and find the correspondences of χA\chi_{A} in BB as χB∗\chi^{*}_{B}. The pixel-to-pixel error (P2P) is then calculated as the average distance between χB\chi_{B} and χB∗\chi^{*}_{B}: P2P=∣∣χB−χB∗∣∣22\text{P2P}=||\chi_{B}-\chi^{*}_{B}||^{2}_{2}. As shown in Tab. 3 and Fig. 6, our method achieves over 2x more accurate correspondences (3.383.38px v.s. 8.468.46px error in a 800×800800\times 800 image) compared to the baselines. We visualize the extracted dense correspondences of our method in Fig. 5, which shows that correspondences across different subjects can be established as long as they share the same canonical pose (T-pose in ZJU-Mocap). Further more, we demonstrate that accurate correspondences can be used for content editing in Fig. 7.

3 Ablation Studies.

Thanks to our model design, we can train a full model with the non-linear deformation Δv\Delta_{v} and the ambient occlusion AO enabled, then strip them out at inference time. Fig. 8 shows a qualitative result to visually demonstrate the impacts on the full model. Notice that without AO, the shading effects are removed during rendering, which produces overall brighter images than the ground-truth. This is an expected effect, but prohibits quantitative evaluation. Furthermore, we ablate these two design decisions during training. For the non-linear deformation, our ablation is to simply disable it during training to see if the LBS is enough to model the deformation. For the AO, our ablation is to compare it with predicting a pose-dependent color by conditioning pose to the color branch of FΘrF_{\Theta_{r}}, and disabling the AO branch FΘaF_{\Theta_{a}}. As shown in Tab. 4, both design decisions contribute to the final model performance. Lastly, we show the two different strategies to deal with root finding failures described in Section 3.4 in Tab. 4. Using the interpolation strategy results in slightly better performance.

Discussion

In this paper, we proposed a volumetric representation for articulated actors based on learned skinning, shape, and appearance. We also model pose-dependent deformation and shading effects. Extensive evaluations demonstrate that our approach consistently outperforms previous methods when generalizing to out-of-distribution unseen poses. Our approach can recover much more accurate dense correspondences across different poses and views than prior works, enabling content editing applications. Moreover, it does not require any body templates, enabling applications for creatures beyond humans. While our approach has clear advantages, there are few limitations. First, our method trains much slower (5 to 8 times) than the baselines, due to the nature of the root finding process for inverse deformation. A future direction could be to use invertible neural networks to avoid the root finding process. Second, even though our forward LBS ensures generalization to unseen poses, pose-dependent non-linear deformation and shading effects are still challenging to estimate correctly for unseen poses. Such effects are fundamentally challenging to model, particularly for lighting-dependent shading. An interesting future direction could be to model these effects across multiple subjects so that information from all subjects can be used to improve non-linear deformation model performance.

Appendix 0.A Inverse Skinning Gradients.

As described in Sec. 3.3, for each sample xv\mathbf{x}_{v} in the view space, we find its canonical correspondence xc\mathbf{x}_{c} through root finding:

In order to optimize the skinning deformation defined by (FΘsF_{\Theta_{s}}, FΘΔF_{\Theta_{\Delta}}), we need to determine the gradients of the overall loss L\mathcal{L} w.r.t the network parameters (Θs\Theta_{s}, ΘΔ\Theta_{\Delta}):

The first term [∂L/∂xc∗\partial\mathcal{L}/\partial\mathbf{x}^{*}_{c}] can be easily calculated through back-propagation. The second terms [∂xc∗/∂Θs\partial\mathbf{x}^{*}_{c}/\partial\Theta_{s}] and [∂xc∗/∂ΘΔ\partial\mathbf{x}^{*}_{c}/\partial\Theta_{\Delta}] can be calculated analytically via implicit differentiation :

Appendix 0.B Dataset Splits and Pose Clustering.

As described in Sec 4.1, to avoid similar poses appearing in both the training and the validation set, we split the dataset by clustering the frames based on pose similarity. Fig. 9 shows an example of the pose similarity matrix on ZJU-Mocap subject 313. It clearly shows that this actor moves with a repetitive motion pattern. Thus the previous way of splitting the dataset into two chunks with consecutive frames will cover similar poses in both sets, which is not suitable for evaluating the pose generalization ability. This motivated us to introduce our new data split protocol based on pose clustering.

Our pose clustering process is as follows: We first disable the global (root) transformation for all poses. Then, the difference of two poses is measured by the Euclidian distance of their corresponding mesh vertices (SMPL mesh for ZJU-Mocap). As we only use pose clustering to construct the dataset splits, the mesh information is considered accessible here. Next the K-Medoids algorithm is adopted to cluster the poses into K=10K=10 clusters (Examples shown in Fig. 10). Finally we calculate the average difference between the K medoids to find the most different one, which we regard as the out-of-distribution poses to form the OOD validation set.

Appendix 0.C Implementation Details.

TAVA employs four MLPs (FΘrF_{\Theta_{r}}, FΘΔF_{\Theta_{\Delta}}, FΘrF_{\Theta_{r}}, FΘaF_{\Theta_{a}}) in our method. Both, FΘrF_{\Theta_{r}} and FΘΔF_{\Theta_{\Delta}}, consist of 4 layers with 128 hidden units, with 4-degree positional encoding on the input coordinates. FΘrF_{\Theta_{r}} is an 8-layer MLP with 256 hidden units and uses 10-degree integrated positional encoding, similar to Mip-NeRF . FΘaF_{\Theta_{a}} is a single-layer MLP with 128 hidden units that connects to the 88-th layer of FΘrF_{\Theta_{r}}. We follow the hyper-parameters in Mip-NeRF for the volume rendering, where 6464 samples are drawn for each ray at both coarse and fine level.

Appendix 0.D Baseline Implementation Details.

As described in Sec. 4.2, for the template-based baselines Animatable-NeRF and NeuralBody , we use their official implementations. For the template-free baselines NARF and A-NeRF , we re-implemented them in our code base for fair comparison. We also carefully adapt their official implementations to the ZJU-Mocap dataset to verify our re-implementation. As shown in Tab. 5, our re-implementation achieves better performance than the official implementations on ZJU-Mocap subject 313.

Appendix 0.E Visualizations for Skinning Weights.

Fig. 11 shows results of our learned skinning weights and canonical geometry for animal subjects and ZJU-Mocap data. The surfaces are extracted with marching cube algorithm with threshold 5.05.0 on the density field.

Appendix 0.F Challenges in ZJU-Mocap Dataset.

The ZJU-Mocap dataset has become an increasingly popular dataset to study human performance capture, reconstruction, and neural rendeirng . Yet we notice that there are a few issues in this dataset that are neither addressed nor mentioned in the previous works, including imperfect camera calibrations and various camera exposures (as shown in Fig. 12). We also get acknowledged from the authors of ZJU-Mocap Dataset on those issues. We discovered these issues after the submission so they are not considered in our designs. Yet they greatly affect both our performance and the baselines’. We believe it is worth to point them out so that they can be considered in the future research.

Appendix 0.G Per-subject Breakdown Comparisons.

We also report a per-subject breakdown of the quantitative metrics against all baseline methods in Tab. 6 and Tab. 7.

References