Neural Articulated Radiance Field

Atsuhiro Noguchi, Xiao Sun, Stephen Lin, Tatsuya Harada

Introduction

In this work, we aim to learn a representation for rendering novel views and poses of 3D articulated objects, such as human bodies, from images. Our approach follows the inverse graphics paradigm of analyzing an image by attempting to synthesize it with compact graphics codes. These codes are typically disentangled to allow for rendering of scenes/objects with fine-grained control over individual appearance properties such as object location, pose, lighting, texture, and shape. For the case of humans, synthesis of novel views and poses can be useful for applications such as movie making, photo editing, virtual clothing and motion transfer .

Various inverse graphics based approaches have been specifically designed for static scenes , rigid objects , blend shapes for keypoints and dense meshes . However, efficient deformation modeling of articulated 3D objects using neural networks remains a challenging task due to the large variance of joint locations (especially for endpoints such as hands), severe self-occlusions, and high non-linearity in forward kinematic transformations . Though work has been done to enable explicit control over the underlying human pose and key point locations , their neural rendering methods are either limited to 2D , which prevents modeling of view-dependent appearance , or based on mesh representations , where rendering quality can be affected by the resolution of the discrete template mesh.

Recent progress on the implicit representation of 3D objects and scenes, such as signed distance functions and occupancy fields , has greatly promoted the development of the inverse graphics paradigm. Such representations are lightweight in model size, continuous, and differentiable, making them highly practical in comparison with the previously-dominant volumetric representations . Particularly, Mildenhall et al. propose the neural radiance field (NeRF) that takes a single continuous 5D coordinate (3D spatial location and 2D viewing direction) as input and outputs the volume density and view-dependent emitted radiance at each spatial location. Combined with a classical differentiable volume rendering technique , it is able to synthesize novel views by learning from a sparse set of input views of static scenes. NeRF completely discards the mesh-based representation and replaces it with a radiance-based model which can effectively and efficiently encode view-dependent appearance, enabling it to reproduce scenes of complex geometry with high fidelity.

In this paper, we extend NeRF to an articulated NeRF, called a Neural Articulated Radiance Field (NARF), to represent articulated 3D objects. Accounting for 3D articulation within the NeRF framework is a challenging problem because a complex, non-linear relationship exists between a kinematic representation of 3D articulations and the resulting radiance field, making it hard to model implicitly in a neural network . In addition, the radiance field at a given 3D location is influenced by at most a single articulated part and its parents along the kinematic tree, while the full kinematic model is provided as input. As a result, dependencies of the output to irrelevant parts may inadvertently be learned, which is known to hurt model generalization to poses unseen in training .

To address these issues, we propose a method that predicts the radiance field at a 3D location based on only the most relevant articulated part. This part is identified using a set of sub-networks that output the probability for each part given the 3D location and the 3D geometric configuration of the parts. The spatial configurations of parts are computed explicitly with a kinematic model, rather than modeled implicitly in the network. A NARF then predicts the density and view-dependent radiance of the 3D location conditioned on the properties of only the selected part. An overview of the method is shown in Fig. 1.

The presented NARF has the following properties:

It learns a disentangled representation of camera viewpoint, bone parameters, and bone pose, allowing these properties to be individually controlled in rendering.

A dense 3D representation is learned from a sparse set of 2D images with pose annotations of the articulated object, which could potentially be obtained through external pose estimation techniques on multi-view images with known camera parameters .

Part segmentation is learned from images with pose annotation. Additional supervision is not needed.

NARF can be trained for articulated objects of various shape and appearance, through the use of an autoencoder that extracts latent shape and appearance vectors which are additionally disentangled.

With this approach, it becomes possible to render both novel views and poses of articulated 3D objects from pose annotated 2D images with little increase in computational complexity.

Related Work

The deformation of articulated objects is traditionally modeled by skinning techniques in which the location of surface mesh vertices is determined from bone transformations controlled by the kinematics . Effective skinning models with subtle pose-dependent and identity-dependent deformation modeling have been developed for the human body and animals . However, the representation capacity of skinning based models is limited to the resolution of the discrete template mesh, and sophisticated shading techniques are usually required for high quality image rendering. In addition, a large amount of 3D scan data and expert supervision are required to prepare a template mesh.

Recently, Deng et al. proposed a neural network based articulated shape representation (NASA). NASA learns the neural indicator/occupancy function of every point in space, conditioned on a latent pose vector that encodes a piece-wise decomposition. NASA provides a continuous and differentiable representation for 3D articulated shapes. However, ground truth occupancy is required to train the network, and NASA does not learn appearance, a critical element for rendering.

Articulated Pose Conditioned Image Generation

The recent advance of image generation models such as variational autoencoders (VAE) and generative adversarial networks (GANs) provides powerful tools for generating realistic-looking images. Image generation for articulated objects (typically persons) conditioned on target poses is an important direction with various applications like movie making, photo editing, virtual clothing and motion transfer . A majority of these works generate the image of a person in a target pose by learning a GAN model from 2D keypoint maps of the target pose. The appearance information of this person is provided by explicit concatenation with an image of this person in the target pose , automatically encoded for a single person or using auto-encoders . These works are limited to 2D, which prevents modeling of view-dependent appearance . Some works leverage the underlying 3D mesh representation and transfer appearance from one mesh to another using aligned mesh triangles. The quality of a mesh based representation is bounded by the resolution of its discrete template mesh, and a 3D template mesh is required.

Implicit 3D representation

Our work builds on the recent success of the implicit 3D representation. This representation is memory efficient, continuous, and topology-free, and has been used for learning 3d shape , 3d texture , static scenes , parts decomposition , articulated objects , deformation , 3d reconstruction from sparse images , and image synthesis .

Early methods required ground truth 3D geometry , but in combination with differentiable rendering, they evolved to learn from 2D images. In particular, neural radiance fields (NeRF) is capable of learning a 3D representation of complex scenes using only multi-view posed images. However, NeRF addresses static scenes and cannot handle deformable objects. Very recently, methods have been proposed to extend NeRF to learn deformations and dynamics . These models have been successful in learning deformable implicit representations using posed video frames . However, these models do not take into account the structure of the object, so they cannot generate images with explicit pose control.

Method

In this section, we present neural articulated radiance field, a novel implicit representation for articulated 3D objects based on NeRF. We start by briefly reviewing the basic NeRF formulation for static scenes in Sec. 3.1. In Sec. 3.2, NeRF is extended to be conditioned on pose via a kinematic model, and a straightforward baseline is derived from this. We reformulate the pose-conditioned NeRF to allow for rigid object transformations as well as global shape variations in Sec. 3.3. In Sec. 3.4, we represent articulated 3D objects as a composition of movable rigid object parts controlled by forward kinematic rules. To achieve constant model complexity with respect to the number of object parts, we propose an efficient Disentangled NARF architecture. The training strategy is then presented in Sec. 3.5.

A neural network is used to represent a Radiance Field such that 3D location x=(x,y,z){\bf x}=(x,y,z) and 2D viewing direction d{\bf d} is converted to density σ\sigma and RGB color value cc. The density σ\sigma acts like a differential opacity controlling how much radiance is accumulated by a ray passing through x{\bf x} .

where γ(p)=[(sin(2lπp),cos(2lπp)]0L\gamma(p)=[(\text{sin}(2^{l}\pi p),\text{cos}(2^{l}\pi p)]_{0}^{L} is a positional encoding (PE) layer that maps an input scalar into a higher dimensional space to represent high-frequency detail of the scene. FΘF_{\Theta} consists of two ReLU MLP networks. Specifically, the volume density σ\sigma is a function of the location x\bf x only, while the RGB color cc is a function of both location x{\bf x} and viewing direction d{\bf d}.

where h{\bf h} is a hidden feature vector.

Classical volume rendering is used to render the color C(r){\bf C}({\bf r}) of a camera ray r(t)=o+td{\bf r}(t)={\bf o}+t{\bf d} with near and far bounds tnt_{n} and tft_{f}, and where o{\bf o} denotes the camera position.

T(t)T(t) denotes the accumulated transmittance along the ray. The integrals are computed by a discrete approximation over sampled points along the ray r{\bf r}.

Here, σj\sigma_{j} and cjc_{j} are the density and color at the jthj^{th} point on the ray r{\bf r}, and δj\delta_{j} is the distance between the jthj^{th} and (j+1)th(j+1)^{th} sample points.

NeRF is trained on images of a single static scene taken from multiple views with known camera parameters. The density and color of each location are trained so that the rendered image for each of the views becomes close to its ground truth. After training, high resolution images can be synthesized from any viewpoint.

2 Pose-Conditioned NeRF: A Baseline

Our goal is to extend the representation capacity of NeRF from static scenes to deformable articulated objects whose configurations can be described by a kinematic model . The radiance field of a 3D location is thus conditioned on the pose configuration. Once this “pose-conditioned NeRF” is learned, novel poses can be rendered in addition to novel views, by changing the input pose configurations.

In this work, we focus on modeling articulated objects without considering the background. Therefore, for compactness, we assume that the backgrounds are pre-cleaned.

Formally, the kinematic model represents an articulated object of P+1P+1 joints, including endpoints, and PP bones in a tree structure where one of the joints is selected as the root joint and each remaining joint is linked to its single parent joint by a bone of fixed length.

Specifically, the root joint J0J_{0} is defined by a global transformation matrix T0{\bf T}^{0}. Let ζi\boldsymbol{\zeta}_{i} be the bone length from the ithi^{th} joint JiJ_{i} to its parent, i∈{1,...P}i\in\{1,...P\}, and θi\boldsymbol{\theta}_{i} denotes the rotation angles of the joint with respect to its parent joint. A bone, considered as a rigid object, defines a local rigid transformation between a joint and its parent. The transformation matrix Tlocali{\bf T}^{i}_{local} is computed as

where Rot and Trans are the rotation and translation matrix, respectively. The global transformation from the root joint to joint JiJ_{i} can thus be obtained by multiplying the transformation matrices along the bones from the root joint to the ithi^{th} joint:

where Pa(i)\text{Pa}(i) includes the ithi^{th} joint and all of its parent joints along the kinematic tree. The corresponding global rigid transformation li={Ri,ti}l^{i}=\{R^{i},{\bf t}^{i}\} for the ithi^{th} joint can then be obtained from the transformation matrix Ti{\bf T}^{i}.

Baseline

The most straightforward way to condition the radiance field at a 3D location x{\bf x} on a kinematic pose configuration P={T0,ζ,θ}\mathcal{P}=\{{\bf T}^{0},\boldsymbol{\zeta},\boldsymbol{\theta}\} is to directly concatenate a vector representing P\mathcal{P} as the model input. Since the forward kinematic computation is a complex non-linear function that is hard to simulate in neural networks, we use the transformations li={Ri,ti}l^{i}=\{R^{i},{\bf t}^{i}\} obtained by the forward kinematics as network inputs.

We refer to this naive approach as Pose-conditioned NeRF (P-NeRF). The implementation details can be found in the supplemental material. Though P-NeRF establishes dependency between the radiance field and pose, generalization with this model is difficult because of the following two reasons.

Implicit Transformations. An articulated object consists of several rigid bodies, and the surface points on the object should move with the rigid transformations of the parts when the pose changes. Therefore, the movement of points can be explicitly described using rigid body transformations of each part, but such transformations may be difficult for a neural network to learn implicitly.

Part Dependency. The density at a 3D location depends only on the parameters of the bone it lies on and its parents along the kinematic tree. However, all the parameters are used to estimate the radiance field of a single location in Eq. 8. As the training on such 3D locations is backpropagated to all the parameters, the network may learn erroneous dependencies that do not physically exist. Correct pose predictions may still be obtained for test poses seen in the training data, but model generalization to novel poses may be degraded .

Towards addressing the above issues, we decompose the articulated object into PP rigid object parts. Each part has its own local coordinate system defined by the rigid transformation li={Ri,ti}l^{i}=\{R^{i},{\bf t}^{i}\}, which is explicitly estimated using forward kinematics, rather than modeled implicitly by a neural network. Then, we show how a rigidly transformed object part can be effectively modeled in a rigidly transformed neural radiance field (RT-NeRF) in Section 3.3. Based on RT-NeRF, we describe in Section 3.4 how to train a single unified NeRF that encodes multiple parts in a manner that avoids the part dependency issue.

3 Rigidly Transformed Neural Radiance Field

Given a rigid transformation l={R,t}l=\{R,{\bf t}\} of an object, we now estimate the radiance field in the object coordinate system where the density is constant with respect to a local 3D location. Formally,

where xl=R−1(x−t){\bf x}^{l}=R^{-1}({\bf x}-{\bf t}) represents the 3D location in the local object coordinate system.

We expect the model to handle certain shape variations. For example, the limb length and thickness of a child should differ significantly from those of an adult. To account for shape variation, we further condition the model on bone parameter ζ\boldsymbol{\zeta}.

Meanwhile, the color cc at a local 3D location may change with a transformation of the object coordinate system, as this may lead to changes in the local lighting condition. Since the RGB color cc at a local 3D location should further depend on rigid transformation ll, we use a 6D vector se(3)\mathfrak{se}(3) representation ξ\boldsymbol{\xi} of transformation ll as a network input.

where dl=R−1d{\bf d}^{l}=R^{-1}{\bf d} is the 2D view direction in the object coordinate system.

Combining Eqs. 10-11, the rigidly transformed neural radiance field (RT-NeRF) defined in the l={R,t}l=\{R,{\bf t}\} space is expressed as

RT-NeRF serves as the basic building block in the neural articulated radiance field, and we will next show how it is utilized to overcome the ’Implicit Transformations’ and ’Part Dependency’ issues.

4 Neural Articulated Radiance Field

The proposed neural articulated radiance field (NARF) is built upon RT-NeRF. We first introduce two basic solutions, Part-Wise NARF and Holistic NARF, and analyze the pros and cons of each. Then, we propose our final solution named Disentangled NARF that shares the merits of both Part-Wise and Holistic NARF. Conceptual figures are visualized in Fig. 2.

Given a kinematic 3D pose configuration of an articulated object {T0,ζ,θ}\{{\bf T}^{0},\boldsymbol{\zeta},\theta\}, we first compute the global rigid transformation {li∣i=1,...,P}\{l^{i}|i=1,...,P\} for each rigid part using the forward kinematics in Eqs. 6-7. To estimate the density and color (σ,c)(\sigma,c) of a global 3D location x{\bf x} from a 2D viewing direction d{\bf d}, we train a separate RT-NeRF, FΘili,ζF_{\Theta^{i}}^{l^{i},\boldsymbol{\zeta}},

for each part using Eq. 12 and combine the densities and colors {σi,ci∣i=1,...,P}\{\sigma^{i},c^{i}|i=1,...,P\} estimated by different RT-NeRFs into one. We denote this approach as Part-Wise NARF (NARFP\text{NARF}_{P}).

Since a surface point of an object can belong to only one of the object parts, only one of the estimates in {σi,ci∣i=1,...,P}\{\sigma^{i},c^{i}|i=1,...,P\} should be nonzero. The density and color (σ,c)(\sigma,c) of a global 3D location x{\bf x} can be determined by taking the estimate with the highest density. However, the max operation is not differentiable, so we instead use the softmax function, which is a differentiable weighted sum over all the estimates:

where τ\tau is the temperature parameter of the softmax function. Volume rendering is then applied using the combined density σ\sigma and color cc to generate the rendered color C(r){\bf C}({\bf r}) by Eq. 5. Since the rendering and softmax operations are both differentiable, the image reconstruction loss can pass gradients to all the RT-NeRF models for effective training.

We note that, in addition to color C(r){\bf C}({\bf r}), the foreground mask M(r){\bf M}({\bf r}) can also be estimated as an integral of the opacity along the camera ray:

We can further render a segmentation image that indicates which RT-NeRF (object part) is used for rendering each pixel:

where sjs_{j} is the index of the part with the greatest density, and Si{\bf S}^{i} denotes a segmentation mask for the ithi^{th} part.

Discussion. The NARFP\text{NARF}_{P} approach models the rigidly transformed parts of an articulated object in separate RT-NeRFs, where each part has a consistent radiance field under different 3D pose configurations. As the rigid transformation of each RT-NeRF is computed explicitly via forward kinematics, rather than implicitly within the network, the issue of implicit transformations and part dependency are avoided. The part dependency issue is also addressed by taking the estimate with the highest density while suppressing the contribution of other parts for a global 3D location in Eq. 15. However, its computation is inefficient for the following reasons.

The computational cost is proportional to the number of object parts, limiting the representation capacity for complex articulated objects.

Training is dominated by the large number of zero density point samples. As a surface point on an object can belong to only one of the object parts, it will be trained as a zero density sample for the remaining parts. Since parts with small densities do not affect the value of the equation very much, it is not really necessary to calculate the density of those parts.

To address the above issues, we present another approach that combines the inputs of the RT-NeRF models in NARFP\text{NARF}_{P} then feeds them as a whole into a single NeRF model for direct regression of the final density and color (σ,c)(\sigma,c). We call this approach Holistic NARF (NARFH\text{NARF}_{H}). Formally,

where Cat denotes the concatenation operator.

Discussion. There is only a single NeRF model trained in NARFH\text{NARF}_{H}. The computational cost is almost constant to the number of object parts and the zero density problem is naturally avoided. However, unlike Part-Wise NARF, NARFH\text{NARF}_{H} does not satisfy Part Dependency, because all parameters are considered for each 3D location. Moreover, object part segmentation masks cannot be generated from Eq. 18 without part dependencies.

We propose Disentangled NARF (NARFD\text{NARF}_{D}) which shares the merits of both NARFP\text{NARF}_{P} and NARFH\text{NARF}_{H} while avoiding their weaknesses by introducing a selector S\mathcal{S}.

The selector S\mathcal{S} identifies which object part a global 3D location x{\bf x} belongs to. S\mathcal{S} consists of PP lightweight sub-networks for each part. For the ithi^{th} part, a sub-network OΓiO_{\Gamma}^{i} takes the local 3D position of x{\bf x} in li={Ri,ti}l^{i}=\{R^{i},{\bf t}^{i}\} and the bone parameter ζ\boldsymbol{\zeta} as input and outputs the probability pip^{i} of x{\bf x} belonging to the ithi^{th} part. Since x{\bf x} should be assigned to only one of the object parts, the softmax activation is used to normalize the selector’s outputs:

It can be seen that OΓiO_{\Gamma}^{i} is actually an occupancy network defined in the local object coordinate system. Comparing this with NASA , NASA’s occupancy networks learn absolute occupancy values to estimate an explicit surface, but our networks learn relative occupancy values to other parts for part selection. For implementation, we use a two-layer MLP with ten hidden nodes for each occupancy net, which is lightweight yet effective.

Disentangled NARF is defined by (softly) masking out the irrelevant parts in the concatenated input using the outputs pip^{i} of the selector.

Note that though we have removed the dependency on irrelevant parts by masking their inputs, the resulting input is still in the form of a concatenation. This is done purposely because all the bones share a single NeRF, which needs to distinguish the different bones in order to generate the corresponding density and color. Different bones are distinguished by dimensions of the concatenated vector that represent them, similar to part identity encoding. The detailed network architecture can be found in Fig. 7 of the supplemental material.

Since the selector outputs the probabilities of a global 3D location belonging to each part, we can generate the segmentation mask by selecting the locations occupied by a specific part followed by Eq. 18:

5 Training Details

The positional encoding dimensions we set for the 3D location x{\bf x} and other parameters are 10 and 4 respectively. During training, at each optimization iteration, we randomly sample a batch of camera rays from the set of all pixels, and then follow the hierarchical volume sampling strategy of the original NeRF to query NN samples for each ray. With known kinematic 3D pose configuration {T0,ζ,θ}\{{\bf T}^{0},\boldsymbol{\zeta},\theta\}, the samples’ densities and colors are estimated by the NARF model. Volume rendering is then used to render the color C(r){\bf C}({\bf r}) and mask M(r){\bf M}({\bf r}) of this ray using Eq. 5 and 16, respectively. The loss is the total squared error between the rendered and true pixel colors and masks.

where R\mathcal{R} is the set of rays in each batch, C^\hat{\bf C} and M^\hat{\bf M} are the ground truth color and foreground mask. In the supplemental material, we empirically show that the extra mask loss helps to learn a cleaner background. Other training details on the learning rate, batch size and optimizer can be found in the supplemental material.

Results of Training on a Single Object

In this section, we evaluate our model in the case of a single articulated 3D object.

We create our own synthetic dataset of human bodies for experimentation. It consists of two persons, one male and one female, selected from the human 3D textured mesh (THUman) dataset . Each person has 56 and 48 different poses, 26 of which are used for training and the others for testing. We render 100 images with various orientations and scaling of each mesh for training and 20 for testing, ending up with 2600 training images for each person. Note that the viewpoint distribution of training and testing sets are the same under this setting. We denote this test data setting as the novel pose/same view setting. All rendered images have a resolution of 128 ×\times 128. Additionally, we introduce three other test settings for a more comprehensive comparison. The same pose/same view setting uses testing images rendered from the same poses and same viewpoint distribution as in training. The novel pose/novel view setting uses novel poses and a different viewpoint distribution than in training. Finally, the same pose/novel view setting uses the same poses but the viewpoint distribution is different from training. The kinematic 3D pose configurations are inferred from the SMPL model parameters provided by the THUman dataset and we use bone length as the bone parameter ζ\zeta in Eq. 10.

Metrics

Three metrics are used to evaluate performance. The peak signal to noise ratio (PSNR) and structural similarity index (SSIM) are two commonly used evaluation metrics for image reconstruction (higher is better). In addition, we introduce the L2 distance error of mask images (Mask), which better describes how close the 3D shape of the rendered object is to the ground truth (lower is better).

Baselines

In addition to the three variants of NARF (NARFP\text{NARF}_{P}, NARFH\text{NARF}_{H} and NARFD\text{NARF}_{D}), three other baselines are included for comparison. The first is a 2D CNN-based method similar to that generates the target subject image from “pose stick figures”. The pose stick figures in our case are generated by projecting the 3D joints into the 2D image (with given camera parameters) then adding lines to connect these 2D keypoints. The second is the P-NeRF method described in Sec. 3.2. The third one is D-NARF, a simple extension of D-NERF to articulated objects. D-NARF aims to learn the mapping Ψ:x→x′\Psi:{\bf x}\rightarrow{\bf x}^{\prime} that transforms a given point to its position in a canonical shape space. In our implementation, a static NeRF model for a canonical pose Pc\mathcal{P}^{c} is learned, then a mapping network Ψ\Psi estimates the deformation field between the scene of a specific pose instance P\mathcal{P} and the scene of the canonical pose Pc\mathcal{P}^{c}. The details of the three baselines can be found in the supplemental material.

Results

Quantitative comparison results are given in Table 1. #Params, #FLOPS, and #Memory denote the number of parameters, floating point operations per ray, and number of elements to preserve during forward propagation per ray, which is proportional to the memory cost.

It can be seen that our method, NARFD\text{NARF}_{D}, outperforms the others under all the evaluation metrics and test data settings (best results shown in bold). Particularly, it exhibits high performance under novel pose and/or novel view settings (slight performance drop on SSIM within 4%) with low computational cost (close to a single NeRF model, P-NeRF). Hence, we can conclude that NARFD\text{NARF}_{D} effectively and efficiently learns the radiance field of an articulated 3D object and the model generalizes to novel poses and views with high fidelity.

In contrast, all the other methods are deficient in one way or another. The CNN-based method fails when tested under novel views (10% performance drop on SSIM) since it is difficult to learn an effective 3D representation from 2D inputs. P-NeRF and D-NARF fail in almost all cases mainly due to both the Implicit Transformations and Part Dependency issues. NARFP\text{NARF}_{P} exhibits good performance and generalization ability but requires much more computation (10×10\times #FLOPS and 17×17\times #Memory of NARFD\text{NARF}_{D}). NARFH\text{NARF}_{H} is less effective when tested on novel poses (8% performance drop on SSIM) due to the Part Dependency issue.

Qualitative results under the novel pose/novel view setting are shown in Fig. 3. Rendered RGB images (first row), depth maps (second row), and part segmentation maps (third row), as well as ground truth RGB images (bottom left corner) are displayed. It can be seen that the NeRF based methods (except for the CNN-based one) can obtain depth images and the “part dependent” methods (NARFP\text{NARF}_{P} and NARFD\text{NARF}_{D}) can obtain segmentation maps. Our final solution NARFD\text{NARF}_{D} generates higher quality RGB, depth and segmentation maps for novel views and poses than the others. Moreover, as shown in Fig. 4, NARFD\text{NARF}_{D} learns a disentangled representation of camera viewpoint, bone parameters and pose, allowing these appearance properties to be individually controlled in rendering.

Appearance Variation with Autoencoder

In this section, we train an autoencoder based on NARF to model shape and appearance variation among multiple articulated objects. The autoencoder consists of an encoder and decoder. First, a 2D CNN-based encoder is used to generate a latent vector z{\bf z} from an input image. The obtained latent vector together with the given camera viewpoint and human pose are fed into our NARF based decoder to reconstruct the input image.

Following the implementation for the NeRF-based generator , we first decompose z{\bf z} into a shape latent vector zs{\bf z}_{s} and an appearance latent vector za{\bf z}_{a}. Then, zs{\bf z}_{s} is concatenated to the density-dependent inputs, namely, the positionally encoded location x{\bf x} and bone parameters ζ\boldsymbol{\zeta}. Meanwhile, za{\bf z}_{a} is concatenated to the color-dependent inputs, namely, the positionally encoded view direction d{\bf d} and the local transformation ξ\boldsymbol{\xi}. Specifically, when combining the autoencoder with the NARFD\text{NARF}_{D} model, we have

The encoder and decoder are trained jointly using the same loss in Eq. 25. For the experiment, we use the best performing NARFD\text{NARF}_{D} by default. Comparison results for other models can be found in the supplemental material.

We create another synthetic human body dataset from THUman for experimentation. All of the males (112 in total) and poses (35 on average for each person) in the THUman dataset are used. We render 10 images for each pose with randomly sampled viewpoints to generate 35450 images for training and 3940 images for evaluation. All rendered images have a resolution of 128 ×\times 128.

Results

Fig. 5 shows the reconstructed RGB images as well as the additional depth images and segmentation maps generated from the input RGB images. A single autoencoder is trained for all objects, indicating that appearance variation is effectively modeled. Fig. 6 shows that the NARF-based autoencoder learns a disentangled representation of camera viewpoint, bone parameters, human pose, and color appearance, allowing these properties to be individually controlled in rendering. For color appearance, it is controlled by replacing the appearance latent vector za{\bf z}_{a} with that of another person. Additional results can be found in the supplemental material.

Conclusion and Future work

In this paper, we propose a method for learning implicit representations for articulated objects. We show that it is possible to learn explicitly controllable representations of viewpoint, pose, bone parameters, and appearance from 3D pose annotated images. Although pose annotation is required to train the model, the model is differentiable and thus may be extended to reduce the required supervisory information, for example, by simultaneously training 3D pose estimation and segmentation with the model. In addition, since the proposed representation provides explicit 3D shape and part segmentation, it may be applied to unsupervised depth estimation and segmentation learning.

Acknowledgement

This work was supported by D-CORE Grant from Microsoft Research Asia and partially supported by JST AIP Acceleration Research Grant Number JPMJCR20U3, and JSPS KAKENHI Grant Number JP19H01115. We would like to thank Sho Maeoki, Thomas Westfechtel, and Yang Li for the helpful discussions. We are also grateful for the GPU resources provided by Microsoft Azure Machine Learning.

References

Appendix A Ablation Studies for AutoEncoder

In Table 1 of the main paper, we quantitatively evaluate our model in the case of a single articulated 3D object, while in Section 5 in the main paper, we only show qualitative results in the case of Autoencoder. Here, we supplement those results with the quantitative ablation study for Autoencoder in Table 2. The same test data settings (namely “same pose/same view”, “novel pose/same view”, “same pose/novel view”, and “novel pose/novel view”), metrics (namely PSNR, SSIM, and Mask), and baseline methods (namely NARFP\text{NARF}_{P}, NARFH\text{NARF}_{H} and NARFD\text{NARF}_{D}, CNN-based, P-NeRF and D-NARF) are used for comparison. The same dataset as in Section 5 of the main paper is used for experimentation.

At testing phase, when extracting the latent shape and appearance vectors (zs{\bf z}_{s} and za{\bf z}_{a}) using the encoder, we use images under the same viewpoint distribution as in the training images as input. Then, images from novel views and poses are rendered by combining zs{\bf z}_{s} and za{\bf z}_{a} with unseen views and poses in the training data.

Quantitative comparison results are given in Table 2. Qualitative results under the novel pose/novel view setting are shown in Fig. 8. Consistent with the case of a single object, our method NARFD\text{NARF}_{D} outperforms the others under all the evaluation metrics and test data settings (best results shown in bold). High quality depth and segmentation images are jointly generated as shown in Fig. 8 (rightmost column). CNN based models cannot represent 3D structure effectively, so the performance drops significantly in the “novel view” settings. Fig. 8 (left-top) shows that in the novel view testing, the CNN-based method produces fuzzy images. Meanwhile, it cannot generate depth and segmentation images. P-NeRF and D-NARF fail in almost all settings due to implicit transformation and part dependency problems. NARFP\text{NARF}_{P} generally performs well, but the computational cost is too high. NARFH\text{NARF}_{H} has poor performance on “novel pose” due to part dependency issues. The performance drop on novel pose in the Autoencoder case is not as significant as in the single object case (shown in Table 1 of the main paper) since the pose diversity in the training data is much larger in the Autoencoder case.

Appendix B Ablation Studies in RT-NeRF

In Section 3.3 of the main paper, we introduced the rigidly transformed neural radiance field (RT-NeRF) to effectively model a rigidly transformed object part. Here, we evaluate the effectiveness of the two most critical design elements in RT-NeRF. The first is the explicit transformation that converts a global 3D location into the local coordinate system and the local 3D location is then used to estimate the density using Eqs. 9 and 10 of the main paper. The second is the pose-dependent color estimation defined in Eq. 11 of the main paper. It takes the 6D vector se(3)\mathfrak{se}(3) representation ξ\boldsymbol{\xi} of transformation ll as a network input to estimate the RGB color cc. To this end, two more baseline methods are introduced accordingly to compare to RT-NeRF. The first is the rigid pose conditioned NeRF (RP-NeRF) that takes the global 3D location and the rigid transformation ξ\boldsymbol{\xi} as network inputs, similar to the P-NeRF defined in Eq. 8 of the main paper. The second is RT-NeRF w/o ξ\boldsymbol{\xi} that estimates the RGB color cc without using the transformation ξ\boldsymbol{\xi} as input in Eq. 11 of the main paper.

We create a synthetic rigid object dataset of a rendered bulldozer using Blender (a software for rendering) for experimentation. In the dataset, the object (a bulldozer) can rigidly transform in the world coordinate system. For each rendered image, both rigid transformation and the camera viewpoint are randomly set. The camera will be translated to point to the center of the object so that the object will appear in the center of the rendered image. The resolution for all rendered images is set to 200 ×\times 200. In total, 480 images are used for training and another 20 images are used for testing. The loss function is the same as in Eq. 25 of the main paper.

Results

The quantitative results are shown in Table 3 and the qualitative results are shown in Fig. 9. The experimental results show that RP-NeRF is unable to learn a good 3D representation and fails to generalize to novel poses and views. In contrast, RT-NeRF effectively models the rigidly transformed object by explicitly transforming the global 3D location into the local coordinate system. In addition, the color estimation without the transformation input is less effective. This is concluded by comparing the results of “RT-NeRF w/o ξ\xi” with the results of “RT-NeRF”. Quantitatively, the performance of “RT-NeRF w/o ξ\xi” drops significantly under the Mask metric in Table 3. Qualitatively, the rendered images from “RT-NeRF w/o ξ\xi” look blurry compared to “RT-NeRF” in Fig. 9.

Appendix C Additional Ablation Studies in NARF

In this section, we provide more ablation studies on the effects of the mask loss (the second item in Eq. 25 of the main paper), temperature parameter for NARFP\text{NARF}_{P} (in Eq. 15 of the main paper), and softmax activation function for NARFD\text{NARF}_{D} (in Eq. 21 of the main paper).

We conduct experiments to evaluate the effect of the mask loss (the second term of Eq. 25 of the main paper) added to the rendered mask image. While the color loss optimizes the final RGB color of the rendered pixels, the gradient from the mask loss directly optimizes the densities of the 3D locations on the camera rays. Therefore, the additional mask loss is helpful in learning 3D shapes efficiently. The quantitative and qualitative results of w/ and w/o mask loss are shown in Table 4 and Fig. 10 respectively. In Table 4, it can be seen that performances of all the three variants of NARF drop significantly without the mask loss, especially under the novel pose/novel view setting. Particularly, in Fig. 10, the NARFH\text{NARF}_{H} model is not able to converge at all without the mask loss. The NARFP\text{NARF}_{P} and NARFD\text{NARF}_{D} models still work without the mask loss but the rendered images get very blurry, especially on the background regions around the object.

We study the effect of the temperature parameter τ∈(0,∞)\tau\in(0,\infty) in NARFP\text{NARF}_{P} (in Eq. 15 of the main paper). The temperature parameter determines how soft the selection is among the multiple RT-NeRFs. When τ\tau is close to , hard selection is performed. Though the Part Dependency prior is strictly satisfied in this case, convergence in the training is difficult since the highest current estimate will completely block the gradient from back-propagating to the others. It is especially worse in the early stage of the training when the highest estimate is almost random. In turn, when τ\tau is close to ∞\infty, averaging is performed. In this case, the gradient is back-propagated to all RT-NeRFs, but the Part Dependency issue arises again, which will harm the generalization ability to novel poses. The quantitative and qualitative results are shown in Table 4 and Fig. 11, respectively. We empirically use the best-performing τ=100\tau=100 setting as shown in Table 4.

In Eq. 21 of the main paper, we use softmax activation, which is also motivated by the Part Dependency prior. Here, we provide the results of using the sigmoid activation as an alternative. Formally, Eq. 21 of the main paper is replaced with Eq. 28:

The quantitative results are shown in Table 4. In conclusion, softmax outperforms sigmoid, especially in the “novel pose/novel view” setting.

Appendix D Implementation Details in P-NeRF

In Eq. 8 of the main paper, P-NeRF takes a global 3D location x{\bf x} and the part transformations {li∣i=1,...,P}\{l^{i}|i=1,...,P\} as input. For implementation, we use the 6D vector se(3)\mathfrak{se}(3) representation ξi{\boldsymbol{\xi}}^{i} of transformation lil^{i}, and concatenate them together with x{\bf x} and the bone length ζ\zeta as the network input. Positional encoding is performed before the concatenation. Formally,

The density and color sub-networks are defined as

Appendix E Implementation Details in D-NARF

D-NERF uses a canonical template and learns the observation-to-canonical deformation

where ω\omega is a deformation latent code.

D-NERF is defined on the deformed position x′{\bf x}^{\prime} in the canonical template,

where ψ\psi is a latent appearance code.

In our case, ω\omega and ψ\psi correspond to the pose configuration P\mathcal{P} and the appearance latent vector za{\bf z}_{a} respectively. Ψ\Psi is implemented using MLP, which seems to suffer from the problem of implicit transformation. In our setup, the pose of each part is given, so it can be implemented more directly. In our implementation, we first use the occupancy network similar to the one defined in Eq. 21 of the main paper to decide which part the input point belongs to.

Then, we calculate the coordinates x′{\bf x}^{\prime} on the canonical shape as

where tcanonicali{\bf t}_{\text{canonical}}^{i} is the origin of a canonical pose’s ithi^{\text{th}} part in the global coordinate system. View direction in the canonical space and the transformation vector are defined as

Then, the D-NARF we have implemented in the experiment is defined as

This implementation is similar to the implementation of NARFD\text{NARF}_{D}, differing only in how the coordinates are input to the model. The results of the experiments show that a concatenation-based NARFD\text{NARF}_{D} model that retains the coordinates for all parts is more effective than a transformation on the input coordinates in D-NARF.

Appendix F Training Details

We used the Adam optimizer with an initial equalized learning rate of 0.010.01. The learning rate is decayed to 0.99995×0.99995\times of the previous iteration. Particularly, P-NeRF and NARFH\text{NARF}_{H} based autoencoders are trained with an initial learning rate of 0.001 since the training will explode if a learning rate of 0.01 is used. The batch size is set to 16 for all experiments. We sample as many camera rays as can be fit in the GPU memory. The training converges at about 100,000 iterations. The training of our method NARFD\text{NARF}_{D} takes 24 hours on 4 V100 GPUs. The code for creating our synthetic datasets is available at https://github.com/nogu-atsu/NARF.

Appendix G Cross Dataset Evaluation on SURREAL Dataset

In order to verify the generalization ability of NARF autoencoder across different datasets, we use a cross dataset evaluation protocol that trains the model on THUman dataset , then tests it on SURREAL dataset . There are several differences between the THUman and SURREAL datasets. First, even though the human samples in SURREAL contain rendered SMPL meshes and textures as in THUman, but unlike in THUman, the meshes do not contain clothing. Second, the camera and shape parameters in SURREAL, including the distance between camera and human, human pose distribution and the range of body size are quite different from that in THUman. Even so, we test images from the test set of SURREAL using the NARF autoencoder trained on THUman. The qualitative reconstruction results are shown in Fig. 12. From left to right, we show the input SURREAL images, their reconstructed images, and the images re-rendered under different pose configurations. Since the THUman dataset is not diverse enough in terms of clothing and body size, the reconstructions of short pants and fat people are less effective. But the novel view and pose rendering results look quite reasonable.

Appendix H Experiment on Real Human Images

In this section, we test our approach on a real human dataset ZJU-MOCAP . ZJU-MOCAP is a multi-view person video dataset. For each frame, SMPL parameters are given. We use the first 1969 frames (90%, 2185 frames in total) of the Taichi class video for training and the remaining 216 frames (10%) for testing (novel pose). The resolution of the image is 512 ×\times 512.

The qualitative results of NARFD\text{NARF}_{D} on this dataset are shown in Fig. 13. The left part of Fig. 13 shows the pose used in the training, but rendered from novel viewpoints, and the right part of Fig. 13 shows the novel pose/novel view testing results. The quality of the rendered images is not as good as testing on our synthetic datasets. This might be caused by the assumption that the parts are rigid objects, which may not be perfectly satisfied for real images. For example, loose clothes may move when a person makes a movement. This issue can be considered in future work, for example, by learning latent variables to account for both pose-dependent and pose-independent deformations similar to Neural Body .

Although the quality of the rendered images for a real person from our method still has room for improvement, we believe that the proposed explicitly controllable representation of viewpoint, pose, bone parameters, and appearance for the articulated object is an important contribution.