Unsupervised Learning of Efficient Geometry-Aware Neural Articulated Representations

Atsuhiro Noguchi, Xiao Sun, Stephen Lin, Tatsuya Harada

Introduction

3D models that allow free control over the pose and appearance of articulated objects are essential in various applications, including computer games, media content creation, and augmented/virtual reality. In early work, articulated objects were typically represented by explicit models such as skinned meshes. More recently, the success of learned implicit representations such as neural radiance fields (NeRF) for rendering static 3D scenes has led to extensions for modeling dynamic scenes and articulated objects. Much of this attention has focused on the photorealistic rendering of humans, from novel viewpoints and with controllable poses, by learning from images and videos.

Existing methods for learning explicitly pose-controllable articulated representations, however, require much supervision, such as videos with 3D pose/mesh annotation and a mask for each frame. Preparing such data involves tremendous annotation costs; thus, reducing annotation is very important. In this paper, we propose a novel unsupervised learning framework for 3D pose-aware generative models of articulated objects, which are learned only from unlabeled images of objects sharing the same structure and a pose prior distribution of the objects.

We exploit recent advances in 3D-aware GAN for unsupervised learning of the articulated representations. They learn 3D-aware image generation models from images without supervision, such as viewpoints or 3D shapes. The generator is based on NeRF and is optimized with a GAN objective to generate realistic images from randomly sampled viewpoints and latent vectors from a prior distribution defined before training. As a result, the generator learns to generate 3D-consistent images without any supervision. We employ the idea for articulated objects by defining a pose prior distribution for the target object and optimizing the GAN objective on randomly generated images from random poses and latent variables. It becomes possible to learn a generative model with free control of poses. We demonstrate this approach by modeling the pose prior as a skeletal distribution , while noting that other models like meshes may bring potential performance benefits.

However, the direct application of existing neural articulated representations to GANs is not computationally practical. While NeRF can produce high-quality images, its processing is expensive because it requires network inference for every point in space. Some methods reduce computational cost by volume rendering at low resolution followed by 2D CNN based upsampling. Although this technique achieves high-resolution images with real-time inference speed, it is not geometry-aware (i.e., the surface mesh cannot be extracted). Recently, a method which we call Efficient NeRF overcomes the problem. The method is based on an efficient tri-plane based neural representation and GAN training on it. Thanks to the computational efficiency, it can produce relatively high-resolution (128 ×\times 128) images with volumetric rendering. We extend the tri-plane representation to articulated objects for efficient GAN training. An overview of the method is visualized in Figure 1. The contributions of this work are as follows:

We propose a novel efficient neural representation for articulated objects based on an efficient tri-plane representation.

We propose an efficient implementation of deformation fields using tri-planes for dynamic scene training, achieving 4 times faster rendering than NARF with comparable or better performance.

We propose a novel GAN framework to learn articulated representations without using any 3D pose or mask annotation for each image. The controllable 3D representation can be learned from real unlabeled images.

Related Work

Articulated 3D Representations. The traditional approach for modeling pose-controllable 3D representations of articulated objects is by skinned mesh , where each vertex of the mesh is deformed according to the skeletal pose. Several parametric skinned mesh models have been developed specifically for humans and animals . For humans, the skinned multi-person linear model (SMPL) is commonly used. However, these representations can only handle tight surfaces with no clothing and cannot handle non-rigid or topology-changing objects such as clothing or hair. Some work alleviates the problem by deforming the mesh surface or using a detailed 3D body scan . Recently, implicit 3D shape representations have achieved state-of-the-art performance in pose-conditioned shape reconstruction. These methods learn neural occupancy/indicator functions or signed distance functions of articulated objects . Photorealistic rendering of articulated objects, especially for humans, is also achieved with 3D implicit representations . However, all these models require ground truth 3D pose and/or object shape for training. Very recently, methods have been proposed to reproduce the 3D shape and motion of objects from video data without using 3D shape and pose annotation . However, they either do not allow free control of poses or are limited to optimizing for a single object. An SMPL mesh-based generative model and image-to-image translation methods can learn pose controllable image synthesis models for humans. However, the rendering process is completely in 2D and thus is not geometry-aware.

Implicit 3D representations. Implicit 3D representations are memory efficient, continuous, and topology free. They have achieved the state-of-the art in learning 3D shape , static and dynamic scenes , articulated objects , and image synthesis . Although early works rely on ground truth 3D geometry for training , developments in differentiable rendering have enabled learning of networks from only photometric reconstruction losses . In particular, neural radiance fields (NeRF) applied volumetric rendering on implicit color and density fields, achieving photorealistic novel view synthesis of complex static scenes using multi-view posed images. Dynamic NeRF extends NeRF to dynamic scenes, but these methods just reenact the motion in the scene and cannot repose objects based on their structure. Recently, articulated representations based on NeRF have been proposed . These methods can render images conditioned on pose configurations. However, all of them require ground truth skeletal poses, SMPL meshes, or foreground masks for training, which makes them unsuitable for in-the-wild images.

Another NeRF improvement is the reduction of computational complexity: NeRF requires forward computation of MLPs to compute color and density for every point in 3D. Thus, the cost of rendering is very high. Fast NeRF algorithms reduce the computational complexity of neural networks by creating caches or using explicit representations. However, these methods can only be trained on a single static scene. Very recently, a hybrid explicit and implicit representation was proposed . In this representation, the feature field is constructed using a memory-efficient explicit representation called a tri-plane, and color and density are decoded using a lightweight MLP. This method can render images at low cost and is well suited for image generation models.

In this work, we propose an unsupervised learning framework for articulated objects. We extend tri-planes to articulated objects for efficient training.

Generative 3D-aware image synthesis Advances in generative adversarial networks (GANs) have made it possible to generate high-resolution, photorealistic images . In recent years, many 3D-aware image generation models have been proposed by combining GANs with 3D generators that use meshes , voxels , depth , or implicit representations . These methods can learn 3D-aware generators without 3D supervision. Among these, image generation methods using implicit functions, thanks to their continuous and topology-free properties, have been successful in producing 3D-consistent and high-quality images. However, fully implicit models are computationally expensive, making the training of GANs inefficient. Therefore, several innovations have been proposed to reduce the rendering cost of generators. Neural rendering-based methods reduce computation by performing volumetric rendering at low resolution and upsampling the rendered feature images using a 2D CNN. Though this enables the generation of high-resolution images at a faster rate, 2D upsampling does not consider 3D consistency and cannot generate detailed 3D geometry. Very recently, a hybrid of explicit and implicit methods has been developed for 3d geometry-aware image generation. Instead of using a coordinate-based implicit representation, this method uses tri-planes, which are explicit 3D feature representations, to reduce the number of forward computations of the network and to achieve volumetric rendering at high resolution.

The existing research is specific to scenes of static objects or objects that exist independently of each other, and do not allow free control of the skeletal pose of the generated object. Therefore, we propose a novel GAN framework for articulated objects.

Method

Recent advances in implicit neural rendering have made it possible to generate 3D-aware pose-controllable images of articulated objects from images with accurate 3D pose and foreground mask annotations. However, training such models from only in-the-wild images remains challenging since accurate 3D pose annotations are generally difficult to obtain for them. In the following, we first briefly review Neural Articulated Radiance Field (NARF) , then propose an adversarial-based framework, named ENARF-GAN, to efficiently train the NARF model without any paired image-pose and foreground mask annotations.

NARF is an implicit 3D representation for articulated objects. It takes a kinematic 3D pose configuration of an articulated object o={lk,Rk,tk}k=1:Ko=\{l_{k},R_{k},{\bf t}_{k}\}_{k=1:K} as input and predicts the color and the density of any 3D location x{\bf x}, where lkl_{k} is the length of the kthk^{\text{th}} part, and RkR_{k} and tk{\bf t}_{k} are its rotation and translation matrices, respectively. Given the pose configuration oo, NARF first transforms a global 3D position x{\bf x} into several local coordinate systems defined by the rigid parts of the articulated object. Specifically, the transformed local location xkl{\bf x}_{k}^{l} for the kthk^{\text{th}} part is computed as xkl=(Rk)−1(x−tk) for k∈{1,...,K}{\bf x}_{k}^{l}=(R^{k})^{-1}({\bf x}-{\bf t}^{k})\text{ for }k\in\{1,...,K\}.

NARF first trains an extra lightweight selector SS in the local space to decide which part a global 3D location x{\bf x} belongs to. Specifically, it outputs the probability pkp^{k} of x{\bf x} belonging to the kthk^{th} part. Then NARF computes color cc and density σ\sigma at the location x{\bf x} from a concatenation of local locations masked by the corresponding part probability pkp^{k}.

where γ\gamma is a positional encoding , Cat is the concatenation operation, and GG is an MLP network. The RGB color C{\bf C} and foreground mask value M{\bf M} for each pixel are generated by volumetric rendering . The network is trained with a reconstruction loss between the generated and ground truth color C^\hat{\bf C} and mask M^\hat{\bf M},

where R\mathcal{R} is the set of rays in each batch. Please refer to the original NARF paper for more details.

2 Unsupervised Learning by Adversarial Training

An overview of this method is illustrated in Figure 2 (b).

However, this training would be computationally expensive. The rendering cost of NARF is heavy because computation is performed for many 3D locations in the viewed space. Even though in supervised training, the time and memory cost of computing the reconstruction loss could be reduced by evaluating it over just a small proportion of the pixels , the adversarial loss in Equation 3 requires the generation of the full image for evaluation. As a result, the amount of computation becomes impractical.

In the following, we propose a series of changes in feature computation and the selector to address this issue. Note that these changes not only enable the GAN training but also greatly improve the efficiency of the original NARF.

3 Efficiency Improvements on NARF

Recently, Chan et al. proposed a hybrid explicit-implicit 3D-aware network that uses a memory-efficient tri-plane representation to explicitly store features on axis-aligned planes. With this representation, the efficiency of feature extraction for a 3D location is greatly improved. Instead of forwarding all the sampled 3D points through the network, the intermediate features of arbitrary 3D points can be obtained via simple lookups on the tri-planes. The tri-plane representation can be more efficiently generated with Convolutional Neural Networks (CNNs) instead of MLPs. The intermediate features of 3D points are then transformed into the color and density using a lightweight decoder MLP. The decoder significantly increases the non-linearity between features at different positions, thus greatly enhancing the expressiveness of the model.

Here, we adapt the tri-plane representation to NARF for more efficient training. Similar to , we first divide the original NARF network GG into an intermediate feature generator (the first linear layer of GG) WW and a decoder network GdecG_{\text{dec}}. Then, Equation 1 is rewritten as follows.

where f{\bf f} is an intermediate feature vector of input 3D location x{\bf x}. However, re-implementing the feature generator WW to produce the tri-plane representation is not straightforward because its input, a weighted concatenation of xkl{\bf x}_{k}^{l}, does not form a valid location in a specific 3D space. We make two important changes to address this issue. First, we decompose WW into KK sub-matrices {Wk}\{W_{k}\}, one for each part, where each one takes the corresponding local position xkl{\bf x}_{k}^{l} as input and outputs an intermediate feature for the kthk^{\text{th}} part. Then, the intermediate feature in Equation 4 can be equivalently rewritten as follows.

where fk{\bf f}^{k} is a feature generated in the local coordinate system of the kthk^{\text{th}} part. Now, Wk(γ(xkl))W_{k}(\gamma({\bf x}_{k}^{l})) can be directly re-implemented by tri-planes FF. However, the computational complexity of this implementation is still proportional to KK. In order to train a single tri-plane for all parts, the second change is to further transform the local coordinates xkl{\bf x}_{k}^{l} into a canonical space defined by a canonical pose oco^{c}, similar to Animatable NeRF .

where RkcR_{k}^{c} and tkc{\bf t}_{k}^{c} are the rotation and translation matrices of the canonical pose. Intuitively, xkc{\bf x}_{k}^{c} is the corresponding point location of x{\bf x} transformed into the canonical space when x{\bf x} is considered to belong to the kthk^{\text{th}} part. Finally, the tri-plane feature FF is learned in the canonical space. The feature extraction for location x{\bf x} is achieved by retrieving the 32-dimensional feature vector fk{\bf f}^{k} on the tri-plane FF at xkc{\bf x}_{k}^{c} for all parts, then taking a weighted sum of those features as in Equation 5,

F∗∗(a)F_{**}({\bf a}) is a retrieved feature vector from each axis-aligned plane at location a{\bf a}.

We estimate the RGB color cc and density σ\sigma from f{\bf f} using a lightweight decoder network GdecG_{\text{dec}} consisting of three FC layers with a hidden dimension of 64 and output dimension of 4. We apply volume rendering on the color and density to output an RGB image C{\bf C} and a foreground mask M{\bf M}.

Although we efficiently parameterize the intermediate features, the probability pkp^{k} needs to be computed for every 3D point and every part. In the original NARF, though lightweight MLPs are used to estimate the probabilities, they are still computationally infeasible.

Therefore, we propose an efficient selector network using tri-planes. Since pkp^{k} is used to mask out features of irrelevant parts, this probability can be a rough approximation of the shape of each part. Thus the tri-plane representation is expressive enough to model the probability. We use KK separate 1-channel tri-planes to represent PkP^{k}, the part probability projected to each axis-aligned plane. We retrieve the three probability values (pxyk,pxzk,pyzk)=(Pxyk(xkc),Pxzk(xkc),Pyzk(xkc))(p_{xy}^{k},p_{xz}^{k},p_{yz}^{k})=(P^{k}_{xy}({\bf x}_{k}^{c}),P^{k}_{xz}({\bf x}_{k}^{c}),P^{k}_{yz}({\bf x}_{k}^{c})) of the kthk^{\text{th}} part by querying the 3D location in the canonical space xkc{\bf x}_{k}^{c}. The probability pkp^{k} that xkc{\bf x}_{k}^{c} belongs to the kthk^{\text{th}} part is approximated as pk=pxykpxzkpyzkp^{k}=p_{xy}^{k}p_{xz}^{k}p_{yz}^{k}.

In this way, the features FF and part probabilities PP are modeled efficiently with a single tri-plane representation. The tri-plane is represented by a (32+K)×3(32+K)\times 3 channel image. The first 96 channels represent the tri-plane features in the canonical space. The remaining 3K3K channels represent the tri-plane probability maps for each of the KK parts. We call this approach Efficient NARF, or ENARF.

4 GAN

To condition the generator on latent vectors, we utilize a StyleGAN2 based generator to produce tri-plane features. We condition each layer of the proposed ENARF by the latent vector with a modulated convolution . Since the proposed tri-plane based generator can only represent the foreground object, we use an additional StyleGAN2 based generator for the background.

We randomly sample latent vectors for our tri-plane generator and background generator: z=(ztri,zENARF,zb)∼N(0,I){\bf z}=({\bf z}_{\text{tri}},{\bf z}_{\text{ENARF}},{\bf z}_{b})\sim\mathcal{N}(0,I), where ztri{\bf z}_{\text{tri}}, zENARF{\bf z}_{\text{ENARF}}, and zb{\bf z}_{b} are latent vectors for the tri-plane generator, ENARF, and background generator, respectively. The tri-plane generator GtriG_{\text{tri}} generates tri-plane feature FF and part probability PP from randomly sampled ztri{\bf z}_{\text{tri}} and bone length {lk}k=1:K\{l_{k}\}_{k=1:K}. GtriG_{\text{tri}} takes lkl_{k} as inputs to account for the diversity of bone lengths.

The ENARF based foreground generator GENARFG_{\text{ENARF}} generates the foreground RGB image Cf{\bf C}_{f} and mask Mf{\bf M}_{f} from the generated tri-planes, and the background generator GbG_{\text{b}} generates background RGB image Cb{\bf C}_{b}.

The final output RGB image is C=Cf+Cb∗(1−Mf){\bf C}={\bf C}_{f}+{\bf C}_{b}*(1-{\bf M}_{f}), which is a composite of Cf{\bf C}_{f} and Cb{\bf C}_{b}. To handle the diversity of bone lengths, we replace Equation 6 with one normalized by the length of the bone: xkc=lkclkRkcxk+tkc{\bf x}_{k}^{c}=\frac{l_{k}^{c}}{l_{k}}R_{k}^{c}{\bf x}_{k}+{\bf t}_{k}^{c}, where lkcl^{c}_{k} is the bone length of the kthk^{\text{th}} part in the canonical space.

We optimize these generator networks with GAN training. We use a bone loss in addition to an adversarial loss on images, R1 regularization on the discriminator , and L2 regularization on the tri-planes. The bone loss ensures that an object is generated in the foreground. Based on the input pose of the object, a skeletal image BB is created, where pixels with skeletons are 1 and others are 0, and the generated mask MM at pixels with skeletons is made close to 1: Lbone=∑r∈R(1−M)2B∑r∈RB\mathcal{L}_{\text{bone}}=\frac{\sum_{r\in\mathcal{R}}(1-M)^{2}B}{\sum_{r\in\mathcal{R}}B}. Additional details are provided in the supplement. The final loss is the linear combination of these losses.

5 Dynamic Scene Overfitting

Since ENARF is an improved version of NARF, we can directly use it for single dynamic scene overfitting. For training, we use the ground truth 3D pose and foreground mask of each frame and optimize the reconstruction loss in Equation 2.

If the object shape is strictly determined by the poses that comprise the kinematic motion, we can use the same tri-plane features for the entire sequence and directly optimize them. However, real-world objects have time or pose-dependent non-rigid deformation such as clothing and facial expression change in a single sequence. Therefore, the tri-plane features should change depending on time and pose. We use a technique based on deformation fields proposed in time-dependent NeRF, also known as Dynamic NeRF. Deformation field based methods learn a mapping network from observation space to canonical space and learn the NeRF in the canonical frame. Since learning the deformation field with an MLP is expensive, we also approximate it with tri-planes. We approximate the deformation in 3D space by independent 2D deformations in each tri-plane. First, a StyleGAN2 generator takes positionally encoded time tt and a rotation matrix of each part RkR_{k} and generates 6-channel images representing the relative 2D deformation from the canonical space of each tri-plane feature. We deform the tri-plane feature based on the generated deformation. Please refer to the supplement for more details. We use constant tri-plane probabilities PP for all frames since the object shares the same coarse part shape throughout the entire sequence. The remaining networks are the same. We refer to this method as D-ENARF.

Experiments

Our experimental results are presented in two parts. First, in Section 4.1, we compare the proposed Efficient NARF (ENARF) with the state-of-the-art methods in terms of both efficiency and effectiveness, and we conduct ablation studies on the deformation modeling and the design choices for the selector. Second, in Section 4.2, we present our results of using adversarial training on ENARF, namely, ENARF-GAN, and compare it with baselines. Then, we discuss the effectiveness of the pose prior and the generalization ability of ENARF-GAN.

Following the training setting in Animatable NeRF , we train our ENARF model on synchronized multi-view videos of a single moving articulated object. The ZJU mocap dataset consisting of three subjects (313, 315, 386) is used for training. We use the same pre-processed data provided by the official implementation of Animatable NeRF. All images are resized to 512×512512\times 512. We use 4 views and the first 80% of the frames for training, and the remaining views or frames for testing. In this setting, the ground truth camera and articulated object poses, as well as the ground truth foreground mask, are given for each frame. More implementation details can be found in the supplement.

First, we compare our method with the state-of-the-art supervised methods NARF and Animatable NeRF . Our comparison with Neural Actor is provided in the supplement. Note that our method and NARF take ground truth kinematic pose parameters (joint angles and bone lengths) as inputs, while Animatable NeRF needs the ground truth SMPL mesh parameters. In addition, Animatable NeRF requires additional training on novel poses to render novel-pose images, which is not necessary for our model.

Table 1 shows the quantitative results. To compare the efficiency between models, we examine the GPU memory, FLOPS, and the running time used to render an entire image of resolution 512×512512\times 512 on a single A100 GPU as evaluation metrics. To compare the quality of synthesized images under novel view and novel pose settings, PSNR, SSIM , and LPIPS are used as evaluation metrics. Table 1 shows that the proposed Efficient NARF achieves competitive or even better performance compared to existing methods with far fewer FLOPS and 4.6 times the speed of the original NARF. Although the runtime of ENARF is a bit slower than Animatable NeRF (0.05s), its performance is superior under both novel view and novel pose settings. In addition, it does not need extra training on novel poses. Our dynamic model D-ENARF further improves the performance of ENARF with little increased overhead in inference time, and outperforms the state-of-the-arts Animatable NeRF and NARF by a large margin.

Qualitative results for novel view and pose synthesis are shown in Figure 3. ENARF produces much sharper images than NARF due to the more efficient explicit–implicit tri-plane representation. D-ENARF, which utilizes deformation fields, further improves the rendering quality. In summary, the proposed D-ENARF method achieves better performance in both image quality and computational efficiency.

Ablation Study To evaluate the effectiveness of the tri-plane based selector, we compare our method against models using an MLP selector or without a selector. A quantitative comparison is provided in Table 1, and a qualitative comparison is provided in the supplement. Although an MLP based selector improves the metrics a bit, it results in a significant increase in testing time. In contrast, the model without a selector is unable to learn clean/sharp part shapes and textures, because a 3D location will inappropriately be affected by all parts without a selector. These results indicate that our tri-plane based selector is efficient and effective.

2 Unsupervised Learning with GAN

In this section, we train the proposed efficient NARF using GAN objectives without any image-pose pairs or mask annotations.

Comparison with Baselines Since this is the first work to learn an articulated representation without image-pose pairs or mesh shape priors, no exact competitor exists. We thus compare our method against two baselines. The first is a supervised baseline called ENARF-VAE, inspired by the original NARF . Here, a ResNet50 based encoder estimates the latent vectors z{\bf z} from images, and the efficient NARF based decoder decodes the original images from estimated latent vectors z{\bf z} and ground truth pose configurations oo. These networks are trained with the reconstruction loss defined in Equation 2 and the KL divergence loss on the latent vector z{\bf z}. Following , ENARF-VAE is trained with images with a black background. The second model is called StyleNARF, which is a combination of the original NARF and the state-of-the-art high-resolution 3D-aware image generation model called StyleNeRF . To reduce the computational cost, the original NARF first generates low-resolution features using volumetric rendering. Subsequently, a 2D CNN-based network upsamples them into final images. Additional details are provided in the supplement. Please note that ENARF-VAE is a supervised method and cannot handle the background, and StyleNARF loses 3D consistency and thus cannot generate high-resolution geometry.

We use the SURREAL dataset for comparison. It is a synthetic human image dataset with a resolution of 128×128128\times 128. Dataset details are given in the supplement. For the pose prior, we use the ground truth pose distribution of the training dataset, where we randomly sample poses from the entire dataset. Please note that we do not use image-pose pairs for these unsupervised methods.

Quantitative results are shown in Table 2. We measure image quality with the Fréchet Inception Distance (FID) . To better evaluate the quality of foreground, we use an extra metric called FG-FID that replaces the background with black color, using the generated or ground truth mask. We measure depth plausibility by comparing the real and generated depth map. Although there is no ground truth depth for the generated images, the depth generated from a pose would have a similar depth to the real depth that arises from the same pose. We compare the L2 norm between the inverse depth generated from poses sampled from the dataset and the real inverse depth of them. Finally, we measure the correspondence between pose and appearance following the contemporary work named GNARF . We apply an off-the-shelf 2D human keypoint estimator to both generated and real images with the same poses and compute the Percentage of Correct Keypoints (PCK) between them, which is commonly used for evaluating 2D pose estimators. We report the averaged PCKh@0.5 metric for all keypoints. Details are provided in the supplement. Qualitative results are shown in Figure 5. Not surprisingly, ENARF-VAE produces the most plausible depth/geometry and learns the most accurate pose conditioning among the three since it uses image-pose pairs for supervised training. However, compared to styleNARF, its FID is worse and the images lack photorealism. styleNARF achieves the best FID among the three, thanks to the effective CNN renderer. However, it cannot explicitly render the foreground only or generate accurate geometry of the generated images. In contrast, our method performs volumetric rendering at the output resolution, and the generated geometry perfectly matches the generated foreground image.

Using Different Pose Distribution Obtaining a ground truth pose distribution of the training images is not feasible for in-the-wild images or new categories. Thus, we train our model with a pose distribution different from the training images. Here, we consider two pose prior distributions. The first uses poses from CMU Panoptic as a prior, which we call the CMU prior. During training, we randomly sample poses from the entire dataset. In addition, to show that our method works without collecting actual human motion capture data, we also create a much simpler pose prior. We fit a multi-variate Gaussian distribution on each joint angle of the CMU Panoptic dataset, and we use randomly sampled poses from the distribution for training. Each Gaussian distribution only defines the rotation angle of each part, which can be easily constructed for novel objects. We call this the random prior.

Quantitative and qualitative results are shown in Table 2 and Figure 5. We can confirm that even when using the CMU prior, our model learns pose-controllable 3D representations with just a slight sacrifice in image quality. When using the random prior, the plausibility of the generated images and the quality of the generated geometry are worse. This may be because the distribution of the random prior is so far from the distribution of poses in the dataset that the learned space of latent vectors too often falls outside the distribution of the actual data. Therefore, we used the truncation trick to restrict the diversity of the latent space, and the results are shown in the bottom row of Table 2. By using the truncation trick, even with a simple prior, we can eliminate latent variables outside the distribution and improve the quality of the generated images and geometry. Further experimental results on truncation are given in the supplement.

Additional Results on Real Images To show the generalization ability of the proposed framework, we train our model on two real image datasets, namely AIST++ and MSCOCO . AIST++ is a dataset of dancing persons with relatively simple backgrounds. We use the ground truth pose distribution for training. MSCOCO is a large scale in-the-wild image dataset. We choose images capturing roughly the whole human body and crop them around the persons. Since 3D pose annotations are not available for MSCOCO, we use poses in CMU Panoptic as the pose prior. Note that we do not use any image-pose or mask supervisions for training. Qualitative results are shown in Figure 6. Experimental results with AIST++, which has a simple background, show that it is possible to generate detailed geometry and images with independent control of viewpoint, pose, and appearance. For MSCOCO, two successful and two unsuccessful results are shown in Figure 6. MSCOCO is a very challenging dataset because of the complex background, the lack of clear separation between foreground objects and background, and the many occlusions. Although our model does not always produce plausible results, it is possible to generate geometry and control each element independently. As an initial attempt, the results are promising.

Conclusion

In this work, we propose a novel unsupervised learning framework for 3D geometry-aware articulated representations. We showed that our framework is able to learn representations with controllable viewpoint and pose. We first propose a computationally efficient neural 3D representation for articulated objects by adapting the tri-plane representation to NARF, then show it can be trained with GAN objectives without using ground truth image-pose pairs or mask supervision. However, the resolution and the quality of the generated images are still limited compared to recent NeRF-based GAN methods; meanwhile, we assume that a prior distribution of the object’s pose is available, which may not be easily obtained for other object categories. Future work includes incorporating the neural rendering techniques proposed in 3D-aware GANs to generate photorealistic high-quality images while preserving the 3D consistency and estimating the pose prior distribution directly from the training data.

This work was supported by D-CORE Grant from Microsoft Research Asia and partially supported by JST AIP Acceleration Research JPMJCR20U3, Moonshot R&D Grant Number JPMJPS2011, CREST Grant Number JPMJCR2015, JSPS KAKENHI Grant Number JP19H01115, and JP20H05556 and Basic Research Grant (Super AI) of Institute for AI and Beyond of the University of Tokyo. We would like to thank Haruo Fujiwara, Lin Gu, Yuki Kawana, and the authors of for helpful discussions.

References

Appendix 0.A ENARF Implementation details

Figure 7 illustrates the learning pipeline of our ENAFR and D-ENAFR models. The decoder network GdecG_{\text{dec}} is implemented with a modulated convolution as in . For the ENARF model, GdecG_{\text{dec}} is conditioned on the time input tt with positional encoding γ(∗)\gamma(*) to handle time-dependent deformations. Following , we use 10 frequencies in the positional encoding. For the D-ENARF model, rotation matrices are additionally used as input to handle pose-dependent deformations. To reduce the input dimension, we input the relative rotation matrices for each part to the root part R1R_{1}. To efficiently learn the deformation field in our D-ENARF model, additional tri-plane deformation Δ\Delta are learned with a StyleGAN2 based generator GΔG_{\Delta},

Each channel represents the relative transformation from the canonical frame in pixels for each tri-planes. The learned deformation field is used to deform the canonical features to handle the non-rigid deformation for each time and pose. The deformed tri-plane feature F′F^{\prime} is formulated as,

where x=[xx,xy,xz]{\bf x}=[{\bf x}_{x},{\bf x}_{y},{\bf x}_{z}]. Intuitively, (xx,xy)({\bf x}_{x},{\bf x}_{y}) is the 2D projection of x{\bf x} on the FxyF_{xy} plane, and Δxy(x)\Delta_{xy}({\bf x}) produces the 2D translation vector for (xx,xy)({\bf x}_{x},{\bf x}_{y}) on the FxyF_{xy} plane. However, this operation requires two look-up tri-planes learned from Δ\Delta and FF for every 3D location x{\bf x}, which is expensive. Therefore, we make an approximation of this by first computing the value of the deformed tri-plane F′F^{\prime} at every pixel grid location using Equation 11, then sampling features from F′F^{\prime} using Equation 7 in the main paper. The increased computational cost is thus proportional to the resolution of F′F^{\prime}, which is much smaller than the number of 3D points x{\bf x}.

We used the coarse-to-fine sampling strategy as in to sample the rendering points on camera rays. Instead of using two separate models to predict points at coarse and fine stages respectively , we use a single model to predict sampled points at both stages. Specifically, for each ray, 48 and 64 points are sampled at the coarse and fine stages, respectively.

A.2 Efficient Implementation

For efficiency, we introduce a weak shape prior for the part occupancy probability pkp^{k} in Section 3.3 of the main paper. Specifically, if a 3D location xkc{\bf x}_{k}^{c} in the canonical space is outside of a cube with one side of 2a2a located at the center of the part pkc{\bf p}_{k}^{c}, we set pkp^{k} to 0, namely, pk←0 if max⁡(∣xkc−pkc∣)>ap^{k}\leftarrow 0\text{ if }\max(|{\bf x}_{k}^{c}-{\bf p}_{k}^{c}|)>a. We set aa to 13\frac{1}{3} meter for all parts. This weak shape prior is used for all NARF-based methods. Since we do not have to compute the intermediate feature fk{\bf f}_{k} (in Equation 7 in the main paper) for the points with part probability pk=0p_{k}=0, the overall computational cost for feature generation is significantly reduced. We implement this efficiently by (1) gathering the valid (pk>0p_{k}>0) canonical positions xkc{\bf x}_{k}^{c} and compute intermediate features fk{\bf f}_{k} for them, (2) multiplying by pkp^{k}, and (3) summing up the feature for each part pk∗fkp_{k}*{\bf f}_{k} with a scatter_add operation.

A.3 Training Details

We use the Adam optimizer with an equalized learning rate of 0.001. The learning rate decay rate is set to 0.99995. The ray batch sizes are set to 4096 for ENARF and 512 for NARF. The ENARF model is trained for 100,000 iterations and the NARF model is trained for 200,000 iterations with a batch size 16. The training takes 15 hours on a single A100 GPU for ENARF and 24 hours for NARF.

Appendix 0.B Ablation Study on Selector (Section 4.1)

To show the effectiveness of the tri-plane based selector, we compared our model with a model without the selector and a model with an MLP based selector. In the model without a selector, we simply set pk=1Kp^{k}=\frac{1}{K}. In the MLP based selector, we used a two-layer MLP with a hidden layer dimension of 10 for each part, as in NARF . The quantitative and qualitative comparisons are provided in Table 1 in the main paper and Figure 8, respectively. The model without a selector cannot generate clear and sharp images because the feature of any location to compute the density and color is evenly contributed by all parts. The MLP based selector helps learn independent parts. However, the generated images look blurry compared to ours. Moreover, it requires much more GPU memory and FLOPS for training and testing. In summary, the proposed tri-plane based selector is superior in terms of both effectiveness and efficiency.

Appendix 0.C Ablation Study on View Dependency

Following Efficient NeRF , our model does not take the view direction as input, i.e. the color is not view dependent. We note that this implementation is inconsistent with the original NeRF model that takes the view direction as an input. Here, we do an ablation study on the view dependent input in our model. Specifically, the positional encoding is added to the view direction γ(d)\gamma({\bf d}) and additionally used as the input of GdecG_{\text{dec}}. Experimental results indicate that additional view direction input leads to darker images and degrades the performance. We thus do not use view direction as input of our models by default. Quantitative and qualitative comparisons are provided in Table 3 and Figure 9, respectively.

Appendix 0.D Comparison with NeuralActor

In this section, we compare ENARF with another state-of-the-art human image synthesis model NeuralActor . Since the training code of NeuralActor is not publically available, we trained ENARF and D-ENARF on two sequences (S1, S2) of the DeepCap dataset following NeuralActor. We then compared them with the corresponding results reported in the NeuralActor paper. NeuralActor uses richer supervision, such as the ground truth UV texture map of SMPL mesh for each frame. The qualitative results on novel pose synthesis are shown in Figure 10. Without deformation modeling, ENARF tends to produce jaggy images and performs the worst due to the enormous non-rigid deformation in training frames. D-ENARF can alleviate the problem by learning the deformation field and can produce plausible results. Compared to NeuralActor, D-ENARF does not generate fine details such as wrinkles in clothing or facial expressions. Still, this gap is acceptable, given that NeuralActor uses GT SMPL meshes and textures for training. We cannot reproduce the quantitative evaluation of NeuralActor (in Table 2) because some implementation details, such as the foreground cropping method, are not publicly available. We thus skip the quantitative comparison with NeuralActor.

Appendix 0.E Implementation Details of ENAFR-GAN

One obvious positional constraint for the foreground object is that it should be generated to cover at least the regions of bones defined by the input pose configuration. This motivates us to propose a bone region loss Lbone\mathcal{L}_{\text{bone}} on the foreground mask MM to facilitate model training. First, we create a skeletal image BB from an input pose configuration. Examples of BB are visualized in Figure 11. The skeletal image BB is an image in which each joint and its parent joint are linked by a straight line of 1-pixel width. The bone region loss Lbone\mathcal{L}_{\text{bone}} is then defined to penalize any overlaps between the background region (1−M1-M) and the bone regions.

We show the comparison results of training with or without Lbone\mathcal{L}_{\text{bone}} in Table 4 and Figure 12. Although the foreground image quality is comparable, Figure 12 shows that the generated images are not well aligned with the input pose without Lbone\mathcal{L}_{\text{bone}}. In Table 4, we can see that PCKh@0.5 metric becomes worse without Lbone\mathcal{L}_{\text{bone}}.

E.2 Training Details of ENARF-GAN

We set the dimension of ztri{\bf z}_{\text{tri}} and zb{\bf z}_{b} to 512, and zENARF{\bf z}_{\text{ENARF}} to 256. We use the Adam optimizer with an equalized learning rate of 0.0004. We set the batch size to 12 and train the model for 300,000 iterations. The training takes 4 days on a single A100 GPU. In the testing phase, a 128×\times128 image is rendered in 80ms (about 12 fps) using a single A100 GPU.

E.3 Shifting Regularization

To prevent the background generator from synthesizing the foreground object, we apply shifting regularization for the background generator. We randomly shift and crop the background images before overlaying foreground images on them. If the background generator synthesizes the foreground object, the shifted images can be easily detected by the discriminator DD, which encourages the background generator to focus only on the background. We generate 128×256128\times 256 background images and randomly crop 128×128128\times 128 images.

Appendix 0.F Pose Consistency Metric

We follow the evaluation metric in the contemporary work GNARF to evaluate the consistency between the input pose and the pose of the generated image. Specifically, we use an off-the-shelf 2D keypoint detector pre-trained on the MPII human pose dataset to detect 2D keypoints in both generated and real images with the same poses and compare the detected keypoints under the metric of PCKh@0.5 . We discard keypoints with low detection confidence (<0.8<0.8) and only compare keypoints that are confident in both generated and real images.

Appendix 0.G Truncation Trick (Section 4.2)

The truncation trick can improve the quality of the images by limiting the diversity of the generated images. Figure 13 shows the results of generating images with the truncation ψ\psi for the tri-plane generator GtriG_{\text{tri}} set to 1.0, 0.7, and 0.4. When truncation ψ\psi is set to 1.0, multiple legs and arms will be generated. Smaller ψ\psi helps generate more plausible appearance and shapes of the object.

Appendix 0.H StyleNARF (Section 4.2)

StyleNARF is a combination of NARF and StyleNeRF . To reduce the computational complexity, it first generates low-resolution features with NARF and then upsamples the features to a higher resolution with a CNN-based generator. Similar to ENARF, StyleNARF can only generate foreground objects, so a StyleGAN2 based background generator GbG_{\text{b}} is used as in ENARF-GAN. First, we sample latent vectors from a normal distribution, z=(zNARF,zb,zup)∼N(0,I){\bf z}=({\bf z}_{\text{NARF}},{\bf z}_{\text{b}},{\bf z}_{\text{up}})\sim\mathcal{N}(0,I), where zNARF{\bf z}_{\text{NARF}} is a latent vector for NARF, zb{\bf z}_{\text{b}} is a latent vector for background, and zup{\bf z}_{\text{up}} is a latent vector for the upsampler. Then the NARF model GNARFG_{\text{NARF}} generates low-resolution foreground feature Ff{\bf F}_{f} and mask Mf{\bf M}_{f}, and GbG_{\text{b}} generates background feature Fb{\bf F}_{b}.

The foreground and background are combined using the generated foreground mask Mf{\bf M}_{f} at low resolution.

The upsampler GupG_{\text{up}} upsamples the feature F{\bf F} based on the latent vector zup{\bf z}_{\text{up}} and generates the final output C{\bf C}.

All layers are implemented with Modulated Convolution .

Appendix 0.I Datasets

We use images at resolution 128×128128\times 128 for GAN training.

We crop the first frame of all videos to 180×180180\times 180 and resize them to 128×128128\times 128 so that the pelvis joint is centered. 68033 images are obtained.

I.2 AIST++

The images are cropped to 600×600600\times 600 so that the pelvis joint is centered and then resized to 128×128128\times 128. We sample 3000 frames for each subject, resulting in 90000 images in total.

I.3 MSCOCO

First, we select the persons whose entire body is almost visible according to the 2D keypoint annotations in MSCOCO. Each selected person is cropped by a square rectangle that tightly encloses the person and it is resized to 128×128128\times 128. The number of collected samples is 38727.

Appendix 0.J Pose Distribution

Two pose distributions are used in our experiments. One is the CMU pose distribution, which consists of 390k poses collected in the CMU Panoptic dataset. We follow the pose pre-processing steps in the SURREAL dataset, where the distance between the person and the camera is randomly distributed in a normal distribution of mean 7 meters and variance 1 meter. The pose rotation around the z-axis is uniformly distributed between 0 and 2π2\pi.

Another pose distribution used in our experiments is a random pose distribution. First, a multivariate normal distribution is fitted to the angle distribution of each joint under the 390k pose samples of the CMU Panoptic dataset. Then, poses are randomly sampled from the learned multivariate normal distribution and randomly rotated so that the pelvis joint is directly above the medial points of the right and left plantar feet.