KeypointNeRF: Generalizing Image-based Volumetric Avatars using Relative Spatial Encoding of Keypoints

Marko Mihajlovic, Aayush Bansal, Michael Zollhoefer, Siyu Tang, Shunsuke Saito

Introduction

3D renderable human representations are an important component for social telepresence, mixed reality, and a new generation of entertainment platforms. Classical mesh-based methods require dense multi-view stereo or depth sensor fusion . The fidelity of these methods is limited due to the difficulty of accurate geometry reconstruction. Recently, neural volumetric representations have enabled high-fidelity human reconstruction, especially where accurate geometry is difficult to obtain (e.g. hair). By injecting human-specific parametric shape models , extensive multi-view data capture can be reduced to sparse camera setups . However, these learning-based approaches are subject-specific and require days of training for each individual subject, which severely limits their scalability. Democratizing digital volumetric humans requires an ability to instantly create a personalized reconstruction of a user from two or three snaps (from different views) taken from their phone. Towards this goal, we learn to generalize metrically accurate image-based volumetric humans from two or three views.

Fully convolutional pixel-aligned features utilizing multi-scale information have enabled better performance for various 2D computer vision tasks , including the generalizable reconstruction of unseen subjects Pixel-aligned neural fields infer field quantities given a pixel location and spatial encoding function (to avoid ray-depth ambiguity). Different spatial encoding functions have been proposed in the literature. However, the effect of spatial encoding is not fully understood. In this paper, we provide an extensive analysis of spatial encodings for modeling pixel-aligned neural radiance fields for human faces. Our experiments show that the choice of spatial encoding influences the reconstruction quality and generalization to novel identities and views. The models that use camera depth overfit to the training distribution, and multi-view stereo constraints are less robust to sparse views with large baselines.

We present a simple yet highly effective approach based on sparse 3D keypoints to address the limitations of existing approaches. 3D keypoints are easy to obtain using an off-the-shelf 2D keypoint detector and a simple triangulation of multi-views . We treat 3D keypoints as spatial anchors and encode relative 3D spatial information to those keypoints. Unlike global spatial encoding , the relative spatial information is agnostic to camera parameters. This property allows the proposed approach to be robust to changes in pose. 3D keypoints also allow us to use the same approach for both human faces and bodies. Our approach achieves state-of-the-art performance for generating volumetric humans for unseen subjects from sparse-and-wide two or three views, and we can also incorporate more views to further improve performance. We also achieve performance comparable to Neural Human Performer (NHP) when it comes to full-body human reconstruction. NHP relies on accurate parametric body fitting and temporal feature aggregation, whereas our approach employs 3D keypoints alone. Our method is not biased to the distribution of the training data. We can use the learned model (without modification) to never-before-seen iPhone captures. We attribute our ability to generalize image-based volumetric humans to an unseen data distribution to our choice of spatial encoding. Our key contributions include:

A simple approach that leverages sparse 3D keypoints and allows us to create high-fidelity state-of-the-art volumetric humans for unseen subjects from widely spread out two or three views.

Extensive analysis on the use of spatial encodings to understand their limitations and the efficacy of the proposed approach.

We demonstrate generalization to never-before-seen iPhone captures by training with only a studio-captured dataset. To our knowledge, no prior work has shown these results.

Related Work

Our goal is to create high-fidelity volumetric humans for unseen subjects from as few as two views.

Classical Template-based Techniques: Early work on human reconstruction required dense 3D reconstruction from a large number of images of the subject and non-rigid registration to align a template mesh to 3D point clouds. Cao et al. employ coarse geometry along with face blendshapes and a morphable hair model to address restrictions posed by dense 3D reconstruction. Hu et al. retrieve hair templates from a database and carefully compose facial and hair details. Video Avatar obtains a full-body avatar based on a monocular video captured using silhouette-based modeling. The dependence on geometry and meshes restricts the applicability of these methods to faithfully reconstruct regions such as the hair, mouth, teeth, etc., where it is non-trivial to obtain accurate geometry.

Neural Rendering: Neural rendering has tackled some of the challenges classical template-based approaches struggle with by directly learning components of the image formation process from raw sensor measurements. 2D neural rendering approaches employ surface rendering and a convolutional network to bridge the gap between rendered and real images. The downside of these 2D techniques is that they struggle to synthesize novel viewpoints in a temporally coherent manner. Deep Appearance Models employ a coarse 3D proxy mesh in combination with view-dependent texture mapping to learn personalized face avatars from dense multi-view supervision. Using a 3D proxy mesh significantly helps with viewpoint generalization, but the approach faces challenges in synthesizing certain regions for which it is hard to obtain good 3D reconstruction, such as the hair and inside the mouth. Current state-of-the-art methods such as NeuralVolumes and NeRF employ differentiable volumetric rendering instead of relying on meshes. Due to their volumetric nature, these methods enable high-quality results even for regions where estimating 3D geometry is challenging. Various extensions have further improved quality. These methods require dense multi-view supervision for person-specific training and take 3–4 days to train a single model.

Sparse View Reconstruction: Large scale deployment requires approaches that allow a user to take two or three pictures of themselves from multi-views and generate a digital human from this data. The use of pixel-aligned features further allows the use of large datasets for learning generalized models from sparse views. Different approaches have combined multi-view constraints and pixel-aligned features alongside NeRF to learn generalizable view-synthesis. In this work, we observe that these approaches struggle to generate fine details given sparse views for unseen human faces and bodies.

Learning Face and Body Reconstruction: Generalizable parametric mesh and implicit body models can provide additional constraints for learning details from sparse views. Recent approaches have incorporated priors specific to human faces and human bodies to reduce the dependence on multi-view captures. Approaches such as H3DNet and SIDER use signed-distance functions (SDFs) for learning geometry priors from large datasets and perform test-time fine-tuning on the test subject. PaMIR uses volumetric features guided by a human body model for better generalization. Neural Human Performer employs SMPL with pixel-aligned features and temporal feature aggregation. In this work, we observe that the use of human 3D-keypoints provides necessary and sufficient constraints for learning from sparse-view inputs. Our approach has high flexibility since it only relies on 3D keypoints alone and thus enables us to work both on human faces and bodies. Prior methods have also employed various forms of spatial encoding for better learning. For example, PVA and PortraitNeRF use face-centric coordinates. ARCH/ARCH++ use canonical body coordinates. In this work, we extensively study the role of spatial encoding, and found that the use of a relative depth encoding using 3D keypoints leads to the best results. Our findings enable us to learn a representation that generalizes to never-before-seen iPhone camera captures for unseen human faces. In addition to achieving state-of-the-art results on volumetric face reconstruction from as few as two images, our approach can also be used for synthesizing novel views of unseen human bodies and achieves competitive performance to prior work in this setting.

Preliminaries: Neural Radiance Fields

where tnt_{n} and tft_{f} define near and far bounds.

Pixel-aligned NeRF. One of the core limitations of NeRF is that the approach requires per-scene optimization and does not work well for extremely sparse input views (e.g., two images). To address these challenges, several recent methods propose to condition NeRF on pixel-aligned image features and generalize to novel scenes without retraining.

Spatial Encoding. To avoid the ray-depth ambiguity, pixel-aligned neural fields attach spatial encoding to the pixel-aligned feature. PIFu and related methods use depth value in the camera coordinate space as spatial encoding, while PVA uses coordinates relative to the head position. However, we argue that such spatial encodings are global and sub-optimal for learning generalizable volumetric humans. In contrast, our proposed relative spatial encoding provides a localized context that enables better learning and is more robust to changes in human pose.

KeypointNeRF

Our method is based on a radiance field function:

where α\alpha is a fixed hyper-parameter that controls the impact of each keypoint. We set this value to 5cm for facial keypoints and to 10cm for the human skeleton.

2 Convolutional Pixel-aligned Features

3 Multi-view Feature Fusion

To model a multi-view consistent radiance field, we need to fuse the per-view spatial encodings sns_{n} (Eq. 3) and the pixel-aligned features Φn\Phi_{n}.

4 Modeling Radiance Fields

The radiance field is modeled via decoupled MLPs for density σ\sigma and color cc prediction.

Density Fields. The density network is implemented as a four-layer MLP that takes as input the geometry feature vector GXG_{X} and predicts the density value σ\sigma.

View-dependent Color Fields. We implement an additional MLP to output the consistent color value cc for a given query point XX and its viewing direction dd by blending image pixel values {Φ(X∣Pn)}n=1N\{\Phi(X|P_{n})\}_{n=1}^{N} similarly to IBRNet . The input to this MLP is 1) the extracted geometry feature vector GXG_{X} that ensures geometrically consistent renderings, 2) the additional pixel-aligned features Φn(X∣Fna)\Phi_{n}(X|F^{a}_{n}), 3) the corresponding pixel values Φn(X∣In)\Phi_{n}(X|I_{n}), and 4) the view direction that is encoded as the difference between the view direction dd and the camera views along with their dot product.

These inputs are concatenated and augmented with the mean and variance vectors computed over the multi-view pixel-aligned features, and jointly propagated through a nine-layer perceptron with residual connections which predicts blending weights for each input view {wn}n=1N\{w_{n}\}_{n=1}^{N}. These blending weights form the final color prediction by fusing the corresponding pixel-aligned color values:

Novel View Synthesis

Given our radiance field function f(X,d)=(c,σ)f(X,d)=(c,\sigma), we render novel views via the volume rendering equation (1), in which we define the near and far bound by analytically computing the intersection of the pixel ray and a geometric proxy that over-approximates the volumetric human and use the entrance and exit points as near and far bounds respectively. For the experiments on human heads, we use a sphere with a radius of 30 centimeters centered around the keypoints, while for the human bodies we follow the prior work and use a 3D bounding box. Similar to NeRF , we employ a coarse-to-fine rendering strategy, but we employ the same network weights for both levels.

The use of the VGG loss for NeRF training was also leveraged by the concurrent methods to better capture high-frequency details. The final loss L\mathcal{L} is minimized by the Adam optimizer with a learning rate of 1e−41e^{-4} and a batch size of one. For the other parameters, we use their defaults. The background from all training and test input images is removed via an off-the-shelf matting network . Additionally for more temporally coherent novel-view synthesis at inference time, we clip the maximum of the dot product (introduced in Sec. 4.4) to 0.8 when the number of input images is two in the supplementary video.

Experiments

In this section, we validate our method on three different reconstruction tasks and datasets: 1) reconstruction of human heads from images captured in a multi-camera studio, 2) reconstruction of human heads from in-the-wild images taken with the iPhone’s camera, and 3) reconstruction of human bodies on the public ZJU-MoCap dataset . As evaluation metrics, we follow prior work and report the standard SSIM and PSNR metrics.

Dataset and Experimental Setup. Our captured data consists of 29 1280×7681280\times 768-resolution cameras positioned in front of subjects. We use a total of 351 identities and 26 cameras for training and 38 novel identities for evaluation. At inference time, we reconstruct humans only from 2–3 input views.

Baselines. As baseline, we employ the current state-of-the-art model IBRNet . In addition, we add several other baselines by varying different types of encoding for the query points in our proposed reconstruction pipeline. Specifically, 1) our pipeline without any encoding, 2) with the camera zz encoding used in , 3) with the encoding of xyzxyz coordinates in the canonical space of a human head that is used in , 4) relative encoding of xyzxyz as the distance between the query point and estimated keypoints, 5) our relative spatial encoding without distance weighing (α→∞\alpha\to\infty in Eq. 3), and 6) the proposed weighted relative encoding as described in the method section (Sec. 4.3). The last three models use a total of 13 facial keypoints that are visualized in Figure 1. All methods are trained with a batch size of one for 150k training steps, except IBRNet which was trained for 200k iterations. For more comparisons and baselines we refer the reader to the supplementary material and video.

Results. We provide novel view synthesis (Fig. 3) results for unseen identities that have been reconstructed from only two input images. The results clearly demonstrate that the rendered images of our method are significantly sharper compared to the baselines and are of significantly higher quality. This improvement is confirmed by the quantitative evaluation (Tab. 1) which further indicates that the proposed distance weighting of the relative spatial encoding improves the reconstruction quality. The third-best performing method is our pipeline without any spatial encoding. However, such a method does not generalize well to other capture systems as we will demonstrate in the next section on in-the-wild captured data.

Robustness to Different Noise Levels. We evaluate the robustness of our relative spatial encoding and the encoding in canonical space proposed by Raj et al. by adding different noise levels to the estimated keypoints and the head center respectively. The results reported in Tab. 2 show that our proposed encoding based on keypoints is significantly more robust compared to modeling in an object-specific canonical space. Note that the canonical encoding requires head template fitting for which we used the ground truth estimation from all views and which, in practice, is very erroneous or even infeasible from two views alone.

Dynamic scenes. Although our model is trained only with a neutral face, it generalizes well to dynamic expressions and outperforms the baseline methods. We evaluate the trained models on 3838 test subjects performing eight different expressions and report results in Tab. 3.

2 Reconstruction from In-the-wild Captures

Setup. To tackle the problem of reconstructing humans in the wild, we acquire a small dataset of subjects by taking several photos with an iPhone and estimate camera parameters. We directly use intrinsic from manufacturing information of iPhone and extrinsic is computed by multi-view RGB-D fitting as in . We evaluate the reconstruction methods trained on the studio captured data in Sec. 6.1 without any retraining.

As input, all methods take three 1920×10241920\times 1024-resolution images of a person and predict a radiance field that is then rendered from novel views. In Figure 4, we display rendered novel views of IBRNet , our method without any spatial encoding, and our method with the proposed spatial encoding. The baseline methods produce significantly worse results with lots of blur and cloudy artifacts, whereas

our method can reliably reconstruct the human heads. This improvement is quantitatively supported by computed SSIM and PSNR on novel held-out views of the visualized four subjects (Tab. 4). This experiment demonstrates that our relative spatial encoding is the crucial component for cross-dataset generalization. Please see the supplementary material for more visualizations.

3 Reconstruction of Human Bodies

Additionally, we demonstrate that our method is suitable for reconstructing full volumetric human bodies without relying on template fitting of parametric human bodies . We use the public ZJU dataset in order to follow the experimental setup used in , so that we could closely compare our method’s ability to reconstruct human bodies to the current state-of-the-art method without changing any experimental variables. We follow the standard training-test split of frames and use a total of seven subjects for training and three for validation. At inference time, all methods use three input views. We compare our method with the generalizable volumetric methods: pixelNeRF , PVA , the current state-of-the-art Neural Human Performer (NHP) , and our method without weighting the relative spatial encoding in Eq. 3. We report results on unseen identities for 438 novel views in Table 5 and side-by-side qualitative comparisons with NHP in Figure 5. The results demonstrate that weighting the spatial encoding benefits reconstruction of human bodies as well. Our method is on par with the significantly more complex NHP, which relies on the accurate registration of the SMPL body model and temporal feature fusion, whereas ours only requires skeleton keypoints.

Conclusion

We present a simple yet highly effective approach for generating high-fidelity volumetric humans from as few as two input images. The key to our approach is a novel spatial encoding based on relative information extracted from 3D keypoints. Our approach outperforms state-of-the-art methods for head reconstruction and better generalizes to challenging out-of-domain inputs, such as selfies captured in the wild by an iPhone. Since our approach does not require a parametric template mesh, it can be applied to the task of body reconstruction without modification, where it achieves performance comparable to more complicated prior work that has to rely on parametric human body models and temporal feature aggregation. We believe that our local spatial encoding based on keypoints might also be useful for many other category-specific neural rendering applications.

Acknowledgments. We thank Chen Cao for the help with the in-the-wild iPhone capture. M. M. and S. T. acknowledge the SNF grant 200021 204840.

References

Appendix 0.A Overview

In this document we provide additional implementation details (Sec. 0.B), information about the baseline methods (Sec. 0.C), more qualitative and quantitative results (Sec. 0.D), and reflect on the limitations of KeypointNeRF and future work (Sec. 0.E).

Appendix 0.B Implementation Details

Density Fields. The geometry feature vector is decoded as density value σ\sigma via a four-layer MLP (64 neurons with Softplus activations).

View-dependent Color Fields. To produce the final color prediction cc for a query point XX, we implemented an additional MLP that predicts blending weights as an intermediate step which are used to blend the input pixel colors. This network follows the design proposed in IBRNet to communicate information among multi-view features by using the mean-variance pooling operator. The per-view input feature vectors (described in Sec. 4.4) are first fused into a global feature vector via the mean-variance pooling operator. Then this feature is attached to the pixel-aligned feature vectors Φ(X∣Fna)\Phi(X|F^{a}_{n}) and propagated through a nine-layer MLP with residual connections and an exponential linear unit as activation to predict the blending weights (Eq. 4).

Appendix 0.C Baseline Methods

We used the publicly released code of MVSNeRF and IBRNet with their default parameters. We re-implemented PVA since their code is not public and we directly used the public results of NHP for the experiments on the ZJU-MoCap dataset .

Appendix 0.D Additional Results

Multi-view studio Capture Results. We further provide qualitative results for two more baseline methods (MVSNeRF and PVA ) for the experimental setup described in Sec. 6.1. The results in Fig. 0.B.1 demonstrate that the best performing baseline (IBRNet) produces incomplete images with lots of blur and foggy artifacts. PVA yields consistent, but overly smoothed renderings, while MVSNeRF does not work well for the widely spread-out input views. For more qualitative results, we refer the reader to the supplementary video.

Keypoint perturbation. To evaluate the sensitivity of our method on a less accurate estimation of keypoints, we perturb them with different Gaussian noise levels (ranging from 1 to 20mm) for unseen subjects from Sec. 6.1 and observe that the rendered images (Fig. 0.D.2) occasionally tend to become blurry around the keypoints (e.g. eyes) for large noise levels (>10>10mm).

The impact of the iPhone calibration for the in-the-wild capture. We evaluate the robustness of KeypointNeRF to a nosier camera calibration by estimating the iPhone camera parameters without the depth term for the experimental setup presented in Sec. 6.2. We observe (Tab. 0.D.1) a negligible drop (PSNR/SSIM by -0.04/-0.5) in performance for our method, demonstrating the robustness of our method under noisy camera calibration.

Convolutional feature encoders. We further measure the impact of the HourGlass feature extractor and compare it with the U-Net encoder that is used by the other baseline methods . We follow the experimental setup from subsections 6.1 and 6.2 and report quantitative results in Tab. 0.D.2 and 0.D.3 respectively. We observe that HourGlass encoder consistently improves the reconstruction quality.

Appendix 0.E Limitations and Future Work

While our method offers an efficient way of reconstructing volumetric avatars from as few as two input images, it still has several difficulties. The image-based rendering formulation of our method parametrizes the color prediction as blending of available pixels, which ensures good color generalization at inference time, however it makes the method sensitive to occlusions. The method itself has also difficulties reconstructing challenging thin geometries (e.g. glasses) and is less robust to highly articulated human motions (see Fig. 0.E.3). As future work we consider addressing these challenges and additionally integrating learnable 3D lifting methods with the proposed relative spatial encoding for more optimal end-to-end network training.