NeuMan: Neural Human Radiance Field from a Single Video

Wei Jiang, Kwang Moo Yi, Golnoosh Samei, Oncel Tuzel, Anurag Ranjan

Introduction

The quality of novel view synthesis has been dramatically improved since the introduction of Neural Radiance Fields (NeRF) . While originally proposed to reconstruct a static scene with a set of posed images, it has since been quickly extended to dynamic scenes and uncalibrated scenes . Recent efforts also focus on animation of these radiance field models of human, with the aid of large controlled datasets, further extending the application domain of radiance-field-based modeling to enable augmented reality experiences.

In this work, we are interested in the scenario where only one single video is provided, and our goal is to reconstruct the human model and the static scene model, and enable novel pose rendering of the human, without any expensive multi-cameras setups or manual annotations. However, even with the recent advancements in NeRF methods, this is far from being trivial. Existing methods require multi-cameras setup, consistent lighting and exposure, clean backgrounds, and accurate human geometry to train the NeRF models. As shown in Table 1, HyperNeRF models a dynamic scene based on a single video, but cannot be driven by human poses. ST-NeRF reconstructs each individual with a time dependent NeRF model from multiple cameras, but the editing is limited to the transformation of the bounding box. Neural Actor can generate novel poses of a human but requires multiple videos. HumanNeRF builds a human model based on a single video with manually annotated masks, but doesn’t show generalization to novel poses. Vid2Actor generates novel poses of a human with a model trained on a single video but cannot model the background. We address these problems by introducing NeuMan, that reconstructs both the human and the scene with the ability to render novel human poses and novel views, from a single in-the-wild video.

NeuMan is a novel framework for training NeRF models for both the human and the scene, which allows high-quality pose-driven rendering as shown in Figure 1. Given a video captured by a moving camera, we first estimate the human pose, human shape, human masks, as well as the camera poses, sparse scene model, and depth maps using conventional off-the-shelf methods .

We then train two NeRF models, one for the human and one for the scene guided by the segmentation masks estimated from Mask-RCNN . Additionally, we regularize the scene NeRF model by fusing together depth estimates from both multi-view reconstruction and monocular depth regression . We train the human NeRF model in a pose independent canonical volume guided by a statistical human shape and pose model, SMPL following Liu et al. . We refine the SMPL estimates from ROMP to better serve the training. However, these refined estimates are still not perfect. Therefore, we jointly optimize the SMPL estimates together with the human NeRF model in an end-to-end fashion. Furthermore, since our static canonical human NeRF cannot represent the dynamics that is not captured by the SMPL model, we introduce an error-correction network to counter it. The SMPL estimates and the error-correction network are jointly optimized during the training.

we propose a framework for neural rendering of a human and a scene from a single video without any extra devices or annotations;

we show that our method allows high quality rendering of human under novel poses, from novel views, together with the scene;

we introduce an end-to-end SMPL optimization and an error-correction network to enable training with erroneous estimates of the human geometry;

our approach allows for the composition of the human and the scene NeRF models enabling applications such as telegathering.

Related Work

As our work is mainly based on neural radiance fields, we first review works on NeRF with a focus on works that aim to control and condition the radiance fields—a necessity for rendering a human in the scene in the context of creating visual and immersive experiences . We also briefly review works that aim to reanimate and perform novel view synthesis of provided scenes.

Since its first introduction , NeRF has become a popular way to model scenes and render them from novel views thanks to its high quality rendering. Representing a scene as a radiance field has the advantage that by-construction you will be able to render the scene from any supported views through volume rendering. Efforts have been made to adapt NeRF to dynamic scenes , and to even edit and compose scenes with various NeRF models , widening their potential application. While these methods have shown interesting and exciting results, they often require separate training of editable instances or careful curation of training data . In this work, we are interested in an in-the-wild setup.

Particularly related to our task of interest, various efforts have been made towards NeRF models conditioned by explicit human models, such as SMPL or 3D skeleton . Neural Body associates a latent code to each SMPL vertex, and use sparse convolution to diffuse the latent code into the volume in observation space. Neural Actor learns the human in the canonical space by a volume warping based on the SMPL mesh transformation, it also utilize a texture map to improve the final rendering quality. Animatable NeRF learns a blending weight field in both observation space and canonical space, and optimize for a new blending weight field for novel poses. ST-NeRF separates the human into each 3D bounding box, and learns the dynamic human within each bounding box. It doesn’t require to estimate the precise human geometry, but it cannot extrapolate to unseen poses since it is dependent on time(frame). However, all these methods require an expensive multi-cameras setup to obtain the ground truth bounding boxes, 3D poses or SMPL estimates. In other words, they cannot be used for our purpose of reconstructing and neural rendering of human from a single video, without extra devices or annotations, and with potential pose estimation errors.

HumanNeRF , a concurrent work, aims to create free-viewpoint rendering of human from a single video. While similar to our work, there are two main differences between HumanNeRF and ours. First, HumanNeRF relies on manual mask annotation to separate the human from the background, while our method learns the decomposition of the human and the scene with the help of modern detectors. Second, HumanNeRF represents motion as a combination of the skeletal and the non-rigid transformations, causing ambiguous or unknown transformations under novel poses, while ours mitigates the ambiguity by using explicit human mesh. Another similar work is Vid2Actor , which builds animatable human from a single video by learning a voxelized canonical volume and skinning weights jointly. Although with similar goals, our method is able to reconstruct sharp human geometry with less than 40 images, comparing to thousands frames are required for Vid2Actor, our method is data-efficient.

Neural Rendering of Humans.

Majority of the literature only consider the problem of reposing a human from a source image to a target image without changing the viewing angle which is essential for enabling new immersive experiences. Grigorev et al tackled a similar problem to ours, namely resynthesizing a human image with a novel pose view given a single input image. They divide the problem into estimating the full texture map of the human body surface from partial texture observations and synthesizing a novel view given the estimated texture map from the first step. They employ two convolutional neural networks (CNN) for each of these steps. With the second one responsible for generating the novel pose consuming the output of the first CNN. This method does not explicitly model the source and target pose, so it is unclear how well it would perform for a novel pose. Sarkar et al additionally consider re-rendering in a novel view as well as novel pose from a single input image. They use a parametric 3D human mesh to recover body pose and shape and a high dimensional UV feature map to encode appearance.

1 Preliminary: Neural Radiance Fields (NeRF)

For completeness, we quickly review the standard NeRF model . We denote the NeRF network as FΘ{\mathcal{F}}_{\boldsymbol{\Theta}} with parameters Θ{\boldsymbol{\Theta}}, that estimates the RGB color c{\mathbf{c}} and density σ\sigma of a given a 3D location x{\mathbf{x}} and a viewing direction d{\mathbf{d}} as

The radiance field function FΘ{\mathcal{F}}_{\boldsymbol{\Theta}} is often implemented with multi-layer perceptrons (MLPs) with positional encodings and periodic activations . The pixel color is then obtained by integrating a discretized ray r{\mathbf{r}} that consists of NN samples from the camera in the view direction d{\mathbf{d}} through the volume given by

and δi\delta_{i} is the distance between two adjacent samples. Here, the accumulated alpha value of a pixel, which represents transparency, can be obtained by α(r)=∑i=1Nwi{\mathbf{\alpha}}\left({\mathbf{r}}\right)=\sum_{i=1}^{N}w_{i}.

Method

An overview of our framework is shown in Figure 2. Our framework is composed mainly of two NeRF networksIn our work, we assume a single human being in the scene, but this can be trivially extended.: the human NeRF that encodes the appearance and geometry of the human in the scene, conditioned on the human pose; and the scene NeRF that encodes how the background looks like. We train the scene NeRF first, then train the human NeRF conditioned on the trained scene NeRF.

The scene NeRF model is analogous to the background model in traditional motion detection work , except it’s a NeRF. For the scene NeRF model, we construct a NeRF model and train it with only the pixels that are deemed to be from the background.

For a ray r{\mathbf{r}}, given the human segmentation mask as M(r){\mathcal{M}}({\mathbf{r}}) where M(r)=1{\mathcal{M}}({\mathbf{r}})=1 if the ray corresponds to the human and M(r)=0{\mathcal{M}}({\mathbf{r}})=0 corresponding the background, we formulate the reconstruction loss for the scene NeRF model as

where C^(r)\hat{{\mathbf{C}}}({\mathbf{r}}) corresponds to the ground-truth RGB color value and Cs(r){\mathbf{C}}_{s}({\mathbf{r}}) corresponds to rendered color value from the scene NeRF model.

As in Video-NeRF , simply minimizing Eq. 3 leads to ‘hazy’ objects floating in the scene. Therefore, following Video-NeRF , we resolve this by adding a regularizer on the estimated density, and forcing it to be zero for space that should be empty—the space between the camera and the scene. For each ray r{\mathbf{r}}, we sample the terminating depth value z^r=Dfuse(r)\hat{z}_{{\mathbf{r}}}={\mathbf{D}}_{fuse}({\mathbf{r}}) and minimize

where α=0.8\alpha=0.8 is a slack margin to avoid strong regularization when the depth estimates are inaccurate. The final loss that we use to train our scene NeRF is

where λempty=0.1\lambda_{empty}=0.1 is a hyper parameter controlling the emptiness regularizer in all our experiments.

Given a video sequence, we use COLMAP to obtain the camera poses, sparse scene model, and multi-view-stereo (MVS) depth maps. Typically, MVS depth maps Dmvs{\mathbf{D}}_{mvs} contain holes, which we fill with the help of dense monocular depth maps Dmono{\mathbf{D}}_{mono} using Miangoleh et al. . We fuse Dmvs{\mathbf{D}}_{mvs} and Dmono{\mathbf{D}}_{mono} together to obtain a fused depth map Dfuse{\mathbf{D}}_{fuse} with consistent scale. In more detail, we find a linear mapping between the two depth maps using the pixels that have both estimates. We then transform the values of Dmono{\mathbf{D}}_{mono} with this mapping to match the depth scale in Dmvs{\mathbf{D}}_{mvs} to obtain a fused depth map Dfuse{\mathbf{D}}_{fuse} by filling in the holes. For retrieving human segmentation maps we apply Mask-RCNN . We further dilate the human masks by 4% to ensure the human is completely masked out. With the estimated camera poses and the background masks, we train the scene NeRF model only over the background.

2 The Human NeRF Model

To build a human model that can be pose-driven, we require the model to be pose independent. Therefore, we define a canonical space based on the 大-pose (Da-pose) SMPL mesh, similar to . In comparison to the traditional T-pose, Da-pose avoids volume collision when warping from observation space to canonical space for the legs.

To render a pixel of a human in the observation space with this model, we transform the points along that ray into the canonical space. The difficulty in doing so is how one expands the transformation of SMPL meshes into the entire observation space to allow this ray tracing in canonical space. Similar to Liu et al. , we use a simple strategy to extend the mesh skinning into a volume warping field.

The error-correction net is only used during training, and is discarded for rendering with validation and novel poses. Since a single canonical space is used to explain all poses, the error-correction network naturally overfits to each frame and makes the canonical volume more generalized.

Due to the nature of the warping field, a straight line in the observation space is curved in the canonical space after warping. Therefore, we recompute the viewing angles by taking into account how the light rays actually travel in the canonical space by looking at where the previous sample is,

where Φ{\boldsymbol{\Phi}} are the parameters of the human NeRF FΦ{\mathcal{F}}_{\boldsymbol{\Phi}}.

To render a pixel, we shoot two rays, one for the human NeRF, and the other for the scene NeRF. We evaluate the colors and densities for the two sets of samples along the rays. We then sort the colors and the densities in the ascending order based on their depth values, similar to ST-NeRF . Finally, we integrate over these values to obtain the pixel using Eq. (2).

To train the human radiance field, we sample rays on the regions covered by the human mask and minimize

where Ch(r){\mathbf{C}}_{h}({\mathbf{r}}) is the rendered color from the human NeRF model. Similar to HumanNeRF , we also use LPIPS as an additional loss term Llpips{\mathcal{L}_{lpips}} by sampling a 32×3232\times 32 patch. We use Lmask{\mathcal{L}_{mask}} to enforce the accumulated alpha map from the human NeRF to be similar to the detected human mask.

where αh{\mathbf{\alpha}}_{h} corresponds to accumulated density over the ray as defined in Sec. 2.1.

To avoid blobs in the canonical space and semi-transparent canonical human, we enforce the volume inside the canonical SMPL mesh to be solid, while enforcing the volume outside the canonical SMPL mesh to be empty, given by

Moreover, we utilize hard surface loss Lhard{\mathcal{L}_{hard}} to mitigate the halo around the canonical human. To be specific, we encourage the weight of each sample to be either 1 or 0 given by,

where ww refers to the transparency where the ray terminates as defined in Sec. 2.1. However, this penalty alone is not enough to obtain a sharp canonical shape, we also add a canonical edge loss, Ledge{\mathcal{L}_{edge}}. By rendering a random straight ray in the canonical volume, we encourage the accumulated alpha values to be either 1 or 0. This is given by,

where αc{\mathbf{\alpha}}_{c} is the accumulated alpha value obtained from a random straight ray in canonical space. Thus, the final loss is given by,

To train, we jointly optimize θf{\boldsymbol{\theta}}_{f}, E{\mathcal{E}}, and FΦ{\mathcal{F}}_{\boldsymbol{\Phi}} by minimizing this loss. We set λlpips=0.01\lambda_{lpips}=0.01, λmask=0.01\lambda_{mask}=0.01, λsmpl=1.0\lambda_{smpl}=1.0, λhard=0.1\lambda_{hard}=0.1, and λedge=0.1\lambda_{edge}=0.1. Since the detected masks are inaccurate, we linearly decay λmask\lambda_{mask} to 0 through the training.

Preprocessing.

We utilize ROMP to estimate the SMPL parameters of the human in the videos. However, the estimated SMPL parameters are not accurate. Therefore, we refine the SMPL estimates by optimizing the SMPL parameters using silhouette estimated from , and 2D joints estimated from as detailed in the supplementary material. We then align the SMPL estimates in the scene coordinates.

Scene-SMPL Alignment.

To compose a scene with a human in novel view and pose, and to train the two NeRF models, we align the coordinate systems in which the two NeRF models lie. This is, in fact, a non-trivial problem, as human body pose estimators operate in their own camera systems with often near-orthographic camera models. To deal with this issue, we first solve the Perspective-n-Point (PnP) problem between the estimated 3D joints and the projected 2D joints with the camera intrinsics from COLMAP. This solves the alignment up to an arbitrary scale. We then assume that the human is standing on a ground at least in one frame, and solve for the scale ambiguity by finding the scale that allows the feet meshes of the SMPL model to touch the ground plan. We obtain the ground plane by applying RANSAC. We show the results of the aligned SMPL estimates in the scene in Figure 4.

Once the two NeRF models are properly aligned we can render the pixel by shooting two rays, one for the human NeRF model, and the other for the scene NeRF model, as describe above. For the near and far planes to generate samples in Eq. 2, we use the estimated scene point cloud to determine them for scene NeRF, and use the estimated SMPL mesh to determine them for the human NeRF, following the strategy in Liu et al. .

Experiments

We introduce our dataset, show qualitative and quantitative results of our method and provide ablation studies. Our method is the first method that can render a human together with a scene with novel human poses and novel views from a single video. We show the importance of our geometry correction and novel loss terms in order to obtain a realistic and sharp human NeRF model.

Existing human motion datasets are not suitable for our experiments. Generally, motion capture data is captured with a static multi-cameras system in a controlled environment which defeats the purpose of reconstructing from a single video. Other video datasets have multiple humans and crowded scenes that are not suitable for our case. Therefore, we introduce NeuMan dataset, a collection of 6 videos about 10 to 20 seconds long each, where a single person performs a walking sequence captured using a mobile phone. Moreover, the camera reasonably pans through the scene to enable multi-view reconstruction. The sequences are named – Seattle, Citron, Parking, Bike, Jogging and Lab. For each video sequence, we first subsample the frames from the video and use them to train our NeRF models. We split frames into 80% training frames, 10% validation frames, and 10% test frames. We provide the details in supplementary material.

2 Qualitative Results

Figure 5 shows the novel view renderings of our scene NeRF models. By reconstructing the background pixels only, our model learns the consistent geometry of the scene, and effectively removes the dynamic human.

Human NeRF Reconstructions.

Our framework learns an animatable human model with realistic details. It captures not only the texture details such as the pattern on the cloth, but also the subject specific geometric details, such as the sleeves, collar, even zipper. Notice that these geometric details are beyond the expressiveness of SMPL model. The learned human models can be reposed to novel driving motions, and produce high quality rendering of the human under novel poses from novel views. Although, the training sequence is as simple as walking, the model can perform stunning cartwheeling motion, which shows the ability to extrapolate to unseen poses. By simple composition, we can render both the human and the scene realistically; see Figure 6.

Telegathering.

The ability to render reposed human together with the scene further allows us to, for example, create telegathering of multiple individuals as in Figure 7. The results show that our framework can facilitate combining human NeRF models in the same scene without any additional training.

3 Novel view synthesis

To compare our method to existing solutions that can do similar tasks, we train NeRF with time(NeRF-T) using the same training data as ours but without the empty space penalty. As HyperNeRF requires a smoothly changing video, we provide it with densely sampled frames from the videos. NeRF-T and HyperNeRF do not perform well at reconstruction of drastically dynamic scenes with humans, while our method is able to do so even with less than 40 training images. We evaluate methods on the test views, and measure PSNR, SSIM , and LPIPS . The results are shown in Table 2 and Figure 8. Our method outperforms across all scenes and metrics. Notice that, unlike HyperNeRF, our method depends on the estimated human pose, and we discard the offset network when rendering validation or test views. Therefore, our reported numbers also reflect the errors in the estimated human pose.

4 Ablation Studies

Our method has 3 components to correct the estimated geometry of the human: the offline SMPL optimization in the preprocessing, the online end-to-end SMPL optimization during the training, and the error-correction network which accounts for the warping errors. Without any geometry corrections components(Ours-GC), the canonical volume overfits to each observations causing averaged color and shape instead of sharp details of the human. However, with geometry correction enabled, the human NeRF model can learn both the textural and geometric details of the human. We show the comparison between Ours-Full and Ours-GC in Figure 9.

When Lsmpl{\mathcal{L}_{smpl}} is disabled, fog appears in the canonical volume around the human, and causes a halo in the final renderings, see Ours-Lsmpl{\mathcal{L}_{smpl}} in Figure 9. By encouraging the volume outside the canonical SMPL mesh to be empty, Lsmpl{\mathcal{L}_{smpl}} effectively removes the fog and suppresses the halos.

Conclusions

We have proposed a novel framework to reconstruct the human and the scene NeRF models that can be rendered with novel human poses and views from a single in-the-wild video. To do so, we use off-the-shelf methods to estimate either 2D or 3D geometry of the scene and the human to provide initialization. Our human NeRF model is able to learn texture details such as patterns on cloth, and geometric details such as sleeves, collar even zipper from less than 40 images.

One major limitation is that the dynamics beyond SMPL cannot be modeled with our static NeRF, those dynamics will degenerate to average shape or color. This is most evident in the hands when the gestures are changing over frames, see Figure 6. We hope to apply more expressive body models to model hand gestures , even garment dynamics , to mitigate this problem.

Additionally, our warping function is a simple extension of the SMPL mesh skinning, it could cause volume collision in some extreme cases. A collision-aware volume warping method or a learned one is required to improve the generalization under extreme poses.

Finally, we assume that the human always has at least one contact point with the ground to estimate the scale relative to the scene. To apply our method to videos with jumping or uneven ground, we need smarter geometric reasoning.

Acknowledgement

We thank Ashish Shrivastava, Russ Webb and Miguel Angel Bautista Martin for providing insightful review feedback.

References

A.1 Dataset Details

A.2 SMPL Refinement

Given an image, we regress the 2D joints j2dj_{2d} and segmentation mask mm of the human using HigherHRNet and DensePose . We further estimate the SMPL mesh M=(V,F)M=(V,F), a collection of vertices and faces using ROMP . The mesh MM is parametrized by SMPL parameters θ\theta such that M=SMPL(θ)M=\text{SMPL}(\theta) and includes the 3D joints j3d{j}_{3d}. The regressed SMPL parameters θ\theta are noisy. Therefore, we use soft-rasterizer , Π\Pi to refine these estimates. Given a mesh, MM and camera θc\theta_{c}, the rasterizer renders a silhouette m^=Π(θc,M)\hat{m}=\Pi(\theta_{c},M). We also project the 3D joints in the image plane using camera matrix j^2d=p(j3d)\hat{j}_{2d}=\mathbf{p}(j_{3d}) where p\mathbf{p} is a projection operator. We obtain the refined SMPL parameters and camera estimates by minimizing

Notice that we use the estimates from DensePose as the target silhouettes in the preprocessing, while using the estimates from Mask-RCNN as the target masks during the training of human NeRF model. It’s because in the preprocessing phase, the rendered masks are from naked SMPL mesh and DensePose is trained on such dense SMPL correspondences. However, we wish to learn extra geometry details beyond the SMPL model with the human NeRF model, and Mask-RCNN better estimates those 2D details such as the hair and clothes.

A.3 Network Architecture

Following , our scene NeRF models consists of a coarse sub-model and a fine sub-model. However, with the prior geometry provided by SMPL estimates, we only use one sub-model for the human NeRF model. Each sub-model and the error-correction network has the same architecture as shown in 10.

A.4 Comparison with previous works

We apply NeuralBody to our dataset in a monocular setting. The results are shown in 11. NeuralBody overfits to the training observations, and produce poor rendering on the back of the subject, while ours generalize better and can faithfully render the back.

We also compare our method with HumanNeRF and NeuralBody on a ZJU Mocap dataset, as shown qualitative comparisons in Figure 12. Our method renders high quality novel view renderings with the ability to extrapolate in pose space.

A.5 Error Correction Network and Scene Model Conditioning

Without the error correction network, the the canonical NeRF lacks details on the cloth and face. Training only the human NeRF in isolation leads to worse performance as the human NeRF model may encode the background pixels into its radiance fields due to segmentation errors. In either case, the canonical NeRF creates fogs around the human to hallucinate the clothing dynamics or the background colors, as shown in 13.