Space-time Neural Irradiance Fields for Free-Viewpoint Video

Wenqi Xian, Jia-Bin Huang, Johannes Kopf, Changil Kim

Introduction

This paper addresses the problem of rendering a video from novel viewpoints. Specifically, we learn a globally consistent, dynamic scene representation that can later be rendered from a novel viewpoint. We learn such a representation from a casually captured single video from everyday devices such as smartphones, without the assistance of multi-camera rigs or other dedicated hardware (which are typically not accessible to casual users).

Free-viewpoint video rendering typically requires a complicated hardware setup consisting of multiple cameras to capture the scene of interest from different viewpoints zitnick2004high; collet2015high; dou2016fusion4d; broxton2020immersive. The multi-camera setup in existing methods is required because conventional 3D reconstruction algorithms (multi-view stereopsis) assume a fully static scene and thus can perform reconstruction using the multiple captured viewpoints of a dynamic scene at any time. Most methods represent the geometric reconstructions as some form of per-frame representation (e.g., depth maps zitnick2004high or meshesbroxton2020immersive). Rendering from a novel viewpoint can then be achieved, e.g., by warping available views using their depth maps to the new viewpoint.

Following the success of single-image depth estimation, recent monocular video depth estimation methods allow for the acquisition of consistent per-frame depth estimates from only a single video Luo-VideoDepth-2020; Yoon-2020-CVPR. While still at an early stage, this line of work opens up new possibilities, where monocular scene depth estimates can be directly used for view synthesis. However, naïve approaches such as per-frame depth-based warping would lead to unnatural stretches and reveal holes in disoccluded regions (even with perfect depth estimates). One can alleviate this by post-processing the incomplete rendering Yoon-2020-CVPR. However, such per-frame processing methods often lead to temporal flickers. The core problem lies in the use of a frame-wise representation (e.g., depth maps associated with the input images), and therefore suffer from issues ranging from temporal inconsistency to high redundancy and thus excessive storage requirements and data transfer bandwidth.

In this work, we build on recent monocular video depth estimation methods and aggregate the entire spatiotemporal aspects of a dynamic scene in a single global representation. While fusing multiple depth maps into a single, global representation has a long tradition, most work on volumetric depth integration has focused on static scenes Curless96; Newcombe11 or geometry alone without textures newcombe2015dynamicfusion. These methods typically use discrete representations such as voxel grids, meshes, or point clouds. Consequently, these methods often suffer from premature hard decisions on geometry estimation and limited resolution due to high storage requirements.

In this paper, we, instead, turn to the recent advances in neural implicit representations, which allow for continuous representations of a scene without resolution loss. Recent work has shown that these representations achieve high-quality view interpolation of complex static scenes while retaining their advantages over discrete representations mildenhall2020nerf; zhang2020nerf++; liu2020neural. The current approaches to learn them, however, require either multiple posed images of a fully static scene mildenhall2020nerf; zhang2020nerf++; liu2020neural or ground truth 3D representations saito2019pifu; saito2020pifuhd. While videos often contain appearances of a scene seen from multiple viewpoints, they only contain exactly one viewpoint at any given time. Combined with the dynamic nature of video, this renders it nontrivial to extend current approaches to learn spatiotemporal representations from a single video.

Specifically, we learn neural irradiance fields as a function of both space and time for each video. We do not model view dependency, hence we use the term irradiance. Using supervision from only color frames of the input video as in mildenhall2020nerf is futile, since the variations between frames can be explained with either a change of appearance or geometry, or a combination of both. We resolve this ambiguity using the per-frame scene depth estimated from monocular video depth estimation. Our depth supervision constrains the scene’s geometry at any moment and disambiguates it from appearance variations. While this enables us to encode physically correct appearance and geometry in a global representation, it fails to fill the holes that could be seen at other time steps in the video. We address this by encouraging the color and volume density to propagate across time whenever spatial locations are not supervised otherwise. The resulting representation allows us to render the video from novel viewpoints and time: our implicit model can be queried at any spatiotemporal location and rendered using standard volume rendering Note that our method can render arbitrary viewpoints at all the observed time steps. We do not extrapolate or interpolate the time steps..

Our technical contributions include the following:

We aggregate frame-wise 2.5D representations into a globally consistent spatiotemporal representation from a single monocular video.

We address the inherent motion–appearance ambiguity using video depth supervision and constrain the disoccluded contents by propagating the color and volume density across time.

We demonstrate a compelling free-viewpoint video rendering experience on various casual videos shot from smartphones, preserving motion and texture details while conveying a vivid sense of 3D.

Related Work

View synthesis for images. Creating novel views from multiple images is a long-standing problem in computer vision and computer graphics. Existing image-based rendering techniques first extract approximated geometric proxy and create novel target views by warping and blending the corresponding contents from multiple source frames kalantari2016learning; penner2017soft; hedman2018deep; riegler2020free. Multi-plane images (MPIs) have been popular in the past few years as a geometric representation zhou2018stereo; mildenhall2019local; flynn2019deepview; li2020crowdsampling; huang2020semantic, allowing for compelling novel view synthesis quality. Many recent works further push the requirement of the number of input images to a narrow-baseline stereo pair zhou2018stereo; srinivasan2019pushing; choi2019extreme or a single image zhou2016view; tucker2020single; niklaus20193d; wiles2020synsin; kopf2020one; shih20203d. Recently, neural implicit representation has shown high-quality view synthesis results by implicitly modeling the volume density and color of the scene using the weights of a multi-layer perceptron mildenhall2020nerf; zhang2020nerf++; liu2020neural. Our work uses NeRF mildenhall2020nerf as our base scene representation for view synthesis. Unlike NeRF that only models static scenes, our focus is on creating new views from arbitrary viewpoint and time for dynamic scenes.

View synthesis for videos. Compared to images, view synthesis for video poses significant challenges due to the need to handle time-varying scene geometry and appearances. Consequently, most of the existing methods typical require laborious multi-camera setup zitnick2004high; collet2015high; orts2016holoportation; dkabala2016efficient; broxton2020immersive, special hardware attal2020matryodshka, or synchronous video captures from multiple viewpoints ballan2010unstructured; bansal20204d. Several methods can reduce the required number of input views by focusing on specific domains such as performance capture carranza2003free; dou2016fusion4d; habermann2019livecap and video re-animation kim2018deep; liu2020neural; liu2018neural; shysheya2019textured; chan2019everybody. In contrast, our work aims to enable view synthesis of a complex dynamic scene at any given viewpoints and time from a single video. Very recently, Yoon et al. Yoon-2020-CVPR also explore the same problem setup. In contrast to Yoon-2020-CVPR that processes each novel view independently (by warping blending multiple images), we can render an interpolation video with temporally smooth transitions across viewpoints. Our method does not assume a simple two-layer (foreground-background) model of the scene and can handle more generic scenes.

Neural implicit representation. Implicit representation has emerged as a powerful tool to overcome conventional limitations of discrete 3D representations such as voxel grids or meshes. The core idea is to use a multilayer perceptron (MLPs) to implicitly model the occupancy mescheder2019occupancy; michalkiewicz2019implicit, signed distance functions park2019deepsdf; atzmon2020sal, object appearance oechsle2019texture, volumetric density eslami2018neural; lombardi2019neural; sitzmann2019deepvoxels; thies2019deferred; sitzmann2019scene; mildenhall2020nerf in the 3D space. Differentiable rendering techniques enable training these models without accessing ground truth data for direct 3D supervision mantiuk2020state; Niemeyer2020DifferentiableVR; mildenhall2020nerf; liu2020neuralvoxel. However, extending the above methods to handle scene dynamics is not trivial due to the motion–appearance ambiguity. Occupancy flow niemeyer2019occupancy achieves 4D reconstruction (3D shape + 1D time) by learning continuous motion fields. Our work builds upon the recent advances in neural implicit representation but focuses on representing a dynamic video. Compared to 4D reconstruction in niemeyer2019occupancy, our method differs in the following two aspects. First, our method does not require direct 3D ground truth training data. Second, in addition to model the time-varying 3D geometry, we also model the appearance of complex scenes.

Video depth estimation. Estimating dense depth from a dynamic video is a challenging task. Existing multi-view stereo (MVS) methods (either geometric-based schonberger2016structure; furukawa2015multi or learning-based huang2018deepmvs; yao2018mvsnet; gu2020cascade; long2020occlusion) assume static scene and thus not suitable for dynamic videos as it often produces erroneous depth for moving objects or untextured regions. Several monocular video depth estimation methods predict depth using the cost volume computed from nearby frames teed2020deepv2d; liu2019neural. Similar to MVS algorithms, these video-to-depth methods have difficulties in handling dynamic scenes well. Very recently, hybrid methods that combine MVS and single-image depth estimation models have been proposed Luo-VideoDepth-2020; Yoon-2020-CVPR; kopf2020robust. Our work leverages the estimated depth from Luo-VideoDepth-2020 to help resolve the ambiguity when learning the spatiotemporal neural radiance fields using a single video. Our approach renders photorealistic views with correctly filled dis-occluded contents compared to view synthesis with per-frame depth-based warping.

Video completion. State-of-the-art video completion methods achieve temporally consistent completion by propagating known contents to missing regions along flow trajectories ilan2015survey; huang2016temporally; xu2019deep; gao2020flow. One may first render a free-viewpoint video according to each frame’s estimated depth, followed by filling the missing pixels (disoccluded regions due to view changes) using video completion algorithms. Our method also produces completed novel views (i.e., with no missing pixels from disocclusion). However, unlike video completion algorithms that inpaint the dis-occluded pixels in the screen space, our approach fills in the dis-occluded content implicitly in the 3D space. Our experiments validate that our approach produces significantly fewer artifacts than the baseline method using video completion.

Depth map fusion. A line of research work focuses on fusing a sequence of RGB-Depth images in a video into a global with voxel-, point-based, signed distance field-based representation zollhofer2018state. Examples include 3D reconstruction for static scenes niessner2013real; izadi2011kinectfusion or dynamic objects from a single newcombe2015dynamicfusion; innmann2016volumedeform or RGB-D cameras dou2016fusion4d; bozic2020deepdeform; innmann2016volumedeform; newcombe2015dynamicfusion; ye2014real; zollhofer2014real. Our work differs from prior 3D reconstruction methods in two aspects. First, we do not assume a fixed, canonical 3D model as in existing dynamic 3D reconstruction methods and, therefore, can naturally handle an entire dynamic scene (as opposed to only individual objects). Second, our approach with neural implicit representations jointly models time-varying geometry and appearance.

Concurrent works. Several concurrent works on extending NeRF mildenhall2020nerf to dynamic scenes from monocular video have been proposed park2020deformable; pumarola2020d; li2020neural; tretschk2020non; du2020neural; gao2021dynamic. These methods either learn a static canonical radiance field with deformation park2020deformable; pumarola2020d; tretschk2020non or a dynamic radiance field directly conditioned on time li2020neural; du2020neural; gao2021dynamic. Our work belongs to the latter; while the other two works li2020neural; du2020neural regularize the training primarily with flow information, ours does so with dynamic scene depth. We refer the readers to these papers for a complete picture.

Background

The color of a pixel can be rendered by integrating the radiance modulated by the volume density along the camera ray r(s)=o+sd\mathbf{r}(s)=\mathbf{o}+s\mathbf{d}, shot from the camera center through the center of the pixel:

is the accumulated transmittance along the ray r\mathbf{r} up to ss.

One can train the MLP using multiple posed images, capturing a static scene from different viewpoints. Specifically, we minimize the photometric loss that compares the rendering through a ray r\mathbf{r} with the corresponding ground truth color from an input image:

where \altmathcalR\altmathcal{R} denotes a set of rays, and C(r)C(\mathbf{r}) and C^(r)\hat{C}(\mathbf{r}) the ground truth and the estimated color, respectively.

In the implementation, the continuous volume rendering of (1) is approximated by numerical quadrature, i.e., computing the color using a finite number of sampled 3D points along a ray and calculate the summation of the radiances, weighted by the discrete transmittance. As this weighted summation process is differentiable, the gradient can propagate backward for optimizing the MLP. We perform the sampling in two steps. First, a ray is sampled uniformly in ss, and then, it is sampled with respect to the approximate transmittance so that more samples are used around surfaces in the scene. The two groups of samples are evaluated in separate coarse and fine networks, and both are used to measure the loss (3).

Space-time Neural Irradiance Fields

A ray r\mathbf{r} at time tt can be determined by a pixel location u\mathbf{u} and the camera calibration \altmathcalPt\altmathcal{P}_{t}: it marches from the camera center through the center of pixel denoted by u\mathbf{u}. Additionally, we parameterize a ray such that the parameter ss denotes the scene depth. This is achieved by setting the directional vector d\mathbf{d} such that its projection onto the principal ray has a unit norm in the camera space.

Color reconstruction loss. To learn the implicit function FF from the input video II, first and foremost, we constrain our representation FF such that it reproduces the original video II when rendered from the original viewpoint for each frame. Specifically, we penalize the difference between the volume-rendered image at each time tt and the corresponding input image ItI_{t}. This amounts to the reconstruction loss of the original NeRF mildenhall2020nerf:

where \altmathcalR\altmathcal{R} is a batch of rays, each associated with a time tt.

Unlike NeRF, for dynamic scenes, we have to reconstruct the time-varying scene geometry at every time tt. However, a single video contains only one observation of the scene at any point in time, rendering the estimation of scene geometry severely under-constrained. That is, the 3D geometry of a scene can be legitimately represented in numerous (infinitely possible) ways since varying geometry can be explained with the varying appearance and vice versa. For example, any input video can be reconstructed with a “a flat TV” solution (with a planar geometry with each frame texture-mapped).

Thus, the color reconstruction loss provides the ground for accurate reconstruction only when the learned representation is rendered from the same camera trajectory of the input, lacking any machinery that drives learning correct geometry. Incorrect geometry would lead to artifacts as soon as we start deviating from the original video’s camera trajectory, as shown in Figure 2a.

Depth reconstruction loss. We resolve this motion–appearance ambiguity by constraining the time-varying geometry of our dynamic scene representation using the per-frame scene depth of the input video (estimated from video depth estimation methods). We estimate the scene depth from the learned volume density of the scene, and measure its difference from the input depth dtd_{t}. It is a non-trivial question on how to define the scene depth of a ray. One possibility is to measure the distance where the accumulated transmittance TT becomes less than a certain threshold. Such an approach, however, involves heuristics and hard decisions. Instead, we accumulate depth values along the ray modulated both with the transmittance and volume density, similarly to the depth composition in layered scene representations tucker2020single.

Our depth reconstruction loss is of the form:

is the integrated sample depth values along the ray and D(r)D(\mathbf{r}).

where u\mathbf{u} denotes the pixel coordinates where r\mathbf{r} intersects with the image plane at tt, dt(u)d_{t}(\mathbf{u}) denote the scene depth for the pixel u\mathbf{u} at time tt.

The empty-space loss combined with the depth reconstruction loss provides geometric constraints for our representation up to and around visible scene surfaces at each frame. The learned representations can thus produce geometrically correct novel view synthesis, as shown in Figure 2b.

Static scene loss. A large portion of spaces that is hidden from the input frame’s viewpoint at any given time is still not constrained, i.e., the MLP has not seen the 3D positions and time as input queries during training. As a result, when these unconstrained spaces are disoccluded due to viewpoint changes, they are prone to artifacts (see Figure 4 for an example). However, there is a high chance that a portion of disoccluded spaces is observed from a different viewpoint at another time. Our idea is to constrain the MLP by propagating these partially observed contents across time. However, instead of explicitly correlating surfaces over time, e.g, using the scene flow, we choose to constrain the spaces surrounding the surface regions. This allows us to avoid misalignment of scene surfaces due to unreliable geometry estimates or other image aberrations commonly seen in captured videos such as exposure or color variations.

We make a simple assumption on unobserved spaces: every part of the world should stay static unless observed not as such. Enforcing this assumption prevents the part of spaces that are not observed from going entirely unconstrained. Our static scene constraint encourages the shared color and volume density at the same spatial location x\mathbf{x} between two distinct times tt and t′t^{\prime}:

where both (x,t)(\mathbf{x},t) and (x,t′)(\mathbf{x},t^{\prime}) are not close to any visible surfaces, and \altmathcalX\altmathcal{X} denotes a set of sampling locations where the loss is measured.

Scene sampling. While we have locations for the color, depth, and free-space supervisions explicitly dictated by quadrature used by volume rendering mildenhall2020nerf, we are free to choose where we apply the static constraints. A straightforward approach would be to use the same sampling locations that are used for other losses. We can then randomly draw another time t′t^{\prime} that is distinct from the current time tt and enforce the MLP to produce similar appearances and volume densities at these two spatiotemporal locations.

However, this still leaves a large part of the scene unconstrained when the camera motion is large. Uniformly sampling in the scene bounding volume would also not be ideal since sampling would be highly inefficient because of perspective projection (except for special cases like a camera circling some bounded volume).

As a simple solution to meet both the sampling efficiency and the sample coverage, we propose to take the union of all sampling locations along all rays of all frames to form our sample pool \altmathcalX\altmathcal{X}. We exclude all points that are closer to any observed surfaces than a threshold ϵ\epsilon (see Figure 3). We randomly draw a fixed number of sampling locations from this pool at each training iteration and add small random jitters to each sampling location. At time t′t^{\prime} the static scene loss is measured against is also randomly chosen for each sample location x\mathbf{x}, while ensuring the resulting location (x,t′)(\mathbf{x},t^{\prime}) is not close to any scene surfaces.

Total loss. Our total loss for training the space-time irradiance fields is a linear combination of all losses presented above:

We validate the effectiveness of these losses in Section 5.

Experimental Results

We first compare our method with baseline approaches using the videos of dynamic scenes. Note that we always render new views at one of the observed times. That is, we do not evaluate our method’s capability of temporal interpolation/extrapolation since our method is not designed to address it. We then provide extensive quantitative ablation studies with the variants of our model where each loss is added one at a time. We urge the readers to watch our supplementary videos in our project webpage (https://video-nerf.github.io), where we provide free-viewpoint rendering of our learned representations.

Datasets. We use the videos of dynamic scenes from the recent consistent video estimation method of Xuan et al. Luo-VideoDepth-2020 along with the camera calibration and the per-frame depth maps provided together. We also intended to use the dataset and video depth of Yoon et al. Yoon-2020-CVPR, but were unable to obtain the depth maps used in their results. Their dataset (denoted by “CVD”) consists of short videos of moving subjects captured by a smartphone.

For quantitative evaluation, we use the synthetic stereo videos from the MPI Sintel dataset Butler12. We select 6 videos that show a variety of characteristics in terms of scene motion, camera motion, and the size of moving subjects. We use the left video to train our models and render them from the right video’s viewpoints for ground truth comparisons.

Baselines. We compare our method against several baseline methods: (“Mesh”) textured mesh representations directly reconstructed from the input depth maps as demonstrated by Xuan et al. Luo-VideoDepth-2020; (“Inpainted”) its inpainted version, where disoccluded empty pixels in the 2D rendered images are inpainted using a recent video inpainting method gao2020flow; and (“NeRF-T”) a version of NeRF with an extra time parameter, which is our model trained with only the color reconstruction loss (4). Note that we do not use the viewing directions for the NeRF baseline.

Qualitative comparisons. Figure 5 presents the comparisons of our model against the baselines using the “CVD” dataset, where we show view synthesis results from novel viewpoints. We used our full method with all losses presented in Section 4 to create these results. Please refer to our project webpage for the full video results.

For quantitative evaluation, we train our models using the Sintel dataset. The left videos are used to train the models rendered from the right videos’ viewpoints and then compared to the ground truth right videos. All metrics aggregate the scores over all frames. Our model always works better when trained with the depth loss. While varying depending on scene types, the static loss and empty space loss help improve the results as well. However, we find that the scene flow loss does not help improve quality. We suspect that this is because casual videos often include strong image-space aberration such as exposure or color changes and monocular depth estimates does not provide as accurate scene depth as stereo-based methods do. For example, this could be addressed by a latent code factoring out such variations, similarly done in NeRF-W martinbrualla2020nerfw. While our “Model-2” works slightly better than our full model (“Model-4”) quantitatively, we have found that our full model works usually the best for real data.

Conclusions

We have presented a simple yet effective algorithm for learning space-time irradiance fields from single casually captured videos. Our core technical contributions are (1) leveraging monocular video depth estimation to constrain the time-varying geometry of our learned neural implicit functions and (2) designing a static scene loss and a sampling strategy to propagate scene contents across time. We extensively validate and justify our design choices both visually and quantitatively on the Sintel dataset. We showcase free-viewpoint video rendering of several challenging dynamic scenes captured with hand-held cellphone cameras.

We thank Ayush Saraf for his help with distributed training. All photos of individuals in this paper are used with permission.

References

Appendix A Additional Details

In this appendix, we provide implementation details (in Section A.1), training details (in Section A.2) and additional quantitative comparison in Section A.3. We include additional qualitative results in our project website https://video-nerf.github.io. In particular, we test on 10 different videos and provide comparison with baselines and different loss configurations.

We implement our framework using PyTorch. In all our experiments, we empirically set the hyper-parameters as α=1\alpha=1, β=100\beta=100, and γ=10\gamma=10.

A.2 Training Details

We used the same MLP architecture as in mildenhall2020nerf, except that we use 1024 activations for the first 8 MLP layers instead of 256. Our models are trained with various combinations of the losses presented in the main paper but otherwise with the same hyperparameters. We used the Adam optimizer with momentum parameters β1=0.9\beta_{1}=0.9 and β2=0.999\beta_{2}=0.999 and a learning rate of 0.0005. We train the MLP for 800k iterations. It takes about 48 hours to train a network with about 100 video frames at the 960×\times540 resolution with 4 NVIDIA V100 GPUs.

A.3 Additional Comparisons to Baseline

In this section, we provide a quantitative comparison to the inpainted mesh method. We rendered the ground-truth mesh provided by the Sintel dataset in 2D and inpaint the missing pixels in disoccluded areas using state-of-the-arts video completion algorithms. Then, we evaluate how well our method handles disoccluded areas, compared with the inpainted mesh method, for which we used the same Sintel GT depth. Table 2 shows the quantitative evaluation measured in PSNR metric. It demonstrates that inpainting is not sufficient to get good disocclusion contents, and validates that our approach produces significantly fewer artifacts than the baseline method using video completion.