Dense Depth Priors for Neural Radiance Fields from Sparse Input Views

Barbara Roessle, Jonathan T. Barron, Ben Mildenhall, Pratul P. Srinivasan, Matthias Nießner

Introduction

Synthesizing realistic views from varying viewpoints is essential for interactions between humans and virtual environments, hence it is of key importance for many virtual reality applications. The novel view synthesis task is especially relevant for indoor scenes, where it enables virtual navigation through buildings, tourist destinations, or game environments. When scaling up such applications, it is preferable to minimize the amount of input data required to store and process, as well as its acquisition time and cost. In addition, a static scene requirement is increasingly difficult to fulfill for a longer capture duration. Thus, our goal is novel view synthesis at room-scale using few input views.

NeRF represents the radiance field and density distribution of a scene as a multi-layer perceptron (MLP) and uses volume rendering to synthesize output images. This approach creates impressive, photo-realistic novel views of scenes with complex geometry and appearance. Unfortunately, applying NeRF to real-world, room-scale scenes given only tens of images does not produce desirable results (Fig. 1) for the following reason: NeRF purely relies on RGB values to determine correspondences between input images. As a result, high visual quality is only achieved by NeRF when it is given enough images to overcome the inherent ambiguity of the correspondence problem. Real-world indoor scenes have characteristics that further complicate the ambiguity challenge: First, in contrast to an “outside-in” viewing scenario of images taken around a central object, views of rooms represent an “inside-out” viewing scenario, in which the same number of images will exhibit significantly less overlap with each other. Second, indoor scenes often have large areas with minimal texture, such as white walls. Third, real-world data often has inconsistent color values across views, e.g., due to white balance or lens shading artifacts. These characteristics of indoor scenes are likewise challenging for SfM, leading to very sparse SfM reconstructions, often with severe outliers. Our idea is to use this noisy and incomplete depth data and from it produce a complete dense map alongside a per-pixel uncertainty estimate of those depths, thereby increasing its value for NeRF — especially in textureless, rarely observed, or color-inconsistent areas.

We propose a method that guides the NeRF optimization with dense depth priors, without the need for additional depth input (e.g., from an RGB-D sensor) of the scene. Instead, we take advantage of the sparse reconstruction that is freely available as a byproduct of running SfM to compute camera pose parameters. Specifically, we complete the sparse depth maps with a network that estimates uncertainty along with depth. Taking uncertainty into account, we use the resulting dense depth to constrain the optimization and to guide the scene sampling. We evaluate the effectiveness of our method on complete rooms from the Matterport3D and ScanNet datasets, using only a handful of input images. We show that our approach improves over recent and concurrent work that uses sparse depth from SfM or multi-view stereo (MVS) in NeRF .

In summary, we demonstrate that dense depth priors with uncertainty estimates enable novel view synthesis with NeRF on room-size scenes using only 18–36 images, enabled by the following contributions:

A data-efficient approach to novel view synthesis on real-world scenes at room-scale.

An approach to enhance noisy sparse depth input from SfM to support the NeRF optimization.

A technique for accounting for variable uncertainty when guiding NeRF with depth information.

Related Work

The ability to synthesize novel views of a scene from a set of observed images and corresponding camera viewpoints is necessary for enabling virtual experiences of real-world environments. In situations where it is feasible to densely sample images of the scene, novel viewpoints can be synthesized with simple light field interpolation . However, when fewer observed views of the scene are available, it becomes increasingly necessary to use information about the scene’s geometry to render new views. A common paradigm for geometry-based view synthesis is to use a triangle mesh representation of scene geometry to reproject observed images into each novel viewpoint and combine them using either heuristic or learned blending algorithms. More recently, these mesh-based geometry models have been replaced by volumetric scene representations such as voxel grids or multiplane images . NeRF popularized an approach that avoids the steep scaling properties of discrete voxel representations by representing a scene as a continuous volume, parameterized by a MLP that is optimized to minimize the loss of re-rendering all observed views of a scene. Since its introduction, NeRF has become the dominant scene representation for view synthesis, and many recent works are built on top of NeRF’s neural volumetric model.

However, in situations where the scene is observed from very few sparsely-sampled viewpoints, NeRF’s high capacity to model detailed geometry and appearance can result in various artifacts, such as “floaters”, i.e., artifacts caused by a flawed density distribution. In this work, we directly address the few-input setting, proposing a strategy that takes advantage of sparse depth to constrain NeRF’s scene geometry and improve rendering quality. This depth data is freely available as a byproduct of running SfM to compute camera poses from the input images (e.g., by using COLMAP ). Our method takes inspiration from techniques that generate complete dense depth maps from sparse depth inputs. These include classic techniques that fuse observed depths into a single 3D reconstruction, typically in the form of a truncated signed distance function , as well as more recent techniques that train deep networks to operate over the sparsely observed geometry in 3D . Although these methods are effective for dense 3D scene reconstruction, the resulting geometry is not ideal for view synthesis since its edges frequently do not align with edges in the observed images. Instead, we leverage recent work on 2D sparse depth completion that directly completes depth maps in image space and extend it to also predict uncertainty.

A few recent works have also proposed incorporating depth observations into NeRF reconstruction. NerfingMVS uses depth from MVS to overfit a depth predictor to the scene. The resulting depth prior guides the NeRF sampling. In comparison, our method employs depth completion on the SfM depth and additionally employs a depth loss to supervise the geometry recovered by NeRF. This way, our novel views achieve significantly better color and depth quality in the few-input setting without relying on the computationally more expensive MVS preprocessing. Concurrent work on depth-supervised NeRF directly uses sparse depth information from SfM in the NeRF optimization. To handle inaccuracy in the sparse reconstruction, the 3D points are weighted according to their reprojection error. In contrast, we learn dense depth priors with uncertainty to more effectively guide the optimization, leading to more detailed novel views, as well as more accurate geometry and higher robustness to SfM outliers.

Method

With the goal of completing sparse depth from SfM, two challenges presented by this input data play a key role in designing the depth prior network. First, sparse reconstructions are noisy and have outliers. As a consequence, dense depth predictions are expected to have spatially varying accuracy, which makes it crucial to know the uncertainty at a per-pixel level. Second, the density of SfM point clouds varies significantly across space, depending on the number of image features. E.g., SfM reconstructions from 18–20 images per ScanNet scene lead to sparse depth maps with 0.04% valid pixels on average. Hence, depth completion must be able to predict dense depth even from largely empty sparse depth maps.

Network Training

Under the assumption that the output is normally distributed, we supervise the network by minimizing the negative log likelihood of a Gaussian:

2 Radiance Field with Dense Depth Priors

Optimization with Depth Constraint

To optimize the radiance field, the color C^(r)\hat{\mathbf{C}}(\mathbf{r}) of each pixel in the batch RR is computed by evaluating a discretized version of the volume rendering integral (Eq. 4 ). Specifically, a pixel determines a ray r(t)=o+td\mathbf{r}(t)=\mathbf{o}+t\mathbf{d} whose origin is at the camera’s center of projection o\mathbf{o}. Rays are sampled along their traversal through the volume. For each sampling location tk∈[tn,tf]t_{k}\in[t_{n},t_{f}] within the near and far planes, a query to Fθ1F_{\theta_{1}} provides the local color and density.

Besides the predicted color of a ray, a NeRF depth estimate z^(r)\hat{z}(\mathbf{r}) and standard deviation s^(r)\hat{s}(\mathbf{r}) are needed to supervise the radiance field according to the learned depth prior (Sec. 3.1). The NeRF depth estimate and standard deviation are computed from the rendering weights wkw_{k}:

Depth-Guided Sampling

In addition to the depth loss function, the depth prior contains valuable signal to guide sampling along a ray. To render one pixel of a room-scale scene, we require the same number of MLP queries as the original NeRF; however, we replace the coarse network used for hierarchical sampling. During optimization, half of the samples are distributed between the near and far planes and the second half are drawn from the Gaussian distribution determined by the depth prior N(z(r),s(r)2)\mathcal{N}(z(\mathbf{r}),s(\mathbf{r})^{2}). At test time, when the depth is unknown, the first half of the samples are used to render an approximate depth z^(r)\hat{z}(\mathbf{r}) and standard deviation s^(r)\hat{s}(\mathbf{r}) that is then used to sample the second half according to N(z^(r),s^(r)2)\mathcal{N}(\hat{z}(\mathbf{r}),\hat{s}(\mathbf{r})^{2}).

Results

We evaluate our method with a baseline comparison (Sec. 4.3) and an ablation study (Sec. 4.4) on the ScanNet and Matterport3D datasets.

We run COLMAP SfM to obtain camera parameters and sparse depth. Specifically, we run SfM on all images to obtain camera parameters. To ensure a clean split between train and test data, we withhold the test images when computing the point cloud used for rendering the sparse depth maps. On average, the resulting depth maps have 0.04% valid pixels. We use three sample scenes each with 18 to 20 train images and 8 test images. This set of images results from excluding video frames with motion blur while ensuring that surfaces are observed from at least one input view. Details are provided in Sec. A.1.

Matterport3D

Using RGB images from the PrimeSense camera, COLMAP SfM struggled to reconstruct complete rooms in Matterport3D, hence, we mimic sparse depth from SfM by sampling and perturbing the sensor depth as described for depth prior training in Sec. 3.1. Sparse depth maps rendered from a SfM point cloud are by nature 3D-consistent. While consistency in 3D is irrelevant for training 2D depth completion, it plays a critical role when optimizing a 3D scene representation with NeRF. Hence, we ensure 3D-consistent sparse depth on the scenes used for NeRF by projecting the sampled and perturbed 3D points to all other views. On average the resulting depth maps are 0.1% complete. The impact of the sparse depth density is studied in Appendix B. We evaluate three example rooms each with 24 to 36 train images and 8 test images.

NeRF Optimization

We process rays in batches of 1024 and use the Adam optimizer with learning rate 0.0005. For fairness, all approaches in the ablation and baseline experiments are configured to use 256 MLP evaluations per pixel, independent of the used sampling approach. The radiance fields are optimized for 500k iterations. Further NeRF and depth prior implementation details are available in Appendix C.

Evaluation Metrics

For quantitative comparison, we compute the peak signal-to-noise ratio (PSNR), the Structural Similarity Index Measure (SSIM) and the Learned Perceptual Image Patch Similarity (LPIPS) on the RGB of novel views as well as the root-mean-square error (RMSE) on the expected ray termination depth of NeRF against the sensor depth in meters. By comparing color values directly, PSNR has limited expressiveness, when the images of the scene have inconsistent color. As shown in Sec. 4.4, the latent codes used to represent view-specific appearance largely help to produce consistent colors across the scene. Still, the color of a rendered image will not necessarily be similar to that of the test view against which it is evaluated. To compensate for this difference, we report an additional PSNR value, which is computed after optimizing for the latent codes on the entire test views. We are unable to use the left/right image split evaluation procedure from NeRF-W , because appearance changes too drastically across the image, so these numbers should be considered an upper bound on performance. This additional metric is listed in parentheses (Tabs. 2 and 3) for all approaches that use a latent code. All other metrics as well as all renderings in the paper are computed by setting the latent code to zero, given that the codes are unknown at test time.

2 Depth Priors

Tab. 1 shows the depth prior accuracy on the three ScanNet and three Matterport3D scenes used for NeRF. These scenes are part of the test sets during depth completion training.

The higher quality, generated sparse depth on Matterport3D leads to more accurate dense depth priors. However, the network interpolates the more noisy sparse depth from SfM on ScanNet without a relevant drop in accuracy.

3 Baseline Comparison

We compare our method to NeRF as well as recent and concurrent work that equally uses sparse depth input in NeRF, namely Depth-supervised NeRF (DS-NeRF) and NerfingMVS . Since DS-NeRF and NerfingMVS rely on SfM and MVS depth, respectively, they are run on ScanNet. NeRF and our method are run on both datasets. The quantitative results (Tab. 2) show that our method outperforms the baselines in all metrics.

“Floaters” are a common problem when applying NeRF approaches in a setting with few input views. By using dense depth priors with uncertainty, our method strongly reduces these artifacts compared to the baselines (example 2 Fig. 3). This contributes to far more accurate depth output and greater detail in color, e.g., visible in the books and the door handle (example 3 Fig. 3). We found that our method is more robust to outliers in the sparse depth input. E.g., erroneous SfM points in the area of the sofa back (example 5 Fig. 3) cause much larger deficiencies in geometry and color of the other approaches. This suggests that dense depth priors with uncertainty focus the optimization on more certain and accurate views, while direct incorporation of sparse depth, as in DS-NeRF, is more error-prone. Besides greater robustness to outliers, dense depth guides NeRF better at object boundaries that are not represented in the sparse depth input. This is observable in example 6 (Fig. 3), where a part of the chair back is missing in DS-NeRF, while it is complete using our method.

NerfingMVS Details The error map calculation in NerfingMVS fails when applied to an entire room as opposed to a local region, causing invalid sampling ranges. The issue is solved as detailed in Sec. C.2. To improve the performance of this baseline, we train its depth predictor 10 epochs longer than was done in the paper. Still, the depth priors on the ScanNet scenes remain at RMSE 0.379m. While our method uses only train images to compute the sparse depth input, NerfingMVS runs COLMAP MVS on train and test images together, which gives them an advantage.

4 Ablation Study

To verify the effectiveness of the added components, we conduct ablation experiments on the ScanNet and Matterport3D scenes. The quantitative results (Tabs. 2 and 3) show that the full version of our method achieves the best performance in image quality and depth estimates. This is consistent with the qualitative results in Fig. 4.

Without Completion Omitting depth completion and supervising with sparse depth only leads to inaccurate depth and color due to “floaters” in areas without depth input. Even in areas with sparse depth points, the results are less sharp than in versions that use completed depth.

Without Uncertainty Removing uncertainty from the optimization causes problems in resolving inconsistency in overlapping areas of the 2D depth priors. This results in wrong edges in RGB and depth (example 2 Fig. 4), duplication artifacts (example 4 Fig. 4) or lacking detail, e.g., in the patterns on the back of the chair (example 1 Fig. 4). The quantitative results on ScanNet (Tab. 2) show that considering uncertainty becomes even more important when using the lower quality sparse depth from SfM.

Without GNLL In this experiment, we replace GNLL with MSE in our depth loss (Eq. 11), and observe that MSE struggles to constrain density behind surfaces. The lack of sharp edges in the density distribution is most visible for novel views looking in tangential direction of a surface, e.g., looking into the corridor (example 3 Fig. 4).

Without Latent Code Omitting the latent code that models per-camera information, leads to incapability to produce smooth and consistent color output across the scene. When rendering novel views, the frustums of training images are clearly visible by causing severe shifts in color intensity (examples 2 and 3, Fig. 4).

5 Limitations

Our method allows for a significant reduction in the number of input images for NeRF-based novel view synthesis while at the same time applying it to larger room-size scenes. However, other NeRF limitations such as long optimization times and slow rendering remain. As a consequence of the drastic reduction in the number of input images, surfaces are typically not observed by more than two other views, hence view-dependent effects are limited. While our approach optimizes NeRF given as few as 18 images, the depth prior network requires a larger training dataset. Although these priors generalize well and only need to be trained once, it would be beneficial if the depth reconstruction could be also learned from a sparse setting.

Conclusion

We have presented a method for novel view synthesis using neural radiance fields (NeRF) that leverages dense depth priors, thus facilitating reconstructions with only 18 to 36 input images for a complete room. By learning a depth prior that generalizes across scenes, our method takes advantage of depth information without requiring depth sensor input of the scene. Instead, the depth prior network relies on the sparse reconstruction, which is available for free after structure from motion (SfM) on the input images. With only a few input views available, we show that our dense depth priors with uncertainty effectively guide the NeRF optimization, thus leading to significantly higher image quality of novel views and more accurate depth estimates compared to other approaches using SfM or multi-view stereo output in NeRF. Overall, we believe that our method is an important step towards making NeRF reconstructions available in commodity settings.

Acknowledgements

This project is funded by a TUM-IAS Rudolf Mößbauer Fellowship, the ERC Starting Grant Scan2CAD (804724), and the German Research Foundation (DFG) Grant Making Machine Learning on Static and Dynamic 3D Data Practical. We thank Angela Dai for the video voice-over.

References

Appendix A Datasets

Motion Blur Detection We consider motion blur when sampling a small subset of images to be used in NeRF: From each window of nn consecutive video frames the sharpest one is selected according to the following metric, where high values indicate sharpness: first, an image is converted to grayscale, then it is convolved with a discrete Laplacian kernel; finally, the variance is computed. nn is set to 10 or 20, depending on how densely the video samples the scene.

Train/Test Image Selection After removing images with severe motion blur, we consider the following criteria: 1) SfM successfully registers the set of images. 2) Surfaces to be reconstructed are observed from at least one input view. In practice, images are removed if their content is visible by other images and the remaining set fulfils 1). This way, 22% of the train pixels are not observed by any other train view, 31% are observed by one other, 47% by two or more. Test views have on average 66% overlap with their most overlapping train view.

Image Resolution The image resolution is 468×\times624 after downsampling and cropping dark borders from calibration.

Test Scenes We ensure that the test scenes are complete, sufficiently large rooms. The following scenes are used for evaluation:

SfM Quality on Few Views Figure 5 shows the mean absolute error (MAE) of the SfM points against the sensor depth. It is computed on the 6291 points from the three ScanNet evaluation scenes. The maximal error is 5.85m. We do not filter the COLMAP SfM output, i.e., all points are projected to the corresponding input views and used as input to the depth completion.

A.2 Matterport3D [2]

Train/Test Image Selection Similar to ScanNet, it is ensured that surfaces are observed from at least one input view. 25% of the train pixels are not observed by any other train view, 45% are observed by one other, 30% by two or more. Test views have on average 67% overlap with their most overlapping train view.

Image Resolution The image resolution is 504×\times630 after downsampling and cropping dark borders from calibration.

Test Scenes We avoid unbounded open space, which is challenging for NeRF approaches. The following scenes are used for evaluation:

Appendix B Impact of Sparse Depth Density

We investigate the impact of the sparse depth density on Matterport3D by decreasing it from 0.1% to 0.05% and 0.01% (Tab. 4).

While reduced sparse depth lowers performance, it clearly shows that depth completion increases the value of very sparse depth input: With just one tenth of the sparse depth our method still performs better, than the version without completion. Despite 0.01% being very sparse—just 32 points per image on average—we expect that using monocular depth estimation is challenging as view-consistent depth is needed.

Appendix C Implementation Details

Depth Completion The depth completion network is based on the architecture from Cheng et al. . We use a ResNet-18 encoder and add a second upsampling branch for uncertainty estimation. It equally consists of up-projection layers with skip connections to the same downsampling layers as the depth prediction branch. To increase performance on very sparse input depth, both branches use a CSPN module, configured to 48 iterations in the depth branch and 24 iterations in the standard deviation branch. The depth completion network is trained at a lower resolution of 256×\times320 on Matterport3D, and 240×\times320 on ScanNet. We use the Adam optimizer with a learning rate of 0.0001 and a batch size of 8. We train for 50 epochs on Matterport3D and 12 epochs on ScanNet. On Matterport3D 80 houses are used for training, 5 houses for validation, and 5 houses for testing. On ScanNet we use the provided data split. We ensure that the scenes used for NeRF are not included during training, and are instead in the test sets.

C.2 NerfingMVS [26]

The error map calculation used by NerfingMVS was not sufficiently robust to by applied to entire rooms, so to improve this baseline’s performance we adapted it as follows:

Original Calculation For each input view an error map is computed by projecting the 3D points according to the depth prior to all other views, where a depth reprojection error is computed and normalized with the projected depth. The mean of the 4 smallest errors are used as values in the error map.

Problem on Entire Rooms When applying the computation on entire rooms as opposed to a local region, the projected 3D points from other views frequently lie behind the camera. As a result the computed mean is often negative. Similarly, the computation of the near and far planes of the scenes is not suited for entire rooms, leading to a negative near plane in our case. Negative near plane and negative error map content lead to invalid sampling ranges, where the far bound lies in front of the near bound. to address this, we set the near and far planes (tnt_{n} and tft_{f}) of each scene such that all depth prior values are contained. In the error map calculation, we assign a maximal error tf−tnt_{f}-t_{n} for projected points that lie behind the camera. Afterwards, the error map values are still computed as the mean of the smallest 4 errors.

C.3 DS-NeRF [10]

We used the same positional encoding frequencies as described for our method in the main paper for this baseline, which improved its performance. A depth loss weight of 0.1 was suitable for the ScanNet scenes.

C.4 NeRF [21]

As in DS-NeRF, we used our own positional encoding frequencies for this baseline, which improved its performance.