Depth-supervised NeRF: Fewer Views and Faster Training for Free

Kangle Deng, Andrew Liu, Jun-Yan Zhu, Deva Ramanan

Introduction

Neural rendering with implicit representations has become a widely-used technique for solving many vision and graphics tasks ranging from view synthesis , to re-lighting , to pose and shape estimation , to 3D-aware image synthesis and editing , to modeling dynamic scenes . The seminal work of Neural Radiance Fields (NeRF) demonstrated impressive view synthesis results by using implicit functions to encode volumetric density and color observations.

In spite of this, NeRF has several limitations. Reconstructing both the scene appearance and geometry can be ill-posed given a small number of input views. Figure 2 shows that NeRF can learn wildly inaccurate scene geometries that still accurately render train-views. However, such models produce poor renderings of novel test-views, essentially overfitting to the train set. Furthermore, even given a large number of input views, NeRF can still be time-consuming to train; it often takes between ten hours to several days to model a single scene at moderate resolutions on a single GPU. The training is slow due to both the expensive ray-casting operations and lengthy optimization process.

In this work, we explore depth as an additional, cheap source of supervision to guide the geometry learned by NeRF. Typical NeRF pipelines require images and camera poses, where the latter are estimated from structure-from-motion (SFM) solvers such as COLMAP . In addition to returning cameras, COLMAP also outputs sparse 3D point clouds as well as their reprojection errors. We impose a loss to encourage the distribution of a ray’s termination to match the 3D keypoint, incorporating reprojection error as an uncertainty measure. This is a significantly stronger signal than reconstructing only RGB. Without depth supervision, NeRF is implicitly solving a 3D correspondence problem between multiple views. However, the sparse version of this exact problem has already been solved by SFM, whose solution is given by the sparse 3D keypoints. Therefore depth supervision improves NeRF by (softly) anchoring its search over implicit correspondences with sparse explicit ones.

Our experiments show that this simple idea translates to massive improvements in training NeRFs and its variations, regarding both the training speed and the amount of training data needed. We observe that depth-supervised NeRF can accelerate model training by 2-3x while producing results with the same quality. For sparse view settings, experiments show that our method synthesizes better results compared to the original NeRF and recent sparse-views NeRF models on both NeRF Real and Redwood-3dscan We also show that our depth supervision loss works well with depth derived from other sources such as a depth camera. Our code and more results are available at https://www.cs.cmu.edu/~dsnerf/.

Related Work

NeRF from few views. NeRF was originally shown to work on a large number of images with the LLFF NeRF Real dataset consisting of nearly 50 images per scene. This is because fitting the NeRF volume often requires a large number of views to avoid arriving at degenerate representations. Recent works have sought to decrease the data-hungriness of NeRF in a variety of different ways. PixelNeRF and metaNeRF use data-driven priors recovered from a domain of training scenes to fill in missing information from test scenes. Such an approach works well when given sufficient training scenes and limited gap between the training and test distribution, but such assumptions are not particularly flexible. Another approach is to leverage priors recovered from a different task like semantic consistency or depth prediction .

Similar to our insight that the primary difficulty in fitting few view NeRF is correctly modeling 3D geometry, MVSNeRF combines both 3D knowledge with scene priors by constructing a plane sweep volume before using a pretrained network with generalizable priors to render scenes. One appeal of an approach that utilizes 3D information is the lack of assumption it makes on the problem statement. Unlike the aforementioned approaches which depend on the availability of training data or the applicability of prior assumptions, our approach only requires the existence of 3D keypoints. This gives depth supervision the flexibility to be used not only as a standalone method, but one that can be freely incorporated into existing NeRF methods easily.

Faster NeRF. Another drawback of NeRF is the lengthy optimization time required to fit the volumetric representation. Indeed Mildenhall et al. trained a single scene’s NeRF model for twelve hours of GPU compute. Many works have found that the limiting factor is not learning the radiance itself, but rather oversampling the empty space during training. Indeed this is a similar intuition to the fact that the majority of the volume is actually empty, but NeRF’s initialization is a median uniform density. Our insight is to apply a supervisory signal directly to the NeRF density to increase the convergence of the geometry and to encourage NeRF’s density function to mimic the behavior of real world surface geometries.

Depth and NeRF. Several prior works have explored ways to leverage depth information for view synthesis and NeRF training . For instance, 3D keypoints have been demonstrated to be helpful when extending NeRFs with relaxed assumptions like deformable surfaces or dynamic scene flows . Other works like DONeRF proposed training a depth oracle to improve rendering speed by directly smartly sampling the surface of a NeRF density function. Similar to DONeRF, NerfingMVS shows how a monocular depth network can be used to induce depth priors to do smarter sampling during training and inference.

Our work attempts to improve NeRF-based methods by directly supervising the NeRF density function. As depth becomes a more accessible source of data, being able to apply depth supervision becomes increasingly more powerful. For example, recent works have demonstrated how depth extracted from sensors like time-of-flight cameras or RGB-D Kinect sensor can be applied to fit implicit functions. Building upon their insights, we provide a probabilistic formulation of the depth supervision, and show this results in meaningful improvements to NeRF and its variants.

Depth-Supervised Ray Termination

We now present our proposed depth-supervised loss for training NeRFs. We first revisit volumetric rendering and then analyze the termination distribution for rays. We conclude with our depth-supervised distribution loss.

To render a 2D image given a pose P\mathbf{P}, we cast rays r\mathbf{r} originating from the P\mathbf{P}’s center of projection o\mathbf{o} in direction d\mathbf{d} derived from its intrinsics. We integrate the implicit radiance field along this ray to compute the incoming radiance from any object that lies along d\mathbf{d}:

where tt parameterizes the aforementioned ray as r(t)=o+td\mathbf{r}(t)=\mathbf{o}+t\mathbf{d} and T(t)=exp⁡(−∫0tσ(s)ds)T(t)=\exp(-\int_{0}^{t}\sigma(s)ds) checks for occlusions by integrating the differential density between to tt. Because the density and radiance are the outputs of neural networks, NeRF methods approximate this integral using a sampling-based Riemann sum instead. The final NeRF rendering loss is given by a reconstruction loss over colors returned from rendering the set of rays R(P)\mathcal{R}(\mathbf{P}) produced by a particular camera parameter P\mathbf{P}.

Ray distribution. Let us write h(t)=T(t)σ(t)h(t)=T(t)\sigma(t). In the appendix, we show that it is a continuous probability distribution over ray distance tt that describes the likelihood of a ray terminating at tt. Due to practical constraints, NeRFs assume that the scene lies between a near and far bound (tnt_{n}, tft_{f}). To ensure h(t)h(t) sums to one, NeRF implementations often treat tft_{f} as an opaque wall. With this definition, the rendered color can be written as an expectation:

Idealized distribution. The distribution h(t)h(t) describes the weighed contribution of sampled radiances along a ray to the final rendered value. Most scenes consist of empty spaces and opaque surfaces that restrict the weighted contribution to stem from the closest surface. This implies that the ideal ray distribution of image point with a closest-surface depth of D\mathbf{D} should be δ(t−D)\delta(t-\mathbf{D}). Figure 3(c) shows that the empirical variance of NeRF termination distributions decreases with more training views, suggesting that high quality NeRFs (trained with many views) tend to have ray distributions that approach the δ\delta-function. This insight motivates our depth-supervised ray termination loss.

2 Deriving depth-supervision

Ray distribution loss. The above equivalence (see our appendix for proof) allows the termination distribution h(t)h(t) to be trained with probabilisitic COLMAP depth supervision:

Experiments

We first evaluate the input data efficiency on view synthesis over several datasets in Section 4.3. For relevant NeRF-related methods, we also evaluate the error of rendered depth maps in Section 4.4. Finally, we analyze training speed improvements in Section 4.5.

DTU MVS Dataset (DTU) captures various objects from multiple viewpoints. Following Yu et al.’s setup in PixelNeRF , we evaluated on the same test scenes and views. For each scene, we used their subsets of size 3, 6, 9 training views. We run COLMAP with the ground truth calibrated camera poses to get keypoints. Images are down-sampled to a resolution of 400×300400\times 300 for training and evaluation.

NeRF Real-world Data (NeRF Real) contains 8 real world scenes captured from many forward-facing views. We create subsets of training images for each scene of sizes 2, 5, and 10 views. For every subset, we run COLMAP over its training images to estimate cameras and collect sparse keypoints for depth supervision.

Redwood-3dscan (Redwood) contains RGB-D videos of various objects. We select 5 RGB-D sequences and create subsets of 2, 5, and 10 training frames for each object. We run COLMAP to get their camera poses and sparse point clouds. To connect the scale of COLMAP’s pose with the scanned depth, we solve a least-squares that best fits detected keypoints to the scanned depth value. Please refer to our appendix for full details.

2 Comparisons

First we consider Local Lightfield Fusion (LLFF) , an MPI-based representation that learns from multiple view points. Next we consider a set of NeRF baselines.

PixelNeRF expands upon NeRF by using an encoder to train a general model across multiple scenes. pixelNeRF-DTU is evaluated using the released DTU checkpoint. For cases where the train and test domain are different, we finetune using RGB supervision for additional iterations on each test scene to get pixelNeRF finetuned.

IBRNet extends NeRF by using a MLP and ray transformer to estimate radiance and volume density.

MVSNeRF initializes a plane sweep volume from 3 views before converting it to a NeRF by a pretrained network. MVSNeRF can be further optimized using RGB supervision.

DS-NeRF (Ours). To illustrate the effectiveness of KL divergence, we include a variant of DS-NeRF with an MSE loss between the SFM-estimated and the rendered depth. Figure 6 qualitatively shows that KL divergence penalty produces views with less artifacts on NeRF Real sequences.

DS with existing methods. As our DS loss does not require additional annotation or assumptions, our loss can be inserted into many NeRF-based methods. Here, we also incorporate our loss when finetuning pixelNeRF and IBRNet.

3 Few-input view synthesis

We start by comparing each method on rendering test views from few inputs. For view synthesis, we report three metrics (PSNR, SSIM , and LPIPS ) that evaluate the quality of rendered views against a ground truth.

DTU. We show evaluations on DTU in Table 2 and qualitative results in Figure 4. We find that DS-NeRF renders images from 6 and 9 input views that are competitive with pixelNeRF-DTU, however metaNeRF-DTU and pixelNeRF-DTU are able to outperform DS-NeRF on 3-views. This is not particularly surprising as both methods are trained on DTU scenes and therefore can fully leverage dataset priors.

NeRF Real. As seen in Table 1, our approach renders images with better scores than than NeRF and LLFF, especially when only two and fives input views are available. We also find that metaNeRF-DTU and pixelNeRF struggle which highlights their apparent weakness. These DTU-pretrained models struggle to perform well outside of DTU. Our full approach is capable of achieving good rendering results because we do not utilize assumptions on the test scene’s structure. We also add our depth supervision loss to other methods like pixelNeRF and IBRNet and find their performances improve, showing that many methods can benefit from adding depth supervision. MVSNeRF has an existing geometry prior handled by its PSV-initialization, thus we did not see an improvement from adding depth supervision.

Redwood. Like NeRF Real, we find similar improvements in performance across the Redwood dataset in Table 3. Because Redwood includes depth measurements collected with a sensor, we also consider how alternative sources of depth supervision can improve results. We train DS-NeRF, replacing COLMAP supervision with the scaled Redwood depth measurements and find that the denser depth helps even more, achieving a PSNR of 20.3 on 2-views.

4 Depth error

We evaluate NeRF’s rendered depth by comparing them to reference “ground truth” depth measurements. For NeRF Real, we use reference depth of test keypoints recovered from running an all-view dense stereo reconstruction. For Redwood , we align their released 3D models with our cameras by running 3dMatch and generate reference depths for each test view. Please refer to our arXiv version for more details regarding depth error evaluation. As shown in Table 4, DS-NeRF, trained with supervision obtained only from depth in training views, is able to estimate depth more accurately than all the other NeRF models. While this is not particularly surprising, it does highlight the weakness of training NeRFs only using RGB supervision. For example, in Figure 5, NeRF tends to ignore geometry and fails to produce any coherent depth map.

RBG-D inputs. We consider a variant of depth supervision using RGB-D input from Redwood. We derive dense depth map for each training view using 3DMatch with RGB-D input. With dense depth supervision, we can render rays for any pixel in the valid region, and apply our KL depth-supervision loss. As shown in Table 3 and Table 4, dense depth supervision produces even better-quality images and significantly lower depth errors.

5 Analysis

Overfitting. Figure 2 shows that NeRF can overfit to a small number of input views by learning degenerate 3D geometries. Adding depth supervision can assist NeRF to disambiguate geometry and render better novel views.

Faster Training. To quantify speed improvements in NeRF training, we compare training DS-NeRF and NeRF under identical settings. Like in Section 4.3, we evaluate view synthesis quality on test views under various number of input views from NeRF Real using PSNR. We can compare training speed performance by plotting PSNR on test views versus training iterations in Figure 7.

DS-NeRF achieves a particular test PSNR threshold using 2-3x less training iterations than NeRF. These benefits are significantly magnified when given fewer views. In the extreme case of only 2-views, NeRF is completely unable to match DS-NeRF’s performance. While these results are given in terms of training iteration, we can translate them into wall time improvements. On a single RTX A5000, a training loop of DS-NeRF takes ∼\sim 362.4 ms/iter while NeRF needs ∼\sim 359.8 ms/iter. Thus in the 5-view case, DS-NeRF achieves NeRF’s peak test PSNR around 13 hours faster, a massive improvement considering the negligible cost.

Discussion. We introduce Depth-supervised NeRF, a model for learning neural radiance fields that takes advantage of depth supervision. Our model uses “free” supervision provided by sparse 3D point clouds computed during standard SFM pre-processing steps. This additional supervision has a significant impact; DS-NeRF trains 2-3x faster and produces better results from fewer training views (improving PSNR from 13.5 to 20.2). While recent research has sought to improve NeRF by exploiting priors learned from category-specific training data, our approach requires no training and thus generalizes (in principle) to any scenes on which SFM succeeds. This allows us to integrate depth supervision to many NeRF-based methods and observe significant benefits. Finally, we provide cursory experiments that explore alternate forms of depth supervision such as active depth sensors. Please see our arXiv version for a discussion on limitations and societal impact of our paper.

Acknowledgments. We thank Takuya Narihira, Akio Hayakawa, Sheng-Yu Wang, Richard Tucker, Konstantinos Rematas, and Michaelu Zollhöfer for helpful discussion. We are grateful for the support from Sony Corporation, Singapore DSTA, and the CMU Argo AI Center for Autonomous Vehicle Research.

References

Appendix

We provide additional implementation details, discussions, and experimental results.

Appendix A Discussion

Limitations. Depth supervision is only as good as the estimates of depth, as such poor SfM or bad depth measurements can result in failure of the optimization process. Next we assume a Gaussian distribution models the uncertainty of the keypoint’s location, but such a simplifying assumption is not necessarily true especially for depth derived from other sources.

Societal Impact. Depth supervision is a technique which empowers NeRF to operate on a wider range of experimental setups. While novel view synthesis is not synthetic media, it can open the door to abuse when generating trajectories through a scene. There may also be privacy concerns as using ubiquitous sensor technology to better render sharper details of a scene could capture personally identifiable information.

Appendix B Derivation Details

In Section 3.1, we claim that h(t)h(t), which is a function that describes a contribution weight from a particular distance tt, is a continuous probability distribution over ray termination. We can verify this by proving that h(t)h(t) is non-negative and the integral of h(t)h(t) over tt’s domain is equal to 1.

We start by assuming that σ(s)\sigma(s) is a real-valued, non-negative (typically a ReLU or softplus activation) function describing the differential density ss units away from the camera origin. Because σ(⋅)\sigma(\cdot) is real-valued, the inner-value integral over σ(s)\sigma(s) must also be real, therefore T(t)T(t) must be non-negative. As a result, h(t)h(t) is the product of two non-negative functions and is therefore non-negative for all values of tt, satisfying the first property of a probability distribution. The next step is to show that the integral of h(t)h(t) over the domain of tt is 1.

To do this we need to make an additional assumption that for every ray cast in a scene will eventually intersect an opaque object: ∀a≥0 ∫a∞σ(s)ds=∞\forall a\geq 0\,\int_{a}^{\infty}\sigma(s)ds=\infty. This is true for most scenes we care about modeling as radiance is emitted by surfaces.

Let u(t)=−∫0tσ(s)dsu(t)=-\int_{0}^{t}\sigma(s)ds. We can compute the derivative du(t)dt=−σ(t)\frac{du(t)}{dt}=-\sigma(t) and rewrite the above equation in terms of u(t)u(t) and du(t)dt\frac{du(t)}{dt}.

This shows that under the above assumptions, h(t)h(t) is guaranteed to be a probability distribution. Due to practical constraints, NeRFs are unable to sample the volume to infinity and instead assume that the scene lies between a near and far bound. To ensure the above assumption still holds true, NeRF implementations will often treat the furthest radiance as an opaque wall.

B.2 Depth-supervision implementation

Depth supervision is implemented by projecting a ray with direction (in local camera coordinates) given by the image coordinates of a detected keypoint and −1-1 in the camera axis. We shoot this ray into a scene and render its depth using the same sampling procedure described in NeRF.

For setups where the training data can fit into GPU memory, the most time-consuming part during training comes from the many forward passes required for a single ray marching rendering step. To gain the benefits of faster training, we must simultaneously train with color supervision and depth supervision with a single ray marching procedure.

To do this, we exploit the fact that image coordinates are actually continuous and the pixels are only samples of the color function at discrete intervals. Therefore we can interpolate RGB supervision for rays corresponding to detected keypoints, allowing us to supervise RGB and depth at the same time. We allocate a portion of the training rays to this.

B.3 COLMAP details

We run the COLMAP with the default configuration on the limited views (the same as NeRF training inputs) (e.g., 2 views). The SfM output is used only during training and is not required for synthesizing novel view.

B.4 metaNeRF and pixelNeRF baselines

For adapting metaNeRF across different domains, we have to deal with different ray bounds and coordinate scales. This is challenging, and we do not directly address this issue as devising approaches to transferring NeRF meta-initialization across different datasets is beyond the scope of this paper. Instead, we used the default scaling and ray bounds provided by previous NeRF works for DTU and NeRF Real .

pixelNeRF-DTU with finetuning and DS. This variation is implemented similarly as pixelNeRF-DTU with finetuning, but the pixelNeRF finetuning stage incorporates our depth supervision loss.

B.5 Dataset Splits

NeRF Real-world. We split each scene from NeRF Real into training views and test views. Because the number of views of each scene varies, we split every eighth image id into the test set (0, 8, 16, 24, …) and construct training views that are evenly distributed over the remaining viewpoint id numbers. This setting gives us sufficient coverage to train different NeRF experiments. We create subsets from these training views of specific sizes to evaluate performance on few-input view synthesis.

Redwood 3d-scan. We selected five test scene from the Redwood 3dscan dataset: table, plant, chair, car, and stool.

Each scene is constructed using 15 frames of RGB and depth. These 15 frames are further sub-divided into training views and test views. We construct training sets with the following viewpoints , truncating when working with a smaller subset (e.g. 2-views uses only 5 and 11). We evaluate the quality of view synthesis on the test views .

B.6 Depth Error Evaluation

To evaluate the depth error of these different baselines, we need to first compute a reference depth of an input scene from test camera poses.

Depth evaluation on NeRF Real. For a given set of training views from a scene, we use COLMAP’s SfM algorithm to get sparse keypoints and camera poses. In addition, we run dense MVS on all training and test views to get a reference depth map from every view and test poses. To align dense MVS depth with the SfM keypoint depth obtained from the training views, we compute a scale aa and shift bb scalar that aligns the keypoint depth visible in a training view to its MVS depth. More specifically, we solve the following least squares optimization on those detected keypoints:

Depth error can be computed by transforming the rendered depth D^\hat{D} from a test camera cc.

Note that dense MVS depth is only used for depth evaluation. It is not used during training and test by any method.

Appendix C Additional Experiments

While pixelNeRF and metaNeRF can achieve reasonable results when given sufficient pre-training data and a small train-test domain gap, many real-world applications cannot rely on the assumptions of similar test domain and sufficient training samples. We further highlight this by showing what happens to these baselines in such a scenario. To construct a more realistic setting for NeRF-based applications, we perform a 4-fold evaluation by splitting NeRF Real into 66 training scenes and 22 test scenes. The 4 test-splits are: fern-horn, flower-fortress, leaves-orchids, and room-trex.

For every split, we use the remaining 66 training scenes to learn NeRF Real priors for pixelNeRF and metaNeRF and evaluate the view synthesis results after fine-tuning the model on the test scenes. We show these baseline results in Table 5 and find that meta-learning NeRF baselines struggle to properly leverage the priors observed during training.

C.2 MPI-based Experiments

We consider Local Light Field Fusion (LLFF) , an MPI-based method, as a baseline on their dataset (referred to as NeRF Real). We have included qualitative comparisons on NeRF Real with LLFF in Figure 8. For quantitative comparison refer to Tab. 1 in the main paper.

We find that while multi-plane images are a powerful representation, they have significant restrictions such as assuming the renderable scene lies within a central frustrum and that depths from different views are related by a homography. As such applying LLFF to scenes like DTU and Redwood is impractical given the complex camera viewpoints compared to NeRF Real.

C.3 Sparseness of keypoints

We report the performance with varying sparseness of keypoint supervision. We ablate DS-NeRF by uniformly removing detected keypoints while fixing the input to 5 views. When given (20%, 50%, and 100%) of possible keypoints, DSNeRF’s PSNR on test view is (21.5, 22.2, 22.6) respectively. Dropping out keypoints reasonably weakens the performance while still outperforming NeRF baseline (18.2).

We additionally consider how the quantity of keypoints degrade as fewer views are provided to COLMAP specifically. On average, there are 1615, 2172, 2621 keypoints for depth supervision per training view when given 2, 5, 10 views respectively. We also show qualitative examples in Figure 9 for different number of input views.