SparseNeRF: Distilling Depth Ranking for Few-shot Novel View Synthesis

Guangcong Wang, Zhaoxi Chen, Chen Change Loy, Ziwei Liu

Introduction

Neural radiance fields (NeRFs) have made tremendous progress in generating photo-realistic novel views of scenes by optimizing implicit function representations given a set of 2D input views. However, in a wide range of real-world scenarios, collecting dense views of a scene is often expensive and time-consuming . Therefore, it is necessary to develop few-shot NeRF methods that can be learned from sparse views without significant degradation in performance.

Learning a NeRF from sparse views is a challenging problem due to under-constrained reconstruction conditions, especially in textureless areas. Directly applying NeRFs to few-shot scenarios suffers from dramatic degradation . Recently, some methods have greatly improved the performance of few-shot NeRF, which can be categorized into three groups. 1) The first group is based on geometry constraints (sparsity and continuity regularizations) and high-level semantics. RegNeRF regularized the geometry and appearance of patches rendered from unobserved viewpoints, and annealed the ray sampling space. InfoNeRF imposed an entropy constraint of the density in each ray and a spatial smoothness constraint into the estimated images. However, since a scene often contains multiple layouts (Figure 1), sparsity and continuity geometric constraints of a few views cannot guarantee the complete 3D geometric reconstruction. 2) The second group resorts to pre-training on similar scenes. For example, PixelNeRF proposed to condition a NeRF on convolutional feature maps to learn high-level semantics from other scenes. 3) The third group exploits depth maps and makes a linearity assumption of depth maps to supervise few-shot NeRFs. For example, DSNeRF exploited sparse 3D points generated by COLMAP or accurate depth maps obtained by high-accuracy depth scanners and the Multi-View Stereo (MVS) algorithm. The depth maps are linearly scaled as supervision to guide the predicted depth of few-shot NeRFs. To use coarse depth maps, MonoSDF uses a local patch-based scale-invariant depth constraint supervised by coarse depth maps instead of global depth maps. However, the scale-invariant depth constraint is strong for real-world coarse depth maps from pre-trained depth models or consumer-level depth sensors due to wide-range depth distances in the wild.

Along the third group, we wish to explore more robust 3D priors from coarse depth maps to complement the under-constrained few-shot NeRF. To address this problem, we present SparseNeRF, a simple yet effective method that distills depth priors from pre-trained depth models or inaccurate depth maps from consumer-level depth sensors (Figure 3), which can be easily obtained from real-world scenes. Deriving useful depth cues from such pre-trained models is non-trivial. In particular, although single-view depth estimation methods have achieved good visual performance, thanks to large-scale monocular depth datasets and large ViT models, they cannot yield accurate 3D depth information due to coarse depth annotations, dataset bias, and ill-posed 2D single-view images. The inaccurate depth information contradicts the density prediction of a NeRF when reconstructing each pixel of a 3D scene based on volume rendering. Directly scaling the coarse depth maps to a NeRF leads to inconsistent geometry against the expected depth of the NeRF.

Instead of directly supervising a NeRF with coarse depth priors, we relax hard depth constraints and distill robust local depth ranking from the coarse depth maps to a NeRF such that the depth ranking of a NeRF is consistent with that of coarse depth. That is, we supervise a NeRF with relative depth instead of absolute depth . To guarantee the spatial continuity of geometry, we further propose a spatial continuity constraint on depth maps such that the NeRF model imitates the spatial continuity of coarse depth maps. The accurate sparse geometry constraints from a limited number of views, combined with relaxed constraints including depth ranking regularization and continuity regularization, finally achieve promising novel view synthesis (Figure 1). It is noteworthy that SparseNeRF does not increase the running time during inference as it only exploits depth priors from pre-trained depth models or consumer-level sensors during the training stage (Figure 2). In addition, SparseNeRF is a plug-and-play module that can be easily integrated into various NeRFs.

The main contribution of this paper is 1) SparseNeRF, a simple yet effective method that distills local depth ranking priors from pre-trained depth models. With the help of the local depth ranking constraint, SparseNeRF significantly improves the performance of few-shot novel view synthesis over the state-of-the-art models (including depth-based NeRF methods). To preserve the coherent geometry of a scene, we propose a spatial continuity distillation constraint that encourages the spatial continuity of NeRF to be similar to that of the pre-trained depth model. Both depth ranking prior and spatial continuity distillation are new in the literature on NeRF. 2) Apart from SparseNeRF, we also contribute a new dataset, NVS-RGBD, which contains coarse depth maps from Azure Kinect, ZED 2, and iPhone 13 Pro. 3) Extensive experiments on the LLFF, DTU, and NVS-RGBD datasets demonstrate that SparseNeRF achieves a new state-of-the-art performance in few-shot novel view synthesis.

Related Work

Neural Radiance Fields. NeRF has made great success in synthesizing novel views of complex scenes due to good representation of neural networks . Block-NeRF , CityNeRF , and Mega-NeRF scaled the standard NeRF up to city-scale or urban-scale scenes. NeRF−−-- , GNeRF , and BARF relaxed the requirements of NeRFs and synthesized novel views without perfect camera poses. Some works attempted to improve NeRFs by considering anti-aliasing , sparse 3D grids with spherical harmonics . Some methods extended NeRFs to dynamic scenes. These methods do not focus on generating novel views with a few views. In this paper, we study sparse-view NeRF to reduce dense capture requirements in the real-world applications.

Few-shot Novel View Synthesis. There are increasing studies on few-shot novel view synthesis. Basically, the existing few-shot NeRF methods can be categorized into three groups. First, some methods exploit continuity constraints on geometry or object semantics. For example, RegNeRF imposed a continuity constraint on geometry and regularized the appearance of patches from unobserved viewpoints with a flow model. InfoNeRF proposed a ray entropy minimization regularization to encourage the density to be as sparse as possible along a ray, and used a ray information gain reduction regularization to constrain the continuous depth of neighbor rays. Second, some methods attempt to pe-train a NeRF on other similar scenes and fine-tune the NeRF on the target scene. For example, PixelNeRF conditioned a NeRF on image inputs in a fully convolutional manner, allowing the model to learn scene priors from other scenes and reduce the requirement of dense views. MVSNeRF leveraged plane-swept cost volumes for geometry-aware scene reasoning, and combined it with physically based volume rendering. Similar to PixelNeRF, MVSNeRF was first trained on other real scenes and was finetuned on target scenes to evaluate its effectiveness and generalizability. Third, depth-based models use available depth information to supervise the training of NeRFs. DSNeRF exploited sparse 3D points generated by COLMAP or accurate depth maps obtained by high-accuracy depth scanners. Different from these depth-based models, SparseNeRF distills robust depth ranking from pre-trained depth models or coarse depth maps from consumer-level sensors. As concurrent works, NeRDi and NeuralLift-360 also use ranking-based methods. However, they mainly focus on single-view setting while SparseNeRF focuses on sparse-view. SparseNeRF uses ranking loss to avoid inconsistent 3D geometry across different views while focuses on a soft geometric regularization on a single view (no cross-view consistency problem during training). Moreover, SparseNeRF introduces a new spatial continuity loss to distill spatial coherence from monocular depth estimators. In addition, some single-view synthesis methods either allow new generative objects from unseen views or focus on specific objects, e.g., face and human . Some NeRF-based methods extended to generation instead of reconstruction. Other AIGC models focus on the 2D generation or category-specific reconstruction , which are different.

Our Approach

We present SparseNeRF to synthesize novel views given sparse view inputs. Single-view depth estimation is a long-standing computer vision task, aiming to predict a depth map given a single image. In this paper, we are interested in mining the depth priors encapsulated in pre-trained models of single-view depth estimation or coarse depth maps captured by consumer-level depth sensors. However, due to coarse annotations of single-view depth maps, (e.g., user clicks , RGB-D , and laser/stereo ), dataset bias, and imperfect depth estimation models, it is challenging to obtain accurate 3D depth estimation given 2D single-view images. As for consumer-level depth sensors, it still struggles to capture accurate depth maps (Figure 3). Driven by these observations, our goal is to make use of the coarse depth maps and distill useful depth priors to guide the learning of a NeRF.

where C^(r)\hat{C}(\mathbf{r}) are rendered color blended by NN samples. C(r)C(\mathbf{r}) is the ground-truth pixel color. We use coarse-to-fine sampling as discussed in the vanilla NeRF . Here we omit the fine rendering for simplification.

Problem Formulation. Vanilla NeRFs aim to learn a mapping function fθ{f_{\theta}} with KdK_{d} dense views by optimizing a color reconstruction loss in Eq. (1). In this work, we aim to study few-shot NeRF when only KsK_{s} sparse views are available (Ks≪KdK_{s}\ll K_{d}). Let HH and WW denote the height and width of an image, we have HWKsHWK_{s} rays as color reconstruction constraints. Without considering view directions d\mathbf{d} and image continuity assumptions (i.e., ci∗\mathbf{c}_{i}^{*} is independent of cj∗\mathbf{c}_{j}^{*} when i≠ji\neq j), suppose we have a Nc×Nc×NcN_{c}\times N_{c}\times N_{c} discrete cubic volume to be optimized, which contains Nc3N_{c}^{3} weighted color variables ci∗\mathbf{c}_{i}^{*}. We have Nc3N_{c}^{3} variables ci∗\mathbf{c}_{i}^{*} and HWKsHWK_{s} constraints. Ideally, we are able to solve a discrete cube of edge length Nc=HWKsN_{c}=\sqrt{HWK_{s}}. That is, we can sample HWKs\sqrt{HWK_{s}} for each edge of a cubic volume. Take H=W=512H=W=512 and Ks=3K_{s}=3 as an example, we have Nc≈92N_{c}\approx 92, which is far from reconstructing a continuous radiance field. To address this under-constrained optimization problem of few-shot NeRF, an intuitive way is to introduce reasonable regularization terms to constrain few-shot NeRFs conditioned on input x\mathbf{x}, d\mathbf{d}, and C(r)C(\mathbf{r}). Considering these constraints, a general formulation is given by

where R\mathcal{R} is a regularization term.

Remark. RegNeRF and InfoNeRF tackles this problem by introducing regularization terms on continuous depth constraint from unobserved viewpoints, sparsity of density on rays, and patch-based semantic constraint on color appearance. Depth-based models, such as DSNeRF and MonoSDF , directly regress depth by leveraging sparse 3D points or deriving ground-truth depth maps with accurate absolute depth maps. Different from them, we propose to use a robust relative depth regularization from coarse depth maps of pre-trained depth models or consumer-level depth sensors.

2 Overview of SparseNeRF

3 Local Depth Ranking Distillation

Single-view depth estimation is a challenging computer vision task, which aims at predicting a depth map of a scene given a single image as input. Due to dataset bias, coarse depth annotations, and imperfect neural models, it is difficult to achieve accurate depth prediction. Depth maps captured by consumer-level sensors are also inaccurate due to wide-range depth distances. Directly using coarse depth maps to supervise a NeRF leads to ambiguous rendered novel view synthesis. To avoid the error of coarse depth maps, we relax the depth constraint and exploit the robust depth ranking prior. Given a pair of pixels in a single image, the depth ranking regularization only considers which point is nearer or farther. However, when the scene is complex, even depth ranking is not accurate. As shown in Figure 4, it is easy for depth estimation models to compare the depth ranking of white and cyan points, but it is hard to estimate the depth ranking of white and red points. It implies that the depth ranking becomes unreliable as the spatial distance increases.

4 Spatial Continuity Distillation

The local depth ranking constraint guarantees that the predicted depth map estimated by NeRF is consistent with the DPT’s depth map. However, it does not constrain the spatial continuity of the depth map. We distill spatial continuity priors from the depth model DPT, which allows large displacement across several depth pixels. If neighbor depth pixels are continuous on the depth map of DPT, we constrain the corresponding depth pixels of NeRF to be continuous. The spatial continuity regularization is given by

Full Objective Loss. The full objective loss of SparseNeRF is formulated by

where λ\lambda and γ\gamma control the importance of the regularization terms. In practice, we set λ=0.2\lambda=0.2, γ=0.02\gamma=0.02, m=m′=1.0×10−4m=m^{{}^{\prime}}=1.0\times 10^{-4}.

NVS-RGBD Dataset

The popular benchmark LLFF on NeRFs is featured with forward-facing and object-centric. Moreover, we focus on consumer-level depth sensors instead of expensive scanners . We study stand-alone depth cameras instead of a complex calibrated structured multiple camera system , and collect from real-world scenes instead of synthesized 3D assets . Few RGBD datasets satisfy all of the requirements. Therefore, we collect a new dataset NVS-RGBD that contains real-world depth maps captured by consumer-level depth sensors, including Azure Kinect, ZED 2, and iPhone 13 Pro. We collect 8 scenes for both Azure Kinect and ZED 2, and 4 scenes for iPhone 13 Pro. The depth maps of Azure Kinect often contain noises around object edges. The depth maps of ZED 2 look smooth but unstable and incorrect when observing time jittering (Figure 3). The poses are computed via COLMAP. We collect indoor scenes for ZED 2 and Kinect. For iPhone, we collect both indoor and outdoor scenes. All the scenes are object-centric. Depth maps obtained from sensors have different artifacts coming from the sensor noises. Please refer to the supplementary material for the details.

Experiments

Datasets. We conduct experiments on the LLFF , DTU , and NVS-RGBD datasets. LLFF contains 8 complex forward-facing scenes. Following RegNeRF , we use every 8-th image as the held-out test set, and evenly sample sparse views from the remaining images for training. Different from LLFF, DTU is an object-level dataset. Following PixNeRF , we use the same 15 scenes in our experiments. On the DTU dataset, the background of these scenes is a white table or a black background, which has few textures. As mentioned in RegNeRF, we mask the background during inference to avoid background bias. We use the same evaluation protocol for a fair comparison. On the NVS-RGBD dataset, for each scene, we use 3 views for training and the rest views for inference. For LLFF and DTU, we use pre-trained depth maps to implement our method. For NVS-RGBD, we use coarse depth maps captured by depth sensors.

Evaluation Metrics. We adopt four evaluation metrics in the experiments, i.e., PSNR, SSIM , LPIPS , and depth error . Depth error is a scale-invariant MSE derived by minimizing ∣∣wd^+b−d∣∣||w\hat{d}+b-d||, where d^\hat{d} and dd denote coarse depth maps and pseudo-ground-truth depth maps predicted by dense-view NeRF. Note that we only use dense views to train a good NeRF to generate pseudo-ground-truth depth maps for evaluation.

Implementation Details. We implement SparseNeRF based on the official JAX . We use the Adam optimizer for the learning of SparseNeRF. We use an exponential learning rate, which decays from 2×10−32\times 10^{-3} to 2×10−52\times 10^{-5}. The batch size is set to 4096. For each scene, we use one GPU-v100 with 32G memory for both training and inference. We also implement the proposed SparseNeRF on GPU-v100 with 16G and RTX 2080 Ti with 11G. We reduce the batch size to 2048 and 1024 for these two types of GPUs, respectively. We use the same backbone and sampled points along rays compared with prior arts.

We train 90k iterations for each scene given three training views, which takes about 2 hours for each scene. For depth maps captured by Kinect and ZED 2, we mask uncertain regions (black regions in depth maps). More precisely, we sample near-far depth pairs in certain regions in local patches to optimize Eqs. 3 and 4 of the paper.

We compare SparseNeRF with state-of-the-art methods on the LLFF dataset, including SRF , PixelNeRF , MVSNeRF , Mip-NeRF , DietNeRF , RegNeRF , DSNeRF and MonoSDF. Among them, SRF, PixelNeRF, and MVSNeRF are pre-trained on other similar scenes to exploit high-level semantics. Mip-NeRF is the state-of-the-art NeRF designed for dense-view training. InfoNeRF and RegNeRF mainly constrain sparsity and continuity of geometry or natural appearance semantics. More precisely, the compared methods can be classified into four groups. In the first group, SRF, PixelNeRF, and MVSNeRF are pre-trained in the large-scale DTU datasets (88 scenes) and are directly tested on the LLFF dataset. In the second group, SRF, PixelNeRF, and MVSNeRF are pre-trained on DTU and fine-tuned on LLFF per scene. The third group includes Mip-NeRF, DietNeRF, and RegNeRF. They uses geometry continuity and semantic constraints. The fourth group distills knowledge from depth maps or sparse points of COLMAP (DSNeRF). The compared results are shown in Table 1. We can see that the proposed SparseNeRF achieves the best performance in PSNR, SSIM, and LPIPS. Note that the first and second groups have to pre-train on lots of other scenes. SparseNeRF achieves better results than DSNeRF and MonoSDF because the depth ranking is more robust than the absolute scale-invariant constraint.

We provide qualitative analysis on LLFF, as shown in Figure 5. It is observed that PixelNeRF and MVSNeRF tend to generate ambiguous pixels. The reason is that they integrate CNN features into NeRFs. CNN features can provide high-level semantics to infer under-constrained textures but suffer pixel-level inference. MipNeRF is designed for dense-view synthesis without considering the under-constrained problem, leading to degraded geometry artifacts. Compared with RegNeRF, SparseNeRF uses robust depth priors from pre-trained depth models, which can handle challenging scene geometry with better performance.

2 Comparisons on DTU

Similar to the LLFF dataset, we conduct four-group comparisons on the DTU dataset, as shown in Table 2. In this first group, given a target scene, SRF, PixelNeRF, and MVSNeRF are pre-trained on other DTU scenes and are tested on the target scene. In the second group, these three methods are further fine-tuned on the target scene. In the third group, we compare Mip-NeRF, DietNeRF, and RegNeRF. Note that for the pre-trained models, they split DTU as a training set (88 scenes) and a test set (15 scenes). Since the training set and the test set contains similar scenes, a test scene shares a similar distribution (e.g., similar backgrounds) with training scenes. Pre-training on other scenes greatly benefits a target scene. Therefore, SRF, PixelNeRF, and MVSNeRF also greatly achieve promising results. Finally, we compare depth-based DSNeRF and MonoSDF . DSNeRF studies two kinds of depth information, i.e., sparse 3D points generated by COLMAP and ground-truth depth maps obtained by high-accuracy depth scanners and the Multi-View Stereo (MVS) algorithm. The depth maps of NVS-RGBD are captured by consumer-level sensors and contain lots of noise, which is hard to be applied to DSNeRF. Therefore, we adopt sparse 3D points for DSNeRF. Compared with these methods, SparseNeRF still achieves the best performance without pre-training on other scenes. We provide a qualitative analysis, as shown in Figure 6. SparseNeRF achieves better visual results than previous state-of-the-art methods.

3 Comparisons on NVS-RGBD

We then implement three state-of-the-art methods on our new NVS-RGBD dataset, including RegNeRF , DSNeRF , and MonoSDF . All of them are optimized per scene without pre-training on the other scenes. RegNeRF adopted continuous geometry assumptions from unobserved viewpoints. DSNeRF uses a global scale-invariant depth constraint to supervise NeRFs. MonoSDF mainly focuses on implicit surface reconstruction. It is assumed that local patches of depth maps from monocular cameras satisfy d′=wd+bd^{{}^{\prime}}=wd+b where dd is predicted by the pre-trained model and d′d^{{}^{\prime}} is then computed by the NeRF. In Table 3, we observe that SparseNeRF achieves better performance compared with these prior arts. We provide a qualitative analysis in Figure 7, demonstrating the superiority of SparseNeRF over previous state-of-the-art methods.

4 Ablation Studies and Further Analyses

Effectiveness of Depth Distillation. To validate the effectiveness of the depth distillation, we isolate the depth constraint and keep other modules unchanged. As shown in the first and third rows of Table 4, it is observed that without the local depth ranking distillation, SparseNeRF degrades on PSNR, SSIM, LPIPS, and depth error, which demonstrates the effectiveness of the depth distillation. We further study the depth maps from unobserved views. The ablation study of visual effect is shown in Figure 8. It shows that depth ranking greatly improves performance in geometry. Spatial continuity regularization further improves the detailed coherence of scenes.

Further Analyses on Depth Ranking and Depth Scaling. To validate the robustness of local depth ranking, we compare depth ranking regularization and depth scaling regularization on LLFF. We use the idea from MonoSDF , which mainly focuses on implicit surface reconstruction. With the specific design for surface shapes, MonoSDF slightly degrades the performance of RGB view synthesis. We re-implement it as depth scaling. Since the single-view depth estimation is not accurate, the linear depth scaling constraint is too strong such that depth distillation misleads the learning of the NeRF. Figure 3 shows two cases of scale-invariant errors. Instead, depth ranking relaxes the constraint and only focuses on depth comparison, which is more robust than depth scaling. As shown in Table 4, depth ranking performs better than depth scaling on LLFF.

Impact of Different Pre-trained Depth Models. We test our method on three depth estimation models, i.e., MiDaS small, DPT Hybrid, and DPT Large. As shown in Table 5, all of the depth models are better than the baseline. DPT Hybrid and DPT Large achieve comparable results, which are better than MiDaS small.

FreeNeRF potentially complements SparseNeRF. FreeNeRF and SparseNeRF improve few-shot NeRF in different ways. FreeNeRF introduces a frequency regularization while SparseNeRF distills depth priors from pre-trained general depth estimators. We implement SparseNeRF on FreeNeRF and the results significantly improve, as shown in Table 6.

User Study. To show that the proposed spatial continuity regularization improves the 3D consistency in 3D space, we conducted a user study. For each scene, we rendered videos for both with and without spatial continuity regularization terms. Users are asked to select the better-quality video according to 3D consistency/coherence and completeness (good geometry, few floaters). We evaluate 8 scenes on LLFF and there are 27 participants. The result shows 88.9% of users prefer “w/ conti” and 11.1% prefer “w/o conti”. To evaluate the advantages of SparseNeRF on 3D consistency and shape completeness, we conduct another user study. Users are asked to evaluate “baseline+RegNeRF’s continuity term” and “baseline+our continuity term”. The 27 users participated in the study and 8 scenes on LLFF are evaluated. The result shows 87.5% of users prefer “baseline+our continuity term” and 12.5% prefer “baseline+RegNeRF’s continuity term”. RegNeRF encourages the overall patch to be as smooth as possible without considering the sharp edges of objects, which conflicts with the reconstruction from sparse views. Our continuity regulation distills the neighbor relationship from depth models is more accurate. We show two examples of depth maps in Figure 9.

Discussion. SparseNeRF could have a wide range of applications due to its promising results. We provide two applications on iPhone 13 Pro. Please refer to the project page for more experiments. Like UNISURF , we also evaluate the SparseNeRF using Chamfer distance, showing the competitive results with MonoSDF and UNISURF, especially for complex scenes.

Conclusion

In this paper, we propose a SparseNeRF framework that synthesizes novel views with sparse view inputs. To tackle the under-constrained few-shot NeRF problem, we propose a local depth ranking regularization that distills the depth ranking prior from coarse depth maps. To preserve the spatial continuity of geometry, we propose a spatial continuity regularization that distills the depth continuity priors. With the proposed depth priors, SparseNeRF achieves a new state-of-the-art performance on three datasets.

Limitation. SparseNeRF significantly improves the performance of few-shot NeRF, but cannot be generalized to occluded views that are unobserved in the training views.

Potential Negative Impact. SparseNeRF focuses on scene synthesis. It might be misused to create misleading content.

Acknowledgement. This study is supported under the RIE2020 Industry Alignment Fund Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner(s). It is also supported by Singapore MOE AcRF Tier 2 (MOE-T2EP20221-0012, MOE-T2EP20221-0011) and NTU NAP Grant.

References

A. Applications

Apple’s latest products (iPhone 12 Pro and higher versions, and iPad Pro now feature built-in LiDAR sensors that can capture depth maps. We use the portrait mode of iPhone 13 Pro to collect RGBD datasets. We train RegNeRF and SparseNeRF with three views (RGB images and depth maps are aligned), and rendered novel views, as shown in Figure 10. RegNeRF does not recover a correct geometry while the proposed SparseNeRF significantly improves the reconstruction of geometry. The main reason is that RegNeRF adds a geometric continuity constraint from unseen viewpoints such that neighbor depth pixels should be as close as possible. This regularization does not fully exploit the depth priors of objects/scenes. Instead, the proposed SparseNeRF distills the depth ranking and the spatial continuity of depth maps captured by an iPhone’s LiDAR, which achieves promising rendered results.

A2. Few-shot NeRF with Simple User-click Depth Annotations

The SparseNeRF significantly improves few-shot novel view synthesis with the help of coarse depth maps, obtained by pre-trained depth models and consumer-level sensors. What if the depth rankings of partial regions of depth maps are incorrect in some cases? Do we have any ideas to fix this issue? We provide an application that fixes incorrect depth rankings of pre-trained depth or sensors when we do not trust these coarse depth maps. As shown in Figure 11 (a), users can easily click near-far depth pairs in a local region, e.g., depth pairs in cyan circles. Green points are nearer while red points are farther. In Figure 11 (b) and Figure 11 (c), we train a few-shot NeRF without and with user-click depth annotations, respectively. It is observed that the performance of the few-shot NeRF significantly improves. In practice, 30∼\sim40 depth pairs of each view lead to significant improvement. In real-world scenes, one may only need to annotate partial regions for refinement. User-click depth annotations provide sparse depth rankings, which tailors for the proposed SparseNeRF. In real-world applications, occlusion often occurs in complex scenes. Even with dense views for training, some regions are observed from a limited number of views. Therefore, this application helps dense-view NeRF to refine novel view synthesis in these cases.

B. Detailed Descriptions of the new Dataset NVS-RGBD

In Section 4, we give a brief introduction to our new dataset NVS-RGBD. In this material, we provide a detailed description. Different from previous methods that derive accurate and expensive ground-truth depth maps by pre-trained NeRF trained with dense views or captured by high-accuracy depth scanners , our goal is to develop a good NeRF with cheap or even free coarse depth maps, i.e., pre-trained single-view depth estimation models and consumer-level depth sensors. To this end, we use Microsoft Azure Kinect, ZED 2, and iPhone 13 Pro to collect a new dataset NVS-RGBD. We collect 4 scenes for iPhone 13, and 8 scenes for ZED 2 and Kinect, respectively. The dataset is similar to the LLFF dataset that focuses on forward-facing scenes. For each scene, we evenly sample 3 views for training and the held-out views are for testing.

The examples of the NVS-RGBD dataset are shown in Figure 12. As for Azure Kinect, the edges of objects often contain lots of noise. Some regions are directly masked by sensors. As for ZED 2, the edges of the depth maps are inaccurate. In Figure 3, we observe that the depth maps of ZED 2 look smooth but unstable and incorrect when observing time jittering. We masked out the uncertain black regions of depth maps for training. Moreover, the scale of these depth maps is not linearly scaled to the ground-truth depth maps predicted by dense-view NeRFs. As for the depth maps of the iPhone, they are more smooth than Kinect and ZED 2. The drawback is that the distance between objects and the iPhone is up to 2.5 meters.

C. More Implementation Details

In Tables 1 and 2, and Figures 5 and 6, we reproduce most of the prior arts as the supplementary Material of RegNeRF. we implement DSNeRF as the official codebase. In Table 7, MonoSDF is mainly designed for surface reconstruction. We apply the depth consistency loss to RegNeRF for novel view synthesis of RGB images for comparison. The depth consistency loss is a scale-invariant loss constrained in local patches. In Table 3, we can see that MonoSDF obtains a lower depth error than RegNeRF, but fails to achieve better novel view synthesis of RGB images.

D. More Visual Comparisons

In Figures 5, 6, and 7 of the main body of the paper, we give two examples of visual comparisons. In this material, Figures 13, 14, and 15 show more visual examples on the LLFF, DTU, and NVS-RGBD datasets, respectively. In particular, we compare five representative methods on LLFF and DTU, respectively. Among them, MipNeRF is a state-of-the-art NeRF method designed for conventional dense-view NeRF, which is optimized by color reconstruction. Without extra geometric regularization, MipNeRF fails to synthesize good novel views. As shown in Figures 13 and 14, it is observed that positions of objects predicted by MipNeRF might be shifted due to the under-constrained problem. PixelNeRF(ft) and MVSNeRF(ft) require pre-training on extra scenes and fine-tuning on a target scene. Thanks to the pre-training, PixelNeRF(ft) and MVSNeRF(ft) greatly improve few-shot NeRFs. Compared with PixelNeRF(ft) and MVSNeRF(ft), RegNeRF and SparseNeRF do not require extra scenes for pre-training. RegNeRF optimizes few-shot NeRF by adding a continuous geometric constraint from unseen viewpoints such that the depth values of neighbor pixels should be as close as possible. RegNeRF achieves a new state-of-the-art performance but fails to guarantee the correct geometry of objects. Unlike RegNeRF, the proposed SparseNeRF exploits coarse depth maps of pre-trained depth models in the experiments. Thanks to the distillation of depth ranking and spatial continuity of pe-trained depth models, SparseNeRF significantly improves few-shot NeRFs.

In Figures 13 and 14, SparseNeRF distills depth priors from pre-trained depth models. In Figure 15, we use depth sensors to collect coarse depth maps. We compare two depth-based methods and the best few-shot NeRF method RegNeRF. Among them, DSNeRF adopts sparse 3D points derived by the COLMAP algorithm . We implement depth-supervised loss of DSNeRF upon RegNeRF. MonoSDF is originally proposed for surface reconstruction. We implement the depth consistency loss upon RegNeRF, which encourages the predicted depth to be scale-invariant to coarse depth maps. The drawback of this method is the strong assumption about the scale-invariant depth constraint. The cheap depth maps from pre-trained single-view depth models or consumer-level depth sensors are incorrect. They only provide coarse geometry information. Compared with the two depth-based NeRFs and the best few-shot NeRF RegNeRF, the proposed SparseNeRF achieves a better visual effect.

E. Visual Comparisons of Predicted Geometry

In Table 3 of the paper, we compare the scale-invariant depth errors of four few-shot NeRF methods. The results show that the SparseNeRF significantly outperforms RegNeRF, i.e., from 5.9×10−3\times 10^{-3} to 3.9×10−3\times 10^{-3} on the scenes captured by Kinect and from 3.5×10−3\times 10^{-3} to 2.4×10−3\times 10^{-3} on the scenes captured by ZED 2. DSNeRF and MonoSDF also improve RegNeRF on depth errors. The depth error indicates the quality of the reconstructed geometry.

In this material, we study the qualitative results of geometry, as shown in Figure 16. The top four rows show the comparisons on the RGBD data of Kinect. The bottom four rows are for ZED 2. Compared with RegNeRF, DSNeRF uses sparse 3D points for depth supervision. MonoSDF imposes a local scale-invariant loss computed by coarse depth maps and predicted depth of NeRFs. Compared with RegNeRF, MonoSDF, and DSNeRF, SparseNeRF yields much better geometry. The sub-optimal results of MonoSDF can be attributed to the fact that coarse depth maps are not able to be linearly scaled to ground-truth depth maps. In the next section, we will analyze the scale-invariant depth errors of coarse depth maps.