FreeNeRF: Improving Few-shot Neural Rendering with Free Frequency Regularization

Jiawei Yang, Marco Pavone, Yue Wang

Introduction

Neural Radiance Field (NeRF) has gained tremendous attention in 3D computer vision and computer graphics due to its ability to render high-fidelity novel views. However, NeRF is prone to overfitting to training views and struggles with novel view synthesis when only a few inputs are available. We term this view synthesis from sparse inputs problem as a few-shot neural rendering problem.

Existing methods address this challenge using different strategies. Transfer learning methods, e.g., PixelNerf and MVSNeRF , pre-train on large-scale curated multi-view datasets and further incorporate per-scene optimization at test time. Depth-supervised methods introduce estimated depth as an external supervisory signal, leading to a complex training pipeline. Patch-based regularization methods impose regularization from different sources on rendered patches, e.g., semantic consistency regularization , geometry regularization , and appearance regularization , all at the cost of computation overhead since an additional, non-trivial number of patches must be rendered during training .

In this work, we find that a plain NeRF can work surprisingly well with none of the above strategies in the few-shot setting by adding (approximately) as few as one line of code (see Fig. 1). Concretely, we analyze the common failure modes in training NeRF under a low-data regime. Drawing on this analysis, we propose two regularization terms. One is frequency regularization, which directly regularizes the visible frequency bands of NeRF’s inputs to stabilize the learning process and avoid catastrophic overfitting at the start of training. The other is occlusion regularization, which penalizes the near-camera density fields that cause “floaters,” another failure mode in the few-shot neural rendering problem. Combined, we call our method Frequency regularized NeRF (FreeNeRF), which is “free” in two ways. First, it is dependency-free because it requires neither costly pre-training nor extra supervisory signals . Second, it is overhead-free as it requires no additional training-time rendering for patch-based regularization .

We consider FreeNeRF a simple baseline (with minimal modifications to a plain NeRF) in the few-shot neural rendering problem, although it already outperforms existing state-of-the-art methods on multiple datasets, including Blender, DTU, and LLFF, at almost no additional computation cost. Our contributions can be summarized as follows:

We reveal the link between the failure of few-shot neural rendering and the frequency of positional encoding, which is further verified by an empirical study and addressed by our proposed method. To our knowledge, our method is the first attempt to address few-shot neural rendering from a frequency perspective.

We identify another common failure pattern in learning NeRF from sparse inputs and alleviate it with a new occlusion regularizer. This regularizer effectively improves performance and generalizes across datasets.

Combined, we introduce a simple baseline, FreeNeRF, that can be implemented with a few lines of code modification while outperforming previous state-of-the-art methods. Our method is dependency-free and overhead-free, making it a practical and efficient solution to this problem.

We hope the observations and discussions in this paper will motivate people to rethink the fundamental role of frequency in NeRF’s positional encoding.

Related Work

Neural fields use deep neural networks to represent 2D images or 3D scenes as continuous functions. The seminal work, Neural Radiance Fields (NeRF) , has been widely studied and advanced in a variety of applications , including novel view synthesis , 3D generation , deformation , video . Despite tremendous progress, NeRF still requires hundreds of input images to learn high-quality scene representations; it fails to synthesize novel views with a few input views, e.g., 3, 6, and 9 views, limiting its potential applications in the real world.

Few-shot Neural Rendering.

Many works have attempted to address the challenging few-shot neural rendering problem by leveraging extra information. For instance, external models can be used to acquire normalization-flow regularization , perceptual regularization , depth supervision , and cross-view semantic consistency . Another thread of works attempts to learn transferable models by training on a large, curated dataset instead of using an external model. Recent works argue that geometry is the most important factor in few-shot neural rendering and propose geometry regularization for better performance. However, these methods require expensive pre-training on tailored multi-view datasets or costly training-time patch rendering , introducing significant overhead in methodology, engineering implementation, and training budgets. In this work, we show that a plain NeRF can work surprisingly well with minimal modifications (a few lines of code) by incorporating our frequency regularization and occlusion regularization. Unlike most previous methods, our approach maintains the same computational efficiency as the original NeRF.

Frequency in neural representations.

Positional encoding lies at the heart of NeRF’s success . Previous studies have shown that neural networks often struggle to learn high-frequency functions from low-dimensional inputs. Encoding inputs with sinusoidal functions of different frequencies can alleviate this issue. Recent works show the benefits of gradually increasing the input frequency in different applications, such as non-rigid scene deformation , bundle adjustment , surface reconstruction , and fitting functions with a wider frequency band . Our work leverages frequency curriculum to tackle the few-shot neural rendering problem. Notably, our approach not only demonstrates the surprising effectiveness of frequency regularization in learning from sparse inputs, but also reveals the failure modes behind this problem and why frequency regularization helps.

Method

Positional encoding.

Directly optimizing NeRF over raw inputs (x,d)({\mathbf{x}},{\mathbf{d}}) often leads to difficulties in synthesizing high-frequency details . To address this issue, recent work has used sinusoidal functions with different frequencies to map the inputs into a higher-dimensional space :

where LL is a hyperparameter that controls the maximum encoded frequency and may differ for coordinates x{\mathbf{x}} and directional vectors d{\mathbf{d}}. A common practice is to concatenate the raw inputs with the frequency-encoded inputs as follows:

This concatenation is applied to both coordinate inputs and view direction inputs.

Rendering.

where c^(r;θ,tK)\hat{{\mathbf{c}}}({\mathbf{r}};\theta,{\mathbf{t}}_{K}) is the final integrated color. Note that the sampled points tK{\mathbf{t}}_{K} are in a near-to-far order, i.e., a point with a smaller index kk is closer to the camera’s origin.

2 Frequency Regularization

The most common failure mode of few-shot neural rendering is overfitting. NeRF learns 3D scene representations from a set of 2D images without explicit 3D geometry. 3D geometry is implicitly learned by optimizing appearance in its 2D projected views. However, given only a few input views, NeRF is prone to overfitting to these 2D images with small loss while not explaining 3D geometry in a multi-view consistent way. Synthesizing novel views from such models leads to systematic failure. As shown on the left of Figure 1, no NeRF model can successfully recover the scene geometry when synthesizing novel views.

The overfitting issue in few-shot neural rendering is presumably exacerbated by high-frequency inputs. shows that higher-frequency mappings enable faster convergence for high-frequency components. However, the over-fast convergence on high-frequency impedes NeRF from exploring low-frequency information and significantly biases NeRF towards undesired high-frequency artifacts (horns and room examples in Fig. 1). In the few-shot scenario, NeRF is even more sensitive to susceptible noise as there are fewer images to learn coherent geometry. Thus, we hypothesize that high-frequency components are a major cause of the failure modes observed in few-shot neural rendering. We provide empirical evidence below.

We investigate how a plain NeRF performs when inputs are encoded by different numbers of frequency bands. To achieve this, we train mipNeRF using masked (integrated) positional encoding. Specifically, we set pos_enc[int(L*x%]):]=0, where LL denotes the length of frequency encoded coordinates after the positional encoding (Eq. 1), and xx is the visible ratio. We briefly demonstrate our observation here and defer the experiment details to §4.1. Figure 2 shows the results for the DTU dataset under the 3 input-view setting. As anticipated, we observe a significant drop in mipNeRF’s performance as higher-frequency inputs are presented to the model. When 10% of total embedding bits are used, mipNeRF achieves a high PSNR of 17.62, while the plain mipNeRF achieves only 9.01 PSNR on its own (at 100% visible ratio). The only difference between these two models is whether masked positional encodings are used. Although removing a significant portion of high-frequency components avoids catastrophic failure at the start of training, it does not result in competitive scene representations, as the rendered images are usually oversmoothed (as seen in Fig. 2 zoom-in patches). Nonetheless, it is noteworthy that in few-shot scenarios, models using low-frequency inputs may produce significantly better representations than those using high-frequency inputs.

Building on this empirical finding, we propose a frequency regularization method. Given a positional encoding of length L+3L+3 (Eq. 2), we use a linearly increasing frequency mask α\boldsymbol{\alpha} to regulate the visible frequency spectrum based on the training time steps, as follows:

where αi(t,T,L)\boldsymbol{\alpha}_{i}(t,T,L) denotes the ii-th bit value of α(t,T,L)\boldsymbol{\alpha}(t,T,L); tt and TT are the current training iteration and the final iteration of frequency regularization, respectively. Concretely, we start with raw inputs without positional encoding and linearly increase the visible frequency by 3-bit each time as training progresses. This schedule can also be simplified as one line of code, as shown in Figure 1. Our frequency regularization circumvents the unstable and susceptible high-frequency signals at the beginning of training and gradually provides NeRF high-frequency information to avoid over-smoothness.

We note that our frequency regularization shares some similarities with the coarse-to-fine frequency schedules used in other works . Different from theirs, our work focuses on the few-shot neural rendering problem and reveals the catastrophic failure patterns caused by high-frequency inputs and their implication to this problem.

3 Occlusion Regularization

Frequency regularization does not solve all problems in few-shot neural rendering. Due to the limited number of training views and the ill-posed nature of the problem, certain characteristic artifacts may still exist in novel views. These failure modes often manifest as “walls” or “floaters” that are located extremely close to the camera, as seen in the bottom of Figure 3. Such artifacts can still be observed even with a sufficient number of training views . To address these issues, proposed a distortion loss. However, our experiments show that this regularization does not help in the few-shot setting and may even exacerbate the issue.

We find most of these failure patterns originate from the least overlapped regions in the training views. Figure 3 shows an example of 3 training views and 2 novel views with “white walls”. We manually annotate the least overlapped regions in the training views for demonstration ((a) and (b) in Fig. 3). These regions are difficult to estimate in terms of geometry due to the extremely limited information available (one-shot). Consequently, a NeRF model would interpret these unexplored areas as dense volumetric floaters located near the camera. We suspect that the floaters observed in also come from these least overlapped regions.

As discussed above, the presence of floaters and walls in novel views is caused by the imperfect training views, and thus can be addressed directly at training time without the need for novel-pose sampling . To this end, we propose a simple yet effective “occlusion” regularization that penalizes the dense fields near the camera. We define:

where mk{\mathbf{m}}_{k} is a binary mask vector that determines whether a point will be penalized, and σK\boldsymbol{\sigma}_{K} denotes the density values of the KK points sampled along the ray in the order of proximity to the origin (near to far). To reduce solid floaters near the camera, we set the values of mk{\mathbf{m}}_{k} up to index MM, termed as regularization range, to 1 and the rest to 0. The occlusion regularization loss is easy to implement and compute.

Experiments

Implementations.

Our FreeNeRF can directly improve NeRF and mipNeRF . To demonstrate this, we use DietNeRF’s codebasehttps://github.com/ajayjain/DietNeRF for NeRF on the Blender dataset and RegNeRF’s codebasehttps://github.com/google-research/google-research/tree/master/regnerf for mipNeRF on the DTU dataset and the LLFF dataset. We disable the proposed components in those papers and implement our two regularization terms on top of their baselines. We make one modification to mipNeRF , which is to concatenate positional encodings with the original Euclidean coordinates (Eq. 2). This is a default step in NeRF but not in mipNeRF, and it helps unify our experiments’ initial visible frequency range. We follow their training schedules for optimization. Please refer to the Appendix for full training recipes.

Hyper-parameters.

Comparing methods.

Unless otherwise specified, we directly use the results reported in DietNeRF and RegNeRF for comparisons, as our method is implemented using their codebases. We also include our reproduced results for reference.

2 Comparison

We compare with state-of-the-art methods in terms of novel view synthesis quality and computation overhead. We show that FreeNeRF outperforms others in synthesis quality while maintaining a much lower cost.

Table 1 shows the image synthesis metrics on the Blender dataset . Our approach outperforms all other methods in the PSNR and SSIM scores, with a comparable LPIPS score to the best one. The improved DietNeRF with fine-tuning still underperforms ours. Note that our direct baseline is “NeRF (repro.)” as we do not use any techniques from DietNeRF . Figure 4 shows two examples for qualitative comparison (see Fig. 1 for plain NeRF’s results). Interestingly, we observe that DietNeRF implicitly distills semantic information from a pre-trained CLIP model into NeRF, which leads to unrealistic and “imaginary” patches that do not exist in the original scenes, such as “ketchup” in the hotdog and rubber-like track-pads in the bulldozer. This behavior is highly correlated to feature distillation and recent developments in 3D object generation that combine NeRF with large pre-trained vision-language models . Although this potentially could be an interesting application, such behavior is undesired in our task and will hamper outputs’ fidelity. In contrast, our method does not require semantics regularization while achieving better performance.

DTU dataset.

Table 2 shows the quantitative results on the DTU dataset. Transfer learning-based methods that require expensive pre-training (SRF , PixelNeRF, and MVSNeRF ) underperform ours in almost all settings, except the full-image PSNR score under 3-view setting. This may be due to the bias introduced by the white table and black background present in many scenes in the DTU dataset, which can be learned as a prior through pre-training. Compared to per-scene optimization methods (mipNeRF , DietNeRF , and RegNeRF ), our approach achieves the best results. Figure 5 shows example novel views rendered by RegNeRF and ours. In the Buddha scene, for instance, piece-wise smoothness imposed by RegNeRF’s geometry regularization leads to the loss of fine-grained details, such as eyes, fingers, and wrinkles. In contrast, our frequency regularization, which can be seen as an implicit geometry regularization, forces smooth geometry at the beginning (due to the limited frequency spectrum) and gradually relaxes the constraint to facilitate the details. In the more challenging scenes (e.g., buildings and the bronze statue in Fig. 5), FreeNeRF produces higher-quality results.

LLFF dataset.

Table 3 and Figure 6 show quantitative and qualitative results, respectively, on the LLFF dataset. We reproduce mipNeRF and obtain better results. Our FreeNeRF is generally the best. Transfer learning-based methods perform much worse than ours on the LLFF dataset due to the non-trivial domain gap between DTU and LLFF. Compared to RegNeRF , our approach predicts more precise geometry and exhibits fewer artifacts. For instance, RegNeRF’s rendered “horns” example (Fig. 6-a) is perceptually acceptable but has poor depth map quality, indicating its incorrect geometry estimation. FreeNeRF, in contrast, renders a less noisy and smoother occupancy field. Also, our approach suffers less from “floaters” than ReNeRF (Fig. 6-b), further demonstrating the efficacy of our occlusion regularization.

Training overhead.

In Table 4, we include the training time of different methods under the same setting. Our method only introduces negligible training overhead (1.02−1.04×1.02-1.04\times) compared to the other approaches (1.62−2.8×1.62-2.8\times). Both DietNeRF and RegNeRF render unobserved patches from novel poses for regularization, which significantly sets back the training efficiency. DietNeRF requires additional forward evaluation of a large model (CLIP ViT B/32, 2242224^{2}, ), and RegNeRF also experiences increased computation due to the use of a normalizing flow model (this part is not open-sourced and therefore not available for our experiments). In contrast, FreeNeRF does not require such additional steps, making it a lightweight and efficient solution for addressing few-shot neural rendering problems.

3 Ablation Study

In this section, we ablate our design choices on the DTU dataset and the LLFF dataset under the 3-view setting. We use a batch size of 1024 for faster training instead of 4096 for the main experiments in Tables 2 and 3.

We investigate the impact of frequency regularization duration TT in Figure 7. Our FreeNeRF benefits more from a longer curriculum in terms of PSNR score across two datasets, with the 90%90\%-schedule being the best. We thus adopt it as our default schedule. However, we notice a trade-off between PSNR and LPIPS where a longer frequency regularization duration can result in higher PSNR but lower LPIPS scores. Fine-tuning the trained model can address this issue and yield better LPIPS scores. More details and discussions are provided in the Appendix.

Occlusion regularization.

Table 5-(a) studies the effect of occlusion regularization. We observe consistent improvements in both datasets when occlusion regularization is included, confirming its efficacy. In contrast, the distortion loss LdistortL_{distort} in worsens the results. Additionally, we find the performance of DTU-3 drops significantly if a large MM is chosen since a large portion of real radiance fields falls in those ranges. The hyper-parameter MM can be set per dataset empirically according to the scene statistics. Further, in Table 5-(b), we show that the way our regularization penalizes points near the camera differs from simply adjusting the near bound. The latter changes the absolute location of the ray starting point, while the occlusion effect remains in the starting area regardless of changes to the near bound.

Limitations.

Our FreeNeRF has two limitations. First, a longer frequency curriculum can make the scene smoother but may decrease LPIPS scores despite achieving competitive PSNR scores. Second, occlusion regularization can cause over-regularization and incomplete representations of near-camera objects in the DTU dataset. Per-scene tuning regularization range can alleviate this issue but we opt not to use it in this paper. Further discussion on these limitations can be found in the Appendix. Addressing these limitations can significantly improve FreeNeRF and we leave them as future work. Still, we consider FreeNeRF to be a simple yet intriguing baseline approach for few-shot neural rendering that differs from the current trend of constructing more intricate pipelines.

Conclusion

We have presented FreeNeRF, a streamlined approach to few-shot neural rendering. Our study unfolds the deep relation between the input frequency and the failure of few-shot neural rendering. A simple frequency regularizer can drastically address this challenge. FreeNeRF outperforms the existing state-of-the-art methods on multiple datasets with minimal overhead. Our results suggest several venues for future investigation. For example, it is intriguing to apply FreeNeRF to other problems suffering from high-frequency noise, such as NeRF in the wild , in the dark , and even more challenging images in the wild, such as those from autonomous driving scenes. In addition, in the Appendix, we show that the frequency-regularized NeRF produces smoother normal estimation, which can facilitate applications that deal with glossy surfaces, as in RefNeRF . We hope our work will inspire further research in few-shot neural rendering and the use of frequency regularization in neural rendering more generally.

References

Appendix A Additional Results

Figure 8 shows more examples to demonstrate the failure mode revealed in Figure 2 that the high-frequency inputs lead to the catastrophic failure of few-shot neural rendering. When taking in 10% of the total embedding bits, mipNeRF can successfully reconstruct scenes despite their over-smoothness. However, with higher-frequency inputs, the scene reconstructions become more unrecognizable and collapse. This experimental finding lies at the heart of FreeNeRF: by restricting the inputs to the low-frequency components at the start of training, NeRF can start from significantly stabilized scene representations at the early stage of training. Upon these stable scene representations, NeRF continues refining the details when high-frequency signals become visible.

A.1 Limitations

In this subsection, we elaborate on the limitations and showcase the failure cases of FreeNeRF.

Figure 7 studies the effect of the duration of frequency regularization on PSNR and LPIPS. From the figure, we observe a trade-off between PSNR and LPIPS that a long-frequency curriculum usually results in a high PSNR score but a low LPIPS score. For example, under the 9 input-view setting, we obtain an object PSNR of 25.59 and an object LPIPS of 0.117 with a 90%-schedule and those of 25.38 and 0.096 with a 50%-schedule. Visually, when the number of input views is relatively sufficient (but still under few-shot settings), results under a shorter schedule usually present more high-frequency details (see the zoom-in patch in Fig. A.1). We thus use 70%-schedule and 50%-schedule for experiments under 6 and 9 input-view settings, respectively. We also found out that training FreeNeRF longer can obtain better LPIPS performance, e.g., 0.182 to 0.167 and 0.308 to 0.290 for DTU-3 and LLFF-3 settings, respectively.

A.2 Depth Evaluation

Here we include results to compare the capability of different methods in depth estimation. As the datasets do not have actual ground truth depth, we utilized depth maps generated by mipNeRFs that were trained on all views as a substitute. FreeNeRF significantly improves its baseline, mipNeRF. RegNeRF, with its patch-based geometry regularization, achieves better performance on the object-centric DTU dataset, while FreeNeRF performs better on the scene-scale LLFF dataset without explicit geometry regularization. This experiment demonstrates the different features of FreeNeRF and RegNeRF, as well as the differences between DTU and LLFF datasets.

A.3 Additional Qualitative Results

Table A.1 provides more numeric results in addition to Table 2 on the DTU dataset. FreeNeRF achieves the best results under the “Average” metrics in most settings. However, we observe less improvement in terms of LPIPS. As we analyze in Section A.1, the slight blurriness introduced by FreeNeRF will result in a low LPIPS score. This is a limitation that could be addressed in the future.

A.4 Additional Visualizations

In Figure A.3, we show more qualitative comparisons between DietNeRF and our FreeNeRF on the Blender dataset. From the zoom-in patches of DietNeRF’s results, we see the generated patches are blurry and do not reflect the same distribution of style as that of ground truth. This is due to implicit semantics distillation behavior done by DietNeRF. In contrast, our FreeNeRF reconstructs scenes closer to the ground truth.

DTU and LLFF.

We provide more rendering results by FreeNeRF in Figures A.4 and A.5 under the 3 input-view setting on the DTU dataset and the LLFF dataset, respectively.

A.5 FreeNeRF for Normal Estimation

We briefly demonstrate a potential FreeNeRF’s application beyond few-shot neural rendering. Specifically, we follow the similar settings in RefNeRF to train a mipNeRF and a FreeNeRF on the “coffee” scene in the Shiny Blender dataset . This dataset aims to benchmark NeRF’s performance on glossy surfaces, where the key challenge is to estimate accurate normal vectors. Figure A.6 shows the comparison between mipNeRF and FreeNeRF. Compared to mipNeRF, FreeNeRF produces more accurate normal estimation and achieves much lower mean angular error (MAE) at no sacrifice of PSNR score. We conjecture that overfitting to high-frequency signals at the start of training is a very common issue in NeRF’s training. However, such partial failure is veiled by good appearance results. We believe these partially degenerated results can be improved with frequency regularization, which makes NeRF’s initial training more stable.

Appendix B Experiment Details

We strictly follow the experimental settings in DietNeRF and RegNeRF to conduct our experiments. We provide some details in the following for completeness.

The Blender dataset has 8 synthetic scenes in total. We follow the data split used in DietNeRF to simulate a few-shot neural rendering scenario. For each scene, the training images with IDs (counting from “0”) 26, 86, 2, 55, 75, 93, 16, 73, and 8 are used as the input views, and 25 images are sampled evenly from the testing images for evaluation. We follow to use a 2×2\times downsampled resolution, resulting in 400×400400\times 400 pixels for each image.

DTU Dataset.

The DTU dataset is a large-scale multi-view dataset that consists of 124 different scenes. PixelNeRF uses a split of 88 training scenes and 15 test scenes to study the “pre-training & per-scene fine-tuning” setting in a few-shot neural rendering scenario. Different from theirs, our method does not require pre-training. We follow to optimize NeRF models directly on the 15 test scenes. The test scan IDs are: 8, 21, 30, 31, 34, 38, 40, 41, 45, 55, 63, 82, 103, 110, and 114. In each scan, the images with the following IDs (counting from “0”) are used as the input views: 25, 22, 28, 40, 44, 48, 0, 8, 13. The first 3 and 6 image IDs correspond to the input views in 3- and 6-view settings, respectively. The images with IDs in serve as the novel views for evaluation. The remaining images are excluded due to wrong exposure. We follow to use a 4×\times downsampled resolution, resulting in 300×400300\times 400 pixels for each image.

LLFF Dataset.

The LLFF dataset is a forward-facing dataset that contains 8 scenes in total. Adhere to , we use every 8-th image as the novel views for evaluation, and evenly sample the input views across the remaining views. Images are downsampled 8×8\times, resulting in 378×504378\times 504 pixels for each image.

Metrics.

B.2 Implementations.

In this codebasehttps://github.com/ajayjain/DietNeRF, a plain NeRF that consists of two MLPs (one coarse MLP and one fine MLP) is used as the baseline. All NeRF models are trained with the Adam optimizer for 200k iterations. The learning rate starts at 5×10−45\times 10^{-4} and decays exponentially with a rate of 0.1. We refer readers to the codebase for more details. In this codebase, the maximum input frequency LL (Eq. 1) used in the position encoding for coordinates is 99. The original coordinates are concatenated with positional encodings by default.

RegNeRF’s codebase.

In this codebasehttps://github.com/google-research/google-research/tree/master/regnerf, a plain mipNeRF is used as the baseline. The maximum input frequency of coordinates is 1616, which is larger than that of the original NeRF . We further concatenate the original coordinates into the positional encodings. All NeRF models are trained with the Adam optimizer with an exponential learning rate decaying from 2⋅10−32\cdot 10^{-3} to 2⋅10−52\cdot 10^{-5} and 512 warm-up steps with a multiplier of 0.01 . Following , we clip gradients by value at 0.1 and then by norm at 0.1 for all experiments. All NeRF models in the main experiments are optimized for 500 epochs with a batch size of 4096. This setting results in around 44k, 88k and 132k training iterations on the DTU dataset for 3/6/9 input views, respectively, and 70k, 140k and 210k training iterations for those on the LLFF dataset, respectively. Note that in the ablation study we use a batch size of 1024 instead of 4096 for faster training.

Occlusion regularization.