RegNeRF: Regularizing Neural Radiance Fields for View Synthesis from Sparse Inputs
Michael Niemeyer, Jonathan T. Barron, Ben Mildenhall, Mehdi S. M. Sajjadi, Andreas Geiger, Noha Radwan
Introduction
Coordinate-based neural representations have gained increasing popularity in the field of 3D vision. In particular, Neural Radiance Fields (NeRF) have emerged as a powerful representation for the task of novel view synthesis, where the goal is to render unseen viewpoints of a scene from a given set of input images.
Though NeRF achieves state-of-the-art performance, it requires dense coverage of the scene. However, in real-world applications such as AR/VR, autonomous driving, and robotics, the input is typically much sparser, with only few views of any particular object or region available per scene. In this sparse setting, the quality of NeRF’s rendered novel views drops significantly (see Fig. 1).
Several works have proposed conditional models to overcome these limitations . These models require expensive pre-training, i.e. training the model on large-scale datasets of many scenes with multi-view images and camera pose annotations, as opposed to test-time optimization which is done from scratch for a given test scene. At test time, novel views can be generated from only a few input images through amortized inference, optionally combined with per scene test time fine-tuning. Though these models achieve promising results, obtaining the necessary pre-training data by capturing or rendering many different scenes can be prohibitively expensive. Moreover, these techniques may not generalize well to novel domains at test time, and may exhibit blurry artifacts as a result of the inherent ambiguity of sparse input data.
One alternate approach is to optimize the network weights from scratch for every new scene and introduce regularization to improve the performance for sparse inputs, e.g., by adding extra supervision or learning embeddings representative of the input views . However, existing methods either heavily rely on external supervisory signals that might not always be available, or operate on low-resolution renderings of the scene that provide only high-level information.
Contribution: In this paper, we present RegNeRF, a novel method for regularizing NeRF models for sparse input scenarios. Our main contributions are the following:
A patch-based regularizer for depth maps rendered from unobserved viewpoints, which reduces floating artifacts and improves scene geometry.
A normalizing flow model to regularize the colors predicted at unseen viewpoints by maximizing the log-likelihood of the rendered patches and thereby avoid color shifts between different views.
An annealing strategy for sampling points along the ray, where we first sample scene content within a small range before expanding to the full scene bounds which prevents divergence early during training.
Related Work
Neural Representations: In 3D vision, coordinate-based neural representations have become a popular representation for various tasks such as 3D reconstruction , 3D-aware generative modelling , and novel-view synthesis . In contrast to traditional representations like point clouds, meshes, or voxels, this paradigm represents 3D geometry and color information in the weights of a neural network, leading to a compact representation. Several works proposed differentiable rendering approaches to learn neural representations from only multi-view image supervision. Among these, Neural Radiance Fields (NeRF) have emerged as a powerful method for novel-view synthesis due to its simplicity and state-of-the-art performance. In mip-NeRF , point-based ray tracing is replaced using cone tracing to combat aliasing. As this is a more robust representation for scenes with various camera distances and reduces NeRF’s coarse and fine MLP networks to a single multiscale MLP, we adopt mip-NeRF as our scene representation. However, compared to previous works , we consider a much sparser input scenario in which neither NeRF nor mip-NeRF are able to produce realistic novel views. By regularizing scene geometry and appearance, we are able to synthesize high-quality renderings despite only using as few as 3 wide-baseline input images.
Sparse Input Novel-View Synthesis: One approach for circumventing the requirement of dense inputs is to aggregate prior knowledge by pre-training a conditional model of radiance fields . PixelNeRF and Stereo Radiance Fields use local CNN features extracted from the input images, whereas MVSNeRF obtains a 3D cost volume via image warping which is then processed by a 3D CNN. Though they achieve compelling results, these methods require a multi-view image dataset of many different scenes for pre-training, which is not always readily available and may be expensive to obtain. Further, most approaches require fine-tuning the network weights at test time despite the long pre-training phase, and the quality of novel views is prone to drop when the data domain changes at test time. Tancik et al. learn network initializations from which test time optimization on a new scene converges faster. This approach assumes that the training and test data are taken from the same domain, and results may degrade if the domain changes at test time.
In this work, we explore an alternative approach which avoids expensive pre-training by regularizing appearance and geometry in novel (virtual) views. Previous works in this direction include DS-NeRF and DietNeRF . DS-NeRF improves reconstruction accuracy by adding additional depth supervision. In contrast, our approach only uses RGB images and does not require depth input. DietNeRF compares CLIP embeddings of unseen viewpoints rendered at low resolutions. This semantic consistency loss can only provide high-level information and does not improve scene geometry for sparse inputs. Our approach instead regularizes scene geometry and appearance based on rendered patches and applies a scene space annealing strategy. We find that our approach leads to more realistic scene geometry and more accurate novel views.
Method
We propose a novel optimization procedure for neural radiance fields from sparse inputs. More specifically, our approach builds upon mip-NeRF , which uses a multi-scale radiance field model to represent scenes (Sec. 3.1). For sparse views, we find the quality of mip-NeRF’s view synthesis drops mainly due to incorrect scene geometry and training divergence. To overcome this, we propose a patch-based approach to regularize the predicted color and geometry from unseen viewpoints (Sec. 3.2). We also provide a strategy for annealing the scene sampling bounds to avoid divergence at the beginning of training (Sec. 3.3). Finally, we use higher learning rates in combination with gradient clipping to speed up the optimization process (Sec. 3.4). Fig. 2 shows an overview of our method.
Here, indicates the network weights and a predefined positional encoding applied to and .
Volume Rendering: Given a neural radiance field , a pixel is rendered by casting a ray from the camera center through the pixel along direction . For given near and far bounds and , the pixel’s predicted color value is computed using alpha compositing:
and and indicate the density and color prediction of radiance field , respectively. In practice, these integrals are approximated using quadrature . A neural radiance field is optimized over a set of input images and their camera poses by minimizing the mean squared error
where indicates a set of input rays and its GT color.
mip-NeRF: While NeRF only casts a single ray per pixel, mip-NeRF instead casts a cone. The positional encoding changes from representing an infinitesimal point to an integration over a volume covered by a conical frustum. This is a more appropriate representation for scenes with varying camera distances and allows NeRF’s coarse and fine MLPs to be combined into a single multiscale MLP, thereby increasing training speed and reducing model size. We adopt the mip-NeRF representation in this work.
2 Patch-based Regularization
NeRF’s performance drops significantly if the number of input views is sparse. Why is this the case? Analyzing its optimization procedure, the model is only supervised from these sparse viewpoints by the reconstruction loss in (3). While it learns to reconstruct the input views perfectly, novel views may be degenerate because the model is not biased towards learning a 3D consistent solution in such a sparse input scenario (see Fig. 1). To overcome this limitation, we regularize unseen viewpoints. More specifically, we define a space of unseen but relevant viewpoints and render small patches randomly sampled from these cameras. Our key idea is that these patches can be regularized to yield smooth geometry and high-likelihood colors.
Unobserved Viewpoint Selection: To apply regularization techniques for unobserved viewpoints, we must first define the sample space of unobserved camera poses. We assume a known set of target poses where
These target poses can be thought of bounding the set of poses from which we would like to render novel views at test time. We define the space of possible camera locations as the bounding box of all given target camera locations
where and are the elementwise minimum and maximum values of , respectively.
To obtain the sample space of camera rotations, we assume that all cameras roughly focus on a central scene point. We define a common “up” axis by computing the normalized mean over the up axes of all target poses. Next, we calculate a mean focus point by solving a least squares problem to determine the 3D point with minimum squared distance to the optical axes of all target poses. To learn more robust representations, we add random jitter to the focal point before calculating the camera rotation matrix. We define the set of of all possible camera rotations (given the sampled position ) as
where indicates the resulting “look-at” camera rotation matrix and is a small jitter added to the focus point. We obtain a random camera pose by sampling a position and rotation:
Geometry Regularization: It is well-known that real-world geometry tends to be piece-wise smooth, i.e., flat surfaces are more likely than high-frequency structures . We incorporate this prior into our model by encouraging depth smoothness from unobserved viewpoints. Similarly to how a pixel’s color is rendered in (2), we calculate the expected depth as:
We formulate our depth smoothness loss as
where indicates a set of rays sampled from camera poses , is the ray through pixel of a patch centered at , and is the size of the rendered patches.
Color Regularization: We observe that for sparse inputs, the majority of artifacts are caused by incorrect scene geometry. However, even with correct geometry, optimizing a NeRF model can still lead to color shifts or other errors in scene appearance prediction due to the sparsity of the inputs. To avoid degenerate colors and ensure stable optimization, we also regularize color prediction. Our key idea is to estimate the likelihood of rendered patches and maximize it during optimization. To this end, we make use of readily-available unstructured 2D image datasets. Note that, while datasets of posed multi-view images are expensive to collect, collections of unstructured natural images are abundant. Our only criterion for the dataset is that it contains diverse natural images, allowing us to reuse the same flow model for any type of real-world scene we reconstruct. We train a RealNVP normalizing flow model on patches from the JFT-300M dataset . With this trained flow model we estimate the log-likelihoods (LL) of rendered patches and maximize them during optimization. Let
and indicates a set of rays sampled from , the predicted RGB color patch with center , and the negative log-likelihood (“NLL”) with Gaussian .
Total Loss: The total loss we optimize in each iteration is
where indicates a set of rays from input poses, and a set of rays from random poses .
3 Sample Space Annealing
For very sparse scenarios (e.g., 3 or 6 input views), we observe another failure mode of NeRF: divergent behavior at the start of training. This leads to high density values at ray origins. While input views are correctly reconstructed, novel views degenerate as no 3D-consistent representation is recovered. We find that annealing the sampled scene space quickly over the early iterations during optimization helps to avoid this problem. By restricting the scene sampling space to a smaller region defined for all input images, we introduce an inductive bias to explain the input images with geometric structure in the center of the scene.
Recall from (2) that are the camera’s near and far plane, respectively, and let be a defined center point (usually the midpoint between and ). We define
where indicates the current training iteration, a hyperparameter indicating how many iterations until the full range is reached, and a hyperparameter indicating a start range (e.g., ). This annealing is applied to renderings from both the input poses and the sampled unobserved viewpoints. We find that this annealing strategy ensures stability during early training and avoids degenerate solutions.
4 Training Details
We build our code on top of of the JAX mip-NeRF codebase.Official codebase available at https://github.com/google/mipnerf. We optimize with Adam using an exponential learning rate decay from to . We clip gradients by value at and then by norm at . We train for pixel epochs, e.g., K, K, and K iterations on DTU for input views respectively (all fewer iterations than mip-NeRF’s default K steps ).
Experiments
Datasets We report results on the real world multi-view datasets DTU and LLFF . DTU contains images of objects placed on a table, and LLFF consists of complex forward-facing scenes. For DTU, we observe that in scenes with a white table and a black background, the model is heavily penalized for incorrect background predictions regardless of the quality of the rendered object-of-interest (see Fig. 3). To avoid this background bias, we evaluate all methods with the object masks applied to the rendered images (full image evaluations in supp. mat.). We adhere to the protocol of Yu et al. and evaluate on their reported test set of scenes. For LLFF, we adhere to community standards and use every -th image as the held-out test set and select the input views evenly from the remaining images. Following previous work , we report results for the scenarios of , , and input views.
Metrics: We report the mean of PSNR, structural similarity index (SSIM) , and the LPIPS perceptual metric . To ease comparison, we also report the geometric mean of , , and LPIPS .
Baselines: We compare against the state-of-the-art conditional models PixelNeRF , Stereo Radiance Fields (SRF) , and MVSNeRF . We re-train PixelNeRF for the view scenarios, leading to better results, and we similarly pre-train SRF with views. We pre-train all methods on the large-scale DTU dataset. The LLFF dataset has been shown to be too small for pre-training and hence serves as an out-of-distribution test for conditional models. We report the conditional models on both datasets also after additional per-scene test time optimization (“ft” for “fine-tuned”). Further, we compare against mip-NeRF and DietNeRF which do not require pre-training, as with our approach. As no official code is available, we reimplement DietNeRF on top of the mip-NeRF codebase (achieving better results) and train both methods for K iterations per scene with exponential learning rate decay from to .
We first compare our model to the vanilla mip-NeRF baseline, analyzing the effect of our regularizers on scene geometry, appearance and data efficiency.
Geometry Prediction: We observe that novel view synthesis performance is directly correlated with how accurately the scene geometry is predicted: in Fig. 4 we show expected depth maps and RGB renderings for mip-NeRF and our method on the LLFF room scene. We find that for input views, mip-NeRF produces low-quality renderings and poor geometry. In contrast, our method produces an acceptable novel view and a realistic scene geometry, despite the low number of inputs. When increasing the number of input images to or , mip-NeRF’s predicted geometry improves but still contains floating artifacts. Our method generates smooth scene geometry, which is reflected in its higher-quality novel views.
To evaluate our gain in data efficiency, we train mip-NeRF and our method for various numbers of input views and compare their performance.Results slightly differ from Tab. 1 as a smaller test set has to be used. We find that for sparse inputs our method requires up to fewer input views to match mip-NeRF’s mean PSNR on the test set, where the difference is larger for fewer input views. For input views, both methods achieve a similar performance (as this work focuses on sparse inputs, tuning hyper-parameters for more input views could result in improved performance for these scenarios).
2 Baseline Comparison
DTU Dataset For input views, our method achieves quantitative results comparable to the best-performing conditional models (see Tab. 1) which are pre-trained on other DTU scenes. Compared to the other methods that also do not require pre-training, we achieve the best results. For and input views, our approach performs best compared to all baselines. As evidenced by Fig. 6, we see that conditional models are able to predict good overall novel views, but become blurry particularly around edges and exhibit less consistent appearance for novel views whose cameras are far from the input views. For mip-NeRF and DietNeRF, which are not pre-trained (like our method), geometry prediction and hence synthesized novel views degrade for very sparse scenarios. Even with or input views, the results contain floating artifacts and incorrect geometry. In contrast, our approach performs well across all scenarios, producing sharp results with more accurate scene geometry.
LLFF Dataset: For conditional models, the LLFF dataset serves as an out-of-distribution scenario as the models are trained on DTU. We observe that SRF and PixelNeRF appear to overfit to the training data, which leads to low quantitative results (see Tab. 2). MVSNeRF generalizes better to novel data, and all three models benefit from additional fine-tuning. For input views, mip-NeRF and DietNeRF are not able to generate competitive novel views. However, with or input views, they outperform the best conditional models. Despite requiring fewer optimization steps than mip-NeRF and DietNeRF and no pre-training at all, our method achieves the best results across all scenarios. From Fig. 7 we observe that the predictions from conditional models tend to be blurry for views far away from the inputs, and the test-time optimized baselines contain errors in predicted scene geometry. Our method achieves superior geometry predictions and more realistic novel views.
3 Ablation Studies
In Tab. 3, we ablate various components of our method. For sparse inputs, we find that the proposed scene space annealing strategy avoids degenerate solutions. Further, regularizing geometry is more important than appearance, and combining all components leads to the best results.
In Tab. 4, we investigate the performance of other geometry regularization techniques. We find that opacity-based regularizers (e.g., enforce rendered opacity values near to either or ) and density or normal smoothness priors (e.g. minimize the distance between neighboring normal vectors in 3D), two strategies often used to enforce solid and smooth surfaces, do not produce accurate scene geometry. Employing the sparsity prior from Hedman et al. leads to better quantitative results, but novel views still contain floating artifacts and the optimized geometry has holes. In contrast, our geometry regularization strategy achieves the best performance. We hypothesize that similar to density-based vs. single surface optimization for coordinate-based methods, providing gradient information along the full ray rather than a single point provides a more stable and informative learning signal.
Conclusion
We have presented RegNeRF, a novel approach for optimizing Neural Radiance Fields (NeRF) in data-limited regimes. Our key insight is that for sparse input scenarios, NeRF’s performance drops significantly due to incorrectly optimized scene geometry and divergent behavior at the start of optimization. To overcome this limitation, we propose techniques to regularize the geometry and appearance of rendered patches from unseen viewpoints. In combination with a novel sample-space annealing strategy, our method is able to learn 3D-consistent representations from which high-quality novel views can be synthesized. Our experimental evaluation shows that our model outperforms not only methods that, similar to us, only optimize over a single scene, but in many cases also conditional models that are extensively pre-trained on large scale multi-view datasets.
Limitations and Future Work: In this work, we do not attempt to hallucinate geometric detail. As a result, our model may lead to blurry predictions in unobserved areas with fine geometric structures (see Fig. 8). We identify incorporating uncertainty prediction mechanisms or generative components as promising future work.