Aug-NeRF: Training Stronger Neural Radiance Fields with Triple-Level Physically-Grounded Augmentations
Tianlong Chen, Peihao Wang, Zhiwen Fan, Zhangyang Wang
Introduction
Neural radiance fields (NeRF) and its variants have demonstrated impressive progresses in learning to represent 3D objects and scenes from images towards photo-realistic novel view synthesis. NeRF leverages a multi-layer perceptron (MLP) to implicitly modeling the mapping from an input 5D coordinates (i.e., 3D coordinates () and 2D viewing directions ()) to volume density and view-dependent emitted radiance color () at the corresponding position in the scene. Then, the obtained continuous 5D function (i.e., MLP) can be utilized to generate novel views with traditional volume rendering mechanisms.
Although NeRF is capable of producing novel views, it unfortunately suffers from inconsistent and non-smooth geometries since the vanilla MLP lacks geometry-awareness. For example, as shown in Fig. 1, the depth maps and 3D geometries of the scene generated by NeRF show obvious discontinuity and outliers, especially around the edge of objects. Considering that the quality of reconstructed geometry plays a central role in view rendering, that might account for NeRF’s limited generalization to unseen views.
To fill in this research gap, a straightforward solution is introducing explicit geometric regularizers like Laplacian or total variation (TV) to enhance the continuity. However, these explicit regularizers are often found to constrain the representation flexibility of MLP too aggressively, resulting in inferior performance. Recent advances in robust data augmentations establish promising successes in image recognition in terms of both improved functional smoothness and generalization.
Motivated by that, we design an Augmented NeRF (Aug-NeRF) training framework, which injects worst-case perturbations to implicitly regularize the NeRF pipeline with physical foundations. Specifically, Aug-NeRF considers to regularize three different levels, including () the input coordinates, where perturbations can imitate the inaccurate camera poses during collecting images; () the intermediate features, in order for a smooth/flat model loss landscape when fitting objects’ 3D geometries that is believed to enhance generalization; () the pre-rendering output, to model potential degradation factors in the image supervision. As presented in Fig. 1, our Aug-NeRF achieves smoother and more consistency reconstructed geometry and improved unseen view synthesis. Additionally, we find Aug-NeRF to show surprising resilience towards severely corrupted supervision images. The main contributions of this paper can be summarized as follows:
We reveal the existence of highly non-smooth geometries in representing scenes as neural radiance fields (NeRF), which we regard as a crucial bottleneck of NeRF’s generalization ability to unseen views.
To address such limitation of NeRF, we propose Aug-NeRF, a triple-level, physically-grounded augmented training pipeline, by leveraging worst-case perturbations to implicated regularize the input coordinate, intermediate feature, and pre-rendering output levels.
Extensive experiments validate the effectiveness of our proposal on diverse scene synthesis tasks, to endow NeRF with smoothness-aware geometry reconstruction, enhanced generalization to synthesizing unseen views, and stronger tolerance of noisy supervisions.
Related Work
It is well-known that deep networks are vulnerable to imperceptible worst-case perturbations . Numerous defense mechanisms have been invented to address the issue, where adversarial training (AT) approaches remains as the de-facto. Although conventional AT enhances model robustness at the price of compromising the standard accuracy , recent studies reveal AT can be harnessed to enhance models’ standard generalization as well . Taking for example, it applies adversarial perturbations to input samples as a form of data augmentation, and shows to improve image classification on the clean dataset. apply worst-case perturbations to the input embedding for natural language understanding, language modeling, and vision-and-language tasks, all successfully boosting their standard generalization. constructed more sophisticated variations of robust augmentations, including both data-driven and heuristic components, to improve model generalization further. However, such robust augmentations on inputs or intermediate features, to our best knowledge, have not been studied in the view synthesis field. This paper explores this possibility by looking into the intrinsic physical grounds.
Neural 3D Representations.
Classic 3D reconstruction approaches utilizes discrete representations such as point clouds , meshes , multi-plane images , depth maps and voxel grids . Neural implicit representations leverage coordinate-based neural networks to approximate visual signals . Such ideas have been successfully applied to both 2D images and 3D objects . Recent advances follow differentiable rendering and end-to-end optimization to reconstruct the neural 3D scene from 2D image supervision . Liu et al. presented the first usage of neural implicit function to infer 3D representation with differentiable rendering. DVR and IDR adopt surface rendering to reconstruct implicit iso-surface by supervising on both images and pixel-accurate object masks.
NeRF pioneered to use differentiable volumetric rendering to optimize a neural radiance field, and achieved more photorealistic and view-consistent results. Many works continue to improve its training and rendering accuracy, efficiency, and generalization. NeRF++ separates two NeRFs to handle foreground and background, respectively. NeRF-W tackles unstructured photos via modeling transient noises and uncertainty. MipNeRF mitigates objectionable aliasing artifacts for NeRF to represent fine details. HyperNeRF introduces topology-aware level-set methods to rectify NeRF geometry especially for dynamics. extend NeRF with lighting and rendering modeling. enhance the underlying geometries reconstructed by NeRF by adopting surface representation in the place of the density volume. leverage multi-view spatial image feature or semi-reconstructed 3D information to reduce input view number and enable generalization to new scenes. free NeRF from accurate camera pose estimation. Acceleration of NeRF training and inference have also been discussed in . Despite so many exciting progresses, studying NeRF’s training stability and data robustness remains an open question.
Preliminaries
NeRF models the underlying 3D scene as a continuous volumetric radiance field of color and density. Formally, a typical radiance field can be written as , where is the spatial coordinate, indicates the view direction, and represent the RGB color and density, respectively. NeRF further parameterizes this 5D-valued function by a composition of Positional Embedding (PE) and the MLP , where is a Fourier feature mapping network , is the network weights. Given a radiance field, NeRF follows the classical volume rendering to render an arbitrary view .
Our goal is to fit a neural radiance from calibrated RGB images captured from multiple views. Suppose we have a set of images with corresponding extrinsic parameters. NeRF simulates the physical imaging process, by casting a ray for each pixel via inverse perspective projection with respect to the camera pose, where denotes the optical center of camera, is the direction of the ray, and is the angular view direction (see Fig. 2). We collect all pairs of rays and pixel colors as the training set , where is the total number of rays, and denotes the ground-truth color of the -th ray. To simulate the color of a ray, NeRF first partitions evenly-spaced bins between the near-far bound along the ray, and then uniformly samples one point within each bin: . Afterwards, NeRF numerically evaluates volumetric ray integration via the following equation:
where , and . With this forward model, NeRF optimizes the expected distance between rendered ray colors and ground-truth pixel colors as follows:
Methodology
NeRF conducts uniform sampling along each ray and interpolates a continuous radiance field via an MLP. However, we argue that the point sampling and the MLP interpolation can never be optimal during training dynamics due to the biased sampling strategy and non-smoothness of MLP. To this end, we propose to train NeRF with a smoothing prior. Sec. 4.1 provides a probabilistic interpretation of this intuition. Different from explicit smoothness modeling, e.g., total variation penalty or low rank prior, we utilize worst-case perturbations as a data-adaptive regularization. We call this training strategy Aug-NeRF.
An overview of our Aug-NeRF is presented in Fig. 2. Following the rendering pipeline of NeRF, Aug-NeRF injects adversarial noises into the following stages: point sampling, intermediate features, and MLP outputs. Each perturbation is searched within a small range to maximize the final loss. It could be treated as a regularization to be jointly minimized with the original training loss (see Sec. 4.2).
1 NeRF as Maximum A Posterior
where is a normalization term, and is the variance.
2 Regularize NeRF with Robust Augmentations
Imposing smoothness onto NeRF can be done in many explicit ways, such as regularizing total variation , Laplacian of surface , etc. However, those regularizers are often not sufficiently data-adaptive, and can constrain the representation flexibility too aggressively, as evidenced in Sec. 5.3. Also, their computation also usually operates on discretized volumetric representations, and needs extra differentiation steps to be added in NeRF.
Recent works suggest a promising alternative by integrating worst-case adversarial perturbations as data augmentations (i.e., AT). AT restricts the change of loss when its input is perturbed, leading to flattening the loss landscape . As a result, the trained network’s intrinsic feature manifold and loss landscape become smoother. Prevailing theories link the generalization ability of deep networks to the geometry of the loss landscape; in particular, a model trained to converge to wide valleys (i.e., flat basins) in loss landscape shows better generalization ability as well as robustness to distributional shifts.
NeRF is trained by given 2D image views (often with known camera poses) and is tested to synthesize novel views from unseen angles. Intuitively, the unsatisfactory novel view synthesis could be seen as a training-testing “generalization gap” issue. This inspires us to incorporate robust augmentations into NeRF to induce a data-adaptive smoothness prior that enhances generalization.
Designing dedicated perturbations for NeRF is far from trivial due to its inherent physics. Unlike conventional deep models, the forward pass of NeRF consists of two white-box simulating stages (point sampling, volumetric rendering) and one black-box network mapping stage. We propose to inject worst-case perturbations into all three levels: coordinates, intermediate features of MLP, and pre-rendering MLP output: all with clear physical meanings. Formally, our approach can be formulated as a min-max game:
where , , and are the perturbations to be learned and injected to the input coordinate, intermediate MLP feature, and pre-rendering RGB- output, respectively, where , , and are the corresponding perturbation search range, is the hidden dimension of the MLP. We elaborate on each perturbation as below.
The original NeRF first randomly samples point along each ray and then conducts importance sampling to simulate the quadrature of the integration. This strategy also mitigates overfitting and produces smoother scene representation . Arandjelovic et al. further proposes an attention-guided sampling scheme to refine this process. However, our insight is that using either coarse-to-fine or learning-based sampling will cause the sampling to overfit the density distribution of the currently rendered ray, which might hold back NeRF when the density field is biased or cannot generalize.
To this end, we propose to produce a worst-case point sampling during training, to simulate a test-time “distributional shift” for NeRF to handle. To be specific, we search a coordinate perturbation following Eqn. 4.2. The coordinate perturbation consists of three parts: 1) the along-ray perturbation shifts point samples along the ray, 2) the point position perturbation is added to the direct input of the NeRF MLP, 3) in addition, we also inject the perturbation to the view direction. Formally, given the perturbation , the input of MLP turns out to be:
The constraint set for is defined as , where is a hyperparamter. The coordinate perturbation lies in a ball to constrain points with a cylinder along the ray. View direction perturbation is restricted within the conical frustum , where is the focal length, and is the pixel size.
Pre-Rendering Output Perturbation.
NeRF next maps points on a ray to the corresponding color and density, then conducts volumetric rendering to compose these point values into the 2D pixel values. As shown by Fig. 1, the reconstructed shape can be noisy and discontinuous. We attribute these artifacts to two reasons: (i) neural implicit functions represented by MLP are not necessarily smooth . When zooming in, we observe the function landscape to be rugged; (ii) the MLP output goes through volumetric rendering to form the RGB output. As the volumetric rendering itself has smoothing effects owing to its point-by-point accumulation, it might “mask” the non-smoothness and noise of the pre-rendering results hence they cannot be effectively eliminated at supervised training.
Inspired by robust training enhancing output smoothness , we propose to intentionally corrupt the output of the MLP with worst-case pertubation, in order to encourage the output smoothness of the MLP, which in turn smooths the NeRF underlying geometry. Given the pre-rendering perturbation , , we perturb the rendering in Eqn. 3 by:
where are outputs by perturbed coordinates, is the transmittance term, is the interval of integral, and correspond to color and density perturbations, respectively. We fix the constraint set as . will be further clamped to make sure lie between $$.
Intermediate Feature Perturbation.
In addition to perturbing per-rendering color and ray density, we also inject adversarial noise into the intermediate features. As revealed by , augmenting intermediate features can further smooth learned functional mappings, more than just augmenting inputs or outputs. To be specific, according to , the backbone MLP can be written as , where (with positional encoding) maps a coordinate to a -dimension feature vector, and project it to RGB color and density, respectively. We crafted the worst-case perturbations as follows:
Intermediate feature perturbation is searched over with a hyperparameter . We also test various injection points of the backbone MLP in Sec. 5.3.
3 Optimization
After incorporating all augmentations, the full training objective is defined as ( as tuned by grid search):
Experiments
Datasets. We evaluate our proposals on public representative datasets of both LLFF and NeRF-Synthetic . Particularly, the face-forwarding scenes {“fern”, “orchids”, “trex”} from LLFF dataset and {“drums”, “ship”, and “chair”} instances in 360° NeRF-Sythetic dataset are adopted in our experiments. To accelerate training, we down-sampled LLFF dataset by 1/8 and 360° NeRF-Synthetic dataset by 1/2.
Training. We employ the same MLP architecture and training recipe with the original NeRF. Aug-NeRF is trained for 500K iterations to guarantee convergence. All hyperparameters are carefully tuned by a grid search and the best configuration is applied to all experiments, as demonstrated in Sec. 5.3. NeRF models are trained on a NVIDIA RTX A GPU with GB memory.
Baseline and Comparison Variants. Our Aug-NeRF is established on the vanilla NeRF . Two groups of current top-performers for view synthesis are compared, including () NeRF-based approaches: NeRF and MipNeRF ; and () classical methods: Neural Volume (NV) , Scene Representation Network (SRN) , and Local Light Field Fusion (LLFF) . For a fair comparison, all above models are trained/tested on the same views of identical scenes.
2 Improved NeRF with Augmentations
In this section, we validate our proposed Aug-NeRF on LLFF and 360° NeRF-Sythetic datasets across six representative scenes. Quantitative comparisons against vanilla NeRF and other top-performing algorithms like {MipNeRF , NV , SRN , LLFF } are provided in Tab. 1 and 2, together with qualitative test views presented in Fig. 3. These results convey several observations:
Aug-NeRF reduces average error by and on the LLFF and 360° NeRF-Sythetic datasets, respectively. It consistently outperforms NeRF on all metrics by a large margin, e.g., {, , , , , } PSNR improvements at scenes {“fern”, “orchids”, “trex”, “drums”, “ship”, “chair”}, showing impressive “generalization” boosts on unseen views thanks to our augmentations.
Compared with recent state-of-the-art MipNeRF and other classical approaches, Aug-NeRF shows a clear advantage, especially in terms of PSNR. In some cases, MipNeRF has a slightly higher SSIM; but Aug-NeRF is able to outperform it in most cases.
Aug-NeRF achieves superior performance in representing fine geometry, as shown in Fig. 3 such as Fern’s and Orchid’s leaves, the skeleton ribs, and railing in T-rex. Both NeRF and MipNeRF reconstruct the low-frequency geometry and color variation, but fail to generate high-quality fine details (see zoom-in).
Depth and Geometry visualization.
The learned depth maps and fitted 3D geometries from NeRFs are provided in and 4 and Fig. 5, respectively. The 3D shapes (Fig. 5) are synthesized by MarchingCube algorithms . We observe that vanilla NeRF suffers from a serrated surface (which overwhelms the fine details), while traditional TV and Laplacian regularizations tend to excessively smoothen the results. Aug-NeRF reduces noises and improves surface smoothness, in a detail- and geometry-preserving manner.
Superior synthesis when trained on noisy data.
As an extra study, we examine Aug-NeRF under supervision images with additive noise corruptions. From Tab. 3 and Fig. 6, compared to the vanilla NeRF, Aug-NeRF shows consistent average error reductions for both Gaussian and Shot noises, while it substantially improves the visual quality of constructed test views (e.g., much fewer noises in the “fern”). We regard it as an additional bonus from enforcing smooth geometry in NeRF training.
3 Ablation Study
To compare the effects of robust augmentations at different levels, we conduct step-wise evaluation as: () NeRF, () NeRF + Feature Aug., () NeRF + Feature & Output Aug., ) NeRF + Feature & Output & Input coordinates Aug., which is our complete Aug-NeRF. Tab. 4 shows that applying robust augmentation to each level brings extra and complementary generalization gains, among which augmenting the pre-rendering output level makes the biggest difference.
Worst-case v.s. random perturbations.
One straightforward baseline for Aug-NeRF is to just use random data augmentation. Particularly, we employ random Gaussian noises to both intermediate features and pre-rendering outputs of NeRFThe vanilla NeRF has already included random noise in coordinates.. As in Tab. 4, Random Aug. obtains moderate performance boosts for all metrics, but are clearly less obvious than our worst-case perturbations.
Effects of augmentation strength and location.
The accuracy gains from Aug-NeRF are largely determined by the strength and location of crafted worst-case perturbations. A comprehensive investigation on three levels of augmentations, i.e., input coordinate, intermediates features, and pre-rendering output, are presented in Fig. 7. When studying one of the factors, we stick to the best configuration for the rest factors. Fig. 7 reveals that: First, NeRF gains the most from {coordinate, features, pre-rendering output} augmentations with {PGD-3, PGD-1, PGD-1} and step size {, , }; Second, applying generated perturbations to the middle layer of NeRF’s MLP contributes the most significantly; Third, too strong (e.g., PGD-10) worst-case perturbations may still deteriorate performance.
Comparison with explicit smooth regularizations.
Conclusion and Broad Impact
In this paper, we have presented Aug-NeRF that addresses the inherent non-smooth geometries of NeRF. Specifically, based on solid physical grounds, Aug-NeRF seamlessly injects worst-case perturbations into three levels of the NeRF pipeline, leading to substantially improved geometry continuity and generalization ability. Extensive quantitative and qualitative results across diverse scenes validate the effectiveness of our proposals. Moreover, the implicit smooth prior induced by triple-level augmentation enables NeRF to recover scenes from noisy supervision images. One limitation is that we only study additive noises (e.g., Gaussian) for corrupted images. We will extend the investigation to other complicated corruptions.
References
Appendix A1 More Technical Details
We summarize the detailed procedures of Aug-NeRF in the Algorithm 1.
Appendix A2 More Experiment Results
We present the constructed test views in Fig. A8 and the learned depth maps in Fig. A9. As shown in Fig. A8, we find that the vanilla NeRF fails to capture the fine-grained details of objects, such as the “ship net”, while Aug-NeRF demonstrates substantially improved visual qualities.
In the meantime, from the depth maps in Fig. A9, NeRF baseline suffers from severe noises. On the contrary, Aug-NeRF enjoys much more smooth depth maps, which suggests that our proposed triple-level robust augmentations indeed enhance the NeRF’s continuity and generate smooth geometry representations.
Benefits in overfitting vs underfitting cases.
We take the scene “chair” as an example, and investigate three combinations of different model sizes and data scales as below Tab. A5: () big NeRF (512) on small images ( Res.); () normal NeRF (256) on small images ( Res.); () small NeRF (128) on large images (Full Res.). We show that in all settings of overfitting / normal case / underfitting, our proposed augmentations are consistently beneficial.
Geometry extraction.
To obtain geometric visualization in Fig. 5, we first query the network with a regular lattice defined over , and export a discretized density field volume. The absolute voxel size is 2/512. Then we employ marching cube algorithm provided in UCSF Chimerahttps://www.cgl.ucsf.edu/chimera/ to extract the surface. We set the threshold to 25 and 1 for chairs and drums, respectively. The step size is chosen as 1. In order to numerically assess the quality of reconstructed geometries, we introduce Chamfer Distance (CD) to measure the difference between reconstructed geometries and ground-truth models:
where and are point sets sampled from the extracted surfaces and ground-truth models, respectively. On scene chair, our AugNeRF achieves CD which is 29.25% lower than vanilla NeRF ().
Different types of noise and inaccurate camera poses.
As shown in Tab. 3 and Fig. 6, we experiment on two kinds of corruptions, i.e., Gaussian and Shot noises. In this paragraph, we add extra results of training with Pepper noise and inaccurate camera poses are collected in below Tab. A6. The results consistently demonstrate the superiority of Aug-NeRF. We note that our main goal is to endow NeRF with smoothness-aware geometry reconstruction, enhanced generalization to synthesizing unseen views, while the improved tolerance of noisy supervisions is a by-product bonus.
Implementation of explicit regularization.
where denotes the density branch of the function . However, evaluating this integral is implausible. Instead, we discretize the integral interval into regular grids and conducting quadrature rule for estimating TV regularization:
where denotes the Laplacian operator.