TermiNeRF: Ray Termination Prediction for Efficient Neural Rendering

Martin Piala, Ronald Clark

Introduction

The idea of using neural networks to represent a 3D scene as a continuous neural field has recently emerged as a promising way to represent and render 3D scenes . In general, this type of representation uses a neural network to learn a function that maps coordinates to optical properties of the scene, such as color and opacity, at each point in space. Unlike any of the traditional representations such as voxel grids, meshes or surfels , neural fields do not require any discretization of the scene.

While this type of representation captures scenes at very high quality, a major limitation is that it takes a long time to train the model and render images. This is because the neural network representing the volume has to be queried along all the viewing rays. For an image of width ww, and height hh, and nn depth samples, the rendering requires Θ(nhw)\Theta(nhw) network forward passes. With NeRF, it can take up to 30 seconds to render a 800×800800\times 800 image on a high-end GPU.

In this paper, we propose a more efficient way to render (and train) a neural field model. Our approach works by jointly training a sampling network along with a color network, with the collective name TermiNeRF. The sampling network estimates where along the viewing direction the surfaces lie, thereby allowing the color network to be sampled much more efficiently, reducing the rendering time to ≤1\leq 1s. We show how this approach can be applied to quickly edit a scene by rapidly learning changes to lighting or materials for a scene.

Related Work

In this section, we give an overview of the related work on raymarching neural network-based 3D scenes and existing attempts to make these approaches more efficient.

Using a neural network that represents a continuous scalar or vector function through space has recently emerged as a promising way to parameterize 3D objects and scenes for rendering. Such models go by various names, including “neural coordinate-based representations” or “neural implicit models”, but perhaps the most appropriate and general term is “neural fields” .

The first approaches in this direction were focused on representing geometric shapes as distance fields and include methods such as DeepSDF , Occupancy Networks , and IM-Net . These approaches are generally trained using a sampled 3D point cloud of the shape as supervision. DVR introduced a “differentiable” volumetric rendering method to train a neural occupancy field using only 2D images as supervision. The main disadvantages of DVR is that it is only able to reconstruct solid surfaces, and view-dependent effects caused by reflections and specularities cannot be captured.

In contrast, neural radiance fields (NeRF) use a neural network to parametrize the color and opacity of each point in space, coupled with volumetric raymarching, to render novel views. The main disadvantage of NeRF is that the volumetric rendering requires sampling the network at multiple positions along a ray projected through each pixel, leading to a prohibitively high computational cost.

More recent works have tried to deal with images captured in uncontrolled conditions with transient distractors and varying lighting conditions . Other works have looked at how the scene can be decomposed into components such as reflectance and incoming radiance so that the scene can be re-lit with new lighting.

PixelNeRF and MVSNeRF proposes a generalized NeRF approach conditioned on image inputs, enabling memory-efficient rendering of various scenes with a small training overhead.

2 Efficient neural radiance fields

In order to improve the efficiency of the volumetric rendering of neural radiance fields, two main types of approaches exist. The first type tries to reduce the number of samples taken along each ray by “intelligently” picking sample locations . utilizes depth supervision for more efficient sample placement during the training. DONeRF takes this idea further and learns a separate depth network that is trained using ground-truth depth maps to predict possible surface locations. This allows them to sample more efficiently in the region near the surface.

uses a transformer or an MLPMixer to predict sampling locations from an input ray. Although being end-to-end trainable, this method only achieves 25%25\% speedup with a significant degradation in quality.

The second approach type tries to improve the efficiency of sampling the field, thus reducing the time taken for each evaluation of the scene model. DeRF decomposes the space into a number of regions and uses a smaller MLP for learning the appearance of each region, which significantly reduces the rendering time. KiloNeRF takes this idea to the extreme and uses a very large collection of many small MLPs. FastNeRF factorizes NeRF into two MLPs and speeds up the rendering process via caching.

replaces the MLP with a “Plen” Octree that stores the direction-dependent radiance at each position. This octree is orders of magnitude faster to sample than an MLP. Instead of learning color and opacity at individual locations and integrating them during volumetric raymarching, Autoint learns a network that can predict the integrated color of ray segments. In so doing, it amortizes the integration and makes rendering noticeably faster.

Method

In this section, we introduce our method for fast rendering and, by extension, training of neural radiance fields models. The key insight of our method is that the only samples that contribute to the final color of a ray are ones which have non-zero opacity (i.e., samples near surfaces or participating media) and are not already occluded. Thus, our method works by learning a distribution along each ray, via a sampling network, that represents the probability of the ray encountering a surface or some participating media at each point. With the information about the distribution of the volume along the ray, we were able to make the rendering pipeline (depicted on fig. 2) more efficient by skipping points that lie in free space.

In this section we describe the two network models used in our approach - the color and sampling network. The color network is a replica of NeRF’s fine network and models color as a function of position. Splitting up a ray into disjunct segments, or bins, the sampling network estimates the corresponding weights to sample points for further evaluation by the color network. With a suitable ray representation, the network is able to do such prediction using only one evaluation per ray, speeding up the rendering process. The weights ww predicted by the network are normalized such that w^i=wi ⁣/∑j=0Nwj\hat{w}_{i}={}^{w_{i}}\!/_{\sum_{j=0}^{N}w_{j}}. We observed that a sampling network of the same size and architecture as a NeRF network (8 hidden layers with 256 units and a skip connection to the fifth layer) has enough capacity to learn good samples.

2 Ray parameterisation

A single ray can be effectively parametrized by infinitely many combinations of the origin and direction (O,\vvd)(O,\vv{d}). The sampling network would not only have to predict the depth but also solve the ambiguity of the ray parameterization. To lift the computational burden from the network, we need to add a structure to the parametrizations and, with additional constraints, establish a bijection between rays and their representations. With an additional assumption that the initial point OO of each ray always lies outside of the sphere circumscribing the scene, the exact origin of the ray is not relevant anymore, and we can choose to replace it with any point on the ray or not to store it altogether.

Sphere intersection: One such way to unify the ray parametrizations is to represent it by two points of intersection with a sphere circumscribing the scene and the direction by the relative order of the points. One issue of such representation is that the line segments delimited by the intersection points are of a different length, depending on the distance of a ray from the center of the sphere. Uneven segment lengths create additional computational pressure on the network.

Additional samples: With the above parametrizations, there is still a considerable amount of complexity left for the network as it would have to solve for intersections in between the two potentially distant points. Similarly to work by Neff et al. we propose additional samples in between the already determined points from the previous sections. With the assumption that the majority of the volume is being located in the middle of the scene, we propose centred-logarithmic sampling, where the points are logarithmically sampled from the middle of the determined line segment towards the edges. The possible boundaries of the segment are described in sections 3.2 and 3.2. Mathematically speaking, if the two boundary points are AA and BB, the placed samples are:

where ss is formed by the concatenation of ll and uu:

The additional points form a neat definition of the ray bin boundaries we wish to predict the weights for. The last bin spans from the last boundary point up to infinity. Using points as bins boundaries additionally means that the two edge points A,BA,B has to lie outside of the rendered object.

We have experimented with various ray representations and combinations of boundary points as well as additional samples. For completeness, to centered-logarithmic sampling we also added equidistant samples. The overview can be found on fig. 3. As the higher number of bins in combination with centered-logarithmic sampling and segment ray representation are favored, we will keep this set-up for all other experiments unless specified otherwise.

3 Sampling distribution learning

NeRF consists of two jointly-trained, coarse and fine, networks. The NeRF model approximates the 5D function F(x,\vvd)→(c,σ)F(x,\vv{d})\rightarrow(c,\sigma) and estimates the color cc and the optical density σ\sigma of each point xx in the corresponding viewing direction \vvd\vv{d}. To train the sampling network, ground truth weights are obtained from the second, fine, network in a similar fashion they are obtained from the coarse network in the original rendering pipeline:

where δi=zi+1−zi\delta_{i}=z_{i+1}-z_{i} is the distance between two neighbouring samples. For each ray, NeRF is evaluated 64+128 times, therefore 196 tuples (zi,wi)(z_{i},w_{i}) are recorded. The zi=∣O−xi∣∣\vvd∣z_{i}=\frac{|O-x_{i}|}{|\vv{d}|} are the z-values capturing the distance of each point from the origin of the ray.

The sampling network estimates NN weights corresponding to the NN bins along the ray. As label weights are obtained from NeRF, inherently, there is a certain amount of noise involved. Therefore we perform Gaussian blurring of the ground truth weights along the ray with a fixed window size, where the window size covers all samples within the given Euclidean distance from the center of the window.

Furthermore, it is not guaranteed that the ground truth weights will be defined in the desired z-values (bin boundaries). Therefore, in order to find the correct label weights for our bins, a distribution matching is necessary as the training would not be possible without correctly placed label weights along the ray.

Given a distribution defined as NN z-value and weight pairs (zi,wi),∀  1≤i≤N(z_{i},w_{i}),\forall\;1\leq i\leq N, we wish to find the new label weights w^j\hat{w}_{j} corresponding with the desired bins. For the correct behavior of the network, certain properties of the distributions have to be preserved:

Any peak in the original distribution, however narrow, has to appear in the resampled variant, too. This problem becomes especially pronounced when the old and new z-values are not spaced out evenly.

Without a further subdivision, two segments of the same size on the ray are equally important for rendering if their largest weights are equal, irrespective of the number of samples in each segment.

To maintain these properties we propose max-resampling. Assuming a bin bjb_{j} is delimited by two consecutive z-values [z^j,z^j+1][\hat{z}_{j},\hat{z}_{j+1}], the weight corresponding to the bjb_{j} is chosen as the largest weight within the bin boundaries from the original distribution as depicted on fig. 5. Mathematically, for each bjb_{j}, the weight

is chosen, where wj′w^{\prime}_{j} is linear interpolation of the weights from the old distribution for the z-value z^j\hat{z}_{j}.

As the sampling network output is normalized, blurred and distribution-matched weights also have to be normalized. Then, using an MSE loss function, the resampled and normalized weights distribution can be used as the supervision for the sampling network.

Experiments

Our primary motivation was to extend NeRF and make the sampling more efficient. We therefore compare our approach with the original implementation of NeRF on multiple datasets that contain both coarse and fine structures, as well as semi-transparent volumes demonstrating the power of multilayer depth prediction. For each dataset we record the average PSNR and LPIPS based on AlexNet over 200 testing images.

Furthermore, we demonstrate the potential of our set-up to train the sampling network jointly with the color network, leading to more accurate depth and color predictions.

Lastly, we demonstrate how the sampling network can be utilized in more efficient retraining and fine-tuning of the color prediction network for scene modifications.

For the most part we use the original NeRF datasets Lego, Mic, Ship and Drums for testing and comparative analysis. We also add a few additional scenes to test the performance of our model in specific scenarios. The first additional scene, called Radiometer, contains multiple layers of volume, as it shows an object inside a glass enclosure.

Additionally, we generate a modified version of the Lego and Ship datasets: Lego Night and Ship Night, which contain the same geometry as the originals but are rendered under darker night-like conditions. These are to demonstrate the fine-tuning ability of our method.

2 Faster Rerendering novel views

The purpose of this experiment is to show that with the use of the sampling network, we can accurately decrease the necessary number of color network evaluations and render the scenes faster.

Experiment set up: For each dataset, we compare the best achieved NeRF results with our method. We report values for multiple rendering time brackets, as there is a natural trade-off between the rendering speed and the image quality as we tweak the number of rendering samples.

After 100 epochs of sampling network training, we select the checkpoint with the lowest validation error for color network fine-tuning. The color network is initialized with the weights of the fine NeRF network (used for depth dataset generation) and is trained on the RGB dataset for 300,000300,000 iterations. The checkpoint with the lowest validation error is chosen for the final evaluation on the test set. During fine-tuning, the sampling network’s parameters are frozen.

For training of both networks, the Adam optimizer with the initial learning rate of 5⋅10−45\cdot 10^{-4} and 5⋅10−55\cdot 10^{-5} with decay was used for the sampling and the color networks.

Figure 6 shows how the TermiNeRF can maintain the image quality of the original NeRF with only a fraction of total network forward passes. Figures 6(c) and 6(g) show how necessary a large number of samples is when we do not have any depth information. The number of samples 8+16 is chosen to match the total number (32) of forward passes through the networks (8 for the coarse network and 8+16 for the fine network) with the TermiNeRF 32.

3 Joint training and fine-tuning

The premise of this experiment is that as the depth dataset is obtained from the NeRF color network, the provided depth data should get better with fine-tuning of the network and, therefore, there is a potential for further training of the sampling network, too. This leads to a positive feedback loop where the sampling network gets more accurate with more precise depth data, and the color network learns to predict the color and volume distribution with better detail with more precisely placed samples.

The experiment set-up stays the same as the previous one, except that at the color network fine-tuning step, both color and sampling networks are updated in alternating iterations, where the depth supervision for the sampling network is provided by the color network. However, we can assign the opacity only to points evaluated with the color network. Assuming the points are suggested only by the predicted distribution, we would never acquire the true density value for regions of the ray that are left unsampled due to low weight, and the sampling network would never learn the additional geometry. To avoid this issue, for the joint fine-tuning, we sample 128 points from the predicted distribution with additional 64 equidistantly sampled points along the whole ray between near and far z-values.

Figure 7 shows that although fine-tuning itself is responsible for the majority of the improvement, we can still improve the image quality if we train the sampling network further.

A false-positive sampling network misprediction poses no issue for the color network, but the color network cannot fill in the gap caused by false-negative misprediction. Even though false-positives are occasional, they still happen and can be fixed only with more precise depth data and further training, which, in this case, happens jointly. Figure 8 focuses on image quality loss due to the sampling network issues.

It is important to note that even though we could train both networks jointly from the beginning, this process would be inefficient as the sampling network would be receiving mostly arbitrary training signal from the color network. Therefore, we first pre-train the color network. Once it converges, we train the sampling network and then jointly fine-tune both.

Fine-tuning for modified scene: Without prior depth information, during the training each ray has to be sampled along its full length from near to far boundaries. This process is just as inefficient as rendering. However, once the sampling network has been trained on one object, the color network can be quickly (fig. 9) retrained for a new scene that shares the same geometric structure, e.g., for a scene with different lighting conditions or different colors or textures.

4 Comparison with other methods

We provide a comparison of our work to the original NeRF implementation, DONeRF and other related methods . A more in-depth comparison with NeRF can be found in the table 2.

DONeRF comparison details: For comparison, we use our own implementation of DONeRF (as no official code is available), which we refer to as DONeRF*.

We replicate DONeRF’s discretization where we first split up the ray into NN bins and then discretize the depth value dsd_{s} along the ray and assign the classification value Cx,y(z)C_{x,y}(z) to bin zz of the ray (x,y)(x,y) as

In addition, we replicate the DONeRF’s blurring approach, where each ray is blurred as follows:

Therefore our re-implementation follows exactly the original DONeRF method, the only difference being that as we are training on 360∘360^{\circ} datasets, and the DONeRF focused on forward facing scenes, we omit the logarithmic sampling and warping from the implementation as we found these do not perform well on 360∘360^{\circ} scenes. The input to the network are therefore equidistantly sampled points and the ray segments (bins) are of a constant length.

Although ground truth depth maps are very precise, the single depth value per ray ultimately lacks the expressivity to capture a more complicated volume that contains multiple layers of depth, such as glass or fog. Such shortcoming can be seen on fig. 11 where the whole distribution of ray-volume intersections are necessary to capture the inside of the radiometer.

Additionally, our method does not model reflections, refractions, or any other effects where the direction of a ray deviates from a straight line. Even though some of the effects can be modeled through the view direction-dependent component of the input, reflections like in the fig. 11 can sometimes be over-expressed in the final render.

Numerical comparison: In this section, we put TermiNeRF into perspective and provide numerical results and comparisons with our best effort DONeRF* implementation, original NeRF, KiloNeRF, PlenOctrees and Autoint on 360∘360^{\circ} Lego and Ship datasets.

Conclusions

Neural Radiance Fields produce high-quality renders but are computationally expensive to render. In this paper, we propose a sampling network to focus only on the regions of a ray that yield a color, which effectively allows us to decrease the number of necessary forward passes through the network, speeding up the rendering pipeline ≈14×\approx 14\times. The network performs one-shot ray-volume intersection distribution estimations, keeping the pipeline efficient and fast. We use only RGB data to train the model, which makes this method not only versatile, but also more accurate and suitable for translucent surfaces.

Even though there are methods that render significantly faster than our approach, the benefits of our method are not limited to rendering/inference but also are applicable during training. For example, TermiNeRF can be trained end-to-end and offers shorter training times when fine-tuning scenes. As we show, this makes our method ideal for quickly adapting to an edited scene (for example, changing the texture, color, lighting or minor geometry modifications).

References