FastNeRF: High-Fidelity Neural Rendering at 200FPS
Stephan J. Garbin, Marek Kowalski, Matthew Johnson, Jamie Shotton, Julien Valentin
Introduction
Rendering scenes in real-time at photorealistic quality has long been a goal of computer graphics. Traditional approaches such as rasterization and ray-tracing often require significant manual effort in designing or pre-processing the scene in order to achieve both quality and speed. Recently, neural rendering has offered a disruptive alternative: involve a neural network in the rendering pipeline to output either images directly or to model implicit functions that represent a scene appropriately . Beyond rendering, some of these approaches implicitly reconstruct a scene from static or moving cameras , thereby greatly simplifying the traditional reconstruction pipelines used in computer vision.
One of the most prominent recent advances in neural rendering is Neural Radiance Fields (NeRF) which, given a handful of images of a static scene, learns an implicit volumetric representation of the scene that can be rendered from novel viewpoints. The rendered images are of high quality and correctly retain thin structures, view-dependent effects, and partially-transparent surfaces. NeRF has inspired significant follow-up work that has addressed some of its limitations, notably extensions to dynamic scenes , relighting , and incorporation of uncertainty .
One common challenge to all of the NeRF-based approaches is their high computational requirements for rendering images. The core of this challenge resides in NeRF’s volumetric scene representation. More than 100 neural network calls are required to render a single image pixel, which translates into several seconds being required to render low-resolution images on high-end GPUs. Recent explorations conducted with the aim of improving NeRF’s computational requirements reduced the render time by up to 50. While impressive, these advances are still a long way from enabling real-time rendering on consumer-grade hardware. Our work bridges this gap while maintaining quality, thereby opening up a wide range of new applications for neural rendering. Furthermore, our method could form the fundamental building block for neural rendering at high resolutions.
This factorized architecture, which we call FastNeRF, allows for independently caching the position-dependent and ray direction-dependent outputs. Assuming that and denote the number of bins for positions and ray directions respectively, caching NeRF would have a memory complexity of . In contrast, caching FastNeRF would have a complexity of . As a result of this reduced memory complexity, FastNeRF can be cached in the memory of a high-end consumer GPU, thus enabling very fast function lookup times that in turn lead to a dramatic increase in test-time performance.
While caching does consume a significant amount of memory, it is worth noting that current implementations of NeRF also have large memory requirements. A single forward pass of NeRF requires performing hundreds of forward passes through an eight layer hidden unit MLP per pixel. If pixels are processed in parallel for efficiency this consumes large amounts of memory, even at moderate resolutions. Since many natural scenes (e.g. a living room, a garden) are sparse, we are able to store our cache sparsely. In some cases this can make our method actually more memory efficient than NeRF.
The first NeRF-based system capable of rendering photorealistic novel views at FPS, thousands of times faster than NeRF.
A graphics-inspired factorization that can be compactly cached and subsequently queried to compute the pixel values in the rendered image.
A blueprint detailing how the proposed factorization can efficiently run on the GPU.
Related work
FastNeRF belongs to the family of Neural Radiance Fields methods and is trained to learn an implicit, compressed deep radiance map parameterized by position and view direction that provides color and density estimates. Our method differs from in the structure of the implicit model, changing it in such a way that positional and directional components can be discretized and stored in sparse 3D grids.
This also differentiates our method from models that use a discretized grid at training time, such as Neural Volumes or Deep Reflectance Volumes . Due to the memory requirements associated with generating and holding 3D volumes at training time, the output image resolution of is limited by the maximum volume size of . In contrast, our method uses volumes as large as .
Subsequent to , the problem of parameter estimation for an MLP with low dimensional input coordinates was addressed using Fourier feature encodings. After use in , this was popularized in and explored in greater detail by .
Neural Radiance Fields (NeRF) first showed convincing compression of a light field utilizing Fourier features. Using an entirely implicit model, NeRF is not bound to any voxelized grid, but only to a specific domain. Despite impressive results, one disadvantage of this method is that a complex MLP has to be called for every sample along each ray, meaning hundreds of MLP invocations per image pixel.
One way to speed up NeRF’s inference is to scale to multiple processors. This works particularly well because every pixel in NeRF can be computed independently. For an pixel image, JaxNeRF achieves an inference speed of 20.77 seconds on an Nvidia Tesla V100 GPU, 2.65 seconds on 8 V100 GPUS, and 0.35 seconds on 128 second generation Tensor Processing Units. Similarly, the volumetric integration domain can be split and separate models used for each part. This approach is taken in the Decomposed Radiance Fields method , where an integration domain is divided into Voronoi cells. This yields better test scores, and a speedup of up to .
A different way to increase efficiency is to realize that natural scenes tend to be volumetrically sparse. Thus, efficiency can be gained by skipping empty regions of space. This amounts to importance sampling of the occupancy distribution of the integration domain. One approach is to use a voxel grid in combination with the implicit function learned by the MLP, as proposed in Neural Sparse Voxel Fields (NSVFs) and Neural Geometric Level of Detail (for SDFs only), where a dynamically constructed sparse octree is used to represent scene occupancy. As a network still has to be queried inside occupied voxels, however, NSVFs takes between 1 and 4 seconds to render an image, with decreases to PSNR at the lower end of those timings. Our method is orders of magnitude faster in a similar scenario.
Another way to represent the importance distribution is via depth prediction for each pixel. This approach is taken in , which is concurrent to ours and achieves roughly 15FPS for images at reduced quality or roughly half that for quality comparable to NeRF.
Orthogonal to this, AutoInt showed that a neural network can be used to approximate the integrals along each ray with far fewer samples. While significantly faster than NeRF, this still does not provide interactive frame rates.
What differentiates our method from those described above is that FastNeRF’s proposed decomposition, and subsequent caching, lets us avoid calls to an MLP at inference time entirely. This makes our method faster in absolute terms even on a single machine.
It is worth noting that our method does not address training speed, as for example . In that case, the authors propose improving the training speed of NeRF models by finding initialisation through meta learning.
Finally, there are orthogonal neural rendering strategies capable of fast inference, such as Neural Point-Based Graphics , Mixture of Volumetric Primitives or Pulsar , which use forms of geometric primitives and rasterization. In this work, we only deal with NeRF-like implicit functional representations paired with ray tracing.
Method
In this section we describe FastNeRF, a method that is 3000 times faster than the original Neural Radiance Fields (NeRF) system (Section 3.1). This breakthrough allows for rendering high-resolution photorealistic images at over 200Hz on high-end consumer hardware. The core insight of our approach (Section 3.2) consists of factorizing NeRF into two neural networks: a position-dependent network that produces a deep radiance map and a direction-dependent network that produces weights. The inner product of the weights and the deep radiance map estimates the color in the scene at the specified position and as seen from the specified direction. This architecture, which we call FastNeRF, can be efficiently cached (Section 3.3), significantly improving test time efficiency whilst preserving the visual quality of NeRF. See Figure 1 for a comparison of the NeRF and FastNeRF network architectures.
In order to render a single image pixel, a ray is cast from the camera center, passing through that pixel and into the scene. We denote the direction of this ray as . A number of 3D positions are then sampled along the ray between its near and far bounds defined by the camera parameters. The neural network is evaluated at each position and ray direction to produce color and transparency . These intermediate outputs are then integrated as follows to produce the final pixel color :
where is the transmittance and is the distance between the samples. Since depends on ray directions, NeRF has the ability to model viewpoint-dependent effects such as specular reflections, which is one key dimension in which NeRF improves upon traditional 3D reconstruction methods.
Training a NeRF network requires a set of images of a scene as well as the extrinsic and intrinsic parameters of the cameras that captured the images. In each training iteration, a subset of pixels from the training images are chosen at random and for each pixel a 3D ray is generated. Then, a set of samples is selected along each ray and the pixel color is computed using Equation (1). The training loss is the mean squared difference between and the ground truth pixel value. For further details please refer to .
While NeRF renders photorealistic images, it requires calling a large number of times to produce a single image. With the default number of samples per pixel proposed in NeRF , nearly 400 million calls to are required to compute a single high definition (1080p) image. Moreover the intermediate outputs of this process would take hundreds of gigabytes of memory. Even on high-end consumer GPUs, this constrains the original method to be executed over several batches even for medium resolution () images, leading to additional computational overheads.
2 Factorized Neural Radiance Fields
Taking a step away from neural rendering for a moment, we recall that in traditional computer graphics, the rendering equation is an integral of the form
where is the radiance leaving the point in direction , is the reflectance function capturing the material properties at position , describes the amount of light reaching from direction , and corresponds to the direction of the surface normal at . Given its practical importance, evaluating this integral efficiently has been a subject of active research for over three decades . One efficient way to evaluate the rendering equation is to approximate and using spherical harmonics . In this case, evaluating the integral boils down to a dot product between the coefficients of both spherical harmonics approximations.
The position-dependent and direction-dependent functions of FastNeRF are defined as follows:
where are -dimensional vectors that form a deep radiance map describing the view-dependent radiance at position . The output of , , is a -dimensional vector of weights for the components of the deep radiance map. The inner product of the weights and the deep radiance map
results in the estimated color at position observed from direction .
3 Caching
Given the large number of samples that need to be evaluated to render a single pixel, the cost of computing dominates the total cost of NeRF rendering. Thus, to accelerate NeRF, one could attempt to reduce the test-time cost of by caching its outputs for a set of inputs that cover the space of the scene. The cache can then be evaluated at a fraction of the time it takes to compute .
For a trained NeRF model, we can define a bounding box that covers the entire scene captured by NeRF. We can then uniformly sample values for each of the 3 world-space coordinates within the bounds of . Similarly, we uniformly sample values for each of the ray direction coordinates with and . The cache is then generated by computing for each combination of sampled and .
The size of such a cache for a standard NeRF model with and densely stored 16-bit floating point values is approximately 5600 Terabytes. Even for highly sparse volumes, where one would only need to keep 1% of the outputs, the size of the cache would still severely exceed the memory capacity of consumer-grade hardware. This huge memory requirement is caused by the usage of both and as input to . A separate output needs to be saved for each combination of and , resulting in a memory complexity in the order of .
Our FastNeRF architecture makes caching feasible. For the size of two dense caches holding and would be approximately 54 GB. For moderately sparse volumes, where 30% of space is occupied, the memory requirement is low enough to fit into either the CPU or GPU memory of consumer-grade machines. In practice, the choice and depends on the scene size and the expected image resolution. For many scenarios a smaller cache of is sufficient, lowering the memory requirements further. Please see the supplementary materials for formulas used to calculate cache sizes for both network architectures.
Implementation
Experiments
We evaluate our method quantitatively and qualitatively on the Realistic 360 Synthetic and Local Light Field Fusion (LLFF) (with additions from ) datasets used in the original NeRF paper. While the NeRF synthetic dataset consists of 360 degree views of complex objects, the LLFF dataset consist of forward-facing scenes, with fewer images. In all comparisons with NeRF we use the same training parameters as described in the original paper .
To assess the rendering quality of FastNeRF we compare its outputs to GT using Peak Signal to Noise Ratio (PSNR), Structural Similarity (SSIM) and perceptual LPIPS . We use the same metrics in our ablation study, where we evaluate the effects of changing the number of components and the size of the cache. All speed comparisons are run on one machine that is equipped with an Nvidia RTX 3090 GPU.
Rendering quality: Because we adopt the inputs, outputs, and training of NeRF, we retain the compatibility with ray-tracing and the ability to only specify loose volume bounds. At the same time, our factorization and caching do not affect the character of the rendered images to a large degree. Please see Figure 3 and Figure 4 for qualitative comparisons, and Table 1 for a quantitative evaluation. Note that some aliasing artefacts can appear since our model is cached to a grid after training. This explains the decrease across all metrics as a function of the grid resolution as seen in Table 1. However, both NeRF and FastNeRF mitigate multi-view inconsistencies, ghosting and other such artifacts from multi-view reconstruction. Table 1 demonstrates that at high enough cache resolutions, our method is capable of the same visual quality as NeRF. Figure 3 and Figure 4 further show that smaller caches appear ‘pixelated’ but retain the overall visual characteristics of the scene. This is similar to how different levels of detail are used in computer graphics , and an important innovation on the path towards neurally rendered worlds.
Cache Resolution: As shown in Table 1, our method matches or outperforms NeRF on the dataset of synthetic objects at cache resolution. Because our method is fast enough to visit every voxel (as opposed to the fixed sample count used in NeRF), it can sometimes achieve better scores by not missing any detail. We observe that a cache of is a good trade-off between perceptual quality, memory and rendering speed for the synthetic dataset. For the LLFF dataset, we found a cache to work best. For the Ship scene, an ablation over the grid size, memory, and PSNR is shown in Table 3.
For the view-filling LLFF scenes at pixels, our method sees a slight decrease in metrics when cached at , but still produces qualitatively compelling results as shown in Figure 4, where intricate detail is clearly preserved.
Rendering Speed: When using a grid resolution of , FastNeRF is on average more than faster than NeRF at largely the same perceptual quality. See Table 2 for a breakdown of run-times across several resolutions of the cache and Figure 2 for a comparison to other NeRF extensions in terms of speed. Note that we log the time it takes our method to fill an RGBA buffer of the image size with the final pixel values, whereas the baseline implementation needs to perform various steps to reshape samples back into images. Our CUDA kernels are not highly optimized, favoring flexibility over maximum performance. It is reasonable to assume that advanced code optimization and further compression of the grid values could lead to further reductions in compute time.
Number of Components: While we can see a theoretical improvement of roughly dB going from to components, we find that or components are sufficient for most scenes. As shown in Table 3, our cache can bring out fine details that are lost when only a fixed sample count is used, compensating for the difference. Table 3 also shows that more components tend to be increasingly sparse when cached, which compensates somewhat for the additional memory usage. Please see the supplementary material for more detailed results on other scenes.
Application
We demonstrate a proof-of-concept application of FastNeRF for a telepresence scenario by first gathering a dataset of a person performing facial expressions for about 20s in a multi-camera rig consisting of 32 calibrated cameras. We then fit a 3D face model similar to FLAME to obtain expression parameters for each frame. Since the captured scene is not static, we train a FastNeRF model of the scene jointly with a deformation model conditioned on the expression data. The deformation model takes the samples as input and outputs their updated positions that live in a canonical frame of reference modeled by FastNeRF.
We show example outputs produced using this approach in Figure 5. This method allows us to render pixel images of the face at 30 fps on a single Nvidia Tesla V100 GPU – around 50 times faster than a setup with a NeRF model. At the resolution we achieve, the face expression is clearly visible and the high frame rate allows for real-time rendering, opening the door to telepresence scenarios. The main limitation in terms of speed and image resolution is the deformation model, which needs to be executed for a large number of samples. We employ a simple pruning method, detailed in the supplementary material, to reduce the number of samples.
Finally, note that while we train this proof-of-concept method on a dataset captured using multiple cameras, the approach can be extended to use only a single camera, similarly to .
Since FastNeRF uses the same set of inputs and outputs as those used in NeRF, it is applicable to many of NeRF’s extensions. This allows for accelerating existing approaches for: reconstruction of dynamic scenes , single-image reconstruction , quality improvement and others . With minor modifications, FastNeRF can be applied to even more methods allowing for control over illumination and incorporation of uncertainty .
Conclusion
In this paper we presented FastNeRF, a novel extension to NeRF that enables the rendering of photorealistic images at 200Hz and more on consumer hardware. We achieve significant speedups over NeRF and competing methods by factorizing its function approximator, enabling a caching step that makes rendering memory-bound instead of compute-bound. This allows for the application of NeRF to real-time scenarios.
References
Overview
We use the supplementary material to provide more detailed results and algorithms. We urge the reader to see our video, in which we show results for all scenes we tested on along with providing an intuition for our method. On a practical note, we also provide guidance on getting smooth results for anyone adapting our method to their own work.
Please note that parameters chosen throughout this work for the rendering algorithm favour the highest possible quality. It is possible to significantly increase rendering speeds by sacrificing a small amount of visual quality for real-world applications.
Effect of modified Fourier feature encoding on metrics: Using an layer hidden unit MLP for position, and a layer MLP for direction, we can obtain an average PSNR of for , for , for on synthetic data, with smoothness of results being reflected in these numbers.
Effect of small direction cache on metrics: Using a smaller direction cache has a negligible impact on metrics, occasionally actually improving them. This shows that the direction-dependent effects required by FastNeRF are low in frequency. For the scenes where we found we needed to use a smaller cache for moving images, the PSNR differences caused by using a smaller directional cache are: Materials () (28.885dB →28.874dB); Drums () (23.745dB →23.836dB); Lego () (32.275dB →32.155dB); Mic () (31.765dB →31.667dB); Hotdog () (34.722dB →34.644dB); Ficus () (27.792dB →28.193dB).
2 Training & Detailed Results
For training, we use the same frequency encoding, noise perturbation and learning rate decay (starting at ) as . For the NeRF 360 Synthetic dataset, we use and samples for the coarse and fine networks, respectively. This becomes and for the LLFF scenes. Note that the coarse networks are discarded and not used in the caching stage - the cached density (or rather, the mesh derived from it) serves as the importance distribution during rendering. We sample random rays per gradient descent step, and use Adam as our optimizer of choice, with , .
In the paper, we compute test metrics on a random subset of 20 images per scene from the test sets of the NeRF 360 Synthetic dataset, and all available test images for the LLFF data. We use the evaluation set only to select the best iteration, which is 300K for all scenes, indicating identical convergence trends as NeRF . For completeness, we show results for all scenes using the full number of test images in Table 4. We note that the results and trends are in agreement with the sub-sampled test set. Please see Figure 7 and beyond for a qualitative comparison in addition to the video. We include further ablations on the number of components in Table 5 and Table 6.
3 Meshing & Rendering
One of the key advantages of working with sparse octrees to store the caches required by our method is that they can serve as a basis for accelerating the rendering process. Using raytracing, we can quickly terminate rays that miss the neurally rendered object(s) all together. For the rays that hit occupied voxels, and so require integrating the volume, we can use raytracing to skip empty space efficiently, saving a significant amount of queries. This matters because caching makes our method memory instead of compute bound. Computing the inner product of the components and weights as well as tracking transmittance for each ray is fast on modern GPUs. Grid look-ups, which require reading from the GPUs RAM on the other hand, are expensive.
While we could use a hierarchical digital differential analyzer such as to determine ray grid intersections, we opt to trace rays against a collision mesh derived from the volume instead, for three reasons. First, this allows us to take advantage of hardware acceleration for the BVH and ray-triangle intersections on modern hardware. Second, we can down/resample or deform the meshes easily using existing tools. Finally, meshes could be used to ‘fake’ shadows, bouncelight and other effects which can be composited over the neurally rendered content.
We derive the meshes by converting the density cache to a sign distance function (thresholding at ), optionally downsampling it before meshing with marching cubes, and optionally simplifying the resulting geometry using standard techniques. We show the resulting geometry for the Lego scene in Figure 6. Note that more aggressive thresholding of the volume or remeshing is easily possible and can lead to significant performance benefits in practise as more rays can be culled early and mesh complexity reduced.
For rendering, we follow the same method as NeRF, using the Beer-Lambert Law to model radiance extinction. Once the contribution of new samples for a ray reaches , we terminate a ray’s rendering kernel. Note that this can be relaxed for greater performance. Another way to accelerate the rendering process is to visit fewer voxels based on radiance or transmittance. For all paper experiments however, we never skip occupied space, or optimise any parameters per scene.
4 Cache size calculation
In this section we show the formulas we used to estimate cache sizes for NeRF and FastNeRF respectively.
where are the sizes of the stored transparency and RGB values in bits and is the inverse volume sparsity, where would indicate a dense volume.
where are the sizes of the stored values and the weights .
For , stored as individual bytes and all other values stored as half-precision floats the cache size would be
Even for highly sparse volumes, where , the cache would take nearly 60TB, which exceeds the memory capacity of consumer-grade hardware. The high memory requirement makes this approach impractical when used with the standard NeRF architecture.
5 Cache size calculation
In this section we show the formulas we used to estimate cache sizes for NeRF and FastNeRF respectively.
where are the sizes of the stored transparency and RGB values in bits and is the inverse volume sparsity, where would indicate a dense volume. Just as in the main paper , correspond to the number of bins per dimension in the position and direction dependent caches respectively.
where are the sizes of the stored values and the weights .
For , stored as individual bytes and all other values stored as half-precision floats the cache size would be
6 Details on Applications
In our proof-of-concept telepresence scenario we use a deformation field network that modifies the input sample positions. This network is similar to the position-dependent network used in FastNeRF, but smaller - a 6-layer MLP with 64 units in each layer. The input is a sample position and an expression vector. The sample position is processed with positional encoding identically to how the FastNeRF input position is processed. The output is an offset that is applied to the input sample position. When training with we add an L2 regularizer that constrains to be an identity transform if a neutral expression is passed as input. We find that this solution stabilizes training and improves results.
In addition to , we also use a non-trainable deformation field that moves the samples in the head region according to the movement of the 3D model bones that represent the shoulders and the neck. This additional deformation accounts for large movements of the head and the shoulders, which reduces the load on and improves quality. The two deformation fields are composed with the position-dependent network as follows:
where is the expression vector. While we could use at test-time as well, we remove it to improve performance.
The training procedure in this scenario is identical to that used in FastNeRF but with 64 samples in the coarse stage and an additional 64 in the fine stage. The position-dependent network has 384 units in each layer, while the view-dependent network uses only 32 units and 2 layers. The view-dependent network is very small as we believe the scene’s illumination is simple and fairly uniform.
The main limitation of this approach in terms of speed is the need to call for all the input samples. In order to mitigate this, at test-time we implement a grid-based sample pruner. For each position in a grid around the face region we compute the density
for 50 randomly chosen samples of . For each grid position we record the minimum and maximum encountered in a sparse volume. At test-time we use this volume to prune the inputs to where the maximum is 0 and where the rays would get saturated, which we estimate using the minimum density values from the volume. This approach reduces the runtime of by more than a factor of 2.
To further reduce the impact of on the runtime we only use 43 samples in the coarse stage and remove the fine stage altogether at test time. To offset this low sample count we add additional samples that are linearly interpolated between the samples generated by the deformation network
The additional samples can be evaluated very cheaply as they can be looked-up in the FastNeRF cache. Note that since the deformation network changes the positions of individual samples along the rays, the rays are no longer straight. This means that in this scenario we cannot use the hardware-accelerated ray tracing procedure described in the main paper, though the pruning method described above serves a similar role.
Even though the steps above significantly reduce the runtime of , we are still limited to rendering pixel images with a reduced sample count if we want to maintain 30FPS with the deformation network. We observed that if this framerate constraint was to be disregarded, the approach described above is able to generate images at significantly better quality. Thus, we believe that a faster deformation approach remains an important goal for future work.