FastNeRF: High-Fidelity Neural Rendering at 200FPS

Stephan J. Garbin, Marek Kowalski, Matthew Johnson, Jamie Shotton, Julien Valentin

Introduction

Rendering scenes in real-time at photorealistic quality has long been a goal of computer graphics. Traditional approaches such as rasterization and ray-tracing often require significant manual effort in designing or pre-processing the scene in order to achieve both quality and speed. Recently, neural rendering has offered a disruptive alternative: involve a neural network in the rendering pipeline to output either images directly or to model implicit functions that represent a scene appropriately . Beyond rendering, some of these approaches implicitly reconstruct a scene from static or moving cameras , thereby greatly simplifying the traditional reconstruction pipelines used in computer vision.

One of the most prominent recent advances in neural rendering is Neural Radiance Fields (NeRF) which, given a handful of images of a static scene, learns an implicit volumetric representation of the scene that can be rendered from novel viewpoints. The rendered images are of high quality and correctly retain thin structures, view-dependent effects, and partially-transparent surfaces. NeRF has inspired significant follow-up work that has addressed some of its limitations, notably extensions to dynamic scenes , relighting , and incorporation of uncertainty .

One common challenge to all of the NeRF-based approaches is their high computational requirements for rendering images. The core of this challenge resides in NeRF’s volumetric scene representation. More than 100 neural network calls are required to render a single image pixel, which translates into several seconds being required to render low-resolution images on high-end GPUs. Recent explorations conducted with the aim of improving NeRF’s computational requirements reduced the render time by up to 50×\times. While impressive, these advances are still a long way from enabling real-time rendering on consumer-grade hardware. Our work bridges this gap while maintaining quality, thereby opening up a wide range of new applications for neural rendering. Furthermore, our method could form the fundamental building block for neural rendering at high resolutions.

This factorized architecture, which we call FastNeRF, allows for independently caching the position-dependent and ray direction-dependent outputs. Assuming that kk and ll denote the number of bins for positions and ray directions respectively, caching NeRF would have a memory complexity of O(k3l2)\mathcal{O}(k^{3}l^{2}). In contrast, caching FastNeRF would have a complexity of O(k3∗(1+D∗3)+l2∗D)\mathcal{O}(k^{3}*(1+D*3)+l^{2}*D). As a result of this reduced memory complexity, FastNeRF can be cached in the memory of a high-end consumer GPU, thus enabling very fast function lookup times that in turn lead to a dramatic increase in test-time performance.

While caching does consume a significant amount of memory, it is worth noting that current implementations of NeRF also have large memory requirements. A single forward pass of NeRF requires performing hundreds of forward passes through an eight layer 256256 hidden unit MLP per pixel. If pixels are processed in parallel for efficiency this consumes large amounts of memory, even at moderate resolutions. Since many natural scenes (e.g. a living room, a garden) are sparse, we are able to store our cache sparsely. In some cases this can make our method actually more memory efficient than NeRF.

The first NeRF-based system capable of rendering photorealistic novel views at 200200FPS, thousands of times faster than NeRF.

A graphics-inspired factorization that can be compactly cached and subsequently queried to compute the pixel values in the rendered image.

A blueprint detailing how the proposed factorization can efficiently run on the GPU.

Related work

FastNeRF belongs to the family of Neural Radiance Fields methods and is trained to learn an implicit, compressed deep radiance map parameterized by position and view direction that provides color and density estimates. Our method differs from in the structure of the implicit model, changing it in such a way that positional and directional components can be discretized and stored in sparse 3D grids.

This also differentiates our method from models that use a discretized grid at training time, such as Neural Volumes or Deep Reflectance Volumes . Due to the memory requirements associated with generating and holding 3D volumes at training time, the output image resolution of is limited by the maximum volume size of 1283128^{3}. In contrast, our method uses volumes as large as 102431024^{3}.

Subsequent to , the problem of parameter estimation for an MLP with low dimensional input coordinates was addressed using Fourier feature encodings. After use in , this was popularized in and explored in greater detail by .

Neural Radiance Fields (NeRF) first showed convincing compression of a light field utilizing Fourier features. Using an entirely implicit model, NeRF is not bound to any voxelized grid, but only to a specific domain. Despite impressive results, one disadvantage of this method is that a complex MLP has to be called for every sample along each ray, meaning hundreds of MLP invocations per image pixel.

One way to speed up NeRF’s inference is to scale to multiple processors. This works particularly well because every pixel in NeRF can be computed independently. For an 800×800800\times 800 pixel image, JaxNeRF achieves an inference speed of 20.77 seconds on an Nvidia Tesla V100 GPU, 2.65 seconds on 8 V100 GPUS, and 0.35 seconds on 128 second generation Tensor Processing Units. Similarly, the volumetric integration domain can be split and separate models used for each part. This approach is taken in the Decomposed Radiance Fields method , where an integration domain is divided into Voronoi cells. This yields better test scores, and a speedup of up to 3×3\times.

A different way to increase efficiency is to realize that natural scenes tend to be volumetrically sparse. Thus, efficiency can be gained by skipping empty regions of space. This amounts to importance sampling of the occupancy distribution of the integration domain. One approach is to use a voxel grid in combination with the implicit function learned by the MLP, as proposed in Neural Sparse Voxel Fields (NSVFs) and Neural Geometric Level of Detail (for SDFs only), where a dynamically constructed sparse octree is used to represent scene occupancy. As a network still has to be queried inside occupied voxels, however, NSVFs takes between 1 and 4 seconds to render an 8002800^{2} image, with decreases to PSNR at the lower end of those timings. Our method is orders of magnitude faster in a similar scenario.

Another way to represent the importance distribution is via depth prediction for each pixel. This approach is taken in , which is concurrent to ours and achieves roughly 15FPS for 8002800^{2} images at reduced quality or roughly half that for quality comparable to NeRF.

Orthogonal to this, AutoInt showed that a neural network can be used to approximate the integrals along each ray with far fewer samples. While significantly faster than NeRF, this still does not provide interactive frame rates.

What differentiates our method from those described above is that FastNeRF’s proposed decomposition, and subsequent caching, lets us avoid calls to an MLP at inference time entirely. This makes our method faster in absolute terms even on a single machine.

It is worth noting that our method does not address training speed, as for example . In that case, the authors propose improving the training speed of NeRF models by finding initialisation through meta learning.

Finally, there are orthogonal neural rendering strategies capable of fast inference, such as Neural Point-Based Graphics , Mixture of Volumetric Primitives or Pulsar , which use forms of geometric primitives and rasterization. In this work, we only deal with NeRF-like implicit functional representations paired with ray tracing.

Method

In this section we describe FastNeRF, a method that is 3000 times faster than the original Neural Radiance Fields (NeRF) system (Section 3.1). This breakthrough allows for rendering high-resolution photorealistic images at over 200Hz on high-end consumer hardware. The core insight of our approach (Section 3.2) consists of factorizing NeRF into two neural networks: a position-dependent network that produces a deep radiance map and a direction-dependent network that produces weights. The inner product of the weights and the deep radiance map estimates the color in the scene at the specified position and as seen from the specified direction. This architecture, which we call FastNeRF, can be efficiently cached (Section 3.3), significantly improving test time efficiency whilst preserving the visual quality of NeRF. See Figure 1 for a comparison of the NeRF and FastNeRF network architectures.

In order to render a single image pixel, a ray is cast from the camera center, passing through that pixel and into the scene. We denote the direction of this ray as d\bm{d}. A number of 3D positions (p1, ⁣⋯ ,pN)(\bm{p}_{1},\dotsi,\bm{p}_{N}) are then sampled along the ray between its near and far bounds defined by the camera parameters. The neural network FNeRF\mathcal{F}_{NeRF} is evaluated at each position pi\bm{p}_{i} and ray direction d\bm{d} to produce color ci\bm{c}_{i} and transparency σi\sigma_{i}. These intermediate outputs are then integrated as follows to produce the final pixel color c^\bm{\hat{c}}:

where Ti=exp⁡(−∑j=ii−1σjδj)T_{i}=\exp(-\sum_{j=i}^{i-1}\sigma_{j}\delta_{j}) is the transmittance and δi=(pi+1−pi)\delta_{i}=\left(\bm{p}_{i+1}-\bm{p}_{i}\right) is the distance between the samples. Since FNeRF\mathcal{F}_{NeRF} depends on ray directions, NeRF has the ability to model viewpoint-dependent effects such as specular reflections, which is one key dimension in which NeRF improves upon traditional 3D reconstruction methods.

Training a NeRF network requires a set of images of a scene as well as the extrinsic and intrinsic parameters of the cameras that captured the images. In each training iteration, a subset of pixels from the training images are chosen at random and for each pixel a 3D ray is generated. Then, a set of samples is selected along each ray and the pixel color c^\bm{\hat{c}} is computed using Equation (1). The training loss is the mean squared difference between c^\bm{\hat{c}} and the ground truth pixel value. For further details please refer to .

While NeRF renders photorealistic images, it requires calling FNeRF\mathcal{F}_{NeRF} a large number of times to produce a single image. With the default number of samples per pixel N=192N=192 proposed in NeRF , nearly 400 million calls to FNeRF\mathcal{F}_{NeRF} are required to compute a single high definition (1080p) image. Moreover the intermediate outputs of this process would take hundreds of gigabytes of memory. Even on high-end consumer GPUs, this constrains the original method to be executed over several batches even for medium resolution (800×800800\times 800) images, leading to additional computational overheads.

2 Factorized Neural Radiance Fields

Taking a step away from neural rendering for a moment, we recall that in traditional computer graphics, the rendering equation is an integral of the form

where Lo(p,d)L_{o}(\bm{p},\bm{d}) is the radiance leaving the point p\bm{p} in direction d\bm{d}, fr(p,d,ωi)f_{r}(\bm{p},\bm{d},\bm{\omega_{i}}) is the reflectance function capturing the material properties at position p\bm{p}, Li(p,ωi)L_{i}(\bm{p},\bm{\omega_{i}}) describes the amount of light reaching p\bm{p} from direction ωi\bm{\omega_{i}}, and n\bm{n} corresponds to the direction of the surface normal at p\bm{p}. Given its practical importance, evaluating this integral efficiently has been a subject of active research for over three decades . One efficient way to evaluate the rendering equation is to approximate fr(p,d,ωi)f_{r}(\bm{p},\bm{d},\bm{\omega_{i}}) and Li(p,ωi)L_{i}(\bm{p},\bm{\omega_{i}}) using spherical harmonics . In this case, evaluating the integral boils down to a dot product between the coefficients of both spherical harmonics approximations.

The position-dependent and direction-dependent functions of FastNeRF are defined as follows:

where u,v,w\bm{u},\bm{v},\bm{w} are DD-dimensional vectors that form a deep radiance map describing the view-dependent radiance at position p\bm{p}. The output of Fdir\mathcal{F}_{dir}, β\bm{\beta}, is a DD-dimensional vector of weights for the DD components of the deep radiance map. The inner product of the weights and the deep radiance map

results in the estimated color c=(r,g,b)\bm{c}=(r,g,b) at position p\bm{p} observed from direction d\bm{d}.

3 Caching

Given the large number of samples that need to be evaluated to render a single pixel, the cost of computing F\mathcal{F} dominates the total cost of NeRF rendering. Thus, to accelerate NeRF, one could attempt to reduce the test-time cost of F\mathcal{F} by caching its outputs for a set of inputs that cover the space of the scene. The cache can then be evaluated at a fraction of the time it takes to compute F\mathcal{F}.

For a trained NeRF model, we can define a bounding box V\mathcal{V} that covers the entire scene captured by NeRF. We can then uniformly sample kk values for each of the 3 world-space coordinates (x,y,z)=p(x,y,z)=\bm{p} within the bounds of V\mathcal{V}. Similarly, we uniformly sample ll values for each of the ray direction coordinates (θ,ϕ)=d(\theta,\phi)=\bm{d} with θ∈⟨0,π⟩\theta\in\langle 0,\pi\rangle and ϕ∈⟨0,2π⟩\phi\in\langle 0,2\pi\rangle. The cache is then generated by computing F\mathcal{F} for each combination of sampled p\bm{p} and d\bm{d}.

The size of such a cache for a standard NeRF model with k=l=1024k=l=1024 and densely stored 16-bit floating point values is approximately 5600 Terabytes. Even for highly sparse volumes, where one would only need to keep 1% of the outputs, the size of the cache would still severely exceed the memory capacity of consumer-grade hardware. This huge memory requirement is caused by the usage of both d\bm{d} and p\bm{p} as input to FNeRF\mathcal{F}_{NeRF}. A separate output needs to be saved for each combination of d\bm{d} and p\bm{p}, resulting in a memory complexity in the order of O(k3l2)\mathcal{O}(k^{3}l^{2}).

Our FastNeRF architecture makes caching feasible. For k=l=1024,D=8k=l=1024,D=8 the size of two dense caches holding {σ,(u,v,w))}\{\sigma,(\bm{u},\bm{v},\bm{w}))\} and β\bm{\beta} would be approximately 54 GB. For moderately sparse volumes, where 30% of space is occupied, the memory requirement is low enough to fit into either the CPU or GPU memory of consumer-grade machines. In practice, the choice kk and ll depends on the scene size and the expected image resolution. For many scenarios a smaller cache of k=512,l=256k=512,l=256 is sufficient, lowering the memory requirements further. Please see the supplementary materials for formulas used to calculate cache sizes for both network architectures.

Implementation

Experiments

We evaluate our method quantitatively and qualitatively on the Realistic 360 Synthetic and Local Light Field Fusion (LLFF) (with additions from ) datasets used in the original NeRF paper. While the NeRF synthetic dataset consists of 360 degree views of complex objects, the LLFF dataset consist of forward-facing scenes, with fewer images. In all comparisons with NeRF we use the same training parameters as described in the original paper .

To assess the rendering quality of FastNeRF we compare its outputs to GT using Peak Signal to Noise Ratio (PSNR), Structural Similarity (SSIM) and perceptual LPIPS . We use the same metrics in our ablation study, where we evaluate the effects of changing the number of components DD and the size of the cache. All speed comparisons are run on one machine that is equipped with an Nvidia RTX 3090 GPU.

Rendering quality: Because we adopt the inputs, outputs, and training of NeRF, we retain the compatibility with ray-tracing and the ability to only specify loose volume bounds. At the same time, our factorization and caching do not affect the character of the rendered images to a large degree. Please see Figure 3 and Figure 4 for qualitative comparisons, and Table 1 for a quantitative evaluation. Note that some aliasing artefacts can appear since our model is cached to a grid after training. This explains the decrease across all metrics as a function of the grid resolution as seen in Table 1. However, both NeRF and FastNeRF mitigate multi-view inconsistencies, ghosting and other such artifacts from multi-view reconstruction. Table 1 demonstrates that at high enough cache resolutions, our method is capable of the same visual quality as NeRF. Figure 3 and Figure 4 further show that smaller caches appear ‘pixelated’ but retain the overall visual characteristics of the scene. This is similar to how different levels of detail are used in computer graphics , and an important innovation on the path towards neurally rendered worlds.

Cache Resolution: As shown in Table 1, our method matches or outperforms NeRF on the dataset of synthetic objects at 102431024^{3} cache resolution. Because our method is fast enough to visit every voxel (as opposed to the fixed sample count used in NeRF), it can sometimes achieve better scores by not missing any detail. We observe that a cache of 5123512^{3} is a good trade-off between perceptual quality, memory and rendering speed for the synthetic dataset. For the LLFF dataset, we found a 7683768^{3} cache to work best. For the Ship scene, an ablation over the grid size, memory, and PSNR is shown in Table 3.

For the view-filling LLFF scenes at 504×378504\times 378 pixels, our method sees a slight decrease in metrics when cached at 7683768^{3}, but still produces qualitatively compelling results as shown in Figure 4, where intricate detail is clearly preserved.

Rendering Speed: When using a grid resolution of 7683768^{3}, FastNeRF is on average more than 3000×3000\times faster than NeRF at largely the same perceptual quality. See Table 2 for a breakdown of run-times across several resolutions of the cache and Figure 2 for a comparison to other NeRF extensions in terms of speed. Note that we log the time it takes our method to fill an RGBA buffer of the image size with the final pixel values, whereas the baseline implementation needs to perform various steps to reshape samples back into images. Our CUDA kernels are not highly optimized, favoring flexibility over maximum performance. It is reasonable to assume that advanced code optimization and further compression of the grid values could lead to further reductions in compute time.

Number of Components: While we can see a theoretical improvement of roughly 0.50.5dB going from 88 to 1616 components, we find that 66 or 88 components are sufficient for most scenes. As shown in Table 3, our cache can bring out fine details that are lost when only a fixed sample count is used, compensating for the difference. Table 3 also shows that more components tend to be increasingly sparse when cached, which compensates somewhat for the additional memory usage. Please see the supplementary material for more detailed results on other scenes.

Application

We demonstrate a proof-of-concept application of FastNeRF for a telepresence scenario by first gathering a dataset of a person performing facial expressions for about 20s in a multi-camera rig consisting of 32 calibrated cameras. We then fit a 3D face model similar to FLAME to obtain expression parameters for each frame. Since the captured scene is not static, we train a FastNeRF model of the scene jointly with a deformation model conditioned on the expression data. The deformation model takes the samples p\bm{p} as input and outputs their updated positions that live in a canonical frame of reference modeled by FastNeRF.

We show example outputs produced using this approach in Figure 5. This method allows us to render 300×300300\times 300 pixel images of the face at 30 fps on a single Nvidia Tesla V100 GPU – around 50 times faster than a setup with a NeRF model. At the resolution we achieve, the face expression is clearly visible and the high frame rate allows for real-time rendering, opening the door to telepresence scenarios. The main limitation in terms of speed and image resolution is the deformation model, which needs to be executed for a large number of samples. We employ a simple pruning method, detailed in the supplementary material, to reduce the number of samples.

Finally, note that while we train this proof-of-concept method on a dataset captured using multiple cameras, the approach can be extended to use only a single camera, similarly to .

Since FastNeRF uses the same set of inputs and outputs as those used in NeRF, it is applicable to many of NeRF’s extensions. This allows for accelerating existing approaches for: reconstruction of dynamic scenes , single-image reconstruction , quality improvement and others . With minor modifications, FastNeRF can be applied to even more methods allowing for control over illumination and incorporation of uncertainty .

Conclusion

In this paper we presented FastNeRF, a novel extension to NeRF that enables the rendering of photorealistic images at 200Hz and more on consumer hardware. We achieve significant speedups over NeRF and competing methods by factorizing its function approximator, enabling a caching step that makes rendering memory-bound instead of compute-bound. This allows for the application of NeRF to real-time scenarios.

References

Overview

We use the supplementary material to provide more detailed results and algorithms. We urge the reader to see our video, in which we show results for all 1616 scenes we tested on along with providing an intuition for our method. On a practical note, we also provide guidance on getting smooth results for anyone adapting our method to their own work.

Please note that parameters chosen throughout this work for the rendering algorithm favour the highest possible quality. It is possible to significantly increase rendering speeds by sacrificing a small amount of visual quality for real-world applications.

Effect of modified Fourier feature encoding on metrics: Using an 88 layer 256256 hidden unit MLP for position, and a 44 layer 128128 MLP for direction, we can obtain an average PSNR of 29.748dB29.748dB for L=1L=1, 29.646dB29.646dB for L=2L=2, 29.449dB29.449dB for L=3L=3 on synthetic data, with smoothness of results being reflected in these numbers.

Effect of small direction cache on metrics: Using a smaller direction cache has a negligible impact on metrics, occasionally actually improving them. This shows that the direction-dependent effects required by FastNeRF are low in frequency. For the scenes where we found we needed to use a smaller cache for moving images, the PSNR differences caused by using a smaller directional cache are: Materials (32332^{3}) (28.885dB →28.874dB); Drums (16316^{3}) (23.745dB →23.836dB); Lego (32332^{3}) (32.275dB →32.155dB); Mic (32332^{3}) (31.765dB →31.667dB); Hotdog (32332^{3}) (34.722dB →34.644dB); Ficus (32332^{3}) (27.792dB →28.193dB).

2 Training & Detailed Results

For training, we use the same frequency encoding, noise perturbation and learning rate decay (starting at 5e−45e-4) as . For the NeRF 360 Synthetic dataset, we use 6464 and 128128 samples for the coarse and fine networks, respectively. This becomes 6464 and 6464 for the LLFF scenes. Note that the coarse networks are discarded and not used in the caching stage - the cached density (or rather, the mesh derived from it) serves as the importance distribution during rendering. We sample 10001000 random rays per gradient descent step, and use Adam as our optimizer of choice, with β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999.

In the paper, we compute test metrics on a random subset of 20 images per scene from the test sets of the NeRF 360 Synthetic dataset, and all available test images for the LLFF data. We use the evaluation set only to select the best iteration, which is 300K for all scenes, indicating identical convergence trends as NeRF . For completeness, we show results for all scenes using the full number of test images in Table 4. We note that the results and trends are in agreement with the sub-sampled test set. Please see Figure 7 and beyond for a qualitative comparison in addition to the video. We include further ablations on the number of components in Table 5 and Table 6.

3 Meshing & Rendering

One of the key advantages of working with sparse octrees to store the caches required by our method is that they can serve as a basis for accelerating the rendering process. Using raytracing, we can quickly terminate rays that miss the neurally rendered object(s) all together. For the rays that hit occupied voxels, and so require integrating the volume, we can use raytracing to skip empty space efficiently, saving a significant amount of queries. This matters because caching makes our method memory instead of compute bound. Computing the inner product of the components and weights as well as tracking transmittance for each ray is fast on modern GPUs. Grid look-ups, which require reading from the GPUs RAM on the other hand, are expensive.

While we could use a hierarchical digital differential analyzer such as to determine ray grid intersections, we opt to trace rays against a collision mesh derived from the volume instead, for three reasons. First, this allows us to take advantage of hardware acceleration for the BVH and ray-triangle intersections on modern hardware. Second, we can down/resample or deform the meshes easily using existing tools. Finally, meshes could be used to ‘fake’ shadows, bouncelight and other effects which can be composited over the neurally rendered content.

We derive the meshes by converting the density cache to a sign distance function (thresholding at 0.00.0), optionally downsampling it before meshing with marching cubes, and optionally simplifying the resulting geometry using standard techniques. We show the resulting geometry for the Lego scene in Figure 6. Note that more aggressive thresholding of the volume or remeshing is easily possible and can lead to significant performance benefits in practise as more rays can be culled early and mesh complexity reduced.

For rendering, we follow the same method as NeRF, using the Beer-Lambert Law to model radiance extinction. Once the contribution of new samples for a ray reaches 0.0010.001, we terminate a ray’s rendering kernel. Note that this can be relaxed for greater performance. Another way to accelerate the rendering process is to visit fewer voxels based on radiance or transmittance. For all paper experiments however, we never skip occupied space, or optimise any parameters per scene.

4 Cache size calculation

In this section we show the formulas we used to estimate cache sizes MNeRF,MFastNeRFM_{NeRF},M_{FastNeRF} for NeRF and FastNeRF respectively.

where sσ,srgbs_{\sigma},s_{rgb} are the sizes of the stored transparency and RGB values in bits and α∈⟨0,1)\alpha\in\langle 0,1) is the inverse volume sparsity, where α=1\alpha=1 would indicate a dense volume.

where suvw,sβs_{uvw},s_{\beta} are the sizes of the stored (ui,vi,wi)(u_{i},v_{i},w_{i}) values and the weights βi\beta_{i}.

For k=l=1024k=l=1024, (r,g,b)(r,g,b) stored as individual bytes and all other values stored as half-precision floats the cache size would be

Even for highly sparse volumes, where α=0.01\alpha=0.01, the cache would take nearly 60TB, which exceeds the memory capacity of consumer-grade hardware. The high memory requirement makes this approach impractical when used with the standard NeRF architecture.

5 Cache size calculation

In this section we show the formulas we used to estimate cache sizes MNeRF,MFastNeRFM_{NeRF},M_{FastNeRF} for NeRF and FastNeRF respectively.

where sσ,srgbs_{\sigma},s_{rgb} are the sizes of the stored transparency and RGB values in bits and α∈⟨0,1⟩\alpha\in\langle 0,1\rangle is the inverse volume sparsity, where α=1\alpha=1 would indicate a dense volume. Just as in the main paper kk, ll correspond to the number of bins per dimension in the position and direction dependent caches respectively.

where suvw,sβs_{uvw},s_{\beta} are the sizes of the stored (ui,vi,wi)(u_{i},v_{i},w_{i}) values and the weights βi\beta_{i}.

For k=l=1024,D=8k=l=1024,D=8, (r,g,b)(r,g,b) stored as individual bytes and all other values stored as half-precision floats the cache size would be

6 Details on Applications

In our proof-of-concept telepresence scenario we use a deformation field network Fdeform\mathcal{F}_{deform} that modifies the input sample positions. This network is similar to the position-dependent network used in FastNeRF, but smaller - a 6-layer MLP with 64 units in each layer. The input is a sample position and an expression vector. The sample position is processed with positional encoding identically to how the FastNeRF input position is processed. The output is an offset that is applied to the input sample position. When training with Fdeform\mathcal{F}_{deform} we add an L2 regularizer that constrains Fdeform\mathcal{F}_{deform} to be an identity transform if a neutral expression is passed as input. We find that this solution stabilizes training and improves results.

In addition to Fdeform\mathcal{F}_{deform}, we also use a non-trainable deformation field Fbone\mathcal{F}_{bone} that moves the samples in the head region according to the movement of the 3D model bones that represent the shoulders and the neck. This additional deformation accounts for large movements of the head and the shoulders, which reduces the load on Fdeform\mathcal{F}_{deform} and improves quality. The two deformation fields are composed with the position-dependent network Fpos\mathcal{F}_{pos} as follows:

where e\bm{e} is the expression vector. While we could use Fbone\mathcal{F}_{bone} at test-time as well, we remove it to improve performance.

The training procedure in this scenario is identical to that used in FastNeRF but with 64 samples in the coarse stage and an additional 64 in the fine stage. The position-dependent network has 384 units in each layer, while the view-dependent network uses only 32 units and 2 layers. The view-dependent network is very small as we believe the scene’s illumination is simple and fairly uniform.

The main limitation of this approach in terms of speed is the need to call Fdeform\mathcal{F}_{deform} for all the input samples. In order to mitigate this, at test-time we implement a grid-based sample pruner. For each position in a grid around the face region we compute the density σ\sigma

for 50 randomly chosen samples of e\bm{e}. For each grid position we record the minimum and maximum encountered σ\sigma in a sparse volume. At test-time we use this volume to prune the inputs to Fdeform\mathcal{F}_{deform} where the maximum σ\sigma is 0 and where the rays would get saturated, which we estimate using the minimum density values from the volume. This approach reduces the runtime of Fdeform\mathcal{F}_{deform} by more than a factor of 2.

To further reduce the impact of Fdeform\mathcal{F}_{deform} on the runtime we only use 43 samples in the coarse stage and remove the fine stage altogether at test time. To offset this low sample count we add additional samples that are linearly interpolated between the samples generated by the deformation network

The additional samples can be evaluated very cheaply as they can be looked-up in the FastNeRF cache. Note that since the deformation network changes the positions of individual samples along the rays, the rays are no longer straight. This means that in this scenario we cannot use the hardware-accelerated ray tracing procedure described in the main paper, though the pruning method described above serves a similar role.

Even though the steps above significantly reduce the runtime of Fdeform\mathcal{F}_{deform}, we are still limited to rendering 300×300300\times 300 pixel images with a reduced sample count if we want to maintain 30FPS with the deformation network. We observed that if this framerate constraint was to be disregarded, the approach described above is able to generate images at significantly better quality. Thus, we believe that a faster deformation approach remains an important goal for future work.