AutoInt: Automatic Integration for Fast Neural Volume Rendering

David B. Lindell, Julien N. P. Martel, Gordon Wetzstein

Introduction

Image-based rendering and novel view synthesis are fundamental problems in computer vision and graphics (e.g., ). The ability to interpolate and extrapolate a sparse set of images depicting a 3D scene has broad applications in entertainment, virtual and augmented reality, and many other applications. Emerging neural rendering techniques have recently enabled photorealistic image quality for these tasks (see Sec. 2).

Although state-of-the-art neural volume rendering techniques offer unprecedented image quality, they are also extremely slow and memory inefficient . This is a fundamental obstacle to making these methods practical. The primary computational bottleneck for neural volume rendering is the evaluation of integrals along the rendered rays during training and inference required by the volume rendering equation . Approximate integration using Monte Carlo sampling is typically used for this purpose, requiring hundreds of forward passes through the neural network representing the volume for each of the millions of rays that need to be rendered for a single frame. Here, we develop a general and efficient framework for approximate integration. Applied to the specific problem of neural volume rendering, our framework improves a tradeoff between rendering speed and image quality, allowing a greater than 10×\times speedup in the rendering process, though with some reduction in image quality.

Our integration framework builds on previous work demonstrating that coordinate-based networks (sometimes also referred to as implicit neural representations) can represent signals (e.g., images, audio waveforms, or 3D shapes) and their derivatives. That is, taking the derivative of the coordinate-based network accurately models the derivative of the original signal. This property has recently been shown for coordinate-based networks with periodic activation functions , but we show that it also extends to a family of networks with different nonlinear activation functions (Sec 3.4 and supplemental).

We observe that taking the derivative of a coordinate-based network results in a new computational graph, a “grad network”, which shares the parameters of the original network. Now, consider that we use as our network a multilayer perceptron (MLP). Taking its derivative results in a grad network which can be trained on a signal that we wish to integrate. By reassembling the grad network parameters back into the original MLP, we construct a neural network that represents the antiderivative of the signal to integrate.

This procedure results in a closed-form solution for the antiderivative, which, by the fundamental theorem of calculus, enables the calculation of any definite integral in two evaluations of the MLP. Inspired by techniques for automatic differentiation (AutoDiff), we call this procedure automatic integration or AutoInt. Although the mechanisms of AutoInt and AutoDiff are very different, both approaches enable the calculation of integrals or derivatives in an automated manner that does not rely on traditional numerical techniques, such as sampling or finite differences.

The primary benefit of AutoInt is that it allows evaluating arbitrary definite integrals quickly by querying the network representing the antiderivative. This concept could have important applications across science and engineering; here, we focus on the specific application of neural volume rendering. For this application, efficiently evaluating integrals amounts to accelerating rendering (i.e., inference) times, which is crucial for making these techniques more competitive with traditional real-time graphics pipelines. However, our framework still requires a slow training process to optimize a network for a given set of posed 2D images.

Specifically, our contributions include the following.

We introduce a framework for automatic integration that learns closed-form integral solutions. To this end, we explore new network architectures and training strategies.

Using automatic integration, we propose a new model and parameterization for neural volume rendering that is efficient in computation and memory.

We improve a tradeoff between neural rendering speed and image quality, demonstrating rendering rates that are an order of magnitude faster than previous implementations , though with a slight reduction in image quality.

Related Work

Over the last few years, end-to-end differentiable computer vision pipelines have emerged as a powerful paradigm wherein a differentiable or neural scene representation is optimized via differentiable rendering with posed 2D images (see e.g., for a survey). Neural scene representations often use an explicit 3D proxy geometry, such as multi-plane or multi-sphere images or a voxel grid of features . Explicit neural scene representations can be rendered quickly, but they are fundamentally limited by the large amount of memory they consume and thus may not scale well.

As an alternative, coordinate-based networks, or implicit neural representations, have been proposed as a continuous and memory-efficient approach. Here, the scene is parameterized using neural networks, and 3D awareness is often enforced through inductive biases. The ability to represent details in a scene is limited by the capacity of the network architecture rather than the resolution of a voxel grid, for example. Such representations have been explored for modeling shape parts , objects , or scenes . Coordinate-based networks have also been explored in the context of generative frameworks .

The method closest to our application is neural radiance fields (NeRF) . NeRF is a neural rendering framework that combines a volume represented by a coordinate-based network with a neural volume renderer to achieve state-of-the-art image quality for view synthesis tasks. Specifically, NeRF uses ReLU-based multilayer perceptrons (MLPs) with a positional encoding strategy to represent 3D scenes. Rendering an image from such a representation is done by evaluating the volume rendering equation , which requires integrating along rays passing through the neural volume parameterized by the MLP. This integration is performed using Monte Carlo sampling, which requires hundreds of forward passes through the MLP for each ray. However, this procedure is extremely slow, requiring days to train a representation of a single scene from multi-view images. Rendering a frame from a pre-optimized representation requires tens of seconds to minutes.

Here, we leverage automatic integration, or AutoInt, to significantly speed up the evaluation of integrals along rays. AutoInt reduces the number of network queries required to evaluate integrals (e.g., using Monte Carlo sampling) from hundreds to just two, greatly speeding up inference for neural rendering.

Integration Techniques.

In general, integration is much more challenging than differentiation. Whereas automatic differentiation primarily builds on the chain rule, there are many different strategies for integration, including variable substitution, integration by parts, partial fractions, etc. Heuristics can be used to choose one or a combination of these strategies for any specific problem. Closed-form solutions to finding general antiderivatives exist only for a relatively small class of functions and, when possible, involve a rather complex algorithm, such as the Risch or Risch-Norman algorithm . Perhaps the most common approach to computing integrals in practice is numerical integration, for example using Riemann sums, quadratures, or Monte-Carlo methods . In these numerical methods, the number of samples trades off accuracy for runtime.

Since neural networks are universal function approximators, and are themselves functions, they can also be integrated analytically. Previous work has explored theory and connections between shallow neural networks and integral formulations for function approximation . Other work has derived closed-form solutions for integrals of simple single-layer or two-layer neural networks . As we shall demonstrate, our work is not limited to a fixed number of layers or a specific architecture. Instead, we directly train a grad network architecture for which the integral network is known by construction.

AutoInt for Neural Integration

In this section, we introduce a fundamentally new approach to compute and evaluate antiderivatives and definite integrals of coordinate-based neural networks.

This equation relates the coordinate-based network Φθ\Phi_{\theta} to its partial derivative Ψθi\Psi_{\theta}^{i} and, hence, Φθ\Phi_{\theta} is an antiderivative of Ψθi\Psi^{i}_{\theta}.

As a result, Φθ\Phi_{\theta} is a network that represents an antiderivative of Ψθi\Psi^{i}_{\theta}. We call this procedure of training Ψθi\Psi^{i}_{\theta} and reassembling θ\theta to construct the antiderivative automatic integration. How to reassemble θ\theta depends on the network architecture used for Φθ\Phi_{\theta}, and is addressed in the next section.

2 The Integral and Grad networks

Coordinate-based neural networks are usually formed from multilayer perceptron (MLP), or fully connected, architectures:

The computational graph of a 3-hidden-layer MLP representing Φθ\Phi_{\theta} is shown in Figure 2. Operations are indicated as nodes and dependencies as directed edges. Here, the arrows of the directed edges point towards nodes that must be computed first.

For this MLP, the form of the network Ψθi=∂Φθ/∂xi\Psi_{\theta}^{i}=\partial\Phi_{\theta}/\partial{x_{i}} can be found using the chain rule

3 Evaluating Antiderivatives & Definite Integrals

To compute the antiderivative and definite integrals of a function ff in the AutoInt framework, one first chooses the specifics of the MLP architecture (number of layers, number of features, type of nonlinearities) for an integral network Φθ\Phi_{\theta}. The grad network Ψθi\Psi_{\theta}^{i} is then instantiated from this integral network based on AutoDiff. In practice, we developed a custom AutoDiff framework that traces the integral network and explicitly instantiates the corresponding grad network while maintaining the shared parameters (additional details in the supplemental). Once instantiated, parameters of the grad network are optimized to fit a signal of interest using conventional AutoDiff and optimization tools . Specifically we optimize a loss of the form

We also note that AutoInt extends to integrating high-dimensional signals using a generalized fundamental theorem of calculus, which we describe in the supplemental.

4 Example in Computed Tomography

and (x,y)(x,y) is on the ray (ρ,α)∈×[0,π)(\rho,\alpha)\in\times[0,\pi) by satisfying x(t)cos⁡(α)+y(t)sin⁡(α)=ρx(t)\cos(\alpha)+y(t)\sin(\alpha)=\rho with α\alpha being the orientation of the ray and ρ\rho its eccentricity with respect to the origin as shown in Figure 3. The measurement ss is called a sinogram, and this particular integral is referred to as the Radon transform of ff .

The inverse problem of computed tomography involves recovering the absorption ff given a sinogram. Here, for illustrative purposes, we will look at a tomography problem in which a grad network is trained on a sparse set of measurements and the integral network is evaluated to produce unseen ones. Sparse-view tomography is a standard reconstruction problem , and this setup is analogous to the novel view synthesis problem we solve in Section 4.

We consider a dataset of measurements D={(ρi,αi,s(ρi,αi)}i<D\mathcal{D}=\{(\rho_{i},\alpha_{i},s(\rho_{i},\alpha_{i})\}_{i<D} corresponding to DD sparsely sampled rays. We train a grad network using the AutoInt framework. For this purpose, we instantiate a grad network Ψθ\Psi_{\theta} whose input is a tuple (ρ,α,t)(\rho,\alpha,t). It is trained to match the dataset of measurements

Thus, at training time, the grad network is evaluated TT times in a Monte Carlo fashion with tj∼U([tn,tf])t_{j}\sim\mathcal{U}([t_{n},t_{f}]). At inference, just two evaluations of Φθ∗\Phi_{\theta^{*}} yield the integral

Results in Figure 3 show that the two evaluations of the integral network Φθ∗\Phi_{\theta^{*}} can faithfully reproduce supervised measurements and generalize to unseen data. Generalization, however depends on the type of nonlinearity used. We show that Swish with normalized positional encoding (details in Sec. 5) generalizes well, and SIREN fits the measurements better but fails to generalize to unseen views.

Note that both the nonlinearity nl and its derivative \textscnl′{\textsc{nl}}^{\prime} appear in the grad network architectures (Eq. (3) and Figure 2). This implies that integral networks with ReLU nonlinearities have step functions appearing in the grad network, possibly making training Ψθ\Psi_{\theta} difficult because of nodes with derivatives that are zero almost everywhere. We explore several other nonlinearities here (with additional details in the supplemental), and show that Swish heuristically performs best in the grad networks used in our application. Yet, we believe the study of nonlinearities in grad networks to be an important avenue for future work.

Neural Volume Rendering

Combining volume rendering techniques with coordinate-based networks has proved to be a powerful technique for neural rendering and view synthesis . Here, we briefly overview volume rendering and describe an approximate volume rendering model that enables using AutoInt for efficient rendering.

Classical volume rendering techniques are derived from the radiative transfer equation with an assumption of minimal scattering in an absorptive and emissive medium . We adopt a rendering model based on tracing rays through the volume , where the emission and absorption along camera rays produce color values that are assigned to rendered pixels.

Rendering from the volume requires integrating the emissive radiance along the ray while also accounting for absorption. The transmittance TT describes the net reduction from absorption from the ray origin to the ray position r(t)\mathbf{r}(t), and is given as

where tnt_{n} indicates a near bound along the ray. With this expression, we can define the volume rendering equation (VRE), which describes the color C\mathbf{C} of a rendered camera ray.

Conventionally, the VRE is computed numerically by Riemann sums, quadratures, or Monte-Carlo methods , whose accuracy thus largely depends on the number of samples taken along the ray.

2 Approximate Volume Rendering for Automatic Integration

Automatic integration allows us to efficiently evaluate definite integrals using a closed-form solution for the antiderivative. However, the VRE cannot be directly evaluated with AutoInt because it consists of multiple nested integrations: the integration of radiance along the ray weighted by integrals of cumulative transmittance. We therefore choose to approximate this integral in piecewise sections that can each be efficiently evaluated using AutoInt. For NN piecewise sections along a ray, we give the approximate VRE and transmittance as

and δi=ti−ti−1\delta_{i}=t_{i}-t_{i-1} is the length of each piecewise interval along the ray. Equation 12 can also be viewed as a repeated alpha compositing operation with alpha values of σˉiδi\bar{\sigma}_{i}\delta_{i}. After some simplification and substitution into Equation 12 (see supplemental), we have the following expression for the piecewise VRE:

While this piecewise expression is only an approximation to the full VRE, it enables us to use AutoInt to efficiently evaluate each piecewise integral over absorption and radiance. In practice, there is a tradeoff between improved computational efficiency and degraded accuracy of the approximation as the value of NN decreases. We evaluate this tradeoff in the context of volume rendering and learned novel view synthesis in Sec. 6.

Optimization Framework

We evaluate the piecewise VRE introduced in the previous section using an optimization framework overviewed in Figure 4. At the core of the framework are two MLPs that are used to compute integrals over values of σ\sigma and c\mathbf{c} as we detail in the following.

Rendering an image from the high-dimensional volume represented by the MLP requires evaluating integrals along each ray r(t)\mathbf{r}(t) in the direction of tt. Thus, the grad network should represent ∂Φθ/∂t\partial\Phi_{\theta}/\partial t, the partial derivative of the integral network with respect to the ray parameter. In practice, the networks take as input the values that define each ray: o\mathbf{o}, tt, and d\mathbf{d}. Then, positions along the ray are calculated as x=o+t d\mathbf{x}=\mathbf{o}+t\,\mathbf{d} and passed to the initial layers of the networks together with d\mathbf{d}. With this dependency on tt, we use our custom AutoDiff implementation to trace computation through the integral network, define the computational graph that computes the partial derivative with respect to tt, and instantiate the grad network.

Grad Network Positional Encoding.

where ωi=2iπ\omega_{i}=2^{i}\pi and LL controls the number of frequencies used to encode each input. We find that using this scheme directly in the grad network produces poor results because it introduces an exponentially increasing amplitude scaling into the coordinate encoding. This can easily be seen by calculating the derivative ∂γ/∂p=(⋯ωicos⁡(ωip),−ωisin⁡(ωip)⋯ )\partial\gamma/\partial p=\left(\cdots\omega_{i}\cos(\omega_{i}p),-\omega_{i}\sin(\omega_{i}p)\cdots\right). Instead, we use a normalized version of the positional encoding for the integral network, which improves performance when training the grad network:

Predictive Sampling.

While AutoInt is used at inference time, at training time, the grad network is optimized by evaluating the piecewise integrals of Equation 13 using a quadrature rule discussed by Max :

We use Monte Carlo sampling to evaluate the integrals σˉi\bar{\sigma}_{i} and cˉi\bar{\mathbf{c}}_{i} by querying the networks at many positions within each interval δi\delta_{i}.

However, some intervals δi\delta_{i} along the ray contribute more to a rendered pixel than others. Thus, assuming we use the same number of samples per interval, we can improve sample efficiency by strategically adjusting the length of these intervals to place more samples in positions with large variations in σ\sigma and c\mathbf{c}. This idea is similar in spirit to accelerated volume rendering techniques using hierarchical sampling or adaptive ray termination .

Fast Grad Network Evaluation.

AutoInt can be implemented directly in popular optimization frameworks (e.g., PyTorch , Tensorflow ); however, training the grad network is generally computationally slow and memory inefficient. These inefficiencies stem from the two step procedure required to compute the grad network output at each training iteration: (1) a forward pass through the integral network is computed and then (2) AutoDiff calculates the derivative of the output with respect to the input variable of integration. Instead, we implemented a custom AutoDiff framework on top of PyTorch that parses a given integral network and explicitly instantiates the grad network modules with weight sharing (see Figure 2). Then, we evaluate and train the grad network directly, without the overhead of the additional per-iteration forward pass and derivative computation. Compared to the two-step procedure outlined above, our custom framework improves per-iteration training speed by a factor of 1.8 and reduces memory consumption by 15% for the volume rendering application. More details about our AutoInt implementation can be found in the supplemental, and our code is publicly availablehttps://github.com/computational-imaging/automatic-integration.

Implementation Details.

In our framework, a volume representation is optimized separately for each rendered scene. To optimize the grad networks, we require a collection of RGB images taken of the scene from varying camera positions, and we assume that the camera poses and intrinsic parameters are known. At training time, we randomly sample images from the training dataset, and from each image we randomly sample a number of rays. Then, we optimize the network to minimize the loss function

where C\mathbf{C} is the ground truth pixel value for the selected ray.

In our implementation, we train the networks using PyTorch and the Adam optimizer with a learning rate of 5×10−45\times 10^{-4}. The networks representing volume density and color each have 8 hidden layers with 256 hidden units, we use a batch size of 4 with 1024 rays sampled from each image, and we decay the learning rate by a factor of 0.2 every 10510^{5} iterations. Training and inference are performed using NVIDIA V100 GPUs. For the sampling network, we evaluate using M=128/NM=128/N samples within each piecewise interval for N∈{2,4,8,16,32,64}N\in\{2,4,8,16,32,64\} (see Figure 5) and find that using 8, 16, or 32 piecewise intervals produces acceptable results while achieving a significant computational acceleration with AutoInt. Finally, for the positional encoding, we use L=10L=10 and L=4L=4 for x\mathbf{x} and d\mathbf{d}, respectively.

Results

We evaluate AutoInt for volume rendering on a synthetic dataset of scenes with challenging geometries and reflectance properties. We demonstrate that the approach allows an improved tradeoff between image quality and rendering speed for neural volume rendering. Rendering times are improved by greater than 10×\times compared to the state-of-the-art , though at slightly reduced image quality.

Our training dataset consists of eight objects, each rendered from 100 different camera positions using the Blender Cycles engine . For the test set, we evaluate on an additional 200 images. We compare AutoInt to two other baselines: Neural Radiance Fields (NeRF) and Neural Volumes . NeRF uses a similar architecture and Monte Carlo sampling with the full volume rendering model, rather than our piecewise approximation and AutoInt. Neural Volumes is a voxel-based method that encodes a deep voxel grid representation of a scene using a convolutional neural network. Novel views are rendered by applying a learned warping operator to the voxel grid and sampling voxel values by marching rays from the camera position.

In Table 1 we report the peak signal-to-noise ratio (PSNR) averaged across all scenes and test images. AutoInt outperforms Neural Volumes quantitatively, while achieving a greater than 10×\times improvement in render time relative to NeRF, though with a tradeoff in image quality. Increasing the number of piecewise sections in the approximate VRE improves render quality at the cost of computation.

We evaluate the effect of the sampling network and the number of sections in the approximate VRE in Figure 5 for the Lego scene. Using the sampling network improves performance and sample efficiency by allocating more sections in regions with large variations in the volume density.

In Table 2 we show the effect of the VRE approximation and grad network architecture on render quality of the Lego scene. Using the full VRE achieves similar performance to the approximate, piecewise VRE with 32 sections. We attribute most of the difference in performance between our method and NeRF to the regularized, tree-like structure of the grad network, which is constrained by weight sharing between the branches (see Figure 2). While evaluating NeRF with fewer samples along each ray reduces computation, rendering quality degrades significantly compared to using AutoInt with the same number of samples.

We also show qualitative results in Figure 6 for the Materials, Hot Dog, Drums, and Lego scenes. Again, the quality of the rendered images improves as the number of sections increases. In the Materials scene (Figure 6), the proposed technique exhibits fewer artifacts compared to Neural Volumes. AutoInt also shows improved modeling of view-dependent effects in the Drums scene relative to Neural Volumes and NeRF (e.g., specular highlights on the symbols). We show additional results on captured scenes from the Local Light Field Fusion and DeepVoxels datasets in the supplemental.

Discussion

In this work, we introduce a new framework for numerical integration in the context of coordinate-based neural networks. Applied to neural volume rendering, AutoInt enables improvements to computational efficiency by learning closed-form solutions to integrals. Although these computational speedups currently come with a tradeoff to image quality, the method takes steps towards efficient learned integration using deep network architectures.

Our approach is analogous to conventional methods for fast evaluation of the VRE; for example, methods based on shear-warping and the Fourier projection-slice theorem . Similar to our method, these techniques use approximations (e.g., with sampling and interpolation) that trade off image quality with computationally efficient rendering. Additionally, we believe our approach is compatible with recent work that aims to speed up volume rendering by pruning areas of the volume that do not contain the rendered object .

A key idea of AutoInt is that an integral network can be automatically created after training a corresponding grad network. We note that grad networks represent a fundamentally different and relatively unexplored type of network architecture compared to conventional fully-connected networks. While using grad networks with certain classes of non-linearities (e.g., Swish) can improve image quality, exploring improved training strategies and more expressive grad network architectures is an important and promising direction for future work. Finally, we believe that AutoInt will be of interest to a wide array of application areas beyond computer vision, especially for problems related to inverse rendering, sparse-view tomography, and compressive sensing.

Julien N. P. Martel was supported by a Swiss National Foundation (SNF) Fellowship (P2EZP2 181817). Gordon Wetzstein was supported by an NSF CAREER Award (IIS 1553333) and a PECASE from the ARO.

References