Neural Fields in Visual Computing and Beyond

Yiheng Xie, Towaki Takikawa, Shunsuke Saito, Or Litany, Shiqin Yan, Numair Khan, Federico Tombari, James Tompkin, Vincent Sitzmann, Srinath Sridhar

Introduction

Visual computing involves the synthesis, estimation, manipulation, display, storage, and transmission of data about objects and scenes across space and time. In computer graphics, we synthesize 3D shapes and 2D images, render novel views of scenes, and animate articulating human bodies. In computer vision, we reconstruct 3D appearance, shape, object pose, and deformation. In human-computer interaction (HCI), we enable the interactive investigation of spacetime data. Beyond visual computing and into adjacent disciplines like robotics, we use the structure of 3D scenes to plan and execute actions. Within these disciplines, fields are widely used to continuously parameterize an underlying physical quantity of an object or scene over space and time. For instance, fields have been used to visualize physical phenomena [Sab88], compute image gradients [SK97], compute collisions [OF03], or represent shapes via constructive solid geometry (CSG) [Eva15].

In visual computing tasks, we solve many different optimization problems, including using machine learning and training data. One class of methods has gained significant attention since 2019: coordinate-based neural networks that represent a field. We refer to these methods as neural fields. Given sufficient parameters, fully connected neural networks can encode continuous signals over arbitrary dimensions at arbitrary resolution. Recent success with neural networks has caused a resurgence of interest in visual computing problems, leading to more accurate, higher fidelity, more expressive, and memory-efficient solutions. This has allowed neural fields to gain traction as a useful parameterization of 2D images [KAL∗21], 3D shape [PFS∗19, MON∗19, CZ19], view-dependent appearance [SZW19, MST∗20], and human bodies and faces[NMOG19, DLJ∗20, SYMB21, YTB∗21, RTE∗21].

Such attention manifests as rapid progress and an explosion of papers (Figure 1), with a need to consolidate the discovered knowledge. However, a shared mathematical formulation to describe related techniques has not yet emerged, making it hard to communicate ideas and train students. Furthermore, there is “selective amnesia” [SC21] of older or even concurrent works, causing research repetition. Finally, rapid progress makes any survey quickly out-of-date, requiring new summarization approaches.

Following a survey of over 250 papers, we address the above issues by defining a neural field via fields of physical quantities, providing a shared mathematical formulation across common techniques, categorizing, describing, and relating many applications, and presenting a living database via a community website.

In Part I, we describe neural field techniques that are common across papers with a consistent notation and vocabulary. For instance, we identify and formalize recent hybrid discrete-continuous representations of neural fields, and techniques for learning priors such as local and global conditioning and meta-learning. Further, neural fields can be combined with a wide variety of differentiable forward maps, and we identify many such maps including surface and volume renderers and partial differential equations.

In Part II, we describe a broad cross-section of applications of neural fields to problems in visual computing. This lets us identify commonalities, connections, and trends across works representing shape and appearance of scenes and objects, including 3D reconstruction, digital humans, generative modeling, data compression, and 2D image processing. We also review works that solve tasks in adjacent communities including robotics (Section 5), medical imaging (Section 1), audio processing (Section 1), and physics-informed problems (Section 2).

Our living database is a companion website that provides search, filter, and visualization features to present the works described in this report (Figure Neural Fields in Visual Computing and Beyond), along with bibliography exporting features. The website allows our community to submit new works in neural fields and automatically updates the databases using keyword categorizations from Part I and II taxonomies. The website requires minimal maintenance, and its source code is open to allow future reports and surveys to provide the same functionality.

In summary, neural field techniques are powerful and widely applicable to visual computing problems and beyond. This state-of-the-art report provides the mathematical formulation of techniques and discussion of applications to make sense of this exciting area, along with a community-driven website to continue to help researchers keep track of new developments in the future.

Following the original definition in physics [FLS65, McM02] we define fields as the following:

A field is a quantity defined for all spatial and/or temporal coordinates.

We can represent a field as a function mapping a coordinate x\mathbf{x} to a quantity, which is typically a scalar or vector. Table 1 provides several examples. The physics community has also studied spinor and tensor fields. Finally, field coordinates are not limited to space and time, such as frequency coordinates in spectrograms. This review will focus on scalar and vector fields defined over spacetime, which are most relevant to visual computing.

In practice, the underlying field generation process may not have a known analytic form. Thus, functions may be described by parameters Θ\Theta that are hand crafted, optimized, or learned. We denote such a field as producing quantity q=Φ(x;Θ)\mathbf{q}=\Phi(\mathbf{x};\Theta). Furthermore, we often index sampled functions using discrete values, such as at camera pixels, or choose discrete function parameterizations using voxels or discretized level sets. However, discrete parameterizations are limited by the Nyquist sampling rate, causing high memory requirements for 3D tasks. Adaptive grids like octrees and k–d trees can reduce memory, but their generation can lead to costly combinatoric optimizations.

A neural network connects many layers of artificial neurons to learn to non-linearly map a fixed-size input to a fixed-size output [LBH15]. A multi-layer perceptron (MLP) neural network can approximate any function through their learned parameters (universal approximation theorem [KA03]). Since 2010, there has been significant interest in using neural networks with corresponding investment in hardware and software to support neural networks. This has significantly lowered the barrier to apply neural networks to a wide range of problems.

Following the universal approximation theorem, any field can be parameterized by an MLP neural network. Thus:

A neural field is a field that is parameterized fully or in part by a neural network.

Neural fields are both continuous and adaptive by construction. Unlike the memory required for discrete parameterizations that scales poorly with spatio-temporal resolution, the memory required for neural fields instead scales with the number of parameters of the neural network—so-called network complexity. While other continuous parameterizations can represent large extents (for example, a Fourier series is a parameterization of a field), it is often difficult to know the required complexity ahead of time for efficient representation. Neural fields help to resolve this problem by using their parameters only where field detail is present. Neural fields are often parameterized as MLPs with activation functions whose gradients are well-defined. By their analytic differentiability, gradient descent, and over-parameterization [FC19], neural fields are effective at regressing complex signals through optimizations for problems that are otherwise ill-posed.

Our definition of a neural field does not include neural networks whose co-domain has a spatial extent, for instance, a network that outputs a grid of voxels. These methods often use convolutional layers and output a 2D or 3D grid of RGB color, occupancy, or latent features [STH∗19, LSS∗19]. While the regular grid output of these architectures can be seen as samplings of fields, the functions parameterized by the neural network are not neural fields as they do not ingest spatio-temporal coordinates.

Within visual computing, neural fields have been called implicit neural representations, neural implicits, or coordinate-based neural networks. In neuroscience, the term neural field may describe theories about the organization and function of the brain [CbGPW14]. The term also describes a regular grid of neurons in a neural network [Lem92]. We exclude works in these two areas as they do not follow from our definition.

2 Related Surveys and Articles

Within visual computing, there are survey papers on specific areas like neural rendering [TFT∗20, TTM∗21] and 3D reconstruction [ZSG∗18], as well as lists of works in scene representations [NeRb, NeRa, Sit, NeRc]. Our survey aims to connect our visual computing communities together and share findings. Finally, neural fields are domain-agnostic and can model arbitrary quantities beyond shape and appearance. Neural fields have been used extensively in computational physics via physics-informed neural networks (PINNs) [RPK19a], with an existing survey [KKL∗21].

Chapter 0 Part I. Neural Field Techniques

A typical neural fields algorithm in visual computing proceeds as follows (Figure 1): Across space-time, we sample coordinates and feed them into a neural network to produce field quantities. The field quantities are samples from the desired reconstruction domain of our problem. Then, we apply a forward map to relate the reconstruction to the sensor domain (e.g. RGB image), where supervision is available. Finally, we calculate the reconstruction error or loss that guides the neural network optimization process by comparing the reconstructed signal to the sensor measurement.

During this process, many problems affect our ability to successfully reconstruct the signal. To better understand and apply neural fields, we identify five classes of techniques (Table 1). We can aid reconstruction from incomplete sensor signals via prior learning and conditioning (Section 1). We can improve memory, computation, and neural network efficiency via hybrid representations using discrete data structures (Section 2). We can supervise reconstruction via differentiable forward maps that transform or project our domain (e.g., 3D reconstruction via 2D images; Section 3). With appropriate network architecture choices, we can overcome neural network spectral biases (blurriness) and efficiently compute derivatives and integrals (Section 4). Finally, we can manipulate neural fields to add constraints and regularizations, and to achieve editable representations (Section 5). Collectively, these classes constitute a ‘toolbox’ of techniques to help solve problems with neural fields.

Suppose we wish to estimate a plausible 3D surface shape given a partial point cloud. This problem arises in reconstructing 3D street scenes from Lidar scans as we only observe points upon surfaces. To accomplish this task, we need a suitable prior over 3D surfaces. One option is to hand-craft a prior via heuristics such as smoothness or sparseness. However, this approach is limited in the complexity of heuristics that we can conceive. Alternatively, we can learn such a prior from data, and these can be encoded within the parameters and architecture (Section 4) of a neural network.

Beyond this, we might wish to adapt our prior based on specific conditions. In our street scene Lidar example, we might vary the expected surface shape of vehicles based on which kind of vehicle it is: a bicycle, a car, or a truck. For neural fields, this is accomplished by conditioning the neural field on a set of latent variables z\mathbf{z} that encode the properties of a specific field. By varying the latent variables, we can then vary the neural field. In this section, we discuss how to pose optimization problems that learn latent variables z\mathbf{z}, how to infer z\mathbf{z} given a set of incomplete observations, and how to condition the neural field on z\mathbf{z} to decode them into different fields.

A conditional neural field lets us vary the field by varying a set of latent variables z. These latent variables could be samples from an arbitrary distribution, or semantic variables describing shape, type, size, color, etc., or come from an encoding of other data types such as audio data [GCL∗21]. For example, by conditioning our neural fields on semantic variables which describe cars, we would like to decode these to a field that represents the shape and appearance of the corresponding car. Instance-specific information can then be encoded in the conditioning latent variable z, while shared information can be encoded in the neural field parameters. If this latent variables are defined on a semantic or smooth space, they can be interpolated or edited.

In the following subsections, we discuss techniques to learn in an unsupervised manner what the latent variables z\mathbf{z} are, how to infer them given (partial) observations O\mathcal{O}, and how to condition neural fields upon z\mathbf{z} to be decoded into a corresponding field. z\mathbf{z} is typically a low-dimensional vector, and is often referred to as a latent code or feature code. We discuss different encoding schemes (Section 1), both global and local conditioning (Section 1), and different mapping functions Ψ\Psi (Section 1).

For feed-forward encoder methods, the conditioning latent code z=E(O)\mathbf{z}=\mathcal{E}(\mathcal{O}) is generated via an encoder E\mathcal{E}, typically a neural network (Figure 2, left). The parameters in E\mathcal{E} can encode priors that can be pre-trained on data or auxiliary tasks. The decoder is the neural field that is conditioned by the latent code. This conditioning method is fast since inference requires only a single forward pass through the encoder and decoder. Examples of encoders for different data types include PointNet [QSMG17] for point clouds, ResNet [HZRS16] for 2D images, or VoxNet [MS15] for voxel grids.

Given the inference-time optimization, auto-decoding is significantly slower than a feed forward encoder approach. However, auto-decoding does not introduce additional parameters and does not make assumptions about observations O\mathcal{O}. For instance, a 2D CNN encoder assumes O\mathcal{O} to be on a 2D pixel grid, whereas an auto-decoder can ingest tuples of pixel coordinates and colors independently of their spatial arrangement. This can lead to robustness in certain out-of-distribution scenarios. For instance, in 3D reconstruction from images when observing a camera pose not in the training set, a convolutional encoder is constrained by the 2D geometry of its kernels whereas an auto-decoder is not [SZW19, LLS∗21]. Due to these benefits, several works have adopted the auto-decoder approach [PFS∗19, GSL∗20, RTE∗21, YWC∗21, LZZ∗21, JA21, TTG∗21, SRF∗21].

These approaches initialize z\mathbf{z} with a forward pass of a parametric encoder E\mathcal{E}, then continue to optimize z\mathbf{z} iteratively via auto-decoding [MQK∗21].

In local conditioning, we introduce multiple latent codes z\mathbf{z}. Each z\mathbf{z} has spatial extent over a local neighborhood in the space of coordinates x\mathbf{x}. Thus, latent codes are a function of x\mathbf{x}: z=g(x)\textbf{z}=g(\mathbf{x}) (Figure 3, right). Example gg are discrete data structures like 2D raster grids [SHN∗19, YYTK21, TY21], 3D voxel grids [PNM∗20, LGL∗20, JSM∗20, CAPM20, CLI∗20], surface patches [TTG∗20], or orthographic 2D projections of 3D grids like floor maps [DBS∗21, PNM∗20]. As each latent code z\mathbf{z} only has to encode information about its local neighborhood, it consequently does not need to store information about the configuration anymore. For instance, imagine that we split a room into small 3D cubes, and in each cube store a latent code that describes the geometry in that cube. These latent codes do not need to encode all possible configurations of the room anymore.

Similarly, for local conditioning, an encoder E\mathcal{E} only has to encode local properties into latent variables z\mathbf{z}. The encoder can preserve the spatial extent of the observation O\mathcal{O}, such as 2D or 3D CNNs to extract features from images or discretized volumes. This can leverage encoder properties such as translation equivariance, granting better out-of-distribution generalization. On the other hand, latent codes z\mathbf{z} do not capture global scene information, which provides fewer constraints and less high-level control.

Global and local conditioning can be combined. For instance, for human face images, to attempt to disentangle a property that is shared across instances (like hair color) from another that is region specific (like skin wrinkle).

Given a set of latent variables z\mathbf{z}, we now wish to use them to parameterize the neural network that represents the corresponding field. Different approaches have been proposed in prior work. Though on first sight, they seem incomparable, all of them follow the same principle of defining a function Ψ\Psi that maps latent variables z\mathbf{z} to a subset of neural network parameters Θ=Ψ(z)\Theta=\Psi(\mathbf{z}) that then parameterize the neural field ΦΘ\Phi_{\Theta}. Different methods of conditioning differ in which parameters Θ\Theta are output by Ψ\Psi, as well as the form of Ψ\Psi itself. These design choices may impact generalization ability, parameter count, and computational cost.

It is not obvious how conditioning via concatenation falls into the framework defined above, where all conditional neural fields are expressed as predicting a subset of the parameters Θ\Theta of a neural field ΦΘ\Phi_{\Theta} via a function Ψ\Psi. However, conditioning via concatenation is equivalent to defining an affine function Ψ(z)=b\Psi(\mathbf{z})=\mathbf{b} that maps latent codes z\mathbf{z} to the vector of biases b\mathbf{b} of the first layer of Φ\Phi [SCT∗20, DPS∗18, MGB∗21]. In other words, in the case of conditioning via concatenation, the subset of parameters that is predicted by Ψ\Psi is only the biases of the first layer of the neural network Φ\Phi, and Ψ\Psi is parameterized as a simple affine mapping.

Hypernetworks [HDL16] parameterize the function Ψ\Psi as a neural network that takes the latent code z\mathbf{z} as input and outputs neural field parameters Θ\Theta via a forward pass [SZW19, SMB∗20, SRF∗21, NWH21, CTT∗21]. We can view this as a general form of conditioning because every other form of conditioning may be obtained from a hypernetwork by outputting only subsets of parameters Θ\Theta, by factorizing parameters of Φ\Phi via low-rank approximations or via additional scales and biases, or by varying Ψ\Psi architectures such as using only a single linear layer. For instance, conditioning via concatenation is a special case of a hypernetwork where Ψ\Psi is an affine transform and only the biases of the first layer of Φ\Phi are predicted. Full hypernetworks provide more complex embeddings of network weights than concatenation [GW20].

Between conditioning via concatenation and full hypernetworks, one can condition an MLP by predicting feature-wise transformations (FiLM) [DPS∗18, CMK∗21, MGB∗21]. To FiLM-condition a neural field Φ\Phi, we use a network Ψ\Psi to predict a per-layer (and potentially per-neuron) scale γ\gamma and bias β\beta vector from latent variables z\mathbf{z}: Ψ(z)={γ,β}\Psi(\mathbf{z})=\{\gamma,\beta\}. The input xi\mathbf{x}_{i} to the ii-th layer Φi\Phi_{i} is transformed as Φi=γi(z)⊙xi+βi(z)\Phi_{i}=\gamma_{i}(\mathbf{z})\odot\mathbf{x}_{i}+\beta_{i}(\mathbf{z}). Yet another trade-off in the size of the subset of neural field parameters predicted is struck by predicting the factors of a low-rank decomposition of the weight matrices of the neural field Φ\Phi [SIE21].

Gradient-based Meta-learning

An alternative to the conditional neural field approach is gradient-based meta-learning [FAL17]. Here, all neural fields in our target distribution are viewed as specializations of an underlying meta-network with parameters θ\theta [SCT∗20, TMW∗21]. Individual instances are obtained from fitting this meta-network to a set of observations O\mathcal{O}, minimizing a reconstruction loss L\mathcal{L} in a small number of gradient descent steps with step size λ\lambda:

Similar to the auto-decoder framework in conditional neural fields, the inference function E\mathcal{E} is implemented via an iterative optimization algorithm. The meta-network can be seen as an initialization that is sufficiently close to all neural fields in the target distribution.

In a conditional neural field, the prior is expressed via the parameters of Ψ\Psi that enforce that the parameters of Φ\Phi lie in a low-dimensional space as defined by latent variable z\mathbf{z}. However, in gradient-based meta-learning, the prior is expressed by constraining the optimization to not move the neural field parameters Θ\Theta too far away from the parameters of the meta-network θ\theta. Gradient-based meta-learning enables fast inference, as only a few gradient descent steps are required to obtain Θ\Theta. As this does not assume a low-dimensional set of latent variables, in principle we retain the full expressivity of the neural field Φ\Phi.

Hybrid Representations

The second category of items in the neural field toolbox are hybrid representations. These combine neural fields with discrete data structures that decompose the space of input coordinates. This allows neural fields to scale up to large signals [RWG∗13] (Figure 4). Discrete data structures are used extensively in visual computing, including regular grids, adaptive grids, curves, point clouds, and meshes. Discrete structures have several important benefits: 1) They typically reduce computation. For instance, bounding volume hierarchies (BVHs) [RW80, MOB∗21] enable fast queries on hardware accelerators [PBD∗10, WWB∗14]. 2) They also allow for more efficient use of network capacity, since large MLP networks have diminishing returns in representation capacity [RJY∗21]. 3) When representing geometry, discrete structures allow for empty space skipping, and so accelerate rendering. 4) Discrete structures are also suitable for simulation such as via finite element methods, and can help in manipulation and editing tasks (Section 5).

The general approach to spatial decomposition is to store some or all of the neural field parameters ΘΦ\Theta_{\Phi} in a data structure gg. Given a coordinate x\mathbf{x}, we query gg to retrieve parameters Θ\Theta of the neural field. Two common approaches to how we map parameters Θ\Theta to data structure gg are network tiling and embedding.

A collection of separate (usually small) neural fields Φ\Phi are tiled across the input coordinate space, covering disjoint regions. The network architecture is shared, but their parameters are distinct for each disjoint region. Given a coordinate x\mathbf{x}, we simply look up the network parameters in the data structure gg:

We store latent variables z\mathbf{z} in the data structure, as in Section 1. Neural field parameters Θ\Theta become a function of the local embedding z=g(x)\mathbf{z}=g(\mathbf{x}) via a mapping function Φ\Phi:

Tiling is a special case of embedding where Ψ\Psi is the identity function and z\mathbf{z} is the parameters Θ\Theta. For design choice details of function Ψ\Psi, please refer to Section 1.

Defining g𝑔g and its Common Forms

We define a discrete data structure gg as a vector field that maps coordinates to quantities using a sum of Dirac delta functions δ\delta:

where coefficients αi\alpha_{i} (scalar or vector) are the quantities stored at coordinates xi\mathbf{x}_{i}. For uniform voxels, the coordinates xi\mathbf{x}_{i} are distributed on a regular grid, whereas for a point cloud the coordinates xi\mathbf{x}_{i} are distributed arbitrarily. For network tiling, αi\alpha_{i} is an entire set of network parameters Θi\Theta_{i}; for embedding, αi\alpha_{i} is latent variable zi\mathbf{z}_{i}.

The discrete data structure gg may use an interpolation scheme to define gg outside the coordinates xi\mathbf{x}_{i} of the Dirac deltas, such as nearest neighbor, linear, or cubic interpolation. If so, instead of Dirac deltas, function δ\delta are basis functions of non-zero, compact, and local support. For instance, voxel grids are regularly combined with nearest-neighbor interpolation, where g(x)g(\mathbf{x}) is defined as αi\alpha_{i} at the xi\mathbf{x}_{i} closest to x\mathbf{x}.

Regular grids, such as 2D pixels and 3D voxels, discretize the coordinate domain with regular intervals. Their regularity makes them simple to index and apply standard signal processing techniques. Although simple, grids suffer from poor memory scaling in high dimensions, and the Nyquist-Shannon theorem requires dense sampling for high-frequency signals. To overcome this, grids can adaptive [MLL∗21] or sparse [CLI∗20] to focus the capacity around higher frequency regions, and can be implemented with data structures like hierarchical trees [LGL∗20, TLY∗21] and textures [SHN∗19, PNM∗20].

Grid tiling discretizes the coordinate domain with a grid and define each local region with smaller neural networks [RWG∗13]. This can help learn larger scale signals [JK20] , make inference faster [RPLG21], and can be suitable for parallel computing [SJK21]. Tiling may increase overfitting given sparse training data [HJKK21], and tile boundary artifacts are possible, though network parameter interpolation can be used to reduce boundary artifacts [DVPL21, MMNM21a].

Grids of embeddings can similarly model larger-scale signals [CLI∗20], enable the use of small neural networks [TLY∗21], and benefit from interpolation [LGL∗20]. Grids of embeddings can also be generated from other neural fields [MLL∗21], generative models [ILK21], images [SSSJ20, TY21, CMK∗21], or user input [HMBL21]. Please see Section 3 for a more detailed discussion.

Irregular grids discretize the coordinate domain with a grid that does not follow a regular sampling pattern (and hence avoding the Nyquist-Shannon sampling limit). These can be morphed to adaptively increase capacity in complex data regions. They may declare connectivity between coordinates explicitly such as in meshes, or implicitly such as in Voronoi cells. They too may also be organized into hierarchies, such as within a BVH or scene graph [GSRN21, OMT∗21].

Point clouds are a collection of sparse discrete coordinates. Each location can hold an embedding [TTG∗20] or a network [RJY∗21]. Although sparse in its support, point clouds can volumetrically define regions through Voronoi cells via nearest neighbor interpolation. For continuous interpolation, we can use Voronoi cells with natural neighbor interpolation [Sib81] or soft-Voronoi interpolation [WPLN∗20].

These are also a collection of points but where each has an orientation and a bounding box or volume [WSH∗19]. Neural field parameters are stored at each point or at each vertex of the bounding volume, and can store embeddings [OMT∗21] or networks [GFWF20, ZLY∗21].

Meshes are a common data structure in computer graphics with well understood properties and processing operations. For triangle meshes, embeddings can be stored on vertices [PZX∗21] and interpolated with barycentric interpolation. For complex polygons, mean-value coordinates [Flo03] and harmonic coordinates [JMD∗07] are options.

Forward Maps

In many applications, the reconstruction domain (how we represent the world) is different from the sensor domain (how we observe the world). Solving an inverse problem recovers the reconstruction from observations obtained from sensors, i.e., finding the parameters Θ\Theta of a neural field Φ\Phi given observations from the sensor Ω\Omega.

We represent the (unknown) reconstruction as a neural field Φ:X→Y\Phi:\mathcal{X}\to\mathcal{Y} that maps world coordinates xrecon∈X\mathbf{x}_{\text{recon}}\in\mathcal{X} to quantities yrecon∈Y\mathbf{y}_{\text{recon}}\in\mathcal{Y}. A sensor observation is often also a field Ω:S→T\Omega:\mathcal{S}\to\mathcal{T} that maps sensor coordinates xsens∈S\mathbf{x}_{\text{sens}}\in\mathcal{S} to measurements tsens∈T\mathbf{t}_{\text{sens}}\in\mathcal{T}.

A forward map is an operator F:(X→Y)→(S→T)F:(\mathcal{X}\to\mathcal{Y})\to(\mathcal{S}\to\mathcal{T}), that is a mapping of functions. The forward map may depend on additional parameters, and may be composed with downstream operators such as sampling or optimization. We call a forward map parameter differentiable if for y=F(Φ(x))\textbf{y}=F(\Phi(\mathbf{x})) we can calculate the derivative ∂y∂θ\frac{\partial\textbf{y}}{\partial\theta}, and input differentiable if we can calculate ∂y∂x\frac{\partial\textbf{y}}{\partial\textbf{x}}.

Given a parameter differentiable forward map, we can solve the following optimization problem to recover the neural field Φ\Phi:

using differentiable programming and an algorithm such as stochastic gradient descent.

In the 2D image example, the forward map can be modeled with equations of radiative transfer, which integrate over a volumetric anisotropic vector field. Neural Radiance Fields (NeRFs) [MST∗20] are one example that recovers isotropic density and anisotropic radiance fields from 2D image measurements. Even though many inverse problems are ill-posed with no guarantee that a solution exists or is unique, empirically forward maps help us to find good solutions in a variety of applications (Part II).

We define a renderer as a forward map which converts some neural field representation of 3D shape and appearance to an image. Renderers take as input camera intrinsic and extrinsic parameters, as well as the neural field Φ\Phi to generate an image. Extrinsic parameters define the translation and rotation of the camera, while intrinsic camera parameters define any other information for the image formation model such as field of view and lens distortion [HZ03]. Renderers often utilize a raytracer which takes as input a ray origin and a ray direction (i.e., a single pixel from the camera model), and returns some information about the neural field. The information may be geometric like surface normals and intersection depth, or they may also aggregate or return some arbitrary features.

If a neural field represents a shape’s surface, we can use several methods to obtain geometric information such as the surface normal and intersection point. For common shape neural fields representations such as occupancy and signed distance fields, ray-surface intersection amounts to a root finding algorithm. One such method is ray marching, which takes discrete steps on a ray to find the surface. Given the first interval, a method such as the Secant method can be used to refine the root within the interval. This method will fail to converge at the right surface if the surface is thinner than the interval which discrete steps are taken. Taking jittered steps can alleviate this by stochastically varying the step length and using supersampling [DW85]. Interval arithmetic can be used to guarantee convergence at the cost of iteration count [FSSV07]. In some cases, a neural network may be employed within ray marching to predict the next step [SZW19, NSP∗21]. If the surface is Lipschitz-bounded (as is the case for signed distance fields), then sphere tracing [Har96] can efficiently find intersections with guaranteed convergence. Segment tracing [GGPP20] can further speed up convergence at the cost of additional segment arithmetic computation at each step.

These methods are all differentiable, but naively back-propagating through the iterative algorithm is computationally expensive. Instead, the surface intersection point alone can be used to compute the gradient [NMOG20, YKM∗20]. If a hybrid representation is used, we can exploit a bounding volume hierarchy to speed up raytracing. In some cases, the data structure can also be rasterized onto the image to reduce the number of rays.

Once the ray-surface intersection point has been retrieved, we can calculate the radiance contribution from the point towards the camera. This is done with a bidirectional scattering distribution function (BSDF) [CT82, BS12] which can be differentiable or be parameterized as a neural field [GN98]. To aggregate lighting contribution to the point, the most phyiscally accurate method is to solve Kajiya’s rendering equation [Kaj86] with a multi-bounce Monte Carlo algorithm [PJH16]. However, this can be especially computationally expensive combined with a neural field.

To overcome this issue, approximations of incident lighting (related to precomputed radiance transfer) can be used. This include cubemaps [Gre86] or the use of spherical basis functions such as spherical harmonics [Gre03]. Neural fields can also be employed to approximate incident lighting [WPYS21]. In the case where the exact material properties or the lighting environment does not need to be modeled, a neural field with position and view direction as input can directly model the radiance towards the camera [MST∗20].

The surface rendering equation cannot model scenes with inhomogeneous media such as clouds and fog. Even for scenes with opaque surfaces, inverse surface rendering has difficulty recovering high frequency or thin geometry such as hair and other meso-scale details because the gradients are only defined at surfaces. Instead of Kajiya’s rendering equation, volume rendering uses the volume rendering integral [KH84] based on the equations of radiative transfer [Cha13], for which the integral can be numerically approximated using quadrature (in practice, stochastic ray marching [PH89, MST∗20]).

In a differentiable setting, under the assumptions of exponential transmittance, integration can be performed with a simple cumulative sum of samples across a ray, making backpropagation efficient and dense with respect to coordinates [MST∗20]. This also propagates gradients throughout space, making the optimization easier. Non-exponential formulations exist [LSCL19, VJK21]; some are not physically accurate but work as an approximation.

One of the important factors in stochastic ray marching-based volume rendering is the number of samples. A higher sample count will mean more accurate models, but at the cost of computational cost and memory. Some approaches to mitigate this include using a coarse neural field model to importance sample [MST∗20], using a classifier network [NSP∗21], or using analytic anti-derivatives [LMW21b].

Many differentiable renderers try to combine the strengths of a opaque surface-based renderer and a volume renderer. Volume rendering is typically under-constrained due to the stochastic samples taken on intervals, resulting in noise near surfaces. Surface rendering only provides gradients at the object surface, which do not smoothly propagate across the spatial domain.

We can combine approaches by re-parameterizing the implicit surface as a density field with soft boundaries using a Laplace distribution [YGKL21], a logistic density function [WLL∗21], a Gaussian distribution, or a smoothed step function. We can also importance sample around the surface intersection point [OPG21].

Physics-informed Neural Networks

Partial differential equations (PDEs) are also powerful forward modules that map network outputs to gradient space supervision. Most of the works thus far have relied on a completely data-driven paradigm where additional constraints can be imposed by the choice of representation or implicit biases and invariance from the network architecture. Another class of methods, known as physics-informed neural networks (PINNs), use learning bias [KKL∗21] which supervises boundary and initial values (from an incomplete simulation or observations) using a loss, and the rest of space by sampling or regularizing with equations of physics, typically partial differential equations (PDEs).

In visual computing, the PINN paradigm is often seen with signed distance functions [GYH∗20]. One of the core properties of an SDF is that they satisfy the Eikonal equation:

. The boundary values are the point cloud X\mathcal{X} which correspond to the 0-level set. In the special case when f(x)=1f(x)=1, u(x)u(x) becomes an SDF. The Eikonal loss is therefore: L=∑x∈X∥∥∇Φ(x)∥−1∥\mathcal{L}=\sum_{\mathbf{x}\in\mathcal{X}}\|\|\nabla\Phi(x)\|-1\|.

PDEs are a natural description of the dynamics in the natural world. As such, a large collection of PDEs have been proposed in the discipline of physics, and naturally many of those PDEs have been used in conjunction with Neural Fields as PINNs. We refer readers to Karniadakis et. al for a much more comprehensive review of PINNs [KKL∗21].

Identity Mapping Function

In some applications, the sensor domain may be the same as the reconstruction domain. In these cases, the forward model is the identity mapping function. The task is simply overfitting a neural field to data. Some examples are: supervising a neural signed distance field directly by ground-truth values [PFS∗19], using neural fields to "memorize" images, or audio signals [SMB∗20].

Network Architecture

The design choices that we make about the structure and components of a neural network have a significant impact on the quality of the field it parameterizes. The most obvious of these choices is the network structure itself, such as how many layers are in the MLP and how many neurons are in each layer. We assume that a reasonable structure with enough learning capacity has been chosen, and focus our discussion on other design decisions that provide inductive biases for effectively learning neural fields.

Real-world signals are complex, making it challenging for neural networks to achieve high fidelity. Furthermore, neural networks are biased to fit functions with low spatial frequency [RBA∗19, HMZ∗21]. Several designs address this shortcoming.

Sinusoidal functions are widely used to equip neural fields with the ability to fit high-frequency signals. Proposed by [ZBDB20a], they can be formally written as γi\gamma_{i} where:

These sinusoidal embeddings, also known as Fourier feature mapping, were subsequently popularized for the task of novel view synthesis [MST∗20, TSM∗20]. In the context of neural tangent kernels [JGH18], sinusoidal positional encoding can be shown to induce a kernel with a spatial selectivity that increases with the frequency of the functions ϕi\phi_{i}. Positional encodings can thus also be seen as controlling the interpolation properties of a neural field.

Wang et al. [WLYT21] showed that the choice of the frequency of γi{\gamma_{i}} biases the network to learn certain, band-width limited frequency content, where lower encoding frequencies result in blurry reconstruction, and higher encoding frequencies introduce salt-and-pepper artifacts. For more stable optimization, one approach is to mask out high-frequency encoding terms at the beginning of the optimization, and progressively increase the high-frequency encoding weights in a coarse-to-fine manner [LMTL21, BHumRZ21]. SAPE proposed a masking scheme [HPG∗21] that also allowed the encoding of spatially varying weights. Finally, alternative positional encoding functions ϕi\phi_{i} have also been proposed [ZRL21, WLYT21, MRNK21]. Zheng et al. [ZRL21] conducted a comprehensive study of various positional encoding functions.

An alternative approach to enable the fitting of high-frequency functions is to replace standard, monotonic nonlinearities with periodic nonlinearities, such as the periodic sine as in SIREN [SMB∗20] enabling fitting of high-frequency content. While a derivation of the properties of the neural tangent kernel is outstanding, some theoretical understanding can be gained by analyzing the relationships of the gradients of the output of the Neural Field with respect to neighboring coordinate inputs—for high frequencies of the sinusoidal activations, these gradients have been shown to be orthogonal, leading to an ability to perfectly fit values at any input coordinates without any interpolation whatsoever. It has been pointed out that positional encoding with Fourier features is equivalent to periodic nonlinearities with one hidden layer as the first neural network layer [BHumRZ21].

Integration and Derivatives

A key benefit of neural fields is that they are highly flexible general function approximators whose derivatives ∇xΦ(x)\nabla_{\mathbf{x}}\Phi(\mathbf{x}) are easily obtained via automatic differentiation [G∗89, BPRS18], making them attractive for tasks that require supervision of derivatives via PDEs. This, however, raises additional requirements in terms of the architecture of the MLP parameterizing the field. Specifically, the derivatives of the network must be nontrivial (i.e., nonzero) to the degree of the PDE (Figure 6, left). For instance, to parameterize solutions of the wave equation — a second-order PDE — the second-order derivatives of the neural field must be nontrivial. This places restrictions on the activation function used in the network. For instance, ReLU nonlinearity has piece-wise constant first derivatives, and zero high-order derivatives, and so cannot generally parameterize solutions to the wave equation. Other common nonlinearities, such as softplus, tanh, or sigmoid, may address this issue, but often need to be combined with positional encodings to parameterize high-frequency content of neural fields. Alternatively, the periodic sine may be used as an activation function [SMB∗20].

In some cases, we are further interested in solving an inverse problem where we measure the derivative quantity of a field, but are later interested in obtaining an expression for an integral over the neural field. In these cases, it is possible to sample the neural field and approximate the integral using numerical quadrature, as is often done for instance in volume rendering [MST∗20]. However, the number of samples needed to approximate the integral may be arbitrarily large for a given accuracy.

While automatic differentiation is ubiquitous in deep learning, automatic integration is not well-explored. The seminal work AutoInt [LMW21b] proposed to directly parameterize the antiderivative of the field as a neural field Φ\Phi. We may then instantiate the computational graph of its gradient network ∇xΦ(x)\nabla_{\mathbf{x}}\Phi(\mathbf{x}), and fit the gradient network to the given derivative values. At test time, a single integral can be calculated in one forward pass, as a consequence of the universal theorem of calculus (Figure 6, right).

Manipulating Neural Fields

While many data structures in visual computing have post-processing tools (e.g. smoothing a polygonal mesh), neural fields have limited tools for editing and manipulation, which significantly limits their use cases. Fortunately, this is an active area of research.

The simplest approach to editing neural fields is to transform its spatial or temporal coordinate inputs. A simple rigid translation of a neural field can be expressed as g(x)=f(x+b)g(x)=f(x+b). We discuss other more sophisticated coordinate remapping methods below (Table 4).

An intuitive solution for controlling shape or appearance neural fields is to leverage explicit shape information. For object modeling problems (Section 1), structural priors are often available in the form of bounding boxes and coarse explicit geometry. Object bounding boxes offer a convenient, explicit handle to add and rigidly transform each object [ZLY∗21, OMT∗21, LGL∗20].

For articulated objects, the kinematic chain offers explicit control over object geometry through its joint angles. At inference time, animating object geometry is achieved by changing joint angle inputs [MQK∗21]. While this approach is effective for a small number of joints, it may introduce spurious correlations in case of long kinematic chains such as human body [LMR∗15]. To avoid this issue, we can represent target shapes as the composition of local neural fields [GCS∗20, DLJ∗20]. Each local field is defined with respect to the joint transformations on a template mesh, and composed by max operation [DLJ∗20] or weighted blending based on relative joint locations [SYZR21, NSLH21]. The primary drawback of articulated compositional neural fields is that each local field is independently modeled, often producing artifacts around joints with unobserved joint transformations. An alternative is to warp an observation space into a canonical space, where the reconstruction is defined, using the articulation of a template model [HXL∗20, SYMB21]. The warping is commonly represented as linear blend skinning (LBS), where the deformations of surface is defined as the weighted sum of joint transformations. The blending weights are computed by a nearest-neighbor query on a template mesh [HXL∗20] or learning as neural skinning fields [SYMB21, MZBT21, CZB∗21, TSTPM21] (Table 4).

Unlike articulated objects with known transformations, modeling general dynamic scenes requires flexible representations that can handle arbitrary transformations. As target geometry and appearance are often modeled with neural fields, a continuous transformation of the field itself is a natural choice. Learning neural fields for spatial transformations is a highly under-constrained problem without 3D supervision [NMOG19]. This motivates the use of regularization loss terms based on physical intuitions:

Smoothness: The first derivative of warp fields w.r.t. spatiotemporal coordinates should be smooth, assuming no sudden movements. This is necessary to constrain unobserved regions [GSKH21, PSB∗21, TTG∗21, DZY∗21].

Sparsity: 3D scenes generally contain large empty space. Thus, enforcing sparsity of the predicted motion fields avoids sub-optimal local minima [GSKH21, XHKK21, TTG∗21, DZY∗21].

Cycle Consistency: If a representation provides both forward and backward warping (e.g. [GSKH21, WELG21, LNSW21], forward/inverse LBS [SYMB21, MZBT21]), we can employ cycle consistency as a loss function. Unlike other regularization terms, cycle consistency does not dampen the prediction as the global optimum should also satisfy this constraint.

Auxiliary Image-space Loss: Image-space information such as optical flows [LNSW21, XHKK21, DZY∗21, GSKH21, WELG21] and depth maps [LNSW21, GSKH21, WELG21] can also be used in auxiliary loss functions.

Conditioning the neural field on temporal coordinates allows time editing such as speed-up/slow-down (Φ(a⋅t)\Phi(a\cdot t), offset (Φ(t+b)\Phi(t+b)), and reversal (Φ(−t)\Phi(-t)) [ZLY∗21, LNSW21].

Editing via Network Parameters

Neural fields can also be edited by directly manipulating the latent features or the learned network’s weights. These methods are task-agnostic, and are applicable to editing geometry, texture, as well as other physical quantities beyond visual computing. A subset of parameters can be selectively modified (e.g., geometry network, texture latent code). Network weight editing and manipulation methods include (see also Table 3):

Latent Code Interpolation/Swapping: For neural fields conditioned on latent codes, interpolation or sampling in the latent space can change properties of the representation [CZ19].

Latent Code/Network Parameters Fine-Tuning: After pre-training, we can fine-tune parameters to fit new, edited observations at test time. Editing may leverage explicit representations, such as sketches on 2D images [LZZ∗21], or moving primitive shapes [HAESB20]. The neural field is coupled to the explicit supervision via differentiable forward maps.

Editing via Hypernetworks: Hypernetworks can learn to map a new statistical distribution (e.g., a new texture style [CTT∗21]) to a pre-trained neural field, by replacing its parameters.

Chapter 1 Part II. Applications of Neural Fields

In Part II of this report, we review neural field works based on their application domain. The recent explosion of interest has seen neural fields used for a wide range of problem in visual computing such as 3D shape and appearance reconstruction, novel view synthesis, human modeling, and medical imaging. Additionally, neural fields are increasingly being used in applications outside of visual computing including in physics and engineering. For each application domain, we limit our discussion to only neural field work while providing pointers to more traditional methods as context.

The reconstruction of representations of 3D scenes from real-world measurements is critical for robotics and autonomous vehicles, and for graphics applications like games and visual effects. Unsurprisingly, the earliest work that we are aware of that uses neural fields was for 3D shape representation [LS04]. 3D scenes have properties including geometry, appearance, materials, and lighting for both static and dynamic parts. Reconstruction is the solution to an inverse problem that maps available observations to a representation. For 3D, available observations are discrete (due to sensors), often sparse (few images), incomplete (partial point clouds), residing in a lower dimension (2D images), or lack vital topological information (point clouds). In this section, we discuss reconstructing, displaying, and editing 3D scenes while identifying which techniques from Part I that they use.

Work on neural fields for geometry reconstruction often focuses on learned priors for reconstruction. AtlasNet [GFK∗18b] proposed to represent 3D shapes via predicting a set of 2D neural fields that lift 2D local patch coordinates to 3D (an “atlas” of patches). Concurrently, FoldingNet [YFST18] proposed a similar idea but is not continuous. These patches were globally conditioned on a latent inferred from either a PointNet [QSMG17] or ResNet [HZRS16] encoder. IM-Net [CZ19] and Occupancy Networks [MON∗19] proposed to represent 3D geometry via an occupancy function. Both proposed to generalize across ShapeNet [CFG∗15] via global conditioning-via-concatenation, where the latent was regressed from either a CNN (from an image) or a PointNet encoder (from a point cloud). Concurrently, DeepSDF [PFS∗19] proposed to represent 3D surfaces via their signed distance function, similarly generalizing across ShapeNet objects with global conditioning-via-concatenation, but performing inference in the auto-decoder framework instead. Figure 1 visualizes the SDF representation. Even though these methods were proposed in conjunction with learning a shape prior, it is possible to overfit a DeepSDF or neural occupancy function on a single 3D shape.

AtlasNet was further extended [DGF∗19] by learning elementary structures in a data-driven manner, and inferring point correspondences across instances. Four concurrent papers proposed to leverage local features stored in voxel grids for local conditioning (Section 1) to improve on prior-based inference. IF-Net [CAPM20], Chibane et al. [JSM∗20], and Convolutional Occupancy Networks [PNM∗20] use 3D CNNs or PointNet operating to process voxelized point clouds into embedding grids to locally parameterize an occupancy network or signed distance function, where conditioning proceeds via concatenation. In contrast, Chabra et al. [CLI∗20] similarly leverages local conditioning from a 3D voxel grid, but infers the latent codes in the voxel grid via auto-decoding. CvxNet [DGY∗20] proposes to represent a 3D shape as a composition of implicitly defined convex polytopes. LDIF [GCS∗20] proposes to represent a 3D shape via a collection of local occupancy functions whose weighted sum represents the global geometry. Littwin et al. [LW19] infer an occupancy function via an image encoder and global conditioning via a hypernetwork.

Many methods aim to improve the quality of learned shapes. MetaSDF [SCT∗20] was the first method to use gradient-based meta-learning (Section 1) for neural fields, using it to infer 3D signed distance functions from point clouds. Yang et al. [YWC∗21] show that optimizing not only the embeddings but also the network weights regularized to be close to the original weights can better resolve ambiguities in unobserved regions. IGR [GYH∗20] uses a neural field to parameterize an SDF, but instead of requiring ground-truth signed distance values in a identity forward map, they leverage a partial differential equation forward map via the Eikonal equation to learn SDFs from a point cloud. Atzmon et al. [AHY∗19] derive analytical gradients of the 3D position of points on a level set with respect to the neural field parameters. SIREN [SMB∗20] similarly learns an SDF with the Eikonal equation from oriented point clouds, but enables encoding of high-frequency detail by leveraging sinusoidal activations. Deep Medial Fields [RLS∗21] proposes to parameterize geometry via its medial field, the local thickness of a 3D shape, enabling faster level-set finding and other applications. NGLoD [TLY∗21] learns multiple level of details for SDFs with a compact hierarchical sparse octree of embeddings. SAL [AL20a] uses a neural field to parameterize SDF, and show that the SDF can be learned by optimizing against the unsigned distance function of the point cloud given suitable initialization. SALD [AL20b] extends this with supervision of normals. Neural Splines [WTBZ21] use as input oriented point clouds and use kernel regression of the neural field to optimize for normal alignment. PIFu [YYTK21] performs prior-based reconstruction of geometry and appearance reconstruction by extracting features from images with a fully convolutional CNN, and, when querying a 3D point, projecting it on the image plane to use image features for local conditioning-via-concatenation. Concurrently, DISN [XWC∗19] enhances single-view reconstruction by using a camera pose estimation that allows to project 3D coordinates onto the image plane and gather local CNN features. Combined with the global feature, the result is a more accurate SDF.

A major breakthrough in 3D reconstruction was the adoption of differentiable rendering (Section 3), which allowed reconstruction of 3D neural fields representing shape and/or appearance given only 2D images, instead of 3D supervision. This has significant implications since 3D data is often expensive to obtain, while 2D images are omni-present. A particularly important social implication is that non-experts can become 3D content creators, without the barrier of specialized hardware or capture rigs. SRNs [SZW19] proposed using a differentiable sphere-tracing based renderer to reconstruct 3D geometry and appearance from only 2D observations. It leveraged global conditioning via a hypernetwork and inference via auto-decoding to enable reconstruction of a 3D neural field of geometry and appearance from only a single image for the first time. Concurrently, Liu et al. [LSCL19] use a CNN to predict an embedding to parameterize occupancy, and a form of differentiable volume rendering to produce a silhouette. Similar to DeepSDF, Occupancy Networks, and IM-Net, SRNs were designed to generalize, although they can be overfit given lots of 2D observations of a single 3D scene.

SDF-SRN [LWL20] enables learning an SRN from only a single observation per object at training time by enforcing a loss on the 2D projection of the 3D signed distance function. Kohli et al. [KSW20] leverages SRNs as a representation learning backbone for self-supervised semantic segmentation. Liu et al. [LZP∗20] similarly reconstruct 3D geometry from 2D images with differentiable sphere-tracing. DVR [NMOG20] represents geometry via an occupancy function, and finds the zero-level set via ray-marching, subsequently querying a texture network for RGB color per ray. Importantly, they derive an analytical gradient of the ray-marcher, which significantly reduces memory consumption at training time, but also requires ground-truth foreground-background masks. They demonstrate the learning of appearance and shape priors across scenes via encoder-based conditioning to enable reconstruction from a single image, as well as overfitting on single scenes for high-quality, watertight 3D reconstruction.

NeuralVolumes [LSS∗19] first proposed differentiable volume rendering, but leverages linearly interpolated voxel grids of color and density as a representation. Though the paper mentions parameterization of color and density functions as neural fields, this was only a part of the ablation studies, and reportedly under-performed a 3D CNN decoder in terms of resolution. cryoDRGN [ZBDB20b] implemented a differentiable volume renderer for the cryo-electron microscopy forward model to reconstruct protein structure from cryo-electron microscopy images, and proposed the positional encoding to allow fitting of high-frequency detail. Here, the neural field is parameterized in the Fourier domain, and the forward model is implemented via the Fourier slice theorem. NeRF [MST∗20, TSM∗20] combined volume rendering with a single ReLU MLP, parameterizing a monolithic neural field, and added positional encodings (Figure 2). By fitting a single neural field to a large number of images of a single 3D scene, this achieved photo-realistic novel view synthesis from only 2D images of arbitrary scenes for the first time. The visual quality and elegant formulation of NeRF has since inspired a large collection of follow-up work. Nerf++ [ZRSK20] improves representation of unbounded 3D scenes via an inverted-sphere background parameterization. Reizenstein et al. [RSH∗21] and Arandjelović et al. [AZ21] propose attention-based accumulation of samples along a ray. DoNeRF [NSP∗21] proposes to jointly train a NeRF and a ray depth estimator for fewer samples and faster rendering at test time. [GKJ∗21, YLT∗21a, RPLG21, HSM∗21, LGL∗20] propose various variants of local conditioning (without generalization, for overfitting a single scene) to speed up the rendering of NeRFs. Mip-NeRF [BMT∗21] proposes to control the frequency of the positional encoding for multi-scale resolution control. NeRF−−, BARF and iNeRF [WWX∗21, LMTL21, YCFB∗20] propose to back-propagate into camera parameters to enable camera pose estimation given a reasonable initialization.

PixelNeRF [YYTK21] and GRF [TY21] perform prior-based reconstruction by extracting features from images with a fully convolutional CNN, and, when querying a 3D point, projecting it on the image plane to use image features for local conditioning-via-concatenation, similar to PIFu and DISN [SHN∗19, XWC∗19]. With more context views, a similar approach can be used for multi-view-stereo-like 3D reconstruction [CXZ∗21, CBLPM21]. NeRF-VAE embeds a globally-conditioned NeRF with encoder-based inference in a VAE-like framework [KSZ∗21]. While volume rendering has better convergence properties than surface rendering and enables photorealistic novel view synthesis, the quality of the reconstructed geometry is worse, due to the lack of an implicit, watertight surface representation. IDR [YKM∗20] leverages an SDF parameterization of geometry, a sphere-tracing based surface-renderer, and positional encodings to enable high-quality geometry reconstruction. Kellnhofer et al. [KJJ∗21] distill an IDR model into a Lumigraph after rendering to enable fast novel view synthesis at test time. Concurrently, UNISURF [OPG21], NeuS [WLL∗21], VolSDF [YGKL21] propose to relate the occupancy function of a volume to its volume density, thereby combining volume rendering and surface rendering, leading to improved rendering results and better geometry reconstruction. Ray marching requires many samples along a ray to faithfully render complex 3D scenes. Even for relatively simple scenes, rendering requires hundreds or even thousands of evaluations of the neural scene representation per ray.

To overcome this limitation, we can parameterize the light field of a scene, which maps every ray to a radiance value. This enables real-time novel view synthesis with a single neural field sample per ray, and geometry extraction from the neural light field without ray-marching. Figure 3 displays the difference between ray-marching and rendering in light fields.

To prevent overfitting to the input views and allow view synthesis, we can learn multi-view consistency via global conditioning, hypernetworks, and inference via auto-decoding [SRF∗21], or through ray embedding spaces [AHZ∗22]. Alternatively, densely sampled rays of a single scene allow easier overfitting for view synthesis. NeuLF [LLYX21] parameterizes forward-facing scenes via light fields, overfitting on single scenes, and addresses multi-view consistency via enforcing similarity of randomly sampled views with the context views in the Fourier domain. Similar to light fields, NeX [WPYS21] parameterizes a set of multi-plane images as a 2D neural field, where the network inputs are pixel locations. Rendering is computationally efficient without ray marching.

Reconstruction of Scene Material and Lighting

The goal of material reconstruction is to estimate the material properties of a surface or participating media from sparse measurements such as images. For opaque surfaces, this may be the parameters of a bidirection scattering distribution function (BSDF). For participating media, this may be phase functions. This is a difficult problem, because materials are diverse and create complex light transport effects. Even just for BSDFs, there is a whole taxonomy [MDH∗20] of characteristic properties. In addition, to accurately estimating materials, we must also perform lighting reconstruction or have prior knowledge of the lighting.

The forward modeling of light transport involves surface rendering [Kaj86] and volume rendering [KH84] for participating media, which both equations having recursive integrals with no closed form solution for forward modeling. Approaches to inversely solve these equations differ in the degree of approximation they make. Because neural networks are general function approximators, they can be useful for estimating arbitrary functions and integrals that are hard to solve. Some methods reconstruct appearance as approximate incident radiance, and other methods attempt to separate appearance into materials via explicit scattering distribution functions and lighting, enabling applications such as relighting.

Many papers use a neural field to parameterize the parameter space of existing material models. Neural Reflectance Fields [BXS∗20a] can reconstruct both the SVBRDF and geometry by assuming a known point light source, and use a neural field which parameterizes density, normals, and parameters of a microfacet BRDF [WMLT07] with a volume rendering forward map with one bounce direct illumination. NeRV [SDZ∗21] extends this to handle more varied lighting setups with an environment map and a one bounce indirect illumination forward map, along with an additional neural field which parameterizes visibility, but assume the environment map is known a priori. NeRD [BBJ∗21] also use a neural field which paramterizes density, normals, and the parameters of a Disney BRDF model [CT82, BS12], but remove any assumptions on the lighting by using spherical gaussians to represent lighting of which the parameters are directly optimized. They do not model visibility or indirect illumination. PhySG [ZLW∗21] also directly optimize spherical gaussians, but uses a neural field to parameterize an SDF and a Ward BRDF [War92] along with a hard-surface-based forward map to improve surface reconstruction accuracy. NeRFactor [ZSD∗21] learns an embedding-based neural field of BRDF parameters as a prior on the material parameter space (Figure 4).

Dynamic Reconstruction

In addition to novel-view synthesis, dynamic scene reconstruction also allows measurement, mixed reality, and visual effects. The challenges of modeling dynamic scenes are that the input data is even sparser in spacetime and often no 4D ground-truth data are available. With careful design choices, regularization loss terms, and the strong inductive bias of neural fields, several works have proposed solutions to this inverse problem. Figure 5 shows one of such approaches. We review the existing approaches based on warp representation and how to embed temporal information. For dynamic reconstruction of humans, please refer to Section 2.

Modeling temporally changing objects or scenes requires additional embedding that encodes frame information [RBZ∗20]. Occupancy Flow [NMOG19] models a dynamic component as a neural field conditioned by normalized temporal coordinates. Similarly dynamic scene reconstruction methods based on differentiable rendering often condition neural fields with temporal coordinates [XHKK21, LNSW21, PCPMMN21, DZY∗21]. Another approach is to jointly optimize per-frame latent code [PSB∗21, LSZ∗21, TTG∗21, PSH∗21, ALG∗21] as embedding. While embedding based on temporal coordinates automatically incorporates temporal coherency as inductive bias, per-frame latent codes can enable the captures of more scene details.

To model dynamic scene from limited input data, we can split the problem into modeling a scene in the canonical space and warping it into each time frame. In Section 1, we provide common warp representations and techniques to regularize the warp fields. While most approaches use the learned warping function as the final scene representation, several works use warp fields only for regularizing the predicted radiance fields [LNSW21, GSKH21]. Unlike other dynamic reconstruction approaches, where non-rigid complex warping is modeled by neural fields, STAR [YLSL21] models motion as a global rigid transformation to primarily focus on tracking of a foreground object.

Digital Humans

Human shape and appearance has received special attention in computer vision and graphics in the last decade, and is one of the most popular application areas of neural fields. The adoption of neural fields has produced unprecedentedly high-quality synthesis and reconstruction of human faces, bodies, and hands. The state-of-the-art continues to evolve quickly.

The data-driven parametric morphable model was introduced by Blanz and Vetter [BV99]. However, these explicit representations lack realism and impose topological limitations making it difficult to model hair, teeth, etc. The expressiveness of fields, such as SDF and radiance fields, has made them an excellent candidate to address these limitations.

Given a collection of high-quality 3D scans, i3DMM [YTB∗21] used coordinate-based neural networks that predict SDF and color to develop a 3D morphable head model with hair. Contrary to mesh-based representation, they learn an implicit reference shape as well as a deformation for each shape instance enabling editing abilities by disentangling color, identity, facial expressions, and hairstyle. Ramon et al. [RTE∗21] use coordinate-based neural networks building upon IDR [YKM∗20] to reconstruct a 3D head. To reduce the required amount of input images, a pre-trained DeepSDF-based latent space is used to regularize the test-time optimization. SIDER [CAMNS21] incorporates coarse geometric guidance via a fitted FLAME model [LBB∗17] followed by implicit differentiation, enabling single-view optimization of facial geometry using SDF parameterized by a coordinate-based network.

While the aforementioned approaches focus on geometric accuracy assuming hard surfaces, NeRF have recently been adopted for photo-realistic view synthesis of human heads. Nerfies [PSB∗21] utilizes a casual video footage captured with a moving hand-held camera to learn radiance fields and deformation fields of a human head. HyperNeRF [PSH∗21] extends Nerfies by incorporating auxiliary hyper dimensions to handle large topological change.

Another line of works enable the semantic control of radiance fields by conditioning on head pose and facial expression parameters obtained from a 3D morphable face model [GTZN21, ASS21]. Supervised by multi-view video sequences, Wang et al. [WBL∗21] use a variational formulation to encode dynamic properties in spatially varying animation codes stored in voxels. Notably, the above methods require per-subject training, and reconstruction from limited observations remains challenging. Several works show promise in few-shot, generalizable reconstruction by exploiting test-time fine-tuning [GSL∗20] or pixel-aligned image features [RZS∗21].

Neural fields have demonstrated efficacy in 3D reconstruction of clothed humans from image inputs [SHN∗19, SSSJ20, HCJS20, HXL∗20, ZYLD21] or point clouds [CAPM20, BSTPM20a]. Due to substantial variations in shape and appearance of clothed human bodies, a global latent embedding does not lead to plausible reconstruction. PIFu [SHN∗19] addresses this by introducing a coordinate-base neural network conditioned on pixel-aligned local embeddings. Its followup works also employ the framework of PIFu for human digitization tasks from RGB inputs [LXS∗20, HCJS20, SJL∗21, YWM∗21, HZJ∗21] or RGB-D inputs [LYP∗20, YZG∗21]. To further improve the fidelity of reconstruction, multi-level feature representations are shown effective [SSSJ20, CAPM20]. Yang et al. [YWM∗21] jointly predict skinning fields and skeletal joints for animating the reconstructed avatars. Also, several works explicitly leverage a parametric body model such as SMPL [LMR∗15] to improve robustness under different poses [ZYLD21] and to enable an animation-ready avatar reconstruction from a single image [HXL∗20, HXS∗21].

Human bodies are dynamic as they are both articulated and deformable. Several works show that providing the structure of human bodies significantly improves the learning of radiance fields [PZX∗21, PDW∗21, CZK∗21, LHR∗21]. Given a fitted body model to images, Neural Body [PZX∗21] diffuses the latent embeddings on the body via sparse convolution, which are inputs to the neural field. Neural Actor [LHR∗21] projects queried 3D points onto the closest point on the fitted body mesh, and retrieves latent embeddings on the UV texture map. A-NeRF [SYZR21] explicitly incorporates joint articulations, and learns radiance fields in the normalized coordinate space. In addition, as shown for face modeling, pixel-aligned local embeddings are highly effective to support novel-view rendering of unseen subjects [SZZ∗21, KKCF21].

Template-mesh registration is another important task in body modeling. IP-Net [BSTPM20a] proposes to jointly infer inner body and clothing occupancy fields given input scans to aid registration. 3D-CODED [GFK∗18a] proposes Shape Deformation Networks to fit a template to a target shape and infer correspondences. Halimi et al. [HIL∗20] show template-based shape completion by conditioning the deformation field on the part and whole encoding. LoopReg [BSTPM20b] presents a self-supervised learning of dense correspondence fields on the predicted implicit surface to a template human mesh, which improves the robustness of surface registration. A continuous local shape descriptor for dense correspondence is proposed by Yang et al. [YLB∗20]. Wang et al. [WGT21] model occupancy fields of input scans in an un-posed canonical space by predicting piece-wise transformation fields (PFT). Since shapes are modeled in the canonical space, we can perform template registration without self-intersection of different body parts.

Lastly, several recent works model a parametric model of human bodies [MZBT21, AXS21], clothed human [SYMB21, TSTPM21, PBT∗21, WMM∗21], hands [KYZ∗20], or clothing [CPA∗21] as neural implicit surfaces. One unique property of human body is the articulation with non-rigid deformations. Occupancy flow [NMOG19] models the non-rigid deformation of human bodies by warp fields. NASA [DLJ∗20] presents articulated, per-body-part occupancy fields to model a pose-driven human body. However, articulated occupancy fields may suffer from artifacts around joints due to discontinuities. Another approach to handle articulation is jointly learning shapes in the un-posed canonical space and transformations from the canonical space to the posed space [SYMB21, MZBT21, PBT∗21, CZB∗21], leading to continuous deformations around body joints. The transformation can be modeled as warp fields (Section 1) in the form of displacements [PBT∗21] or Linear Blend Skinning weights [SYMB21, MZBT21, CZB∗21, TSTPM21, WMM∗21].

Generative Modeling

Assuming a dataset of samples drawn from a distribution y∼D\mathbf{y}\sim\mathcal{D}, generative modeling defines a latent distribution Z\mathcal{Z}, such that every sample y\mathbf{y} can be identified with a corresponding latent z∼Z\mathbf{z}\sim\mathcal{Z}. The mapping from Z\mathcal{Z} to samples from D\mathcal{D}, is performed via a learned generator g\mathbf{g}, parameterized as a deep neural network, g(z)=y\mathbf{g}(\mathbf{z})=\mathbf{y}.

Neural fields provide greater flexibility for generative modeling because their outputs can be sampled at arbitrary resolutions, and because input samples can be shifted by simply applying transforms to the input coordinates. Each sample y\mathbf{y} is associated with a neural field Φ\Phi that can be densely queried across the coordinate domain to obtain the sample y\mathbf{y}. Consistent with the notation in Section 1, latent variables z\mathbf{z} can either globally or locally condition the neural field Φ\Phi, yielding a conditional neural field Φ(x,z)\Phi(\mathbf{x},\mathbf{z}).

GRAF [SLNG20] first adapted this via a neural radiance field and volume rendering forward model, globally conditioned via concatenation. Chan et al [CMK∗21] improved upon GRAF by using sinusoidal activation functions [SMB∗20] in the MLP, and global FiLM conditioning. StyleNeRF [Ano22] combine StyleGAN-style FiLM conditioning with accelerated rendering techniques such as low-resolution rendering and 2D upsampling. By leveraging local conditioning via object-centric compositions (see Section 2), GIRAFFE [NG21b] generates an output image as the composition of multiple neural fields, which allows control over shape, appearance and scene layout (Figure 6). CAMPARI [NG21a] trains an additional camera generator along with the 3D volume generator to generalizes to complex pose distributions.

An advantage of neural fields is that they are, in principle, agnostic to the signal they parameterize. Du et al. [DCTS21] combine global conditioning, the auto-decoder framework, and latent space regularization to learn multi-modal (such as audio-visual) manifolds.

2D Image Processing

A compelling feature of 2D neural fields is the ability to represent continuous images. The first neural network to parameterize an image was demonstrated by Stanley et al. [Sta07]. These networks were not fit via gradient descent, instead relying on architecture search in a genetic algorithm framework, and could thus not represent images with fine detail. Fitting natural images with neural fields in a modern deep-learning framework was first demonstrated by SIREN [SMB∗20] and FFN [TSM∗20], and by [SCA19] who decomposed the image into continuous vector layers. Unlike grid-based convolutional architectures, continuous images can be sampled at any resolution. As a result, they have been used for a variety of image processing tasks described below.

These techniques take an input image and map it to another image that preserves some representation of the content. Common tasks include image enhancement, super-resolution, denoising, inpainting, semantic mapping and generative modeling [IZZE17]. Since this task requires learning a prior from data, an encoder-decoder architecture is often used, where the encoder is a convolutional neural network, and the decoder is a locally conditioned 2D neural field (Section 1).

Chen et al. [CLW21a] propose such a locally conditioned neural field with a convolutional encoder for the task of image super-resolution, naturally leveraging the resolution independence of the neural field decoder. Shaham et al. [SGZ∗21] leverage a convolutional encoder to produce a low-resolution, 2D feature map from an input image, which is then upsampled with nearest-neighbor interpolation to locally condition a decoder neural field. They demonstrate speedups for a variety of image-to-image translation tasks, such as segmentation and segmentation-to-RGB image as compared to the fully convolutional baseline. Neural Knitworks [CCA∗21] discretize the 2D space into patches to introduce the appropriate receptive field, for inpainting, super-resolution, and denoising (Figure 7). CIPS [ADK∗21] proposes an image synthesis architecture whose input pixels coordinates are conditionally-independent given latent vector z. Global information is given by z, which is used by a hypernetwork to modulate the neural field weights. INR-GAN [SIE21] similarly uses a hypernetwork and latent code z to modulate the linear layer weights and biases in the neural field, for image generation. Henzler et al. [HMR20] learns 2D texture exemplars and maps them into 3D by sampling random noise fields at desired positions as the neural texture field input. PiCA [MSS∗21] applies a lightweight SIREN to predict human facial texture over a guide mesh, where the 2D coordinate input lie in the UV space. Li et al. [LTW∗21] use a 2D neural field to parameterize deformation, for recovering images distorted by turbulent refractive media.

X-Fields [BMSR20] parameterizes the image as a 2D neural field conditioned on time and illumination to enable time and illumination interpolation. Alternatively, Nam et al. [NBB21] represents dynamic images (video) with 2D neural fields, with additional homography warp, optical flow, or occlusion operations to model dynamic changes. Kasten et al. [KOWD21] represent dynamic images as time-dependent 2D neural fields, where individual foreground components are segmented as atlas, and alpha composited to obtain the final rendering.

Robotics

Robotics requires complex perception systems that allow agents to efficiently infer, reason about, and manipulate representations of real-world scenes. Many robotics problems share similarities with visual computing, such as using 3D reconstruction for robot navigation, and so neural fields have recently been used in robotics too. We discuss neural fields for robot perception, planning, and control.

Cameras observe the 3D world via 2D images. The projection of a world location onto the image plane is obtained through the extrinsic and intrinsic matrices, where the extrinsic matrix [R∣t][\mathbf{R}|\mathbf{t}] defines the 6DoF transformation between the world coordinate frame and the camera coordinate frame, while the intrinsics matrix K\mathbf{K} describes the projection of a 3D point onto the 2D image plane. Commonly, these camera parameters are obtained via SFM and keypoint matching with off-the-shelf tools like COLMAP [SF16, SZPF16]. Since neural rendering is end-to-end differentiable, camera parameters can be jointly estimated with the neural field making them useful for Simultaneous Localization and Mapping (SLAM) [SLOD21] and Absolute Pose Regression (APR) [CWP21].

There exist numerous ways to parameterize camera intrinsic and extrinsic matrices. Representing the extrinsic matrix, particularly rotation, however, is a long-standing challenge. Since the 3-by-3 rotation matrix lies on SO(3)SO(3), continuity is not guaranteed [ZBL∗19, LEC∗20]. Consequently, alternative parameterizations were used in neural field literature: exponential coordinates [YCFB∗20], Rodriguez formula [WWX∗21], continuous 6D representation [MCL∗21, ZBL∗19], Euler angle[AMBG∗21], and homography warp (for planar scenes) [MCL∗21, YCFB∗20].

Furthermore, joint reconstruction and registration is a long-standing chicken-and-egg problem: camera parameters are needed to reconstruct the scene, and a reconstruction is needed to estimate camera parameters. One simplification is to assume known reconstruction and solve only the registration problem. iNeRF [YCFB∗20] estimates camera poses given pre-trained neural fields, thus inverting the problem statement. Due to the strong assumption, iNeRF has a limited use case of registering new, un-posed images to an already-reconstructed scene. Furthermore, the optimization problem is known to be non-convex. A coarse initialization of camera parameters alleviates this challenge. As such, several works refine camera parameters during reconstruction [AMBG∗21, LMTL21, JAC∗21, WWX∗21, CWP21]. The chicken-and-egg nature of the problem is reflected by NeRF– [WWX∗21], in which the authors jointly optimizes the scene and cameras, but retrain scene reconstruction with the optimized camera parameters for better reconstruction quality.

The more challenging problem of estimating unknown camera parameters, while jointly reconstructing the scene has been looked into by several works [LMTL21, WWX∗21]. However, the scene is assumed to be forward-facing and the camera pose initialization relies on hand-crafted rules (e.g., cameras centered at origin, facing the -z axis, with focal lengths equal to the reference image size [WWX∗21]). GNeRF [MCL∗21] further relaxes these assumptions, and supports inward-facing 360∘360^{\circ} scenes. The optimization between reconstruction and registration proceeds iteratively, rather than jointly, and are softly coupled via a discriminator. Approaches to overcome local minima in optimization and preserve details include coarse-to-fine training [LMTL21, JAC∗21, CWP21], high-frequency positional encoding weights [CWP21, LMTL21], and curriculum learning [JAC∗21] for optimizing focal lengths followed by intrinsics and lens distortion.

In the examples above, supervision is provided via appearance. Additional signals such as depth map from RGB-D sensor could provide strong signals for camera parameter optimization [SLOD21]. Similarly, time-of-flight image may also reduce the problem complexity [ALG∗21]. A closely related topic is object pose estimation. While object pose can be modelled explicitly, a different approach is to predict a probability distribution of pose over SO(3)SO(3) [MEJ∗21]. The probability distribution is itself a neural field, mapping rotation to probability. The approach is especially useful for symmetric objects with multiple valid solutions.

Planning

Planning in robotics is the problem of identifying a sequence of valid robot configuration states to achieve an objective. This includes path planning for navigation, trajectory planning for grasping or manipulation, or planning for interactive perception [BHS∗17]. Given the spatial nature of planning problems, neural fields have been used as a representation in many solutions.

Among the first attempts to use neural networks for robot path planning was by Lemmon [Lem91, Lem92]. While this work refers to a 2D grid of neurons as a “neural field,” the neurons are mapped one-to-one to a 2D map and is thus a (non-continuous) coordinate-based network. Their method finds the variational solution of Bellman’s dynamic programming equation [Bel54] used in path planning. Other coordinate-based representations have been used for planning, including estimating affordance maps [KS15, QMGR20]. Zhang et al. parameterize a Fisher Information Field [ZS20] via neural network. World constraints can be specified during planning, for example, by specifying constraints as a level set in a high-dimensional space [SFE∗20] which is learned as a neural field.

Grasping and manipulation problems require knowing 2D or 3D position of a robot gripper relative to the surface of objects. Some approaches model the shape of the object as a collection of points [BKP11], and learn potentially good points for grasping or manipulation using a Markov Random Field (MRF). This approach can handle points in any continuous 3D position as long as an MRF can be built. Similarly, other data-driven grasping methods often use coordinate-based representations [BMAK14]. ContactNets [PHP20] models contact between objects via learned neural fields. Neural fields can also synthesize human grasps [KYZ∗20]. GIGA [JZS∗21] uses locally-conditioned neural fields encoding quality, orientation, width, and occupancy for grasp selection.

Control

Controllers are responsible for realizing plans, while ensuring that physical constraints and mechanical integrity are preserved. Control can be achieved either by relying on a planner or directly from observations. Neural fields have been used for this task by learning an obstacle barrier function approximated by an SDF [LQCA20]. Similarly, Bhardwaj et al. [BSM∗21a] solve the robot arm self-collision avoidance task by using neural fields to predict the closest distance between robot links, given its joint configuration. In visuomotor control, control is driven directly by visual observations. Li et al. [LLS∗21] use NeRF to facilitate view-invariant visuomotor control to achieve robot goal states specified via a 2D goal image. This is achieved by auto-decoding a dynamics model to support future prediction and novel view synthesis.

Lossy Compression

The goal of lossy data compression is to approximate a signal as best as possible with as few bits as possible. These opposing forces naturally form a tradeoff which can be characterized as a Pareto frontier: the rate-distortion curve. In practice, signals are often stored as discrete sequences of data which are transformed into an alternate basis such as the discrete cosine transform which help to decorrelate the signal (making downstream tasks like quantization and entropy coding more effective). Many standards exist, such as JPEG [Wal92] for images or HEVC [SOHW12] for videos. Recent work has explored the potential of neural fields as an alternate signal storage format which directly represents the continuous signal with a parameteric, continuous function. Compression may be achieved in one of two ways. First, by leveraging the inductive bias of the network architecture itself, and simply overfitting a neural field to a signal, i.e., finding parameters Θ\Theta of a neural field Φ\Phi. Only the parameters and architecture of the neural field need to be stored. Second, prior-based compression schemes achieve compression via learning a space of low-dimensional latent code vectors z\mathbf{z} that may be decoded into neural field parameters, where the storage cost of the decoder is amortized over many latent codes.

Encoding 3D geometry with neural fields presents an alternative to conventional mesh representations that may enable a significantly reduced memory footprint. Subsequently, SIREN [SMB∗20] demonstrated fitting a wide array of signals with neural fields: audio, video, images, and 3D geometry (including large-scale scenes). Lu et al. [LJLB21] show that SIREN in conjunction with scalar weight quantization can compress dense volumetric data better than state-of-the-art approaches at the cost of higher encoding and decoding latency. Davies et al. [DNJ21] compare neural field 3D geometry with decimated meshes and found better reconstruction quality with the same memory footprint. Takikawa et al. [TLY∗21] show that a hierarchical tree structure can be used to learn multiresolution signals that can perform level-of-detail more effectively than decimated meshes. Dupont et al. [DGA∗21] compared the memory use of neural fields parameterizing images, and found neural fields can outperform JPEG [Wal92] but not state-of-the-art image compression techniques. ACORN [MLL∗21] use an adaptive quadtree data structure to fit high resolution images with neural fields, but do not outperform traditional image compression methods. For dynamic 3D scenes, DyNeRF [LSZ∗21] uses 28 MB of memory for a 10-second 30 FPS 3D video sequence. Light Field Networks [SRF∗21] allow a large reduction in memory used over a classical discrete light field. General network compression and quantization techniques can further reduce network size [Isi21, BBSC21]. Bird et al. apply entropy penalization to NeRF [BBSC21], and obtain higher compression rates compared to standard HEVC video encoding [SOHW12] and LLFF [MSOC∗19] for forward-facing scenes.

The above methods study a variety of continuous signal modalities, and use different datasets, architectures, and metrics, making comparisons difficult. Few rigorous comparisons exist between neural field and conventional compression schemes. Many works lack ablation of design choices and the ability to make direct queries to the without decoding. While neural fields for compression remains in its infancy at the time of writing, it is nonetheless a valuable perspective to consider signals as functions and neural networks as a data format.

Beyond Visual Computing

Visual computing problems are a subset of all inverse problems which can be parameterized by neural fields. These problems often share the same challenges involving incomplete observations and the need for a flexible parameterization. In this section we will highlight some of the emerging research directions in neural fields beyond visual computing.

Most of the works surveyed so far have been concerned with modeling the imaging process of consumer cameras, which measure the visible electromagnetic radiation via optical lenses, using sensors that digitize irradiance into intensity over a 2D raster grid. Nonetheless, neural fields can also model alternative signal modalities such as non-line-of-sight imaging [SWL∗21], non-visible x-rays for computed tomography (CT) [SPX21, ZIL∗21, SLX∗21], magnetic resonance imaging (MRI) [SPX21], pressure waves for audio [RBBJ], chemiluminescence [PXZ∗21], time-of-flight imaging [ALG∗21], as well as volumetric light displays [ZBW∗20].

In CT and MRI, raw sensor measurements are the Radon and Fourier transformation of spatially-varying density, respectively [SPX21]. The sensor domains are not human-readable, while the reconstructed density volume (whose 2D slices are called the image domain) is. Reconstructing the density volume is an ill-posed problem [SPX21], and classical techniques are sensitive to measurement noise [ZIL∗21]. Reconstruction and NVS in medical imaging is also limited by capture constraints such as scene movement, finite sensor resolution, limited viewing angle, and sparse views (to reduce X-ray exposure, and speed up procedure).

Neural fields can either parameterize the sensor domain [SLX∗21], or the density domain directly [SPX21, ZIL∗21]. In the former, the neural field maps sensor coordinates to predicted sensor activations, and can augment real measurement data, before applying classical reconstruction technique (filtered back-projection [KS01]) [SLX∗21]. In the latter the neural field directly predicts the density value at a 3D spatial coordinate, and is supervised by mapping its output to the sensor domain via Radon (CT) or Fourier (MRI) transform [SPX21, ZIL∗21] (Figure 9).

Similarly, in cryo-electron microscopy (cryo-EM), the 2D sensor measures the convolution between a point spread function and electron density. In CryoGRGN [ZBDB20b], the neural field maps a 3D coordinate to the deconvolved electron density.

Audio signals can be represented as raw waveforms or spectrograms, and each can be parameterized by neural fields [SMB∗20, GCM∗21]. Raw waveform is a continuous 1D field function that maps temporal coordinates to amplitude, often stored as a collection of discrete samples. A spectrogram is the Fourier transform of a waveform, which is a 2D field function mapping time and frequency to amplitude. Neural field methods fit all time-dependent waveforms, which means that the techniques are applicable to all waves mechanical (e.g., seismic waves, ocean waves, acoustic waves) or electromagnetic (e.g., wave optics, radio waves).

Synthetic aperture sonar (SAS) also exhibits the common problems in inverse problems: noisy sensor domain and ill-posed problem. Reed et al. [RBBJ] use a neural field mapping a 2D position to a distribution of point scatter (reconstruction domain), and obtain supervision by mapping to the sensor domain via convolution.

Physics-informed Problems

Physics-informed problems have solutions that are restricted to a set of partial differential equations (PDEs) based on laws of physics. The solutions are often continuous in spatio-temporal coordinates. Neural fields are therefore a natural parameterization of the solution space, given that neural networks are continuous, differentiable, and universal function approximators. These neural fields are also referred to as physics-informed neural networks (PINNs), whose use was first popularized by [RPK19a]. Problems constrained by nonlinear PDEs often require arbitrarily-small step sizes for traditional methods, which often require prohibitive computation resources. Parameterizing the solution via neural networks re-formulates these problems as optimization, rather than simulation, which is more data-efficient [RPK19a]. Since many physical processes in nature are governed by PDEs (boundary conditions), supervising (or regularizing) the neural network via its gradients is a common technique. In fact, the Eikonal regularization for SDF is an example in visual computing. Physics-informed problems include topology optimization [ZLCT21], geodesy estimation [IG21], collision dynamics estimation [PHP20], solving the Schrodinger Equation [RPK19a], Navier-Stokes equation [RPK19a], and Eikonal equation [SAR20].

Chapter 2 Discussion & Conclusion

As we have seen, neural fields have a rich but evolving technical ‘toolbox’, and a rapidly increasing space of applications both within and outside of visual computing. In our view, there are several factors that have resulted in this progress.

First, the idea of parameterizing a continuous field using an MLP without the need to use more complex neural network architectures has simplified the training of fields and reduced the entry barrier [GN98]. Neural fields provide a fundamentally different approach to signal processing that is no longer discrete and more faithful to the original continuous signal. Second, techniques such as positional encoding and sinusoidal activations have significantly improved the quality of neural fields leading to large leaps in applications focused on quality. Finally, applications in novel view synthesis and 3D reconstruction have been particularly important in popularizing neural fields because of the visually appealing nature of these applications [MST∗20, SZW19, GFK∗18b, PFS∗19]. Third, researchers have realized that differentiable volume and voxel rendering commonly used in novel view synthesis methods [MSOC∗19, YFKT∗21] can be useful in solving others tasks like 3D reconstruction [YGKL21] and even semantic segmentation [ZLLD21, VRG∗21].

Despite the progress, we believe that neural fields have only started to scratch the surface and there remains great potential for continued progress in both techniques and applications. In terms of techniques, a common limitation of many neural fields is their inability to generalize well to unseen data. We believe that integration of stronger priors can enable these methods to generalize better and be data efficient. Other inductive biases in the form of task-specific heuristics, laws of physics, or network architecture can further help generalization. Building a common framework for incorporating these priors is a fruitful direction for future work. Furthermore, the rapid progress has been at the expense of methodical evaluation and analysis of common techniques. We need shared datasets and benchmarks on which different techniques can be evaluated and compared.

In terms of applications, a majority of neural field methods have thus far been used to solve “low-level” (e.g., image synthesis) and “mid-level” (e.g., 3D reconstruction) tasks. The application of neural fields for “high-level” semantic tasks remains an open problem. Examples of these problems include understanding 3D scene layout [WYN21], 3D scene interaction [LMW∗21a], and grouping of data into more meaningful entities [TLV21]. Furthermore, neural fields have focused on a single data modality, but exploring the fusion of multiple modalities could be a fruitful topic of research. For example, synthesizing fields based on language input [JMB∗21], or joint modeling of image and text or audio input could foster closer connections with the NLP community. Finally, future work should consider moving beyond supervised learning of fields and consider weakly- or self-supervised learning as an alternative. Building upon advances in 3D deep learning such as transformation equivariance [SPJ∗22, SSM∗20, STD∗21] could allow neural fields to be data efficient and generalize better.

Finally, the neural fields community must become more self-aware to build a culture that promotes scholarship, sustainable growth, inclusivity, and diversity. We must pay careful attention to past work and avoid repetition of work through “selective amnesia” [SC21]. We must avoid becoming a “scientific bandwagon” [Sha56] by acknowledging limitations of neural fields and by collaborating with domain experts to identify shortcomings. While the explosion of work in neural fields helps us make important advances, we must be aware of its impact on the mental health and well-being of researchers [SC21].

As the applications of neural fields increase, so will their societal impact, especially in domains where they make a significant difference. In generative modeling of audio, imagery, 3D scenes, etc., neural fields have enabled the generation of realistic content for the purpose of deception, for instance, impersonating the image and voice of actors without their consent. Recent work explores the detection of digitally-altered imagery [TVRF∗20]. Similarly, improving the capabilities of inverse-problem solvers, such as denoising and blur-removal algorithms, has implications for privacy. For instance, this may enable de-anonymizing recordings that were previously assumed to not contain sufficient information to allow identification of participating actors. It may also decrease the cost of surveillance, making it more widely available as the hardware requirements may decrease. Finally, neural fields could also have negative environmental impact as significant computational resources are spent optimizing neural networks with GPUs.

Neural fields can also positively impact society. Democratizing the generation of photo-realistic imagery helps more artists and content creators to tell stories. The ability to aggregate information from low-dimensional supervision (e.g., 2D images) relaxes hardware constraints for 3D content creation. Neural fields for computer vision may help build robotic automation that helps people.

Conclusion

This report provides an overview of the flourishing research direction of neural fields. From a review of over 250 papers, we have summarized five classes of techniques using a shared mathematical formulation, including prior learning and conditioning, hybrid representations, forward maps, network architectures, and manipulation methods. Then, we have surveyed applications across graphics, vision, robotics, medical imaging, and computational physics. Neural fields research will continue to grow. To aid in continued understanding, we have established a living report as a community-driven website, where authors can submit their own papers and classify them using the taxonomy developed in this report. Neural fields will be a key enabler for progress across many areas of computer science, and we look forward to the research yet to come.

Chapter 3 Authors

Yiheng Xie (https://yxie20.github.io/) is a final-year undergraduate student at Brown University and a researcher at Unity Technologies. His research interests include 3D reconstruction, physically-based rendering, and robotics. He is a member of Brown Visual Computing (BVC), advised by Professor Srinath Sridhar, and Brown Humanity Centered Robotics Initiative (HCRI), advised by Professor Michael Littman.

Towaki Takikawa (https://tovacinni.github.io/) is a Ph.D. student at the University of Toronto with Prof. Sanja Fidler and Prof. Alec Jacobson. He is also a Research Scientist at NVIDIA in the Hyperscale Graphics Systems group. He received his bachelors in Computer Science at the University of Waterloo.

Shunsuke Saito (http://www-scf.usc.edu/~saitos/) is a Research Scientist at Meta Reality Labs in Pittsburgh. He finished his Ph.D. at University of Southern California, where he worked with Prof. Hao Li. Prior to USC, he spent one year at University of Pennsylvania as Visiting Researcher. He obtained B.Eng. and M.Eng. in Applied Physics at Waseda University in 2013 and 2014.

Or Litany (https://orlitany.github.io/) is a Research Scientist at NVIDIA. Prior, he was a postdoc at Stanford University working under Prof. Leonidas Guibas, and a postdoc at FAIR working under Prof. Jitendra Malik. Previously, he was a postdoc at the Technion and a research intern at Microsoft, Intel and Google. He received his PhD from Tel-Aviv University, advised by Prof. Alex Bronstein. He received my B.Sc. in Physics and Mathematics from the Hebrew University under the auspices of “Talpiot.”

Shiqin Yan (https://player-eric.com/about) is a masters student in Computer Science at Brown University. He is passionate about solving real-world problems with large-scale web-based systems.

Numair Khan (http://cs.brown.edu/~nkhan6/) is a PhD student at Brown University where his research focuses on methods and representations for scene reconstruction, differentiable rendering, and novel-view synthesis.

Federico Tombari (https://federicotombari.github.io/) is a Research Scientist and Manager at Google Zurich (Switzerland), where he leads an applied research team in Computer Vision and Machine Learning. He is also affiliated to the Faculty of Computer Science at TU Munich (Germany) as lecturer (Privatdozent).

James Tompkin (https://jamestompkin.com/) is an assistant professor in visual computing at Brown University. He was a PhD student at UCL, with postdocs at the Max Planck Institute and Harvard. His lab develops visual understanding techniques for camera-captured media to remove barriers from image and video creation, editing, and interaction. This requires image and scene reconstruction techniques, especially from multi-camera systems.

Vincent Sitzmann (https://vsitzmann.github.io/) is an incoming assistant professor at MIT EECS. Currently, he is a Postdoc at MIT’s CSAIL with Joshua Tenenbaum, William Freeman, and Frédo Durand. Previously, he finished his Ph.D. at Stanford University. His research interest lies in neural scene representations — the way neural networks learn to represent information on our world. His goal is to allow independent agents to reason about our world given visual observations, such as inferring a complete model of a scene with information on geometry, material, lighting etc. from only few observations, a task that is simple for humans, but currently impossible for AI.

Srinath Sridhar (https://srinathsridhar.com/) is an assistant professor of computer science at Brown University. His research interests are in 3D computer vision and machine learning. Specifically, he focuses on visual understanding of 3D human physical interactions with applications ranging from robotics to mixed reality. He has won several fellowships (e.g., Google Research Scholar) and awards (e.g., Eurographics Best Paper Honorable Mention) for his work, and has previously spent time at Stanford, Max Planck Institute for Informatics, Microsoft Research Redmond, and Honda Research Institute.

Chapter 4 Acknowledgements

This work was supported by NSF CNS-2038897 and the Google Research Scholar Program. We thank Sunny Li for their help in designing the website, Jayden Yi for a conceptual readthrough, and Alexander Rush and Hendrik Strobelt for the Mini-Conf project.

References

Appendix 4.A Variable Naming Conventions

Appendix 4.B Implicit Surface Representations

Since neural fields can store arbitrary quantities, they offer a flexible way to represent geometry and other data. We will summarize the popular output types for neural fields that can represent geometry. In addition to geometry, other types of output include radiance, BRDF parameters, deformation/warping parameters, classification weights, etc. We have discussed these output types throughout Part II. In this appendix, we focus on implicit surface representations.

Distance Functions represent the distance to the nearest surface in some metric. These distances are useful for tasks like path planning, computational fabrication, approximating occlusion, and more. A special case of the distance function is the signed distance function (SDF), where the sign of the distance encodes whether the surface is inside or outside. These signed distance functions can be efficiently visualized using an algorithm like sphere tracing [Har96]. A necessary but not sufficient condition of the signed distance function is the Eikonal property, which has been utilized for neural fields in many different contexts.

Occupancy represents whether a point is considered to be inside the object or outside the object, usually with a binary 0,10,1 value. Since the neural field output is continuous, the binary 0,10,1 values is often approximated with a continuous function, and an appropriate isosurface value b∈b\in needs to be selected. These occupancy fields can be visualized using raymarching.

Volume Density represents the density of particles that exist. Unlike occupancy, this is not a binary value and are continuous values with no upper bound. While distance functions and occupancy can only represent hard surfaces, volume density is useful for representing physically-realizable volumetric scenes like clouds, fog, or hair, where the geometric structures are too fine to be efficiently modeled as hard surfaces, and a distribution of particles can effectively approximate the true geometry. These densities can be visualized through volume rendering [NGHJ18].

Medial Fields [RLS∗21] represent the local thickness of the geometry which can be derived from the medial axis. Similarly to SDFs, these quantities can be used for rendering and approximating quantities like ambient occlusion.