Implicit Neural Representations with Periodic Activation Functions

Vincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell, Gordon Wetzstein

Introduction

We are interested in a class of functions Φ\Phi that satisfy equations of the form

A continuous parameterization offers several benefits over alternatives, such as discrete grid-based representations. For example, due to the fact that Φ\Phi is defined on the continuous domain of x\mathbf{x}, it can be significantly more memory efficient than a discrete representation, allowing it to model fine detail that is not limited by the grid resolution but by the capacity of the underlying network architecture. Being differentiable implies that gradients and higher-order derivatives can be computed analytically, for example using automatic differentiation, which again makes these models independent of conventional grid resolutions. Finally, with well-behaved derivatives, implicit neural representations may offer a new toolbox for solving inverse problems, such as differential equations.

For these reasons, implicit neural representations have seen significant research interest over the last year (Sec. 2). Most of these recent representations build on ReLU-based multilayer perceptrons (MLPs). While promising, these architectures lack the capacity to represent fine details in the underlying signals, and they typically do not represent the derivatives of a target signal well. This is partly due to the fact that ReLU networks are piecewise linear, their second derivative is zero everywhere, and they are thus incapable of modeling information contained in higher-order derivatives of natural signals. While alternative activations, such as tanh or softplus, are capable of representing higher-order derivatives, we demonstrate that their derivatives are often not well behaved and also fail to represent fine details.

To address these limitations, we leverage MLPs with periodic activation functions for implicit neural representations. We demonstrate that this approach is not only capable of representing details in the signals better than ReLU-MLPs, or positional encoding strategies proposed in concurrent work , but that these properties also uniquely apply to the derivatives, which is critical for many applications we explore in this paper.

To summarize, the contributions of our work include:

A continuous implicit neural representation using periodic activation functions that fits complicated signals, such as natural images and 3D shapes, and their derivatives robustly.

An initialization scheme for training these representations and validation that distributions of these representations can be learned using hypernetworks.

Demonstration of applications in: image, video, and audio representation; 3D shape reconstruction; solving first-order differential equations that aim at estimating a signal by supervising only with its gradients; and solving second-order differential equations.

Related Work

Recent work has demonstrated the potential of fully connected networks as continuous, memory-efficient implicit representations for shape parts , objects , or scenes . These representations are typically trained from some form of 3D data as either signed distance functions or occupancy networks . In addition to representing shape, some of these models have been extended to also encode object appearance , which can be trained using (multiview) 2D image data using neural rendering . Temporally aware extensions and variants that add part-level semantic segmentation have also been proposed.

Periodic nonlinearities.

Periodic nonlinearities have been investigated repeatedly over the past decades, but have so far failed to robustly outperform alternative activation functions. Early work includes Fourier neural networks, engineered to mimic the Fourier transform via single-hidden-layer networks . Other work explores neural networks with periodic activations for simple classification tasks and recurrent neural networks . It has been shown that such models have universal function approximation properties . Compositional pattern producing networks also leverage periodic nonlinearities, but rely on a combination of different nonlinearities via evolution in a genetic algorithm framework. Motivated by the discrete cosine transform, Klocek et al. leverage cosine activation functions for image representation but they do not study the derivatives of these representations or other applications explored in our work. Inspired by these and other seminal works, we explore MLPs with periodic activation functions for applications involving implicit neural representations and their derivatives, and we propose principled initialization and generalization schemes.

Neural DE Solvers.

Neural networks have long been investigated in the context of solving differential equations (DEs) , and have previously been introduced as implicit representations for this task . Early work on this topic involved simple neural network models, consisting of MLPs or radial basis function networks with few hidden layers and hyperbolic tangent or sigmoid nonlinearities . The limited capacity of these shallow networks typically constrained results to 1D solutions or simple 2D surfaces. Modern approaches to these techniques leverage recent optimization frameworks and auto-differentiation, but use similar architectures based on MLPs. Still, solving more sophisticated equations with higher dimensionality, more constraints, or more complex geometries is feasible . However, we show that the commonly used MLPs with smooth, non-periodic activation functions fail to accurately model high-frequency information and higher-order derivatives even with dense supervision.

Neural ODEs are related to this topic, but are very different in nature. Whereas implicit neural representations can be used to directly solve ODEs or PDEs from supervision on the system dynamics, neural ODEs allow for continuous function modeling by pairing a conventional ODE solver (e.g., implicit Adams or Runge-Kutta) with a network that parameterizes the dynamics of a function. The proposed architecture may be complementary to this line of work.

Formulation

Our goal is to solve problems of the form presented in Equation 1. We cast this as a feasibility problem, where a function Φ\Phi is sought that satisfies a set of MM constraints {Cm(a(x),Φ(x),∇Φ(x),...)}m=1M\{\mathcal{C}_{m}(\mathbf{a}(\mathbf{x}),\Phi(\mathbf{x}),\nabla\Phi(\mathbf{x}),...)\}_{m=1}^{M}, each of which relate the function Φ\Phi and/or its derivatives to quantities a(x)\mathbf{a}(\mathbf{x}):

This problem can be cast in a loss function that penalizes deviations from each of the constraints on their domain Ωm\Omega_{m}:

We parameterize functions Φθ\Phi_{\theta} as fully connected neural networks with parameters θ\theta, and solve the resulting optimization problem using gradient descent.

We propose siren, a simple neural network architecture for implicit neural representations that uses the sine as a periodic activation function:

Interestingly, any derivative of a siren is itself a siren, as the derivative of the sine is a cosine, i.e., a phase-shifted sine (see supplemental). Therefore, the derivatives of a siren inherit the properties of sirens, enabling us to supervise any derivative of siren with “complicated” signals. In our experiments, we demonstrate that when a siren is supervised using a constraint Cm\mathcal{C}_{m} involving the derivatives of ϕ\phi, the function ϕ\phi remains well behaved, which is crucial in solving many problems, including boundary value problems (BVPs).

We will show that sirens can be initialized with some control over the distribution of activations, allowing us to create deep architectures. Furthermore, sirens converge significantly faster than baseline architectures, fitting, for instance, a single image in a few hundred iterations, taking a few seconds on a modern GPU, while featuring higher image fidelity (Fig. 1).

2 Distribution of activations, frequencies, and a principled initialization scheme

We present a principled initialization scheme necessary for the effective training of sirens. While presented informally here, we discuss further details, proofs and empirical validation in the supplemental material. The key idea in our initialization scheme is to preserve the distribution of activations through the network so that the final output at initialization does not depend on the number of layers. Note that building sirens with not carefully chosen uniformly distributed weights yielded poor performance both in accuracy and in convergence speed.

Hence, we propose to draw weights with c=6c=6 so that wi∼U(−6/n,6/n)w_{i}\sim\mathcal{U}(-\sqrt{6/n},\sqrt{6/n}). This ensures that the input to each sine activation is normal distributed with a standard deviation of 11. Since only a few weights have a magnitude larger than π\pi, the frequency throughout the sine network grows only slowly. Finally, we propose to initialize the first layer of the sine network with weights so that the sine function sin⁡(ω0⋅Wx+b)\sin(\omega_{0}\cdot\mathbf{W}\mathbf{x}+\mathbf{b}) spans multiple periods over $.Wefound. We found\omega_{0}=30$ to work well for all the applications in this work. The proposed initialization scheme yielded fast and robust convergence using the ADAM optimizer for all experiments in this work.

Experiments

In this section, we leverage sirens to solve challenging boundary value problems using different types of supervision of the derivatives of Φ\Phi. We first solve the Poisson equation via direct supervision of its derivatives. We then solve a particular form of the Eikonal equation, placing a unit-norm constraint on gradients, parameterizing the class of signed distance functions (SDFs). siren significantly outperforms ReLU-based SDFs, capturing large scenes at a high level of detail. We then solve the second-order Helmholtz partial differential equation, and the challenging inverse problem of full-waveform inversion. Finally, we combine sirens with hypernetworks, learning a prior over the space of parameterized functions. All code and data will be made publicly available.

We demonstrate that the proposed representation is not only able to accurately represent a function and its derivatives, but that it can also be supervised solely by its derivatives, i.e., the model is never presented with the actual function values, but only values of its first or higher-order derivatives.

An intuitive example representing this class of problems is the Poisson equation. The Poisson equation is perhaps the simplest elliptic partial differential equation (PDE) which is crucial in physics and engineering, for example to model potentials arising from distributions of charges or masses. In this problem, an unknown ground truth signal ff is estimated from discrete samples of either its gradients \gradientf\gradient f or Laplacian Δf=\gradient⋅\gradientf\Delta f=\gradient\cdot\gradient f as

Poisson image editing.

2 Representing Shapes with Signed Distance Functions

Inspired by recent work on shape representation with differentiable signed distance functions (SDFs) , we fit SDFs directly on oriented point clouds using both ReLU-based implicit neural representations and sirens. This amounts to solving a particular Eikonal boundary value problem that constrains the norm of spatial gradients ∣∇xΦ∣|\nabla_{\mathbf{x}}\Phi| to be 11 almost everywhere. Note that ReLU networks are seemingly ideal for representing SDFs, as their gradients are locally constant and their second derivatives are 0. Adequate training procedures for working directly with point clouds were described in prior work . We fit a siren to an oriented point cloud using a loss of the form

Here, ψ(x)=exp⁡(−α⋅∣Φ(x)∣),α≫1\psi(\mathbf{x})=\exp(-\alpha\cdot|\Phi(\mathbf{x})|),\alpha\gg 1 penalizes off-surface points for creating SDF values close to 0. Ω\Omega is the whole domain and we denote the zero-level set of the SDF as Ω0\Omega_{0}. The model Φ(x)\Phi(x) is supervised using oriented points sampled on a mesh, where we require the siren to respect Φ(x)=0\Phi(\mathbf{x})=0 and its normals n(x)=\gradientf(x)\mathbf{n}(\mathbf{x})=\gradient f(\mathbf{x}). During training, each minibatch contains an equal number of points on and off the mesh, each one randomly sampled over Ω\Omega. As seen in Fig. 4, the proposed periodic activations significantly increase the details of objects and the complexity of scenes that can be represented by these neural SDFs, parameterizing a full room with only a single five-layer fully connected neural network. This is in contrast to concurrent work that addresses the same failure of conventional MLP architectures to represent complex or large scenes by locally decoding a discrete representation, such as a voxelgrid, into an implicit neural representation of geometry .

3 Solving the Helmholtz and Wave Equations

The Helmholtz and wave equations are second-order partial differential equations related to the physical modeling of diffusion and waves. They are closely related through a Fourier-transform relationship, with the Helmholtz equation given as

Here, f(x)f(\mathbf{x}) represents a known source function, Φ(x)\Phi(\mathbf{x}) is the unknown wavefield, and the squared slowness m(x)=1/c(x)2m(\mathbf{x})=1/c(\mathbf{x})^{2} is a function of the wave velocity c(x)c(\mathbf{x}). In general, the solutions to the Helmholtz equation are complex-valued and require numerical solvers to compute. As the Helmholtz and wave equations follow a similar form, we discuss the Helmholtz equation here, with additional results and discussion for the wave equation in the supplement.

Results are shown in Fig. 5 for solving the Helmholtz equation in two dimensions with spatially uniform wave velocity and a single point source (modeled as a Gaussian with σ2=10−4\sigma^{2}=10^{-4}). The siren solution is compared with a principled solver as well as other neural network solvers. All evaluated network architectures use the same number of hidden layers as siren but with different activation functions. In the case of the RBF network, we prepend an RBF layer with 1024 hidden units and use a tanh activation. siren is the only representation capable of producing a high-fidelity reconstruction of the wavefield. We also note that the tanh network has a similar architecture to recent work on neural PDE solvers , except we increase the network size to match siren.

Neural full-waveform inversion (FWI).

In many wave-based sensing modalities (radar, sonar, seismic imaging, etc.), one attempts to probe and sense across an entire domain using sparsely placed sources (i.e., transmitters) and receivers. FWI uses the known locations of sources and receivers to jointly recover the entire wavefield and other physical properties, such as permittivity, density, or wave velocity. Specifically, the FWI problem can be described as

where there are NN sources, r samples the wavefield at the receiver locations, and ri(x)r_{i}(x) models receiver data for the iith source.

We first use a siren to directly solve Eq. 7 for a known wave velocity perturbation, obtaining an accurate wavefield that closely matches that of a principled solver (see Fig. 5, right). Without a priori knowledge of the velocity field, FWI is used to jointly recover the wavefields and velocity. Here, we use 5 sources and place 30 receivers around the domain, as shown in Fig. 5. Using the principled solver, we simulate the receiver measurements for the 5 wavefields (one for each source) at a single frequency of 3.2 Hz, which is chosen to be relatively low for improved convergence. We pre-train siren to output 5 complex wavefields and a squared slowness value for a uniform velocity. Then, we optimize for the wavefields and squared slowness using a penalty method variation of Eq. 8 (see the supplement for additional details). In Fig. 5, we compare to an FWI solver based on the alternating direction method of multipliers . With only a single frequency for the inversion, the principled solver is prone to converge to a poor solution for the velocity. As shown in Fig. 5, siren converges to a better velocity solution and accurate solutions for the wavefields. All reconstructions are performed or shown at 256×256256\times 256 resolution to avoid noticeable stair-stepping artifacts in the circular velocity perturbation.

4 Learning a Space of Implicit Functions

and use a ReLU hypernetwork , to map the latent code to the weights of a siren, as in :

We replicated the experiment from on the CelebA dataset using a set encoder. Additionally, we show results using a convolutional neural network encoder which operates on sparse images. Interestingly, this improves the quantitative and qualitative performance on the inpainting task.

At test time, this enables reconstruction from sparse pixel observations, and, thereby, inpainting. Fig. 6 shows test-time reconstructions from a varying number of pixel observations. Note that these inpainting results were all generated using the same model, with the same parameter values. Tab. 1 reports a quantitative comparison to , demonstrating that generalization over siren representations is at least equally as powerful as generalization over images.

Discussion and Conclusion

The question of how to represent a signal is at the core of many problems across science and engineering. Implicit neural representations may provide a new tool for many of these by offering a number of potential benefits over conventional continuous and discrete representations. We demonstrate that periodic activation functions are ideally suited for representing complex natural signals and their derivatives using implicit neural representations. We also prototype several boundary value problems that our framework is capable of solving robustly. There are several exciting avenues for future work, including the exploration of other types of inverse problems and applications in areas beyond implicit neural representations, for example neural ODEs .

With this work, we make important contributions to the emerging field of implicit neural representation learning and its applications.

Broader Impact

The proposed siren representation enables accurate representations of natural signals, such as images, audio, and video in a deep learning framework. This may be an enabler for downstream tasks involving such signals, such as classification for images or speech-to-text systems for audio. Such applications may be leveraged for both positive and negative ends. siren may in the future further enable novel approaches to the generation of such signals. This has potential for misuse in impersonating actors without their consent. For an in-depth discussion of such so-called DeepFakes, we refer the reader to a recent review article on neural rendering .

Acknowledgments and Disclosure of Funding

Vincent Sitzmann, Alexander W. Bergman, and David B. Lindell were supported by a Stanford Graduate Fellowship. Julien N. P. Martel was supported by a Swiss National Foundation (SNF) Fellowship (P2EZP2 181817). Gordon Wetzstein was supported by an NSF CAREER Award (IIS 1553333), a Sloan Fellowship, and a PECASE from the ARO.

References