Modulated Periodic Activations for Generalizable Local Functional Representations
Ishit Mehta, Michaël Gharbi, Connelly Barnes, Eli Shechtman, Ravi Ramamoorthi, Manmohan Chandraker
Introduction
Functional neural representations using Multi-Layer Perceptrons (MLPs) have garnered renewed interest for their conceptual simplicity and ability to approximate complex signals like images, videos, audio recordings , light-fields and implicitly-defined 3D shapes . They have shown to be more compact and efficient than their discrete counterparts . While recent contributions have focused on improving the accuracy of these representations, in particular to model complex signals with high-frequency details , it is still challenging to generalize them to unseen signals. Recent approaches typically require training a separate MLP for each signal . Previous efforts sought to improve generalization by imposing priors on the functional space spanned by the MLP parameterization , using hypernetworks , or via meta-learning . But multi-instance generalization still causes significant degradations in quality.
We introduce a neural functional representation that simultaneously achieves high-reconstruction quality and generalizes to multiple instances. Our approach can encode functional representations for multiple discrete signals using a single model. Unlike previous works, which train a model for each signal, it can do so in a single feed-forward pass (Fig. 1). We represent each signal using a low-dimensional latent code. These codes serve as conditioning variables in a functional mapping that uses two MLPs: a modulator and a synthesis network. The synthesis network implements a mapping from coordinates (e.g. spatial position) to signal values (e.g. RGB colors). It uses the sine function as activation, which enables accurate reconstructions of high-frequency content , but also makes naive conditioning strategies ineffective (§ 3.2). The modulator is the key to generalization. It consumes latent code and outputs, at each layer, parameters that modulate the amplitude, phase and frequency of periodic activations in the synthesis network. The modulator uses ReLU activations. Our model can either be used as an autoencoder, where the latent codes are produced by a third network (the encoder); or as an auto-decoder, where the latent codes are optimized jointly with all the network parameters.
As we show in Figure 2, the quality of functional representations that fit images as-a-whole, degrades as we increase the target resolution. High-resolution images typically have a broad power spectrum, thereby requiring more expensive models to represent them functionally. But images are usually much simpler locally: simple edges and textures re-occur commonly across images that are otherwise quite distinct at the global level. This motivates our strategy to exploit locality. We partition the signal domain into a regular tiling, and assign each tile a latent code (Fig. 1). By “zooming in” on the local structure, computing functional approximations that generalize becomes more tractable , because simple parts exhibit fewer variations than complete objects . Some recent work has explored locality, but they focus on relatively simple, low-frequency signals like signed distance fields , which can locally be well approximated using a single linear decision boundary—well in the purview of ReLU-based MLPs. For more complex signals like images and videos, even local patterns contain high-frequency components that a standard ReLU-MLP fails to reconstruct (Figure 3). We show that locality, together with our model architecture, makes it possible to obtain functional representations of large, complex signals.
Compared to previous methods, ours produces qualitatively and quantitatively superior functional representations, with improved generalization capabilities. In summary, our contributions are as follows:
A local neural-functional representation that enables generalization and achieves high fidelity. We use a set of local functions defined on a tiling of the input domain that combine to reconstruct the target signal.
A new network architecture, which uses modulation and synthesis sub-networks for high-fidelity functional neural representations of images, shapes and videos.
A novel conditioning mechanism in which a ReLU-MLP modulates the amplitude, phase, and frequency of periodic activations in the synthesis sub-network.
Related Work
Our work builds upon the extensive use of Multi Layer Perceptrons (MLPs) to encode images , videos , shapes and 3D scenes . Once trained, these models yield continuous representations that can be queried at arbitrary locations in the signal’s input domain. They have had significant impact in view-synthesis and other interpolation problems . Similar approaches have been used for end-to-end differentiable texture mapping and volumetric rendering .
Lapedes and Farber show the earliest use of periodic activations in neural networks. They observe that networks with more than one layer with periodic activations are difficult to train, and often converge to undesired local minima. This problem is formalized further in . For small datasets, Sopena et al. show compelling results using sine activations in the first layer and monotonic functions in the others. This is similar to preconditioning the input using a Fourier basis, which was shown to be useful in feature visualization and image synthesis . More recently, Tancik et al. proposed Fourier Feature Networks (FFN), where they encode the MLP’s input into a high-dimensional space using a random sampling of Fourier basis functions. Concurrently, Sitzmann et al. showed that, with careful network initializations, sine activations can be used in all layers. They demonstrate regressions of small, single images and videos, as well as more complex shapes. However, as we show in Section 4, these networks struggle with larger datasets, or individual instances when the complexity is increased. In Figure 2, we show how quality degrades with these networks as we regress a progressively higher-resolution image, or a longer video. Our work lifts these limitations by exploiting locality, and introduces an effective modulation mechanism to enable generalization.
A major limitation of current implicit representations, is that they need to be optimized for each test signal individually, unlike more established models than only require a forward pass at test time, having been trained on large datasets. Building implicit models with similar generalization properties is typically done via a conditioning mechanism, using latent variables. Conditioning by contatenating the latent code with an MLPs spatial input coordinates has been successful in signed-distance fields regression tasks . Schwarz et al. use the same strategy for MLPs encoding radiance fields, although with limited resolution. In Section 3.2, we show why this conditioning-by-concatenation approach is inadequate for MLPs with sine activations, and limits reconstruction quality. Conditional hypernetworks achieves similar goals. A hypernetwork estimates all the parameters of a hyponetwork from the latent code. Hypernetworks are prohibitively expensive, in both compute and memory, thereby limiting the resolution of the reconstructed signal in practice. We propose a new approach to modulate the implicit function based on the conditioning variable, which is inspired from attention mechanisms .
Local methods have been used largely to process complex systems in the form of KD-trees for real-time fluid simulation , as regular grids for photon mapping and popularly for fast ray tracing . Local representations have also been used to compress surface light fields and for pre-computed radiance transfer . Our method also relates to recent work on voxelized implicit models for 3D representation . We show that our approach is general, and can be used for a variety of applications.
Method
Concretely, we define our functional representation as a continuous conditional mapping.
The synthesis network defines a continuous function from the spatial coordinates of a discrete signal like an image to its output domain (e.g. color). It is a composition of hidden layers, with hidden features . Each layer uses a periodic-nonlinear activation function and is defined recursively as:
1.2 Modulation Network
We modulate the activations of the synthesis network with a second MLP using ReLU activations, which acts on the latent code corresponding to the target signal. It is defined recursively as:
where and , are the weights and biases of the MLP. We establish an explicit relationship between and , by feeding in at every layer in the modulation network in the form of a skip connection. As can be seen from Equations 4 and 2, the latent codes can modulate the amplitude of the sine activations of each hidden layer in the synthesis network, through the modulation parameters Furthermore, expanding Equation 2, we get
which shows the latent codes also indirectly control the frequency and phase shift of the sinusoids in subsequent layers. Figure 4 illustrates the expressivity of our modulation mechanism visually.
2 Expressivity of Modulation
A simpler alternative to using a separate modulation network would be to concatenate the latent codes and the input coordinates, and use the resultant vector as a single input to the synthesis network. This strategy has shown to be fruitful for ReLU-based synthesis networks for encoding signed distance fields . However, we find it consistently fails with sine activations (see Section 4.1 for details).
In this alternative conditioning mechanism, the network takes the concatenation as input, so the first layer can be rewritten as:
where and are submatrices of , corresponding to and respectively. The latent codes can therefore only act as a phase shift, , on the first layer. This severely limits expressivity, in contrast to our model, where the latent code modulates the amplitude, frequency and phase-shift of the functional representation at all layers, via the . Figure 4 illustrates the difference in expressivity between our model and a concatenation-based MLP.
3 Local Functional Representations
The ability to generalize to more than one signal gives us an additional opportunity. Rather than computing a single neural function for the entire signal, we decompose the domain into a regular grid, and calculate a local continuous representation for each tile (illustrated in Figure 5). Concretely, we assign each tile a latent code , so that the entire signal is represented by a codebook , and the corresponding neural functions , whose first argument is the normalized local coordinates in the tile .
In practice, to eliminate visual discontinuities at the tile boundaries, the images are split into a set of overlapping tiles. When evaluating the continuous representation, the contribution of overlapping tiles is weighted -linearly according to the distance between the point and the tile centers (Fig. 5).
4 Training procedure
We present two modes of training of our model. In the auto-encoder setting, (§ 3.4.1), the latent codes are estimated using a discrete encoder. In the auto-decoder configuration (§ 3.4.2), the latent codes are randomly initialized and optimized with the network parameters as in .
Unless otherwise specified, we use our model in an auto-encoder configuration. Auto-encoding lets us to build a continuous representation, from discrete input signals, using an auxiliary encoder network (shown in Figure 1). This could be useful in spatial super-resolution (images, videos), frame interpolation (videos), or reconstruction problems from sparse samples (lightfields, compression).
4.2 Auto-decoder configuration
In the auto-decoder configuration, we jointly optimize the network parameters , and the latent codes for all the training signals. That is, we do not use the optional encoder of Figure 1. We use this configuration in our shape reconstruction experiments, as proposed by . After training, we obtain a functional representation for new, unseen test signals by sampling a new latent code for the unseen signal, and optimizing it with the same objective used during training, but this time keeping the network parameters constant. We initialize all latent codes as Gaussian random vectors with .
Experiments
We demonstrate two classes of experiments, on three domains (images, videos, 3D shapes). First, we demonstrate the generalization capabilities of the proposed model (§ 4.1) in a global setting. That is, we compute functional representations for many discrete signals, each of which represented (as a whole) by a latent code (i.e., without the tiling procedure described in Section 3.3). Second, we show how our model can be used to learn local functional representations of discrete signals (§ 4.2), with high reconstruction quality. In this set of experiments, each signal is defined using a latent codebook (one code per tile of the input signal). We also show our model can be applied to other multi-domain tasks, such as image relighting (§ 4.3), where the function’s input is a D pixel coordinate and a D lighting direction.
We compare to state-of-the-art MLP-based functional baselines, illustrated in Figure 6, together with our model. These are:
ReLU/FFN a standard MLP with two inputs—latent code and sample coordinates. In case of FFN , the sample coordinates are transformed using a random fourier gaussian matrix with scale .
SIREN+ A single MLP with sine activations adapted from , with an additional input for the latent code , concatenated with the coordinates .
HyperNet-SIREN A SIREN whose weights are conditionally predicted using a hypernetwork , as described in . Instead of a modulator sub-network, a hypernetwork takes the latent code as input, and predicts all the parameters of the synthesis network (the , ).
We run our image experiments on the CelebA and CIFAR-10 datasets separately. We use our model in an auto-encoder setting (§ 3.4.1). A convolutional encoder estimates a latent code for each image. From the latent codes, we decode a functional representation for each image. All the images are resampled to resolution for training.
The parameters of the modulator, synthesizer and encoder are trained simultaneously to minimize the sum of a reconstruction loss,
We train all the models for 1000 epochs. Our train/test splits contain 167K/33K and 60K/10K images respectively. We use center-crops for CelebA and entire image for CIFAR-10 as ground truh. The images are resampled to and using bicubic sampling.
We evaluate generalization by sampling at pixel centers, at the input image resolution () and computing the PSNR. Additionally, we evaluate continuity by sampling more finely, at () resolution, and comparing to ground-truth resampled to the same resolution; we do not train the models with these higher resolution targets. Table 1, summarizes our result on the CelebA dataset, and Table 2 on CIFAR-10 . SIREN+ struggles with generalization, and in the case of CIFAR-10, does not even converge. We hypothesize this is due to the higher image variability in CIFAR-10, in comparison to CelebA where faces are aligned. We observe a similar behavior with HyperNet-SIREN. As shown in Tables 1 and 2, the scaling parameter of FFN is critical for continuity: PSNR drops with the recommended value. Since in the case of HyperNet-SIREN, the last layer predicts all the parameters of the hyponetwork, it makes the last layer highly over-parameterized. This leads to slow training, unstable convergence and inefficient memory usage. We show reconstructions on CelebA test images in Figure 7.
Generative modeling of 3D shapes has recently been driven by implicit neural representations trained to regress a shape’s signed distance field (SDF), by sampling discrete locations in the 3D space. The shape can be reconstructed from the learned SDF using sphere tracing or marching cubes. We show that our model is a powerful replacement for the conditional ReLU-MLPs typically used for this application; it can encode SDFs more accurately. For this experiment, we sample K points for each shape in the cars category of ShapeNet . Half these points are sampled close to the surface, the remaining are randomly sampled inside the unit sphere encompassing the shapes . We use a similar training objective as in case of images (Eq. 7), but we repalce the loss with an penalty in the fidelity term. Following , all conditional models arer trained in as auto-decoder for this experiment (§ 3.4.2). Table 3 shows quantitative comparisons in terms of bi-directional Chamfer distance, computed between the ground-truth shapes and the reconstructions. We show renderings in Figure 8. Compared to DeepSDF , we produce higher-quality reconstructions, with finer details. As for images, we found SIREN+ does not converge. In this comparison, we do not include recent improvements that are orthogonal to our contribution, e.g., improvements to the spatial sampling , training procedure or loss functions These improvements would benefit our method as well as the baselines.
For videos, we train our model on 90K videos from the Vimeo-90k septuplet dataset . Each video is frames long and has a spatial resolution of . During training, we randomly crop tiles from the videos. We use a 3D convolutional encoder to predict the latent codes. At test time, the videos are structured in a grid and each tile is reconstructed with the estimated latent code. The train-test split is used as provided in the dataset. We show quantitative comparisons in Table 4.
2 Local Functional Representations
Our dual-MLP model can also be used to generalize to high-resolution implicit functions. We train our model on images from Div2K . Each image has a long-side resolution of K and split into overlapping tiles. We found tile size to provide good reconstruction accuracy as well as good interpolation properties. The total number of tiles in the training set is M. Each tile is encoded using our method as an auto-encoder. At test time we sample unseen images, and encode them using the trained model. Since other baselines do not generalize, we train a separate MLP for each of the images individually for global methods (i.e. ReLU, SIREN, FFN). Reconstruction PSNR at resolution is reported in Table 5. Additionally, we perform an ablation on our model by using a standard ReLU MLP with our local parameterization.
For this experiment, we collect 18 high-resolution (2M triangles) shapes from the the ThreedScans project . These shapes are split into two Scenes A and B, each of which have 9 shapes. We compute a ground truth signed-distance-field (SDF) as in the global shapes experiment (§ 4.1) for supervision. We use our model in an auto-decoder configuration. We train it on voxels extracted from Scene A. and we evaluate reconstruction accuracy on Scene B, where we only optimize the latent codes. For our global baselines, we overfit the models individually for each scene. We extract meshes from the learned neural SDFs using marching cubes , and report the chamfer-distance from the ground truth to the reconstructions in Table 6. In Figure 10 we show Scene A renders using all the baselines and our method.
Similar to experiments shown in , we encode high-resolution videos using our method. We use videos (pexels.com) with resolution and downsample them to . Each video is split into a grid of tiles for the local model. The reconstruction PSNR is reported in Table 7. Our model is pre-trained on Vimeo-90k (§ 4.1) and tested on the collected videos. It is able to achieve similar reconstruction accuracy as to the one obtained by previous methods while being 1000 faster. We found that SIRENs struggle to reconstruct high-frequency content for complex and varied video signals, in both the spatial and time dimensions as shown in Figure 9.
3 Image-based Relighting
Conclusion
We propose a novel method for representing signals using multi-layer perceptrons (MLPs). We show that partitioning the signal domain into tiles simplifies the signal locally. This leads to representing images, videos and shapes using MLPs with high-quality reconstructions. MLPs with ReLU activations fail to reconstruct high-frequency components of the signals. Instead, we use sine activations which we show to work with a wider frequency spectrum. Using local models requires MLPs to be conditioned on latent codes. We show that concatenating latent codes with the input hinders expressivity. Our method uses a dual-MLP architecture instead. The proposed model also generalizes to multiple instances of these signals. Our local parameterization is general enough to be applied in other applications that use implicit neural functions . We merge local functions using -linear blending which mitigates perceptual discontinuities for the tasks that we explored; however, it is still unclear if that strategy can be applied to other function domains.
Acknowledgements
This work was funded in part by ONR grant N000142012529, ONR grant N000141912293, NSF-Chase CI, NSF CAREER 1751365 and Adobe.